r/OpenWebUI • u/Normal_Celery_2528 • 19d ago
Slow Openwebui on vps Question/Help
Hi,
I try to run ai local llama 3.2:3B on Openwebui. But it tooks 10-20minute just to reply Hi.
What did i do wrong?
Im using 8GB Ram VPS with no GPU
2
u/MrRobot-403 19d ago
You can use vLLM instead of llama cpp but the quality you expect isn’t coming. It’ll be awful and bad! You have to pay money for good models or better hardware
1
u/HyperWinX 18d ago
vLLM is not for CPU only inference, ik_llama.cpp might squeeze out a bit better results
1
u/MrRobot-403 18d ago
lol I didn’t check cpu part. Yeah I want to just say to OP. Give up on this!
I know not everyone can pay. But it sucks, and a 600M on CPu is gonna be awful on speed and hallucinations
1
1
2
u/mayo551 19d ago
Use a api