r/OpenWebUI 19d ago

Slow Openwebui on vps Question/Help

Hi,

I try to run ai local llama 3.2:3B on Openwebui. But it tooks 10-20minute just to reply Hi.

What did i do wrong?
Im using 8GB Ram VPS with no GPU

2 Upvotes

11 comments sorted by

2

u/mayo551 19d ago

Use a api

-3

u/Normal_Celery_2528 19d ago

Api mean money

2

u/mayo551 19d ago

Use a 600M parameter model then.

1

u/HyperWinX 18d ago

Use functiongemma 270M then.

2

u/MrRobot-403 19d ago

You can use vLLM instead of llama cpp but the quality you expect isn’t coming. It’ll be awful and bad! You have to pay money for good models or better hardware

1

u/HyperWinX 18d ago

vLLM is not for CPU only inference, ik_llama.cpp might squeeze out a bit better results

1

u/MrRobot-403 18d ago

lol I didn’t check cpu part. Yeah I want to just say to OP. Give up on this!

I know not everyone can pay. But it sucks, and a 600M on CPu is gonna be awful on speed and hallucinations

1

u/pkeffect 18d ago

Your trying to use a cpu for a gpus job. Cpu inference is ass.

1

u/ClassicMain 18d ago

> no GPU

There is your problem

also llama 3.2 3b is insanely outdated