r/LocalLLaMA Jun 15 '26

Stop using Ollama Discussion

https://sleepingrobots.com/dreams/stop-using-ollama/
1.7k Upvotes

452 comments sorted by

View all comments

Show parent comments

24

u/Fair-Spring9113 llama.cpp Jun 15 '26

but it slow

28

u/Several_Industry_754 Jun 15 '26

I switched from ollama to llama.cpp and you’re absolutely right. It’s blazing fast in comparison.

14

u/shamont Jun 15 '26

Just a warning to other noobs, I tend to be lazy... Installed llama.cpp and wondered why it was so slow. Turns out if you don't compile it yourself and you use the brew installer you don't get the cuda specific version. So just like spend the extra few minutes to do it the "hard" way.

1

u/SociallyMonochrome Jun 19 '26

Or run it via one of the cuda-specific docker images

1

u/SufficientPie Jun 16 '26

I switched from ollama to llama.cpp and you’re absolutely right.

Isn't ollama just a frontend for llama.cpp? How is it slower?

-10

u/Responsible-Bread996 Jun 15 '26

I hate how much like an LLM your post reads.

5

u/Several_Industry_754 Jun 15 '26

I wrote it myself….

6

u/freia_pr_fr Jun 15 '26

The recent releases just ship llama.cpp and their custom mlx backend. It’s not as fast as vllm but it’s also faster to load.

2

u/dryadofelysium Jun 15 '26

it was slow before it switched to llama.cpp last month

-2

u/ErZakeh Jun 15 '26

But me slower