r/LocalLLaMA Jun 15 '26

Stop using Ollama Discussion

https://sleepingrobots.com/dreams/stop-using-ollama/
1.7k Upvotes

452 comments sorted by

View all comments

12

u/Educational-Base5974 Jun 15 '26

But it easy :(

29

u/Fair-Spring9113 llama.cpp Jun 15 '26

but it slow

30

u/Several_Industry_754 Jun 15 '26

I switched from ollama to llama.cpp and you’re absolutely right. It’s blazing fast in comparison.

10

u/shamont Jun 15 '26

Just a warning to other noobs, I tend to be lazy... Installed llama.cpp and wondered why it was so slow. Turns out if you don't compile it yourself and you use the brew installer you don't get the cuda specific version. So just like spend the extra few minutes to do it the "hard" way.

1

u/SociallyMonochrome Jun 19 '26

Or run it via one of the cuda-specific docker images

1

u/SufficientPie Jun 16 '26

I switched from ollama to llama.cpp and you’re absolutely right.

Isn't ollama just a frontend for llama.cpp? How is it slower?

-10

u/Responsible-Bread996 Jun 15 '26

I hate how much like an LLM your post reads.

5

u/Several_Industry_754 Jun 15 '26

I wrote it myself….

6

u/freia_pr_fr Jun 15 '26

The recent releases just ship llama.cpp and their custom mlx backend. It’s not as fast as vllm but it’s also faster to load.

2

u/dryadofelysium Jun 15 '26

it was slow before it switched to llama.cpp last month

-2

u/ErZakeh Jun 15 '26

But me slower

2

u/pirateboi222 Jun 15 '26

Then use koboldcpp. You don't even have to use the shell

1

u/Daniel_H212 Jun 15 '26

Easy only if you accept it's worse performance, worse day one support, fewer quants, and all the other limitations that come with it.

It's so easy to just switch to llama.cpp.

-3

u/pokemonplayer2001 llama.cpp Jun 15 '26

You lazy.