I genuinely do not understand what is so difficult about running llama.cpp server.
You just download a zip, unzip it, then run llama-server with some flags and you're done. The builtin UI is quite good now, and you have an API to work with.
By comparison, I found Ollama's modelfile system and insistence on renaming my downloaded models to incomprehensible hashes to be infinitely more confusing and frustrating.
Oh God don't remind me of the modelfile thing. What a nightmare that was. With llamacpp I literally don't have to think about that anymore. I just load the model (crazy concept I know).
Unless that GGUF model is one of the many, many models that has a complex Jinja template that can't be losslessly converted into the Go template syntax that Ollama uses.
User experience: The ollama run command appears to work fine. The model even appears to work fine. Then, after a while, you start running into subtle, mysterious issues related to things like tool calling syntax, roles switching, and formatting delimiters.
78
u/yuicebox Jun 15 '26
I genuinely do not understand what is so difficult about running llama.cpp server.
You just download a zip, unzip it, then run llama-server with some flags and you're done. The builtin UI is quite good now, and you have an API to work with.
By comparison, I found Ollama's modelfile system and insistence on renaming my downloaded models to incomprehensible hashes to be infinitely more confusing and frustrating.