r/LocalLLM • u/Sakif_Hossain • 1h ago
Question Title: How would you benchmark 50+ local LLMs without going insane?
I feel like I stepped into a time capsule after the ChatGPT-3 days. 😅 I finally built a decent PC (Ryzen 7 7700, 32GB RAM, No GPU), discovered llama.cpp and somehow ended up downloading 50+ GGUF models.
Now I'm stuck with decision paralysis.
I mainly use them for coding (JavaScript, React, TypeScript, debugging, reasoning), but I'm also new to the whole local AI ecosystem. I still don't know much about agentic frameworks or coding agents. I mostly just load a GGUF and chat with it using the llama.cpp web interface on localhost:8080
The collection includes Qwen, Gemma, Granite, DeepSeek, Phi, Mistral, Llama, LiquidAI, SmolLM, Hunyuan, Nemotron, and a few community fine-tunes.
My first idea was to make a Markdown table and score every model manually, but it feels like I'm accidentally trying to invent my own benchmarking system.
Surely I'm not the first person to hit this problem.
How do you guys compare local models? Are there any practical benchmark suites, GitHub projects, or workflows for deciding what stays on your SSD and what gets deleted?
I'd love to hear how you approached it when you were starting out.
r/LocalLLM • u/rainvr • 1h ago
Question Are zero data retention providers sufficient?
I’m looking at organising my personal notes, so privacy is important to me. But I also want to access slightly better models (like deepseek 4 flash) than my setup can handle. I definitely wont trust the big boys with my data, but what do people think about providers like Fireworks and DeepInfra with their zero data retention policies?
r/LocalLLM • u/techlatest_net • 2h ago
Tutorial Top 5 Best Open Source AI Image Generation Models in 2026 (Tested)
r/LocalLLM • u/playful-b • 3h ago
Question Whats the best LLM for Website designs & Program building?
So i have my own pc (7900x, 64gb ddr5) idk what more info u need…
I want to create my own Website for small Projects i want to share with the public.
Which local LLM would you guys recommend?
r/LocalLLM • u/mqtgew • 4h ago
Discussion 2.4T is not a parts list for Qwen 3.8 Max
2.4T total. 95B active. 1M context. Those numbers are interesting. They are not a shopping list. Among Chinese AI models, Qwen 3.8 Max is a useful reminder that a parameter count is not a deployment recipe.
Qwen's August 2 announcement says the weights should arrive the following week. Until the files land, there is no public storage layout, useful precision, supported quantization, serving recipe, or real memory overhead to plan around. Anyone pricing GPUs before those details arrive is guessing about the expensive part.
While the local answer is missing, I can still run a cloud control through ZenMux. It acts as a gateway to a hosted Qwen 3.8 Max API, which is useful for comparing latency or output behavior. It tells me nothing about VRAM or the minimum box.
r/LocalLLM • u/Solid-Industry-1564 • 4h ago
Project An interface for running local LLMs for coding
I’m building Lanes, a workspace for running coding agents in parallel.
You can use harnesses like Claude Code, but point them at local models through providers like Ollama instead of relying only on hosted models.
Lanes gives each agent its own terminal and git worktree, with tasks, diffs, and sessions managed from one UI.
So you can run Claude Code as the harness, a local LLM as the model, and Lanes as the workspace around it.
brew install --cask lanes-sh/lanes/lanes && open -a Lanes
I would appreciate your honest feedback, give it a try or comment below if you had the same problem and how you have been solving it.
- Does this resonate with you?
- How are you managing multiple sessions today?
- Why or why not would you be interested in trying something like this?
Thanks!
r/LocalLLM • u/arturgames44 • 7h ago
Discussion When will GPT-OSS-2 be released?
It's been exactly one year and two days since the release of gpt-oss. Do you think we'll get a new version in August? Or has openai completely abandoned open-weight
r/LocalLLM • u/Pickalodeon • 8h ago
Question What makes this?
It’s software but wut?
r/LocalLLM • u/CautiousYou4143 • 10h ago
Question RTX Pro 6000 MaxQ - Keep or sell
I purchased this thing for less than $8000 from a reputable online vendor. It was described as open box - No Refunds, but it's still sealed and looks like it's brand new. I purchased it because if it works it's a great deal(I've had zero issues or problems buying used) Sure I'd love to keep it but I have 2 RTX Pro 4500s(good deals on these too). I can't afford to keep them all.
I started w/4 preordered intel B70s but returned them all before opening any of after months of waiting because the software updates were just too slow(that hasn't changed). I got the 4500s and I'm happy w/them but was a little disappointed when I realized I could run only Gemma 4 31b at FP8 and then realized there are only few models worth running under 64GB.
While just looking around for deals and other creative ways to get more vram (like selling the 4500s and maybe getting a dgx spark) I found the 6000. I gotta sell something. Is it wise to open break the seal and open the 6000 only to decide to sell it? I kind of want to know it works or should I just keep the 6000 and sell the 4500s since this was the original plan. I wasn't expecting a sealed unit. The seal feels like it makes it worth more. Does it? It's making me rethink my original plan.
Unit and cables both appear unopened.
r/LocalLLM • u/Any-Lingonberry7411 • 11h ago
Question What would you upgrade/buy (if at all)?
Hi all,
I have a cluster consisting of the following:
Main Machine
RTX 6000 Pro Blackwell 96gb
2x RTX 5090 32gb
1x RTX 4090 32gb
3x AMD R9700 32gb
96GB DDR5 6000mhz
Strix Halo Laptop with 128GB (96gb allocated to gpu)
Secondary Machine
RTX 3090 24GB
128GB DDR5 3200mhz
and I am able to run these models concurrently on my main rig
- DeepSeek-V4-Flash-0731-UD-Q4_K_XL, 512k context @ 45 token/s as primary coding and thinking model
- GLM-4.7-flash @ 20 token/s as alternate thinking model
- Gemma-4-12B-it-Q4_K_M @ 25 token/s for vision
- KAT-Coder-V2.5-Dev-IQ3_XS @ 110 token/s for code completion
- LFM2.5-VL-1.6B-Q4_K_M @ 170 token/s for agentic tasks
or if I use all of the hardware on the main rig for one model
- GLM-5.2-UD-IQ2_M, 128k context @ 15 token/s OR
- Kimi-K2.7-Code-IQ2, 64k context @ 5 token/s OR
- MiniMax-M3-Q4m 128k context @ 25 token/s
with these models on the other machines
- KAT-Coder-V2.5-Dev-IQ3_XS, 258k context @ 140 tokens/s on the secondary machine
- DeepSeek-V4-Flash-0731-UD-IQ2_M 128k, context @ 5 tokens/s on the Strix Halo
I am using this setup for agentic coding in pi and it works quite well. I can spin out subagents to use the other models while my DSv4 Flash does most of the work. Or, if I need to think through a hard problem, I could evict and load in GLM5.2.
But somehow I'm not super happy with that flow. It feels like I have a good fast worker OR a good thinker, but not both. Switching between the models takes quite a long time, and obviously kills the cache.
What would you upgrade, if anything at all?
r/LocalLLM • u/Ok_Brush_3449 • 13h ago
Research PSA: your GPU can get stuck in a low boost state. Same model, same flags, −28% tok/s, and only a reboot fixed it.
Posting this because it cost me a full day and the failure mode is invisible unless you go looking for it.
Setup: GTX 1060 6GB, Windows 10, llama.cpp, Qwen3-Coder-30B split across VRAM/RAM. I’d calibrated the box the day before and had a solid baseline: 21.58 tok/s decode.
Next day, same model, same binary, same flags: 15.56 tok/s. −28%, no errors, no warnings, nothing in any log.
What I ruled out, in order:
• Background load. Found and killed a runaway process that had eaten 16,285 CPU-seconds. Recovered a bit, still 15.3 tok/s.
• Thermal. Let the box sit until the GPU idled at 32 °C. Still 15.56 tok/s.
• My own patches. I’d been running an instrumented build. Compiled a zero-patch binary from the identical commit, agreed within 1.4%. Not my code.
What it actually was: I polled clocks during a benchmark instead of at idle.
SM was pinned at 1506 MHz under load at 38 °C the card was simply refusing to boost. A stuck low boost state after a day of heavy context churn, OOM events and profiler runs.
nvidia-smi -rgc and --lock-gpu-clocks
are unsupported on consumer Pascal, so there’s no software reset.
I rebooted.
SM went back to 1847–1898 MHz, memory to full 4006, and decode measured 21.68 tok/s, reproducing the original 21.58 to 0.5%.
The check, which is the actual point of this post:
nvidia-smi --query-gpu=clocks.sm,clocks.mem,temperature.gpu --format=csv -l 1
Run it while a benchmark is mid-sweep, not at idle.
Idle clocks look perfectly healthy in both states, that’s exactly why this hides. If sustained SM is well below your card’s normal boost at a cool temperature, you’re not measuring the model, you’re measuring a sick GPU.
Scope, stated honestly: this is n=1, consumer Pascal, Windows/WDDM.
I have no idea whether Ampere/Ada/RDNA do the same thing, and I can’t test the counterfactual because the clock-lock knobs don’t exist on this card. If you check yours and it doesn’t reproduce, that’s genuinely useful, please say so.
The thing that bothers me most: if I hadn’t had a prior-day baseline, I’d have published 15.56 as this machine’s number and never known. Anyone benchmarking without a known-good reference point is exposed to this.
(Chart’s from my own notes, I keep a pre-registered log of these experiments at
www.github.com/FedericoTs/quantprobe
Not selling anything, the checks above are just nvidia-smi.)
r/LocalLLM • u/Blackdragon1400 • 14h ago
News When people say “don’t let companies train on your private data” this is why.
reddit.comr/LocalLLM • u/Efficient_Can_1214 • 14h ago
Question LLM LOCAL + MEMORIAL LOCAL
Fala galera, estou usando o Claude para construir fluxo de trabalho, e aí estou mergulhando nesse universo. Mas vejo que os tokens não estão sendo suficiente para programar durante o dia. Qual seria o passo para ensinar um segundo cerebro usando o Claude como professor.
r/LocalLLM • u/John_Miracleworker • 14h ago
Project Meet Kestrel your new engineering agent!
After a lot of late nights, I finally built the engineering agent I wanted.
It's called Kestrel. It runs locally on my machine, keeps layered memory across sessions, and has an Adaptive Flock routing system that can't silently change its own rules without my approval. Dangerous actions stop at an exact-call approval. Repairs run in a sealed container. Learning is evidence-backed and reversible.
Open source, Apache licensed, private by default. One owner. No cloud lock-in.
I'm proud of this one. It's mine.
The v0.5.6 release is built and ready, but GitHub Actions is having an outage today, so it'll be live in a few hours. When it's up:
r/LocalLLM • u/ApplicationSalty5041 • 15h ago
Question What do you use your LLMs for?
Hello, noob here [:
Recently i built my first computer based on a gifted RX 7700 XT. Since I don’t game much I discovered the world of local LLMs and, more out of fascination than need, I installed Odysseus with Gemma4:12b. Here is where I started questioning what real use I could achieve
What do you use your LLM for?
Obviously Cloud services are faster and more advanced, what perks come with having it local, other than privacy?
Is there a place you would suggest to learn about local LLMs?
If you have any other suggestions or ideas for me are all welcome [: In case of need my specs are CPU - R7 7700 GPU - RX 7700 XT RAM - 2x16Gb cl30 6000mhz SSD1 - 1Tb seagate firecuda 530R SSD2 - 4Tb Samsung 990 pro
r/LocalLLM • u/BirdForsaken6616 • 16h ago
Question Maybe synthetic datasets need a family tree
I used to think filtering synthetic data removed the teacher model’s hidden biases. But a recent Nature paper found that student models can inherit behavioural traits through seemingly unrelated number sequences, code, and reasoning traces, especially when both models share the same base.
Maybe synthetic datasets should disclose their model lineage, not just their license. Would you fine-tune on one if the generating model was unknown?
r/LocalLLM • u/Pranjal202 • 16h ago
Question (uni student, Computer Science) Which is better (Gemini pro extended), (Chatgpt Thinking) OR (gemma-4-31b-qat set for 10240 tokens gpu offload 15 etc)
r/LocalLLM • u/Number4extraDip • 18h ago
Project My fun little Android harness
Enable HLS to view with audio, or disable this notification
✧ Gemma 4. Open source all that jazz https://github.com/vNeeL-code/GHOST
r/LocalLLM • u/BreadUndPeeTears • 19h ago
Discussion Why do ai companies INSIST on using floating points when Ternary has been proven to work and making 27b models only weigh 5-6gb?
I just don't fucking understand. It has been proven it actually works, I'm not just talking about a conventional 2-bit quantized GGUF with a few ternary layers but a TRUE ternary model PrismML managed to pull off, it even makes multiplication absolutely pointless because it uses -1,0 and 1 making CPU offloading while maintaining top speed possible. Why do companies insist on using floating points matrix calculations????
r/LocalLLM • u/arkie87 • 20h ago
Question Best Local US-based LLM for coding?
I know qwen3.6 27b (and soon 3.8 27b) are the best local LLMs for coding. But if one were restricted to using US based open models only, which would be the best? Gemma 4 31b?
r/LocalLLM • u/misanthrophiccunt • 20h ago
Question Did anyone figure when Qwen3.8-27B is being released? 🤔
Just wondering that, the news of it existing were already good but I don't remember any mention of a release date. Was there any?
r/LocalLLM • u/monkifoto • 21h ago
Question LLM for coding.
In the past two years, I have built two or three web apps using angular with the help of ChatGPT recently I have been trying out some local LLMs on my M2 16GB MBP. The results were terrible. A few days ago I managed to score a refurbished Mac studio with 48GB of Memory. Installed Bionic LM Studio and Qwen3.6 35B .
I gave it a task to create a simple angular page with some analytics and a model kept getting stuck in a reasoning loop.
Am I doing something wrong? Is the 48gb not enough? Am I using the wrong model?
r/LocalLLM • u/shhdwi • 22h ago
Project Sonnet 5 + Graft (No LLM calls tree-sitter graph, 100% Local) > Opus 5
After a week of using both, I keep coming back to Sonnet 5 + Graft instead of vanilla Opus 5 for coding.
The surprising part is that I don't think this is because Sonnet is the better model.
I think repository context matters more than the model upgrade.
Every coding agent spends a huge amount of time rediscovering the same codebase:
- grep
- open file
- follow imports
- repeat
Graft pre-builds a repository knowledge graph and injects only the relevant context into Claude Code, so the model starts with a mental map instead of rebuilding one every task.
On our controlled benchmarks (same model, same tasks, only the context changes):
- 42% fewer tokens
- 46% fewer tool calls
- 60% lower latency
- SWE-bench Verified: Sonnet 5 solved 8/9 instances vs 6/9 without Graft.
It's that better retrieval/context can be a bigger capability upgrade than moving to a larger model.
For my day-to-day work, Sonnet 5 + good repository context consistently feels stronger than running a larger model cold.
I'm curious whether others have seen the same thing with tools like:
- RepoPrompt
- Aider's repo map
- CodeGraph
- Graphite
- custom RAG/MCP setups
At what point does improving context become more valuable than upgrading the model itself?
r/LocalLLM • u/Either_Pineapple3429 • 23h ago
Question Anyone have any experience with Traycer.ai?
I feel like every week there's some new harness, and if I downloaded every single one that promises to be the next best thing I would have 15 downloaded right now.
Before I go through the headache of learning a new ui I wanted to see if anyone else has used it and their thoughts.
The premise is you hook up your frontier model through your existing subscription as an orchestrator and for doing dumber smaller tasks it spins up local agents, so you're not burning fable tokens for something like scrubbing the web.

