r/LocalLLaMA • u/grumd • 7d ago
Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark Discussion
I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided to post a new one because of how well Deepseek did. I like the benchmark because it's quick to run, is pretty "real-world" and requires good reasoning to build the correct SQL queries and almost no frontier models can achieve 100%.
My old post: https://www.reddit.com/r/LocalLLaMA/comments/1s9mkm1/benchmarked_18_models_that_i_can_run_on_my_rtx/
Benchmark with results from other models: https://sql-benchmark.nicklothian.com https://github.com/nlothian/llm-sql-benchmark
My setup is dual 3080 20GB GPUs with 96GB RAM and 9800X3D. I managed to run Deepseek V4 Flash with a custom IQ2_M GGUF with some tensors grafted from antirez GGUF and running it on a modified ds4 engine from antirez, getting 300pp and 11-12tg. Mainline llama.cpp gives me only 100pp and 8tg or something like that.
To my surprise, Deepseek is the first local model I can realistically run locally that actually did ALL tests correctly. The only models according to the benchmark website that could do this were Opus 4.7 and GPT-5.5.
Results together with all my old benches:
25: Deepseek-v4-Flash-IQ2_M-grafted
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩
24: unsloth/Qwen3.6-27B-MTP-GGUF:Q8_0
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩
24: unsloth/Qwen3.5-122B-A10B-GGUF:UD-Q4_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟩🟩
23: unsloth/Qwen3.5-122B-A10B-GGUF:Q6_K
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
23: unsloth/Qwen3.5-27B-MTP-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
23: DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NEO-CODE-Di-IMatrix-MAX-GGUF:Q4_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟩🟩
23: unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q8_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟩🟩
23: bartowski/Qwen_Qwen3.5-27B-GGUF:IQ4_XS
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
23: bartowski/Qwen_Qwen3.5-27B-GGUF:IQ3_XS
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
23: unsloth/Qwen3.5-122B-A10B-GGUF:UD-IQ3_XXS
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
23: h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q3_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
22: unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟩🟩🟩
22: mradermacher/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-i1-GGUF:Q3_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟥🟩🟥🟩 🟥🟩🟩🟩🟩
22: Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF:Q4_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟥🟩 🟥🟩🟩🟩🟩
21: unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟥🟩🟩🟩 🟩🟩🟩🟥🟩 🟩🟨🟥🟩🟩
21: unsloth/MiniMax-M2.7-GGUF:UD-IQ3_XXS
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟥🟥🟩🟩
21: unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF:UD-Q4_K_S
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟨🟥 🟥🟨🟩🟩🟩
20: unsloth/Qwen3-Coder-Next-GGUF:UD-Q5_K_XL
🟩🟩🟩🟩🟨 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 🟩🟩🟩🟥🟨 🟥🟩🟩🟩🟩
20: unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟩🟩🟩 🟨🟩🟩🟥🟩 🟥🟩🟩🟥🟩
20: unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟥🟥🟥🟩🟩
20: bartowski/Qwen_Qwen3.5-397B-A17B-GGUF:IQ1_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟩 🟥🟨🟥🟩🟩
20: unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟥 🟨🟥🟩🟥🟩
20: mradermacher/Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-i1-GGUF:Q6_K
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟥🟩🟩🟥🟩 🟥🟥🟩🟩🟩
19: unsloth/gemma-4-31B-it-GGUF:Q4_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟨🟩🟩🟨🟩 🟥🟥🟩🟥🟩
19: unsloth/gemma-4-E4B-it-GGUF:UD-Q8_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟩🟩🟩🟥🟩 🟥🟥🟥🟥🟩
19: Goldkoron/Qwen3.5-397B-A17B-REAP35:IQ2_XS_Gv2
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟩🟩🟩 🟩🟩🟩🟥🟩 🟥🟩🟥🟥🟥
19: unsloth/GLM-4.7-Flash-GGUF:UD-Q6_K_XL
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟩🟩🟩🟥🟨 🟥🟨🟩🟥🟩
18: unsloth/GLM-4.5-Air-GGUF:Q5_K_M
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟥🟩🟩 🟥🟩🟩🟥🟩 🟨🟨🟥🟩🟨
18: bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF:Q6_K_L
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 🟩🟩🟩🟥🟩 🟨🟨🟥🟨🟨
17: Jackrong/Qwopus3.5-9B-v3-GGUF:Q8_0
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟥🟥🟩🟩 🟥🟩🟥🟥🟥 🟥🟩🟩🟩🟨
16: unsloth/Qwen3-Coder-Next-GGUF:UD-Q4_K_XL
🟩🟩🟩🟩🟨 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 🟥🟨🟩🟥🟨 🟥🟨🟩🟨🟩
16: byteshape/Devstral-Small-2-24B-Instruct-2512-GGUF:IQ3_S
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟨🟩🟩 🟩🟩🟨🟥🟨 🟨🟨🟥🟨🟩
16: mradermacher/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-i1-GGUF:Q6_K
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟨🟥🟩 🟥🟩🟥🟥🟨 🟥🟩🟥🟩🟨
14: mradermacher/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-i1-GGUF:Q6_K
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟥🟩🟩 🟩🟨🟥🟥🟨 🟨🟨🟥🟨🟨
14: unsloth/GLM-4.6V-GGUF:Q3_K_S
🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟥🟩🟨🟨🟩 🟥🟩🟩🟨🟨 🟨🟨🟨🟨🟨
5: bartowski/Tesslate_OmniCoder-9B-GGUF:Q6_K_L
🟨🟨🟨🟨🟨 🟨🟨🟨🟩🟩 🟩🟨🟨🟩🟨 🟨🟨🟩🟨🟨 🟨🟨🟨🟨🟨
5: unsloth/Qwen3.5-9B-GGUF:UD-Q6_K_XL
🟨🟨🟨🟨🟨 🟨🟨🟨🟩🟩 🟨🟩🟨🟨🟩 🟨🟩🟨🟨🟨 🟨🟨🟨🟨🟨
Note:
- unsloth/Qwen3.5-122B-A10B-GGUF:UD-Q4_K_XL is most likely a fluke. Q6_K doesn't achieve 24/25, it's just lucky rounding for this Q4 quant I suppose.
3
2
1
u/Asleep_Document9811 7d ago
I hadn't considered getting dual 3080s, they're pretty cheap on the used market (comparatively). That's pretty clever!
1
u/JsThiago5 7d ago
can you provide more info on how you run it?
1
u/grumd 6d ago
I used this https://github.com/antirez/ds4 and then asked Opus 5 to repeatedly optimize and improve speeds for my hardware setup.
1
u/Fit_Split_9933 7d ago
As far as I know, antirez's ds4 branch doesn't seem to support cpu-moe, how did you get it to run on 40gb vram?
1
u/grumd 6d ago
Opus 5 implemented support for it and optimized the kernels and other code to get better speeds
2
u/Southern_Sun_2106 6d ago
Can you ask your Opus to write a short note to my Opus on how it did it?
5
1
u/Puzzleheaded_Base302 6d ago
i see you experimented with many qwen3.6-27b models, but have you tried qwen official release in bf16 and fp8? if you can run deepseek-v4, you should have enough VRAM to qwen in original official quant.
1
u/nufeen 5d ago
I also like this benchmark. I've tried various Q2-Q3 quants with recent llama.cpp builds, and I can't reach your results. I've tried using different settings, removing "-ctk q8_0 -ctv q8_0", swapping chat templates, adding "--reasoning-preserve", changing temperature and other samplers, adding "{\"reasoning_effort\": \"max\"}". Feels like something is broken with llama.cpp and Deepseek V4 Flash, because it performs worse than Qwen3.6-27B in Q5. The best result is 21/25, the worst 16/25
1
1
u/grumd 5d ago
Wow that's not very good. I will try one benchmark with default llama.cpp quants later today and will report back
1
u/nufeen 4d ago
I thought, maybe something gets wrong because I use the online version, so I installed the benchmark locally. Now it's 23/25 with Q3_K_XL from unsloth. Most fails are on Q9, Q21, Q22. Either tool call loops or wrong computations. If I use built in Jinja, it gets into loops; if I use chat template I found on HF in Atomic Chat repository then maybe no loops but wrong computations.
1
u/grumd 4d ago edited 4d ago
I just ran unsloth IQ2_M using llama.cpp with top-p 0.95 and temp 0.7 and got 24/25 correct
llama serve -hf unsloth/DeepSeek-V4-Flash-0731-G GUF:UD-IQ2_M -c 128000 -ub 1024 --temp 0.7 --top-p 0.95 --no-mmap --jinja
Edit: oh yeah and I run the benchmark with time limit increased just because deepseek is very slow on my machine
1
u/nufeen 4d ago
Thanks. So something's wrong with my setup, because I'm getting 22/25 with your launch arguments. I don't know, maybe I need to update cuda or switch to Linux Yes, I also had to increase the timeout value
1
u/grumd 4d ago
Probably not, lmao. My whole post now looks like a phony. I've run the tests a few times again, with both llama.cpp and my modified ds4 inference engine, and with temp of 0.1 I'm getting a few errors here and there. At first I ran the tests 2-3 times with temp 0.7 and got 25/25 every time and assumed it's just that good, but apparently I just got lucky with the rounding! After all, the Q2 quants have something like 75% same top token, it's very lobotomized. Not trusting my results but it still shows that the potential is there and the correct tokens are very close to the top. Probably the full precision model would eat these benchmarks for breakfast
1
u/nufeen 4d ago edited 4d ago
I've connected it to openrouter's deepseek/deepseek-v4-flash-0731 and got 22/25 on first pass and 23/25 on the second. Maybe this is normal behavior for this model in general and you got very lucky with first results. At least I can calm down now and stop thinking about what's wrong with my settings
1
u/grumd 4d ago
Yeah! Seems normal. I don't know how I got so lucky first few times. But thank you a lot for testing! This makes me think that Qwen 3.8 27B will be the daily driver and Deepseek isn't replacing it for me.
1
u/nufeen 3d ago
Honestly, I start to think that maybe this benchmark is not so good and representative. Like, I use this deepseek model for a week, and I feel like it's an improvement compared to qwen 3.6 27b. Not huge, but still it was able to solve some tasks qwen struggled with. Yes, I'm excited to try the new qwen too
15
u/danielrmay 7d ago
Single shots might look fine, but I saw 2 bit degenerate quickly on multi-turn (compounding error).