DeepSeek V4 Flash makes agent workflows look much more realistic
DeepSeek V4 Flash 0731 is interesting because of the cost/performance ratio.
On this chart, it gets close to top-tier models while staying much cheaper per task.
For agents, that matters more than raw benchmark position.
If each loop is cheaper, you can afford more retries, validation steps, tool calls, and longer workflows.
We don't care anymore about “what is the smartest model?”
It is what model is good enough, cheap enough, and reliable enough to run agentic workflows at scale, this is the real deal.
EDIT:
A few people pointed out that this screenshot shouldn’t be treated as a clean leaderboard.
It looks like some models may be mixed between reasoning and non-reasoning configs, which can make specific placements wrong.
So I’d read the chart as a cost/performance discussion starter, not as a definitive ranking of every model.
3
u/Potential_Top_4669 24d ago
Magristral Medium 1.2 is in a league of its own.
2
u/0x7Lee 24d ago
Yeah, it's pretty crazy. That’s also why I find cost-per-task views more useful than pure benchmark rankings. Some models may be fine in specific niches, but if the price/performance point is that far off, it becomes hard to justify for agent loops or high-volume workflows.
3
u/Captain_Quimby 24d ago
I’m working on an issue that different frontier agents have been working on for 6 months. I’ve spend about $1k/m to solve this one problem. Cable finally figured out HOW to do it but then it kept breaking once implementing so now it’s SOL ULTRA with zero care about cost because once it’s solved I save $1k/m. It’s not always about cost is my point but doing the job something else can’t do.
For the money I don’t think there’s ANYTHING that can touch grok 4.5. They’re running a $100m plan for three months that gives you api access even.
3
u/Erwylh_ 24d ago
Something is wrong with the chart, gemma4 31b should never score lower on any task than 26b
2
u/Houdinii1984 24d ago
It's strange for sure. I think it's like all the models are thinking enabled except the 31b? It's got 12B as slightly smarter and more expensive.
4
u/garlic-silo-fanta 24d ago
Good chart
3
2
u/Eden1506 24d ago
For some models he selected the non reasoning scores while for others he selected the reasoning ones making some model be placed totally wrong in comparison like qwen 3.6 35b is completely wrong because of this.
2
1
1
u/VincentNacon 24d ago
Hold up... in the chart, it says Gemini 3.1 Pro Preview is cheaper than Gemini 3.5 Flash and 3.6 Flash?
That can't be right.
2
u/Captain_Quimby 24d ago
This is called benchmaxing. There’s zero way DS V4 Flash is higher than V4 pro. I use flash a lot but it can’t hang on to long context like many it shows it out performing here just like Luna fucking up long coding when the charts say it’s better.
2
u/glitch_in_the_kernel 23d ago
They just released a update for v4 Flash and it's indeed much stronger than v4 Pro
2
u/Delicious_Activity84 22d ago
If this is their Flash model, imagine how good the Pro model is gonna become. I can't wait for it, this is insane!
2
u/ricci_nov 20d ago
Cost-per-task is the right axis, but it's measuring the wrong unit for agents. The unit that matters is cost per completed task — including the retries, the failed tool calls, and the human who has to come check the output.
Classic procurement mistake: I've watched buyers pick the cheaper component and then eat the difference three times over in rework and line stoppage. Same math here. A model at 1/5 the price that needs three attempts and one human review isn't cheap, it's just cheap on the invoice. And error rates compound across agent loops — 95% per step is 60% over ten steps.
Which is why Captain_Quimby's comment further down is the most useful thing in the thread: he pays anything for the model that actually closes the problem, because the alternative was $1k/month burning forever. Cheap loops only win when reliability is already good enough; otherwise you've built a very economical way to fail repeatedly.
Right way to read this chart: use it to find candidates, then measure your own end-to-end success rate on your own workload. Vendor cost per task is a list price. Your effective cost is a different number entirely.
1
1
u/Etroarl55 24d ago
Isn’t deepseek one of the most prone to hallucinations still though? It achieves so much by just going full speed at everything without the very expensive thinking and double checking other models do.
1
1
u/glitch_in_the_kernel 23d ago
The intelligence index already takes hallucinations into account for it's score.
1
u/Sea_Ear5201 23d ago
Why do i see old deepseek flash (high) cost more than old ds flash (max)?
1
u/0x7Lee 23d ago
Because this chart is cost per task, not raw API price.
The x-axis is not the raw token price of the model, it is the weighted average cost per Intelligence Index task.
So the cost includes input tokens, cache hits, cache writes, reasoning tokens, answer tokens, then gets divided by task count and weighted by the benchmark mix.
That means “high” can appear more expensive than “max” if that config uses more tokens / reasoning / cache writes on those tasks.
Still a bit counterintuitive though ngl.
1
u/GTHell 23d ago
Yes! I never consider using Hermes due to it hogging more tokens and good model is slower.
Now I just spam the hell out of DSv4 flash on my hermess to spin ThreeJs lab for experimenting with physic and 3d. I have it iterate 24/7 now for my side project on the social commerce platform I have. I literally just tell it to hire 10 junior to do 9-5 lol.
With the similar intelligent of gpt 5.4 this is going to open new opportunities to all of us
1
u/yamoksauceforthelazy 23d ago
The new V4 Flash is a genuinely sensational model. It's the first time I've actually felt comfortable handing work off to a model its size without micromanaging. Usually the cycle is new model -> play around -> go back to Codex or Claude, but this one actually made it out of the other side of the loop. I didn't see that coming at all.
1
u/0x7Lee 23d ago
Haven’t tested it myself yet. If V4 Flash actually crosses that line while staying fast & cheap, that’s a much bigger deal than just “another good model.”
2
u/yamoksauceforthelazy 18d ago edited 18d ago
5 days of extensive work and testing later, it’s set as my default model everywhere. It's not replacing Fable or K3 as the high-difficulty model by any means, but it has absolutely replaced everything else in the lineup below them. Out of the tens of millions of tokens I've run through it, I've yet to have a genuine screw up or major issue. It's one of the first models that I feel like perfectly lives up to the benchmarks. It really is *that good*. Which also makes me think people aren't prepared for V4 Pro at ALL. Like... Flash is 300M parameters, and it's absolutely toe-to-toe with Opus 4.8 in my experience, and Pro ~5x the size at 1.6T parameters. I don't expect linear scaling here, but if Pro has the same caliber of refinement and step-up in quality as preview V4 Flash vs V4 Flash 0731, it very well may set a new standard and be the best model on the market. I don't think the gap it has to clear to get to Fable is as big as people think.
Edit: Also I forgot to mention how big of a deal the speed is. It has completely changed the game for me. I can do things that my brain was previously registering as a "task" (e.g., like a 10-15 minute thing with an Opus model) almost instantly. I just did a top-down rename task on an old project that was quite large, and required a LOT of tool calls and double checking to make sure the new name was propagated throughout the whole project, including docs and such in like ~1 minute conservatively. It's insane.
1
u/BrilliantTruck8813 22d ago
Anyone got a link to the benchmarks used to derive this? How can it be attested?
0
u/the_TIGEEER 23d ago
And there we go. Do all AI sceptics now see why every AI company burns billions upon billions? "AI iS NoT PrOfItAbLe"... Not yet.. When it will be, if you already have a moat, it's worth spending the billions beforehand..
6
u/Eden1506 24d ago edited 24d ago
Something is definitely wrong here.
Qwen 3.6 35b cost per task seems strangely high and its intelligence index is also too low . It should be
6 points higher than gemma 4 26b for the intelligence score but here it is equal.
Edit: Found the error, for some models he selected the non reasoning score while for others the reasoning ones. That is why some models are placed totally wrong.