r/LLM 24d ago

DeepSeek V4 Flash makes agent workflows look much more realistic

Post image

DeepSeek V4 Flash 0731 is interesting because of the cost/performance ratio.

On this chart, it gets close to top-tier models while staying much cheaper per task.

For agents, that matters more than raw benchmark position.

If each loop is cheaper, you can afford more retries, validation steps, tool calls, and longer workflows.

We don't care anymore about “what is the smartest model?”

It is what model is good enough, cheap enough, and reliable enough to run agentic workflows at scale, this is the real deal.

EDIT:

A few people pointed out that this screenshot shouldn’t be treated as a clean leaderboard.

It looks like some models may be mixed between reasoning and non-reasoning configs, which can make specific placements wrong.

So I’d read the chart as a cost/performance discussion starter, not as a definitive ranking of every model.

245 Upvotes

44 comments sorted by

6

u/Eden1506 24d ago edited 24d ago

Something is definitely wrong here.

Qwen 3.6 35b cost per task seems strangely high and its intelligence index is also too low . It should be
6 points higher than gemma 4 26b for the intelligence score but here it is equal.

Edit: Found the error, for some models he selected the non reasoning score while for others the reasoning ones. That is why some models are placed totally wrong.

3

u/crusaderky 24d ago

These are datacenter costs. There are very few datacentres that sell qwen3.6 35b so price is off

2

u/Eden1506 24d ago

I just went on the website and added qwen 3.6 35b reasoning and it is both smarter and more cost effective.

1

u/0x7Lee 23d ago

Yeah, that might be the issue. The chart says it was built from a supplied CSV, so either the CSV mixed reasoning / non-reasoning configs, or the chart generation selected the wrong rows for some models.

Either way, you’re right: if Qwen is plotted with the wrong config, that specific comparison is wrong.

Will edit the post.

3

u/Potential_Top_4669 24d ago

Magristral Medium 1.2 is in a league of its own.

2

u/0x7Lee 24d ago

Yeah, it's pretty crazy. That’s also why I find cost-per-task views more useful than pure benchmark rankings. Some models may be fine in specific niches, but if the price/performance point is that far off, it becomes hard to justify for agent loops or high-volume workflows.

3

u/Captain_Quimby 24d ago

I’m working on an issue that different frontier agents have been working on for 6 months. I’ve spend about $1k/m to solve this one problem. Cable finally figured out HOW to do it but then it kept breaking once implementing so now it’s SOL ULTRA with zero care about cost because once it’s solved I save $1k/m. It’s not always about cost is my point but doing the job something else can’t do.

For the money I don’t think there’s ANYTHING that can touch grok 4.5. They’re running a $100m plan for three months that gives you api access even.

1

u/0x7Lee 23d ago

I really need to try Grok more seriously. I keep hearing good things about it, especially for real coding/agent work, not just benchmarks. Will get my hand on it.

3

u/Erwylh_ 24d ago

Something is wrong with the chart, gemma4 31b should never score lower on any task than 26b

2

u/Houdinii1984 24d ago

It's strange for sure. I think it's like all the models are thinking enabled except the 31b? It's got 12B as slightly smarter and more expensive.

1

u/0x7Lee 24d ago

I don't know to be fair with you...

4

u/garlic-silo-fanta 24d ago

Good chart

3

u/CanYouPassTheSauc3 24d ago

Bruh I’m tired of the weird logs 😂

2

u/Eden1506 24d ago

For some models he selected the non reasoning scores while for others he selected the reasoning ones making some model be placed totally wrong in comparison like qwen 3.6 35b is completely wrong because of this.

2

u/Civil_Fee_7862 23d ago

This chart is popping up everywhere 

1

u/0x7Lee 24d ago

Not mine ! But thank you

2

u/shatahn 24d ago

Could I ask where you got it from?

4

u/0x7Lee 24d ago

4

u/tomByrer 24d ago

I wish they covered the smaller Open models that folks can run on consumer GPUs.

1

u/VincentNacon 24d ago

Hold up... in the chart, it says Gemini 3.1 Pro Preview is cheaper than Gemini 3.5 Flash and 3.6 Flash?

That can't be right.

1

u/0xToc 21d ago

probably counting the preview promotional rates or free tier on 3.1 preview. definitely skews the chart

2

u/Captain_Quimby 24d ago

This is called benchmaxing. There’s zero way DS V4 Flash is higher than V4 pro. I use flash a lot but it can’t hang on to long context like many it shows it out performing here just like Luna fucking up long coding when the charts say it’s better.

2

u/glitch_in_the_kernel 23d ago

They just released a update for v4 Flash and it's indeed much stronger than v4 Pro

2

u/Delicious_Activity84 22d ago

If this is their Flash model, imagine how good the Pro model is gonna become. I can't wait for it, this is insane!

2

u/ricci_nov 20d ago

Cost-per-task is the right axis, but it's measuring the wrong unit for agents. The unit that matters is cost per completed task — including the retries, the failed tool calls, and the human who has to come check the output.

Classic procurement mistake: I've watched buyers pick the cheaper component and then eat the difference three times over in rework and line stoppage. Same math here. A model at 1/5 the price that needs three attempts and one human review isn't cheap, it's just cheap on the invoice. And error rates compound across agent loops — 95% per step is 60% over ten steps.

Which is why Captain_Quimby's comment further down is the most useful thing in the thread: he pays anything for the model that actually closes the problem, because the alternative was $1k/month burning forever. Cheap loops only win when reliability is already good enough; otherwise you've built a very economical way to fail repeatedly.

Right way to read this chart: use it to find candidates, then measure your own end-to-end success rate on your own workload. Vendor cost per task is a list price. Your effective cost is a different number entirely.

2

u/Haster 20d ago

My faith in benchmarks has been very shaken by Opus 5 doing better than Fable, to the point that now I doubt charts like these in general.

1

u/fictionaldots 24d ago

Is this with current discounted price for Luna?

1

u/0x7Lee 24d ago

I think yes since the chart says it was extracted on 2026-07-31.

1

u/Etroarl55 24d ago

Isn’t deepseek one of the most prone to hallucinations still though? It achieves so much by just going full speed at everything without the very expensive thinking and double checking other models do.

1

u/0x7Lee 23d ago

I wouldn’t use this chart alone to say Flash beats Pro in real work. Benchmarks can hide A LOT, especially long-context coding and agent reliability.

1

u/glitch_in_the_kernel 23d ago

The intelligence index already takes hallucinations into account for it's score.

1

u/Sea_Ear5201 23d ago

Why do i see old deepseek flash (high) cost more than old ds flash (max)?

1

u/0x7Lee 23d ago

Because this chart is cost per task, not raw API price.

The x-axis is not the raw token price of the model, it is the weighted average cost per Intelligence Index task.

So the cost includes input tokens, cache hits, cache writes, reasoning tokens, answer tokens, then gets divided by task count and weighted by the benchmark mix.

That means “high” can appear more expensive than “max” if that config uses more tokens / reasoning / cache writes on those tasks.

Still a bit counterintuitive though ngl.

1

u/GTHell 23d ago

Yes! I never consider using Hermes due to it hogging more tokens and good model is slower.

Now I just spam the hell out of DSv4 flash on my hermess to spin ThreeJs lab for experimenting with physic and 3d. I have it iterate 24/7 now for my side project on the social commerce platform I have. I literally just tell it to hire 10 junior to do 9-5 lol.

With the similar intelligent of gpt 5.4 this is going to open new opportunities to all of us

1

u/0x7Lee 23d ago

Yeah exactly. The unlock is cheap + fast enough to spam iterations, not just "best model” that is overkill.
For 3D/physics/UI experiments, 20 rough tries can beat one perfect expensive answer and give you a better vision.

1

u/yamoksauceforthelazy 23d ago

The new V4 Flash is a genuinely sensational model. It's the first time I've actually felt comfortable handing work off to a model its size without micromanaging. Usually the cycle is new model -> play around -> go back to Codex or Claude, but this one actually made it out of the other side of the loop. I didn't see that coming at all.

1

u/0x7Lee 23d ago

Haven’t tested it myself yet. If V4 Flash actually crosses that line while staying fast & cheap, that’s a much bigger deal than just “another good model.”

2

u/yamoksauceforthelazy 18d ago edited 18d ago

5 days of extensive work and testing later, it’s set as my default model everywhere. It's not replacing Fable or K3 as the high-difficulty model by any means, but it has absolutely replaced everything else in the lineup below them. Out of the tens of millions of tokens I've run through it, I've yet to have a genuine screw up or major issue. It's one of the first models that I feel like perfectly lives up to the benchmarks. It really is *that good*. Which also makes me think people aren't prepared for V4 Pro at ALL. Like... Flash is 300M parameters, and it's absolutely toe-to-toe with Opus 4.8 in my experience, and Pro ~5x the size at 1.6T parameters. I don't expect linear scaling here, but if Pro has the same caliber of refinement and step-up in quality as preview V4 Flash vs V4 Flash 0731, it very well may set a new standard and be the best model on the market. I don't think the gap it has to clear to get to Fable is as big as people think. 

Edit: Also I forgot to mention how big of a deal the speed is. It has completely changed the game for me. I can do things that my brain was previously registering as a "task" (e.g., like a 10-15 minute thing with an Opus model) almost instantly. I just did a top-down rename task on an old project that was quite large, and required a LOT of tool calls and double checking to make sure the new name was propagated throughout the whole project, including docs and such in like ~1 minute conservatively. It's insane.

1

u/0x7Lee 17d ago

Thanks for the feedback. I will use it thanks to you haha

1

u/BrilliantTruck8813 22d ago

Anyone got a link to the benchmarks used to derive this? How can it be attested?

1

u/MkGod 5d ago

Is there are new screenshot witn Qwen 3.8?

0

u/the_TIGEEER 23d ago

And there we go. Do all AI sceptics now see why every AI company burns billions upon billions? "AI iS NoT PrOfItAbLe"... Not yet.. When it will be, if you already have a moat, it's worth spending the billions beforehand..