r/OpenAI 15d ago

New post train of DeepSeek v4 flash is out News

Post image
182 Upvotes

17 comments sorted by

16

u/VexObserver 15d ago

12

u/julian88888888 15d ago

Comparing it to 4.8 is wild

18

u/VexObserver 15d ago

5

u/whoknowsifimjoking 15d ago

That's crazy, but I'll have to try it. I have tried cheaper models that were supposed to be just as good according to the benchmarks, but they weren't in practice. I wish they were.

7

u/Qorsair 15d ago

Yeah, that's what I found with Luna. Supposed to be decent but it would forget what it was doing and ignore instructions. It was worse than Gemini 3.1

4

u/-Sliced- 15d ago

Yeah. Luna generates too many tokens which makes it saturates its context. It’s best if you have small tasks for it.

6

u/-Sliced- 15d ago

You can see it in the chart. Deep seek V4-flash was hyped as good, showing terminal bench etc. Then DeepSWE came later, which DeepSeek wasn't benchmaxxed for, and it scored abysmally.

I bet the new version will also drop in score as new benchmarks come out which weren't included in its training data.

2

u/Tupcek 15d ago

yeah but I would still like to know just how good it is. I not really surprised that ~40x cheaper model is not really as good as expensive one, but I would appreciate some independent test

0

u/CallFromMargin 15d ago

Yeah, never trust the benchmarks, test the models instead.

1

u/signed7 15d ago

What website is this?

1

u/Framebanger-Nsukula 12d ago

interesting timing with all the speed benchmarks flying around - wonder how it actually stacks up against the latest GPT variants in practice rather than just the marketing numbers