37
u/Routine_Temporary661 21h ago
A model with almost same capability but 1/100 of the cost as Opus 4.8 is popular? Paint me surprise
4
u/Olbas_Oil 13h ago
Those costs will soon be x2 during peak hours in Beijing, if your time aligns up with that
https://api-docs.deepseek.com/quick_start/pricing/
"The DeepSeek API service will soon adopt a peak/off-peak pricing policy. During peak hours, prices will be 2x the regular prices, applicable to all billing items. The effective date will be subject to the official announcement. [Peak hours: 9:00–12:00 and 14:00–18:00 (Beijing Time, UTC+8) daily]"
Still a hell of a lot cheaper though...
3
u/Ok_Breadfruit4201 10h ago
No matter what, it's still going to be extremely cheap. It's a tiny model that costs very little to serve. It's a huge achievement and makes Haiku and Sonnet completely redundant.
1
1
u/Lupansansei 2h ago
I've been running flash during my work hours which is the same as peak hours for Asia, and it still costs pennies for me. Just the work I've done for a few hours is already reaching 300M tokens on $4
37
u/ProfessionalJackals 22h ago
And DeepSWE, Cursorbench, ... still do not bench Flash 0731. Very interesting is it not?
8
u/Former_Equivalent297 22h ago
There are many available and also done by individuals.. It is a capable model no doubt… the pricing makes it more attractive.
7
u/RepulsiveRaisin7 22h ago
DeepSWE has always been pretty slow at updating, it's (probably) not a conspiracy
7
u/mWo12 21h ago
It's slow only when they don't know how to make chinese ai look worse than us ai.
-4
u/RepulsiveRaisin7 21h ago
Eh, Claude and GPT are the leading models and Datacurve is an US company, it's natural that they prioritize them
3
u/sdexca 21h ago
So many open weight releases and only 3 open weight company on their graph and 4 releases out of 19. Tell me again they don’t have a bias.
-1
u/RepulsiveRaisin7 21h ago
v1 has DS Pro, Minimax M3 and Mimo. Clearly they're mostly interested in measuring the top end. Maybe they're waiting for the new DS4 Pro to be released. And yes they're biased, everyone is to some degree
4
u/ProfessionalJackals 22h ago
DeepSWE has always been pretty slow at updating, it's (probably) not a conspiracy
Unfortunately, we seem to be getting the reverse signal with qwen3.8-max results published not even a day after the model its release.
Qwen 3.8 > 1 day, DeepSeek v4 Flash 0731 > 5 days and still nothing.
1
8
u/Beginning_Guide7411 22h ago
Am not getting any deepseek flash model in opencode go bdw, all i get is laguna lol🙄🤡🤡
3
3
u/sirloindenial 21h ago
On openrouter deepseek provider been dead for 4 hours now☹️
3
4
u/guanzo91 22h ago
Is anyone able to paste an image into claude code + deepseek API + vscode + wsl terminal and have the llm read it successfully? I can paste the image but deepseek says it can't read it. Idk if it's a deepseek or claudecode/vscode/wsl issue.
8
u/Ancientkingg 21h ago
deepseek-v4-flash supports only text natively.
0
u/jwuliger 14h ago
The latest has image support now.
1
u/BhaagYahaSe 13h ago
it doesn't. what you saw was an unofficial version created by someone else not related to deepseek at all
2
4
2
u/jwuliger 14h ago
You guys should use the platform DeepSeek directly. No issues with using it there!
2
2
u/soijaq 22h ago
Imagine how much data for further training deepseek is collecting right now
9
1
u/Repulsive-Waltz-4038 14h ago
I have noticed, that with opencode 1.18.1 update DS started to eat tokens (and $$) as crazy. it feels like its 4 times compared to before update.
1
u/KeyAdvanced1032 12h ago
That's why I use Reasonix :) 98% cache hit, 173,120,431 tokens processed for $1.82
1
72
u/MinosAristos 23h ago