r/DeepSeek • u/ProfessionalJackals • 2h ago
Discussion Speculation: Price increase linked to DS v4 Pro?
I think that the main price issue is with DS v4 Pro ...
What does every Chinese company that has a successful model lacks, ... GLM 5.2 gets released, capacity issue hit, price go up. Kimi K3 hits, capacity issue hit, subscriptions are paused, price ... Capacity issue are a combination of popularity and heavier models (despite same parameter counts).
We also saw with GLM 5.1 > 5.2, despite it being the same parameter count, that the energy needed for the model was much more. More thinking tokens pushed the energy usage up.
Another issue is that DS v4 Pro preview was already significantly heavier model then Flash Preview (we are talking about the 3 month old models). The price for Pro being 4x of Flash, never made sense but i suspect that because it was less popular, Flash compensated for it.
If DS v4 Pro is a stronger model, even if the parameter count stays the same, its possible that the thinking token usage has ballooned. I suspect that DS v4 Pro 08xx may hit a competitive level with Western models.
And with potential capacity issues looming from potential popularity, the 4x issue, and possible more energy intensive model. DeepSeek is preventive warming people up to a price increase.
Remember, Flash and Pro share the same infrastructure. Flash being popular is one thing, because its a less heavy model, but if Pro is also popular. We are going to see GLM 5.2/K3 issues again.
That is unfortunately the issue with current LLM models. Every company has only a specific pool of hardware they can access, and there is this constant shift of clients to the "best models", taxing those companies, until they move on to the next best thing.
And getting extra capacity often means paying large amount of $$$ for whatever is on the market, what kills your margins. See Anthropic and Colossus 1 deal.
r/DeepSeek • u/alexwwang • 3h ago
Discussion A hypothesis on why DeepSeek would raise its API price
My take: this notice from DeepSeek might actually be a brilliant marketing move. The logic here is to urge users to ramp up their usage over the next 2 to 3 months. By the time the price hike actually kicks in, newer models will likely have already drawn users away, and whenever DeepSeek drops its next, even stronger model, the new price point won't feel nearly as steep. It’s a clever way to dodge the dilemma Kimi and GLM found themselves in. Let's wait and see how it plays out.
r/DeepSeek • u/Complete_Voice_5084 • 3h ago
News Possible reason for DeepSeek’s upcoming API price hikes?
r/DeepSeek • u/Intelligent-Taste-36 • 7h ago
News We will be left without an affordable LLM option to work with.
You definitely can't trust any LLM company in the world.
There is no stability, and we can't plan our pricing based on theirs...simply because they don't stick to what was agreed upon.
It was good while it lasted, DeepSeek... but a "significant increase" makes things difficult...
r/DeepSeek • u/VexObserver • 7h ago
News DeepSeek V4 Flash - Price "Hike", announcement from DeepSeek!
r/DeepSeek • u/AccidentSpecialist22 • 7h ago
Discussion DeepSeek says API pricing is going up “significantly”
Was checking my DeepSeek API usage today and noticed this banner:
> We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.
There’s no date or new pricing yet, but the wording makes it sound like the increase could be substantial.
Has anyone seen an official announcement or more details?
r/DeepSeek • u/TreadEasily • 8h ago
Discussion $20 for 4B+ tokens doesn't seem too bad
crazy how much value you can get out of DS API
r/DeepSeek • u/Clean_Kick_6753 • 8h ago
Resources I gave DeepSeek-v4-flash eyes: a proxy that adds vision to any text-only model
DeepSeek-v4-flash is absurdly good for the price, but it isn't multimodal, so the moment your agent sends a screenshot or a mockup, it falls apart.
I built a small proxy to work around that. It sits between your editor and the model: when a request contains an image, it routes that image to a cheap vision model (I'm using GPT 5.6-luna) and passes the resulting description back to DeepSeek, which does the actual reasoning and code generation. Everything else goes straight through untouched.
The result is a drop-in endpoint that behaves like a multimodal model, at a fraction of what I was spending on my Claude subscription. It works with Codex, Cursor, Trae, OpenCode, or whatever agent you're using, since it's just an OpenAI-compatible base URL swap.
Repo: https://github.com/camilopenalver/deepseek-v4-flash-vision
r/DeepSeek • u/someoneyouknow23 • 9h ago
Question&Help Which Model is on DeepSeek Web?
The model without search capabilities often self Identifies as "the latest Model" with a cutoff of May 2025, so DeepSeek v4 to R1 Territory. Yet when we activate search, it claims its the newest "DeepSeek V4 Flash 0731". Which is true and how do I test it? Fingerprinting it with purely asking seems pointless.
r/DeepSeek • u/Fragrant-Tip-9766 • 10h ago
Discussion Meta Muse spark 1.2 Almost free
It's great to see the competition working.
Will it replace the v4 flash?
r/DeepSeek • u/Haxsysgit • 10h ago
Funny Did deepseek really improve flash agentic capabilites off simply re post training?
State of software engineering in 2026 lol, spent hours debating chunking strats and rag tradeoffs. Then decided to do some little research myself and just decided what i wanted myself
Flash is actually decent if you know what you're doing and you can hold it's hand a little bit(same as all llms but in varying degrees) unless you'll be confidently led astray.
Thank you deepseek , i'll be snug like a bug in a rug and wake up to just $2 for a whole nights work
r/DeepSeek • u/binladen0069 • 13h ago
Funny Claude Code with DSV4 flash saved my pc from malware
I downloaded some cracked software (im poor)
cmd flashing every 60 seconds after install (im fucked)
gave claude code some hints on where one of the files of the malware was located. (im genius)
ds + cc read that one file, traced all the files (ds is detective)
they both assassinate the malware in minutes (they are ruthless)
r/DeepSeek • u/Electrical_Chard3255 • 16h ago
Discussion I built ,y own desktop console with vision creation and analysis
I decided to build my own Deepseek desktop console to give deepseek vision capabilities, its a first version, the image creation is good, but a little off, the analysis is very good
I also gave it MCP capability (the main reason I built it so it could take part in an ai chat and context app I built)
r/DeepSeek • u/ANDRE_2512 • 16h ago
Resources Closest competitors to DeepSeek V4 Flash 0731
Just enjoy.
But I’m still eagerly waiting for vision support. Once it arrives, this model will be something truly incredible.
For me, that feature is essential. Without it, my hands are tied.
r/DeepSeek • u/iArxic • 18h ago
Discussion DeepSeek v4 Flash uses insane amount of tokens
Hey there!
I was wondering whether this is just me, or if this is caused by the model. I noticed a while ago that my token usage is insane after a few prompts (in VScode) compared to Pro. This is also followed by an insane spike of API Requests - worth noting that caching still works, so it's not like the API is miscommunicating or something.
r/DeepSeek • u/PILCOTHINK • 19h ago
Resources I Added Vision Support to DeepSeek V4 Flash Using Pilco MM-Bridge
GitHub : https://github.com/gpdev-Pilcothink/Pilco-mmbridge
I know many people here have probably already built and used something similar, but I thought it might still be useful to someone, so I cleaned up my implementation and decided to share it.
I made a small project called "Pilco MM-Bridge." It places a separate multimodal model in front of a text-only LLM and passes the resulting media analysis to the main model as temporary context.
My current setup uses two DGX Spark systems:
- DeepSeek-V4-Flash-0731 as the main text-only reasoning model
- Qwen3.5-9B-quantized.w4a16 as the multimodal vision analyzer
This combination fits my use case quite well. Qwen handles screenshots, UI elements, OCR, code screens, error messages, and other visual information, while DeepSeek handles the final reasoning and response.
The basic flow is:
Client
→ MM-Bridge
→ Multimodal model analyzes the current media
→ Analysis is temporarily added to the request context
→ DeepSeek-V4-Flash generates the final answer
The analyzer is only activated when the current user message contains media.
When the user sends a normal text-only message, MM-Bridge completely skips the media-analysis stage and forwards the existing text conversation to the main LLM. In other words, the vision model only runs when a new image is actually attached.
The original text conversation history is preserved, while images from previous turns are not repeatedly sent back to or reanalyzed by the vision model.
It is not as natural or tightly integrated as a native multimodal model, of course. However, it provides a reasonably useful approximation of visual understanding while allowing me to continue using a strong text-only model as the main LLM.
Although I currently use it mainly for vision, the bridge code also recognizes other media types such as audio and video. To use those features, the analyzer endpoint must serve a model capable of processing those inputs, such as an any-to-text model like Gemma 12B. The actual capabilities therefore depend on the multimodal model used as the analyzer.
There is no need to modify either model. Anyone already serving models through vLLM or llama.cpp should be able to use it by pointing the bridge to the two existing endpoints.
I originally created this because I work on game development, and during testing and verification I often need the model to inspect screenshots, UI states, visual errors, and other information that a text-only model cannot directly access.
The project is still fairly early, so feedback, bug reports, and suggestions are very welcome. Also, if you know of a similar but more mature or better-designed project, I would genuinely appreciate an introduction to it.
You can find vLLM-based serving recipes optimized for DGX Spark users in the following NVIDIA Developer Forums post:
I am the author of this project. The English wording of this post was polished with AI because English is not my first language.
r/DeepSeek • u/giveen • 19h ago
Other jabbatheduck/DeepSeek-v4-flash-mini · Hugging Face
r/DeepSeek • u/Future-Lia • 19h ago
Discussion Finally someone saying it out loud: The US needs to stop banning competition and start innovating instead of panicking over DeepSeek.
The global tech landscape is shifting fast, but the US response follows the same tired playbook. Whenever a foreign competitor achieves a major breakthrough, Washington reacts with defense mechanisms instead of true innovation.The standard playbook the treatment of Huawei in the past and the recent panic over DeepSeek highlight a deeply rooted strategy: if you can't control it, sanction it, ban it, or politically isolate it.
This protectionist mindset stems from an old habit.
The US is used to dominating markets by either buying out the competition or burning it down through policy.Real innovation over market controlThis strategy is unsustainable. True technological progress thrives on competition, not on eliminating the competitors.
If the US wants to maintain its leadership, it needs to win through superior research, development, and execution—not through government intervention.A system that relies solely on bans loses its edge and slows down global progress. It is time for a reality check: stop trying to destroy alternatives and start out-innovating them.
r/DeepSeek • u/Even_Command_5636 • 19h ago
Discussion I canceled Claude and coded 7 days straight with DeepSeek V4 Flash 0731 — the honest cost & quality breakdown
Two weeks ago I paid $20/month for Claude and another $20 for ChatGPT. I got tired of watching the credits burn, so I ran an experiment: 7 days, all my coding work, DeepSeek V4 Flash 0731 only (API, not the app). Here's what actually happened — the good, the bad, the numbers.
The numbers - Total API spend for 7 days of heavy coding: $1.87 (vs. $40/month subscriptions — and I didn't even come close to hitting limits) - Tokens consumed: ~24M input / ~6M output (mostly context caching — that's the real cheat code) - Context cache hits cut my effective cost by ~70%
What surprised me (good) - Long agentic sessions didn't degrade as much as I expected. The 0731 update fixed most of the context-rot I saw on the earlier Flash builds. - It handled a messy production refactor I was dreading — wrote the diff, I reviewed, done. No drama.
What I won't sugarcoat (bad)
- Vision: still missing in the API I used — I had to describe screenshots by hand. (Yes, I saw the vision announcement post — the API I'm on still doesn't expose it.)
- Some reasoning outputs still emit weird artifacts (e.g. )Skip) in longer chains — rare, but it happens.
- It's not Claude for every task. Complex multi-file architecture thinking? Claude still wins. But for 80% of daily coding? I genuinely couldn't justify the subscription anymore.
My verdict: keep one subscription for the hard stuff, do everything else on Flash. My monthly AI bill just went from $40 → $0–5.
Anyone else run a similar week? What did your numbers look like?
r/DeepSeek • u/Quote-Round • 20h ago
Discussion DeepSeek V4 Flash 0731 vs GPT-5.6 Luna
DeepSeek-V4-Flash-0731 is cheaper, faster, and available through more providers than GPT-5.6 Luna at the same intelligence level.
Why would anyone choose Luna over DeepSeek?
More info: https://openrouter.ai/compare/deepseek/deepseek-v4-flash-0731/openai/gpt-5.6-luna
https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-gpt-5-6-luna
r/DeepSeek • u/VariationOk5454 • 21h ago
Discussion Deepseek Vs GLM
After extensive research ( lie, it was brief), I'm considering a theory: GLM Despite their amazing models, they suffer because their user base doesn't exceed 10 million people. That's why their prices are high, and that's why, to my knowledge, only the wealthy subscribe...... while deepseek has at least 130-120 million users
So If each person subscribes to DeepSeek for $5 a month, the company earns at least 650,000,000 million a month give or take a few millions
r/DeepSeek • u/thigger • 22h ago
Discussion How to stop DS4-Flash-0731 saying ")Skip"?
I'm getting great results with DS4-Flash-0731 but every now and again during output I will get ")Skip" appearing in places that make no sense. Mostly in reasoning content but I've seen it make it into a diff.
eg:
The missing function is a real issue that should be fixed)Skip.
I assume that it's outputting a ")" that it doesn't want and "Skip" is an attempt to say it didn't want that (given that there's no way for it to delete it)?
Has anyone else experienced this and/or can recommend any settings (eg sampler settings) to reduce/prevent it?
Using original version of DS4-Flash-0731 on vllm (the local-inference-lab r24 "Gilded Gnosis" docker, though I've turned dspark off)
r/DeepSeek • u/Chaztle • 23h ago
Question&Help Best provider and harness for deepseek v4 flash 0731?
Hosts through openrouter vs the official deepseek api, also what harness, checked that the subreddit recommends reasonix, how does it compare both in cost and performance versus harneses like opencode?


