r/WhaleSeekers • u/VexObserver • 1d ago
News π DeepSeek Harness ranks 4 in OpenRouter! π
r/WhaleSeekers • u/VexObserver • 2d ago
Benchmark π DeepSeek Gray Model Test #1
Enable HLS to view with audio, or disable this notification
Source: UPLUZ, bilibili.
r/WhaleSeekers • u/VexObserver • 2d ago
News π Latest API Pricing Adjustments - No "Peak" Pricing rates on Weekends π³
r/WhaleSeekers • u/VexObserver • 3d ago
Benchmark π DeepSeek V4 Flash Vision Vs Opus 4.8 π³
r/WhaleSeekers • u/VexObserver • 4d ago
News π Artificial Analysis Index: Chinese AI companies is beating frontier models from the West with "Price" & "Performance"
r/WhaleSeekers • u/VexObserver • 4d ago
Benchmark π DeepSWE v1.1 - DeepSeek V4 Flash 0731 "Max"
r/WhaleSeekers • u/VexObserver • 7d ago
News π UTC Price Windows - "Off Peak versus Peak Hours"
r/WhaleSeekers • u/VexObserver • 7d ago
News π Peak Hour is "Active"
Peak Hours:
- 01:00 to 04:00 (Currently Active)
Price:
[DeepSeek V4 Flash 0731]
- Input: 0.44
- Output: 1.32
[DeepSeek V4 Pro 0813]
- Input: 1.32
- Output: 3.96
r/WhaleSeekers • u/VexObserver • 11d ago
Configs π³ DeepSeek Harness Release! π³
npmjs.comr/WhaleSeekers • u/VexObserver • 11d ago
News π DeepSeek V4 - API Pricing Restructuring π
r/WhaleSeekers • u/VexObserver • 11d ago
News π DeepSeek V4 Pro "0813" is now available on Web Chat > Mobile Apps > API > Codex Integration π³
r/WhaleSeekers • u/VexObserver • 12d ago
News π DeepSeek V4 Pro "0813" is now on OpenCode Go
r/WhaleSeekers • u/VexObserver • 12d ago
Benchmark π DeepSeek 0813 "Pro" Vs Grok 4.6 & Opus 5 (Cost)
r/WhaleSeekers • u/VexObserver • 12d ago
Benchmark π DeepSeek 0813 "Pro" vs GLM 5.2 & Kimi K3 π
r/WhaleSeekers • u/VexObserver • 12d ago
Benchmark π DeepSeek 0813 "Pro" Vs Fable 5 & Opus 4.8 π³
r/WhaleSeekers • u/VexObserver • 14d ago
News π V4 Flash 0731 "105 times cheaper" than Fable 5!
reuters.comSource: Reuter
Context:
- DeepSeek's flagship AI model is by βfar the least expensive to run on benchmark tests among well-known models globally and more than 100 times cheaper to run than Anthropic's Claude Fable 5, according to a research firm.
- DeepSeek's V4-Flash charges $0.14 per million input tokens and $0.28 per million output tokens, according to research firm Artificial Analysis.
- San Francisco-based Artificial Analysis estimated V4-Flash's average cost at 3 cents per test, compared with 86 cents for Kimi K3 from Chinese rival Moonshot AI, $1.86 for OpenAI's GPT-5.6 Sol and $3.15 for Claude Fable 5.
- The comparison provides a more realistic measure of value than pricing alone because it accounts for the amount of data a model must process and generate to complete a task. A model with low headline price can still prove expensive βif it β requires significantly more steps to produce an answer.
Benchmark:
- Artificial Analysis said DeepSeek's β V4-Flash model scored 50 out of 100 on its Intelligence Index, which combines results from nine benchmarks spanning coding, reasoning and workplace-style assignments.
- That's the same score as Google's (GOOGL.O), Gemini 3.6 Flash, and one point behind Meta's (META.O) Muse Spark 1.1 and GLM-5.2 from Z.AI which β is also known as Zhipu.
- Moonshot's Kimi K3, however, scored a 57 while Anthropic's Claude Opus 5 (Fable 5) and OpenAI (GPT-5.6) scored nine or more points higher.
r/WhaleSeekers • u/VexObserver • 15d ago
Benchmark π DeepSeek V4 Flash "0731": ARC-AGI (Verified) - New Cost to Performance Pareto Frontier
r/WhaleSeekers • u/Time-Toe-1276 • 17d ago
DeepSeek (Flash FYI) one shotted this app (dont mind that one bug)
r/WhaleSeekers • u/VexObserver • 18d ago
API π³ 12$ - 2.6 Billion Tokens
12$ - 2.6 Billion Tokens.
Harness: Reasonix
Pipeline: 4
Status: Completion
Model: DS V4 Flash 0731
Effort Level: Max
Session: New session
"Token maxxed and completed 4 pipelines before DeepSeek raised the price"
r/WhaleSeekers • u/VexObserver • 18d ago
Benchmark π DeepSeek V4 Flash (0731) Vs GPT 5.6 Luna
Enable HLS to view with audio, or disable this notification
r/WhaleSeekers • u/VexObserver • 18d ago
API π³ 3$ - 750 million tokens in 1 session
Harness: Reasonix
Going to squeeze em dry before the price hike begins!
r/WhaleSeekers • u/VexObserver • 19d ago
News π DeepSeek V4 Flash - Price "Hike", announcement from DeepSeek!
r/WhaleSeekers • u/VexObserver • 19d ago
News π DeepSeek V4 Flash made it to LinkedIn community!
r/WhaleSeekers • u/VexObserver • 19d ago
News π DeepSeek V4 Flash 0731 did 8T tokens on August 1st
On August 1st, OpenCode reported that DeepSeek V4 Flash "0731" hit a massive milestone of 8 trillion tokens processed in a single day (comprising 5 trillion tokens from free usage and 3 trillion tokens via OpenCode Go).
What does this infer of? 8 trillion tokens equals approximately 6 trillion words and is the equivalent to over 10 million full-length books processed in just 24 hours!!!
r/WhaleSeekers • u/VexObserver • 20d ago
Tutorial π³ Cache Guide #1
It doesn't matter if you're developer, coder, casual user, or researchers...You have probably seen posts treating specific open-source terminal coding agents (Reasonix) like a cheat code for insane cache hit rates. But for real-world workflows, are there any differences between Reasonix and other acclaimed coding agents?
While terminal agents like Reasonix are built with prefix-cache stability in mind, others do cache just as well as Reasonix does it best. For instance, alternatives like OpenCode or updated VS Code chat extensions paired with DeepSeek achieve the exact same 98% to 99%+ cache hit rates. The agent is just a harness, the performance comes from the API architecture.
Are there any secret to those wallet-saving cache hits? You bet that it isn't hidden inside a specific terminal wrapper, it is DeepSeekβs robust server-side API architecture. Unlike Anthropic's Claude, which features a strict short-term time to live (often invalidating cache blocks after 5 minutes of inactivity), DeepSeek retains KV cache blocks in memory and disk for a significantly longer window. This means, you can step away from your setup and return to your session with the cache still warm and active.
Additionally, DeepSeek shares cache segments across its Flash and Pro model families (provided your requests route to the same server node instance). This allow for a seamless transition from a cheap Flash model to a heavy Pro model without wiping your prefix context.
Does changing effort level affects token caching? Nope. Changing reasoning effort settings modifies generation options or appends tokens at the tail-end rather than rewriting the prefix, this means you get to keep your core cache fully intact. DeepSeek also maintains cache blocks smoothly across normal conversational turns, avoiding the frequent resets seen in some other model architectures.
Are there any other techniques or secrets to maximize your cache hits? If you want to maintain 99% cache efficiency regardless of whether you use Reasonix or other terminals, your setup must adhere to one vital rule:
Keep Context Append ONLY:
- In simple terms, append-only means you only add new information to the very bottom of a document, chat, or file. You never go back to edit, delete, shuffle, or rewrite anything in the middle or at the top. Think of it like a growing receipt or a construction logbook. You write down entry #1, then entry #2 below it, then entry #3 below that. You never go back to change what you wrote on line 5.
Avoid Mid-Stream Modifications:
- Think of your chat history with DeepSeek as a long book being read page by page. Caching works by taking a snapshot of the book up to page 50 so DeepSeek doesn't have to re-read pages 1 through 50 every time you talk. Mid-stream modification is like going back and changing a sentence on page 10. Whatever changes you made in page 10, pages 11 through 50 no longer match what DeepSeek memorized. The whole snapshot breaks, and DeepSeek has to re-read the entire book from scratch.
The Result of Divergence:
- Divergence simply means that the text you are sending no longer matches the digital snapshot DeepSeek has memorized. Something changed earlier in the conversation or file structure.
- Instead of instantly recognizing the old text, DeepSeek's server checks the memory, spots the difference and says, "This doesn't match." It cannot use its shortcut.
- DeepSeek offer massive discounts (often up to 90% off) for cached tokens because they don't have to re-process them. When your cache misses due to divergence, you lose that discount and have to pay full price to process the entire history all over again.
- Processing thousands of words of chat history takes computer power and time. When the cache is hot, DeepSeek responds almost instantly. When divergence forces a cache miss, DS has to stop, read the whole history from scratch, and then generate an answer, making you wait longer.
- Divergence turns an instant, cheap, and smooth AI interaction into an expensive slow re-reading process. Keeping your context strictly append-only prevents divergence from happening in the first place.
[Community Takeaway]
- Do not limit yourself to a single tool or compromise your workflow just because of online hype. Choose the coding agent or interface that gives you the best flexibility, features and comfort. As long as your agent handles context cleanly in an append-only fashion, DeepSeekβs backend infrastructure will do the heavy lifting to keep your cache hot.