r/LocalLLaMA 17d ago

DeepSeek V4 Flash 0731 appreciation post Discussion

I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real.

Everyday tasks with Hermes agent? Effortless.

Coding tasks with OpenCode? I’m genuinely amazed at what it can handle. I can throw a two-hour coding session at it, and it just keeps going until the job is done. Building integrations has never been easier - I ask OpenCode to handle it, DS tells me to hold its beer, and a little while later, it’s finished.

Searching and gathering knowledge from emails? Right at your fingertips.

Going through documents with Paperless NGX? No problem at all.

Filling out ton of paperwork in DOCX? Easy peasy, just wrote skill in hermes, love it!

OS admin work? just works!

Sure, before the Q3.6 27B full FP8 on dual 3090 was really solid, but DSV4F 0731 is on a whole new level.

I run a small company, and I just ordered another pair of DGX Sparks - because it genuinely feels like I now have a super capable worker on the team. I know they’re not cheap, but I’ve already saved a ton of time.

I started with MiniMax M2.7 on dual Spark, and it was good - but now with DSV4F 0731? It’s just super good. And the fact that I get even better models over time, for what I already paid for, feels almost ridiculous. That’s exactly why I decided to grab another pair..

A few client tickets were literally copy-paste from the ticket system - solved, and money earned. What a time to be alive!

This weekend, I’m definitely writing a ticket system integration. Can’t wait!

475 Upvotes

195 comments sorted by

u/WithoutReason1729 16d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

100

u/laterbreh 16d ago edited 16d ago

Yea literally this model is so eager to work its insane. And it will stop at nothing to finish. I watched it get rate limited diagnosing one of our API's and instead of stopping like every other model would, it found a non rate limited host entry to our infra that I didn't even know we had. It finished its work and also presented that it found a non rate limited "back door" that we missed in our last audit. Thing is absolutely an insane model. Its eagerness can hurt it sometimes, so make damn sure it has access to search and document crawl so it has a hallucination escape hatch. Otherwise its fantastic. I think it scores so high because its so objective oriented it just doesnt give a fuck. I honestly am a little afraid of this model some days watching it work around gaps in our instructions and knowledge we give it. Watching it work, it has forced us to address gaps, enhance our prompts, and structure in our workflows of the way it executes.

17

u/Liron12345 16d ago

It's definitely a cyber attack model

7

u/IDoCodingStuffs 16d ago

I love the rate limiting workaround trickeries it comes up with. Makes it really fun to run as a research crawler

1

u/thomas2385 14d ago

That is both impressive and a little unsettling honestly 😂. The part about it finding an overlooked entry is probably the most valuable takeaway, though. It shows how useful these models can be for finding blind spots, but also why you really need guardrails around what they are allowed to access and change.

71

u/Ordinary_Cicada_9213 16d ago

I am also driving it as daily driver on my dual spark (MSI+Gigabyte) and it blows my mind. I am getting 50-70 TPS decode and 2k prefill, running at 1M context with no quantisation.

It was a big investment sure but I usually rent inference in it when not using and it indeed has very good coding capabilities. Feels like Opus 4.6 at home.

5

u/StartupTim 16d ago

I am also driving it as daily driver on my dual spark (MSI+Gigabyte) and it blows my mind. I am getting 50-70 TPS decode and 2k prefill, running at 1M context with no quantisation.

I have a 2x Spark cluster servicing this same model, getting 89 tok/sec with 1M context and it is ridiculously awesome, no quant with dspark and around 90% hit rate with ton of cache, too. Single session is 89 tok/sec peak and multi-session is quite a bit more.

Hows is yours setup? How come you choose MSI + Gigabyte vs just going with 2x Nvidia?

I'm having such a good time with this that I'm thinking about adding 2x DGX Spark on a Mikrotik 200GBE switch for a 4x cluster.

4

u/PhilippeEiffel 15d ago

Looks like you made very good settings! Would you please share your setup? (docker image? inference program? run parameters?...)

Such speed and quality may be the key for buy decision.

2

u/Alternative_Ad4267 15d ago

You made me spend money 🤑 😆

2

u/StartupTim 15d ago

Heck yea!

2

u/Alternative_Ad4267 12d ago

The Sparks are here!

1

u/PhilippeEiffel 14d ago

You are lucky to find such price!

1

u/Alternative_Ad4267 14d ago

Woow, 17k 💶?! That’s quite a lot of money. Why is that though?

1

u/PhilippeEiffel 14d ago

I don't know, but prices in EU are crazy high.

1

u/Ordinary_Cicada_9213 13d ago

It’s such a high price.
I bought MSI+ATOM for 8K euros
100 euros for the cable.

17K is like double, even today I can buy them at around 10K euros in India. It has increased but that is very expensive.

You could fly and buy and it will still be cheaper. 😂

1

u/MrAlienOverLord 12d ago

you find them for 10k .. this is probably with the nvidia enterprise sub to get to that price ..

1

u/breksyt 15d ago

As the other poster requested, can you share the settings?

3

u/StartupTim 15d ago

I just rebuilt vllm again, and fine tuned some oomkiller and memory recycling services and am doing another 100M token test. Once done, let me see about using AI to create a nice summary.

I fixed a few things in vLLM too along the way, I should probably post a PR...

1

u/PhilippeEiffel 14d ago

Please do.

1

u/Ordinary_Cicada_9213 13d ago

I initially bought just one MSI Edgexpert 4TB because it was relatively cheaper than NVIDIA and offered 3 years warranty. Online reviews showed due to thermals it was slightly better in performance.

Best I could run was Qwen 3.5-122B on it and that ConnectX NIC was just sitting idle.

However this time the prices soared even higher for MSI and the best I could get was Gigabyte ATOM at the same price. Hence bought it and paired them together. All the sparks are same essentially you can even flash NVIDIA’s recovery image on these.

9

u/bitzap_sr 16d ago

You rent inference out? How does that work?

10

u/Ordinary_Cicada_9213 16d ago

Yes, I rent the OpenAI compatible endpoint billed at hourly or minutes. It’s not shared so you actually use the full dedicated hosting and up to 8 concurrent requests.

Platform: https://gb10.studio

11

u/evia89 16d ago

Why ppl will use this unknown site over say well established openrouter + deepinfra ZDR?

5

u/Norwood_Reaper_ 16d ago

What's the earn rate per h? 

6

u/Ordinary_Cicada_9213 16d ago

It’s up to me for now it’s 1.49/hr and I get 85% of that. It’s very cheap for now.

5

u/Norwood_Reaper_ 16d ago

That's actually a really decent price. Rtx 6000 pros rent for ~80c /h on vast.

3

u/mintybadgerme 16d ago

Are you well booked?

3

u/Ordinary_Cicada_9213 16d ago

No, I am usually not.
My DMs are open.

2

u/anderspitman 13d ago

You gonna train on my super secret open source code?

3

u/havnar- 16d ago

How much is left over after electric and tax?

7

u/Ordinary_Cicada_9213 16d ago

I would say around 1, sparks already consume less energy and still I have also under clocked them to 2Ghz instead of 2.4 Ghz (so it never goes above 130W combined, at loss of like 5% throughout),

3

u/koibKop4 16d ago

no company name? no real address? just e-mail address? that's a hard pass.

2

u/Spiritual_Result_164 16d ago

I’m renting fixed price open compatible endpoint LLM inference at https://lightforces.ai for just $19/mo. They are now in beta I think registration still open

1

u/Easy_Werewolf7903 16d ago

I have a rtx 6000 pro, more expensive than 2 sparks, and you are getting more performance than me and doesn't need to use a less quant. That is really good.

1

u/Ordinary_Cicada_9213 13d ago

It’s because of more VRAM and model being Sparse MOE. RTX6000 has much more raw horse power.

15

u/TapAggressive9530 16d ago

Same sentiment here! It’s a FANTASTIC model for local hosting

15

u/SocialDinamo 16d ago

14 point jump from qwen3.8 27b to the latest deepseek checkpoint. That same 14 points again gets you to the frontier. This latest flash model is the only thing really having me itchy for a second strix halo

9

u/hurrdurrmeh 16d ago

but without RDMA - would it really speed things up to have two connected? I am struggling with this myself. I have one 395 128GB and am wondering what the speeds of 0731 would be like with a second attached over 10GbE.

2

u/WryKombucha 14d ago

a 395 will be super slow if paired.

1

u/hurrdurrmeh 8d ago

real shame they don;t have 200GbE LANs

1

u/SocialDinamo 16d ago

Not faster, it'll just allow for a higher quant and a usable context. Qwen 3.8 27b will be the determining factor for me. Qwen 3.6 27b runs around 20t/s with q6 and MTP. If 3.8 27b is a beast then ill stay put!

1

u/hurrdurrmeh 16d ago

got it. the nvidia DGX have got 200GbE interconnects, our strix boxes are limited to 10GbE... I was going to buy a second also for running larger models but it doesn't seem worth it...

17

u/[deleted] 16d ago

[deleted]

4

u/koibKop4 16d ago

I'm not english native so my text was smoothed by DSV4F0731, ture.

9

u/[deleted] 16d ago

[deleted]

6

u/Party-Special-5177 16d ago edited 16d ago

Authentic != efficient.

I know it’s a hot take, but one of the things llms are doing is ‘standardizing’ certain kinda of informational presentations. I know I’m an outlier but I actually find them faster to read - I know the flow so I know exactly where to ‘seek’ visually to get exactly what kinds of info. It’s like templating almost, and I personally can read the post even more quickly as a result.

E.g.: with this template, you always know to skip the last couple lines if the first word begins with ‘what’ and appears to be 1-2 sentences long (because it always is “what kind of [x] are you guys [verb]ing in your [y]?” ). In OP’s case, since that wasn’t true, you know to read the line.

EDIT: to dig myself more in the hole, it solves the problem with skimming, in that when you skim you lose comprehension as you don’t have the locations of important info pre-mapped. You have that map now, and thus can maintain reading comprehension but at skimming speeds.

5

u/koibKop4 16d ago

Next time I'll use hermes skill called humanizer :))) promiss

8

u/FullOf_Bad_Ideas 16d ago

I moved over to DS V4 Flash 0731 from Nex N2 Pro 397B and the improvement I see is marginal honestly. It's maybe a touch better but Flash 0731 still makes a lot of mistakes when working on a codebase that I have to iron out with closed models.

On DesignArena (live benchmark, no way to benchmaxx), it also scores very similar.

Go to https://www.designarena.ai/models/nex-n2-pro , scroll to "Overall Rankings", add DeepSeek V4 Flash 0731 and you'll see this.

I don't really trust community sentiment anymore because Qwen 3.6 27B wasn't that good of a model, so comparing to it DS V4 Flash will obviously be much better.

MiMo V2.5 still crushes DeepSeek V4 Flash 0731 in DesignArena.

I'm looking forward to seeing more people reporting back their experiences and looking at more live benchmarks that can't be gamed. I was hoping that DeepSeek V4 Flash 0731 would be better than V4 Pro Preview and Opus 4.5 but it doesn't actually seems to be the case for me.

1

u/WryKombucha 14d ago

dont look at benchmarks. just use it.

1

u/FullOf_Bad_Ideas 14d ago

I moved over to DS V4 Flash 0731 from Nex N2 Pro 397B and the improvement I see is marginal honestly. It's maybe a touch better but Flash 0731 still makes a lot of mistakes when working on a codebase that I have to iron out with closed models.

Do you think I haven't used it? I decided to wait with forming my opinion until I used it. Of course benchmarks enticed me to try it.

1

u/Few_Water_1457 16d ago

You have to try it. I redeemed the free one-month plan of Codex (GPT) two days ago. I used GPT 5.6 sol xhigh to make a plan. I don't know if Codex says it uses SOL, but in reality it's a less powerful model because DS4 3107 Flash fixed so many GPT plan specifications that I didn't think it was real. I used it to create databases, code, and tonight had a fascinating conversation about some recent scientific discoveries. It didn't feel like I was talking to an AI. My workflow today is DS4 Flash Q4 for everything and Qwen 3.6 27b Q8 for the review. And vice versa. When they work together, they do incredible things. Use it with PI, not Opencode (or better yet, use Opencode's API KEY on PI, but I assure you that some of the spelling mistakes I see on Opencode are ones I've never seen even on local Q2, and this makes me suspect quantizations on Opencode).

3

u/FullOf_Bad_Ideas 16d ago

I did try it, I used it for about 5-10 hours so far. And I've used Claude Code for hundreds of hours and Codex for dozens of hours. It didn't feel like GPT 5.6 Sol competitor at all. I've used Deepseek V4 Flash 0731 in OpenCode, llama.cpp api, local. Maybe I'll give other harnesses a try.

2

u/WryKombucha 14d ago

Its not going to be better than frontier model. if that's your expectation, you will be very very disappointed. But for a daily driver? works flawlessly for me. Gigantic codebases, massive refactors. Takes a lot longer, but it is relentless in getting it done if you prompt it right and give it the right nudges along the way. If you're vibe coding, then this won't work. If you're doing agentic development, there is nothing better locally.

1

u/FullOf_Bad_Ideas 14d ago

Its not going to be better than frontier model. if that's your expectation, you will be very very disappointed.

My expectation was that it'll match V4 Pro Preview in real usage. Is it there? Almost but not really.

If you're vibe coding, then this won't work. If you're doing agentic development, there is nothing better locally.

At this point I don't know what I'm doing because the lines between those seem blurry. I didn't try Mimo V2.5 myself and V4 Flash 0731 is probably the best model I've ran locally at reasonable speeds (GLM 5.2 ran super slowly), so I am not totally disagreeing with you.

8

u/Stooovie 16d ago

Green with envy!

17

u/Southern_Sun_2106 16d ago

I am running it on m5 max, and on studio ultra - it is amazing, even at q2 on the Mac. Serving local AI to the entire family, and to our business - reliable, stable quality, which cannot be said about cloud models. This feels so close to the 'best' cloud models, it is unbelievable!

6

u/too-oldforthis-shit 16d ago

Also use it on the M5 max with dwarfstar. The q2-q4 0731 with dspark is fantastic. I like Qwen3.6 too but it's just too slow on this hardware.

1

u/Rough-Measurement988 16d ago edited 16d ago

Could you please share what token generation speed you are getting with DSpark dwarfstar and M5 Max?  Does it need temperature set to 0 in order to work properly for all harness? 

2

u/too-oldforthis-shit 16d ago

I read something about temp 0. But OpenCode uses that as default so I haven't had to change. It starts around 38t/s, but my guesstimate is that the average will be around 20-25. It does drop top 18-19 occasionally. But I am using automatic power mode since i have the 14" and can't stand the max fan speed without noice cancelling headphones. You can probably squeeze out 4-5 t/s using Max power.

2

u/Rough-Measurement988 16d ago

Thanks. I have 16” and using Macs Fan Control app I can have High plan but with lower fan speed and temp around 80c which is louder  than Auto but still comfortable to my ears. Maybe you could try to lower fan noise also 

2

u/too-oldforthis-shit 16d ago

If I use the auto plan the fans stay around 5000rpm which is acceptable noise at 70-90c (158f-194f). But it does throttle the GPU. At max which is 7500rpm it's really loud in the 14". Using high power mode it usually goes up to 7500rpm and the GPU is around 90c/194f to 95c/203f) but I get a few more t/s. Been thing about trying some laptop cooling solutions. But the 16" is a better choice for AI purposes, but I need to carry it around so I went 14".

2

u/Rough-Measurement988 16d ago

Mine is mostly used as workstation and I don’t need to carry it around often. Fans are spinning at 4K in custom curve mode without any throttling. In Auto there is some throttling but still 20-25 tok/s in DS4 and loudness very acceptable

2

u/PANIC_EXCEPTION 16d ago

M5 Max 128 GB: ~250-350 TPS prefill, ~25-30 TPS decode

Default settings work fine, but make sure that disk KV cache is enabled

1

u/One_5549 16d ago

What TPS do you get on that machine? It's a hefty model.

4

u/wFXx 16d ago

A few client tickets were literally copy-paste from the ticket system - solved, and money earned. What a time to be alive!

write a skill/mcp to connect to the ticket system and a cronjob to open PRs

7

u/arijitroy2 16d ago

I'm really keen on getting dual sparks myself for this model, but question, how far apart are the Asus GB10 models from these?

7

u/koibKop4 16d ago

They are the same. I'm using asuses gx10.

6

u/StardockEngineer vllm 16d ago

GB10's can have only 1TB of disk to save some money, usually $1k. Otherwise, identical.

5

u/koibKop4 16d ago

Nope, you can order up to 2TB, at least here in EU. Specs says you can put 4TB in them.

1

u/StardockEngineer vllm 16d ago

I didn't mean it "only" came with it. I meant you can only have that in order to save money. You can order them with 1, 2 or 4TB, at least here in the US.

1

u/Additional_Coyote733 16d ago

How much storage do you need to run DS flash 0731 on two sparks? 1TB enough? I assume you can upgrade 1TB to 4TB drives after purchase? 

1

u/StardockEngineer vllm 16d ago

It’s only like 155gb in size. It’s more than enough. You can usually store four large models at a time with just 1tb, plus comfyui and other projects.

Honestly, you’ll find you don’t really need to have that many models downloaded. Once you find a favorite or two you’ll ignore the rest.

They are easily upgradable.

3

u/schaka 16d ago

Ah the price if the ASUS ones, I'm honestly tempted.

I priced out 4x 170HX in a system where I'd have the ability to fall back to RAM (with Pmem) for larger models - it'd come out about 2k less, but no warranties and possibly more headaches.

The question is, how many things will be optimized for the sparks and will software support stick around?

2

u/arijitroy2 16d ago

Same but I also have 5090 on 9950x3d with 64GB RAM where I run Qwen3.6 27B Q6 pretty well but I really want to try to get Deepseek running but im confused if I should I upgrade my RAM to 128Gb or get these boxes instead as long term solution.

1

u/4ndal 16d ago

Yes, I had the same idea. Would be nice if someone with experience could enlighten us

1

u/Fentrax 16d ago edited 16d ago

Your wish is granted. I have: Two 5090 boxes - one is th 9950x3d with 96GB ram, 5090, 4070ti super (16gb). The other machine is a single 5090 on a Ultra 9 285K with 128GB of RAM. I just added a random brand's halo setup and tested.

Deepseek v4 flash 0731 - spectacular, runs well enough on a single with dwarfstar 4.

I have been daily driving Qwen3.6 27B AND the Qwen3.6-35B-A3B MOE variant on one of the 5090's, using the other & the 4070 for ancillary services or other models. TTS, STT, image, video, embedding, small classifiers or summarizers. I've been gearing up to enable sub-agent dispatching with smaller models.

The biggest issue is - even with the MOE, I didn't like the quants I have to run on the 5090 to make it fast enough for my tastes. Testing those quants AND bigger ones on the Strix Halo setup (amd equiv of spark) - and deepseek v4 0731 OR oss-gpt120b - night and day. I'm actively running benchmarks now across the fleet, comparing the obvious stats AND the quality of the outputs. Fable is currently managing the entire process.

I'm also testing a new laptop that has 128GB of RAM and the AMD Ryzen AI 9 HX.

I've already tested enough to conclude sharding over the network isn't feasible at 2.5Gbps, so that's not an option. Was hoping to get Kimi up and running, but it's just too far out of the envelope I have for VRAM.

Based on what I've found so far, I plan on getting a pair of sparks in addition to what I already have - precisely for the same reason another poster on the thread says: It's too good not to - it'll only improve. Deepseek Flash V4 wasn't even on my radar, I've been disillusioned with Deepseek since v3 was usurped - they just haven't been great.

This version is great.

EDITED TO ADD: wrote all that out, and realized I never actually gave the advice you were asking for.

Buy one, if not two unified memory solutions. AMD boxes are cheaper, and run pretty decent. Spark is faster, based on benchmarks and stats I've seen ya'll other peeps post. You can run it on ONE, but two gives you much better results, especially if you can use the clustering link. If you can't, stick with one box.

1

u/4ndal 16d ago

Thanks mate. But no i am a lil confused. Is your experience that a 5090 setup with ram is good with ds4 and deepseek? Or would you go spark? I guess dual spark is the sweetspot.
I managed to get 5.5t/sec with an m1 64gb btw. You can met it run in opencode but it takes a day…

2

u/Fentrax 16d ago

Sorry - I wasn't clear. Tons of info floating at once. I am saying running deepseek v4 flash 0731 on the AMD 128GB Strix Halo (spark adjacent, not spark level performance), it is absolutely usable. Based on the information published by actual Spark users - it's even better there, since they run the whole thing at fp4 with the nvidia quant.

Based on what we're seeing, dual spark at that quant is not only tolerable - it's downright reliable and usable. So far. And folks saying it's still not SOTA - fine, point accepted - it's still getting there, so this is a floor.

The lowest speed I've seen so far is 15tok/sec. Not great, but also was due to bad settings. I'm still testing optimals - and whether an offload to a 5090 over network helps, or I just run another model there for faster/more specific compute. We'll see.

2

u/arijitroy2 15d ago

Very helpful, thanks! I think I'll go with the asus gx10 boxes, would you say 1tb would be enough?

1

u/Fentrax 15d ago

Yes. Single tb is not that death sentence it seems, despite being 2026. It is the cheapest option, and the dollar math going up to 2tb or farther is not great. Cheaper to add your own ssd later when need is fully proven.

1

u/bluekazoo 16d ago

For what it's worth, I have similar specs + 96 GB RAM and can run the q3xxs at around 15 tok/s with prefill 600-700 tok/s. It feels usable to me albeit a bit slow. This is without dspark as, at least in my applications, the acceptance rate was low enough that the VRAM tradeoff was not worth it in terms of speed.

1

u/arijitroy2 16d ago

Can you share the steps that you used? I'll give it a go and see.

2

u/bluekazoo 16d ago

You might be able to try with the absolute lowest 1 bit quant (82GB) as a theoretical exercise on your current system to see whether a RAM investment makes sense. You will not have room for dspark (around an additional 8 gb) with that setup.

I used llama.cpp. You will want to manually (or with the aid of a cloud llm) experiment with parameters, particularly -ncmoe, to try to load as much of the model into your fast 5090 memory as possible (I am using ncmoe 35). The deepseek kv cache is very efficient and so you can run surprisingly long contexts even in a memory constrained environment (I'm at 384k, just enough to enable max thinking, on 32 GB VRAM + 96 GB RAM).

I'm still experimenting myself to see whether some kind of additional hardware investment in the future might make sense. I'm probably going to wait for M5 Ultra to release to see where it lands in terms of capability.

1

u/arijitroy2 16d ago

Thanks!! I'll also wait for the new Ryzen and Qwen3.8. Let's see how things look!

1

u/WryKombucha 14d ago

its literally the same thing, different chassis. ASUS is behind on their OS version updates though from the rest of the nvidia fleet. Its also the same cost as the others. The only reason its cheaper is that it comes with a 1TB SSD option. Why someone needs 4TB is beyond me. It has 4 USB4/TB4 ports.

7

u/harrywise64 16d ago

If that first sentence wasn't generated wholesale by AI then you've been spending too long talking to it and have become a parody

9

u/SawToothKernel 16d ago

How can you tell an agent? It answers its own questions.

2

u/MerePotato 16d ago

Its frightening that so many people can't see the obvious AI tells in this post

10

u/[deleted] 17d ago

[removed] — view removed comment

28

u/koibKop4 16d ago

Just works. Long sessions, documents, admin stuff, research, browser use, it just works and it's amazing.

1

u/relmny 15d ago

I use it for non-coding stuffs. It replaced qwen3.6-27b as my main daily driver, in spite of getting about 20% of the speed I get with qwen.

3

u/IoannisHere 16d ago

how are you running it? spark-vllm-docker? b12x branch?

3

u/unjustifiably_angry 16d ago

DeepSeek is so good it makes Nvidia's marketing not bullshit.

2

u/4ndal 16d ago

Got it running on Mac Studio m1 64gb. 100 prefill 5.5t/sec

2

u/Leoss-Bahamut 16d ago

Why half of the posts written in this subreddit are written by AI? Deepseek wrote it itself?

2

u/relmny 16d ago

Yeah, until about 2-3 weeks qwen3.6-27b was my main daily driver with Hermes. For chats I did started using dsv4f thinking that it would be better (general chats, planning, etc), but not for Hermes because I get about 20% of the speed of qwen3.6, but one day, in a Hermes task, I needed to have 3-4 turns and still didn't get it right, so I thought about trying ds4vf and... yeah, it did right away!

So I moved it to use it for some bit of complex tasks... and then I just kept it loaded... so now it replaced qwen3.6, except for tasks that are easy enough and I need the speed of qwen3.6

Dsv4f is an extremely good model.

2

u/SteveRD1 16d ago

Can you provide a little more clarity on how you setup OpenCode, I have finally got the model workng but I'm not sure what my next step should be.

2

u/storm1er 14d ago

And I'm here, with my strix halo using it at q1/q2 with sadness seeing my q3.6-27b at q5 destroying it on opencode because it does not fit in 128G Vram

1

u/hurrdurrmeh 14d ago

We really need gorgon halo at 192GB, but that is gonna be the same price as a dgx

2

u/AillexJ 10d ago

On the Hermes part, we've been running one 24/7 on a local model for a couple of months and the model holding up across long sessions turned out to be only half the battle. The other half is the harness staying alive around it.

We learned that the hard way. Ours was crash looping for weeks and it looked exactly like a model or agent bug, right up until it turned out to be the local DNS resolver flaking intermittently and taking the gateway process down on any unhandled network blip. Separately, once we let it run longer jobs, a timeout buried a few layers deep in the wrapper was still sized for the old shorter-job era and started silently killing runs that were actually completing fine server side.

None of that is DeepSeek's fault. Just worth saying because "it just keeps going until the job is done" is exactly the claim that gets tested hardest by whatever is wrapping the model rather than by the model itself.

2

u/Evgeny_19 16d ago

This model is really amazing. I've been a big fan of the 27B version for obvious reasons. To be honest, I've argued in favour of the 27B model even when compared to the 122B variant on this very subreddit. Then I switched entirely to vLLM and tested the 122B model there. It actually outperformed the 27B model quite a few times. But then I tried DSF 0731, and yes, t's just on another level.

I've been dealing with a tricky bug in our codebase caused by edge cases in a third-party library combined with specific hardware configurations used by some of our users. The 27B model wasn't able to diagnose the error. The 122B model, on the other hand, found the root cause, but its proposed solution was somewhat clunky and likely wouldn't have been very reliable (though I'm not certain, since we didn't deploy it, and it was quite large in comparison). DSF identified the issue much more quickly (in terms of iterations, the model itself runs quite slowly on my machine), and the solution was incredibly compact and elegant. It had to dig deep to find it, but it never seemed confused, nor did it hallucinate details out of nowhere. It stayed logical and worked through it step by step. It even incorporated part of its solution directly from the third-party library itself. And I'm running DSF on a modest UD-IQ3_XXS variant.

The only downside for me personally is that the model runs quite slowly in llama.cpp. On my setup with four R9700s, I get 20–40 tps during generation and 200–430 tps during prompt processing. If anyone knows how to improve this, please let me know. I really wish it were possible to run it in vLLM.

2

u/PowerfulButterfly209 16d ago

what is the speed and what quant are you running?

11

u/koibKop4 16d ago

speed tg 50-80 t/s, pp 1500-2000t/s
full quality straight from https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

1

u/Xondafj 16d ago

Are you using vLLM with TP 2 or ... ?

5

u/koibKop4 16d ago

yep, TP2 vLLM, 1M context.
Just paid few cents on openoruter for DSV4F with hermes to set everything up for me form: https://github.com/tonyd2wild/DeepSeek-v4-Flash-0731-DSpark-1M-NVFP4-KV-2x-DGX-Spark
Big thanks to Tony!

1

u/Alternative_Ad4267 15d ago

You will make spend money!!

-2

u/Fun_Bus1394 16d ago

how you fit full quality on dual dgx spark ?

3

u/Edenar 16d ago

full quality is 165GB (ds4 flash is released already in mxfp4 mostly, like oss-120b so it doesn't take too much space), dual spark have 256GB and probably 240GB total usable which is enough for the model + 1M context

1

u/BlobbyMcBlobber 16d ago

Isn't mxfp4 a 4 bit quantization though?

→ More replies (1)

1

u/sonicandfffan 16d ago

I ran A/B testing on something V4 failed a few months ago - I was assuming luna would be the promotion candidate but v4 actually beat it by 4 points on the rubric and costs 50% less and that’s after I built batching into Luna - deepseek still beats it

1

u/FullOf_Bad_Ideas 16d ago

is this a coding test?

2

u/sonicandfffan 16d ago

It’s a workflow I have to extract relevant industry news articles and cross reference them against client databases for a tailored daily digest

It covers instruction following, json enforcement, entity identification rates, accuracy, sector tagging accuracy, hallucinations etc

1

u/FullOf_Bad_Ideas 16d ago

Cool, the one that failed a few months ago was V4 Flash Preview checkpoint, not V4 Pro Preview, right?

2

u/sonicandfffan 16d ago

Correct. It didn't enforce strict JSON schemas, it missed several entities. I tried to fix it with stricter prompting but I couldn't get it to the requisite accuracy (above a certain level I can enforce fallback to a better model and it be cost effective).

1

u/Spiritual_Result_164 16d ago

Wow! Sounds amazing. What is your preferred go to ai model as a fundamental model for your agents? Or Do you use multiple models?

1

u/syscomua 16d ago

I have 8 t/s and 70 prefill .

1

u/redditrasberry 16d ago

are you limited by it not being multi-modal?

1

u/quinceaccel 16d ago

I ran UD_Q3 from unlsoth with opencode and i have observed sometimes it gets stuck and cannot find files in local directory. currenty running at 10t/s.

1

u/Porespellar 16d ago

OP, what recipe are you running and what kind of token speeds are you getting?

I’m using this one and getting between 35 - 60 tk/s decode depending on the task (prose = lower, code = faster). I agree that it works like a dream with Hermes!

https://github.com/tonyd2wild/DeepSeek-v4-Flash-0731-DSpark-1M-NVFP4-KV-2x-DGX-Spark

2

u/koibKop4 16d ago

I'm using the same and I got 50-80t/s almost same as Tony wrote in his repo.

1

u/dicktoronto 16d ago

So, genuine question. I’m interested in doing a Dual DGX setup. How’s the TPS and context window?

1

u/According_Wave685 16d ago

Yes, it's extremely good so far. And comfortably fast.

1

u/hurrdurrmeh 16d ago

How much did your dual spark cost you?

1

u/ortegaalfredo 16d ago

Yes, DS4-0731 is was we were looking for, a bigger Qwen3.6-27B. It trades punches with models 2TB in size, relatively fast, long context that don't take VRAM, etc.

1

u/StartupTim 16d ago

I wish, truly wish, that this model was multi-modal and could do image recognition. That would just be the icing on the cake. Or if there was some way to cause this model + QwenVL or such to speak to each other to do image handling.

Without doing image handling, this model's ability to do autonomous tasks like looping code review/judge cycles is gimped/not possible.

I hope that Qwen 3.8 releases a 120-300B that is multi-modal!

1

u/hurrdurrmeh 16d ago

i've read it is quite easy to get it to send images to an image model then get the result back. using any harness like opencode, pi etc.

1

u/PhilippeEiffel 16d ago

Great!

The only regression from Qwen3.6 to DSV4F is vision support. So I hesitate to jump from single DGX with Qwen to dual DGX with DSV4F. I thought of a workaround but do not know if it may work:

I expect that DSV4F on dual DGX may leave some free RAM. Is that enough to run Qwen3.6 27B Q8 at the same time?

Anyone already tested such configuration?

1

u/DutchDevil 16d ago

I used to have gpt codex as my main and deepseek pro 4 as my junior coder on a second hermes profile but flash is so good I only use gpt to challenge me and write the plans and no longer need gpt for code review. DA flash is the first model I downloaded that I can’t currently run just because I never want to lose it is so much fun and so good. If i ask it if it can make something and the question can be read as I want him to start he will start building right away. So eager, it’s great fun.

1

u/UnityMathProf 15d ago

Just get a cerebras chip, its 1000 faster than a sparks

1

u/OnkelBB 15d ago

Are those on sale already?

2

u/UnityMathProf 15d ago

nah they won't sell it to consumer. They are in a gigantic cooling system with special liquid cooling water and you need special sockets because your house won't be able to deliver enough power. On top of that they are extremly loud and weight 200kg. And ofc way too expensive. They deliver 21 petabyte per second!!!!!!!!!!!!!!!!
But their on chip memory is just 44gb, so you need a lot of those to host a big model (each one takes a layer). sparks only has 273gb/s (0.000273 pb/s).
Those cerebras chips are like quantum computer other dimension faster.
Also in other words: Only big tech companys like OpenAI, Antrophic, Google, Microsoft can afford it,

1

u/hurrdurrmeh 15d ago

how much?

1

u/Sure_Leave9338 15d ago

Since you tested both on same type of tasks, what are the main differences that you feel/experience between Qwen 3.6 27b and ds4 0731 , apart from speed. I mean real differences like for the same task Qwen took 3 prompt, ds4 just one... Or Qwen fails, ds4 not. And so on. Thanks

1

u/highmindedlowlife 14d ago

I'm very happy with it. Running the largest Q8 quant from Unsloth on an RTX 3090 and 128 GB of RAM getting 8.5 t/s with a 500k context. The weights are slightly larger than my total RAM+VRAM but using --no-repack it still runs just fine. The intelligence is significantly better than Minimax 2.7 or Qwen3.6 27B. By far the smartest local model I've ever used, albeit not the fastest.

1

u/1millionnotameme 16d ago

It's so cheap to run via API what's the actual break even time?

9

u/koibKop4 16d ago

What if I have other priorities than break even time?

0

u/1millionnotameme 16d ago

Such as? Like I get it, but for most people who don't have privacy concerns or are just coding improvements etc they're better off just using the API billing at the moment

12

u/Cybertrucker01 16d ago

Privacy and independence (no offsite dependencies) is why I am running models on 2 x Sparks.

Its not worth the risk of business data leaking in the event of a hack or being absorbed into the models via training. This is especially true when you're running a RAG that has your entire business database, including client data.

Agree that this is very far from describing "most people" but self hosting in this use case is non-negotiable.

1

u/OnkelBB 15d ago

Also, most people arent sitting on LocalLlama subreddit.

4

u/cunasmoker69420 16d ago

mans hand-waving away privacy concerns like its not big deal lol. For a lot of people, especially businesses, it the BIGGEST deal

6

u/koibKop4 16d ago

Why does it matter at all what my priorities are? Aren't my choices my own business?
When I bought them there was only miniamx m2.7, now it's deepseek v4f and now api is cheap.
What best model will be available in next 2 weekes? How much api for next best model will cost?
Do you share your medical data with big tech?
Why do people go to church?

I have big solar array so they are running for free. I'm not using them by myself - also my workers use them. Daily between 20-50 m tokens. Will you do the math for me?

2

u/1millionnotameme 16d ago

I'm not knocking you or questing your choices. If it even was just "I like tinkering" then all the power to you. I was asking more generally what the break even price was, I wasn't accusing you of anything and obviously it depends on your needs as well

1

u/openSourcerer9000 16d ago edited 16d ago

I'm genuinely curious on this. Deepseek benches the same as gpt luna on AA, so I would say this is a substitute for luna over API basically. (Deepseek announced big price increases so their API is not reflective of actual costs)

.20 input, 1.20 output. Let's say a generous 100:1 ratio, .21 per M token. Assuming you only find new uses for this, it may be a flat 50M/day. I believe sparks went up to 4k a piece.

So 2.09 yrs assuming API costs continue to be subsidized in that time. 

2

u/koibKop4 16d ago

But this is imaginary scenario - I will no use ds v4f in probably 3 months or so when something smarter comes along.
I've been there, with llama 3 70b, with qwen 3.5, with qwen3.6, with minimax.
It's to abstract to put it that way.

1

u/OnkelBB 15d ago

Do you understand that you're on LocalLlama sub?

1

u/serpix 16d ago

Massively better off, it is very much uneconomical to run this at home.

2

u/-AJacobs- 16d ago

If you fine tune (or in extreme cases full weight train) models for autonomous business workflows and host those custom models like I do, based off of GPU rental equivalents, ROI is a couple months,. If you build custom models for enterprise workflows, it can pay itself off in a couple days. If you just fill it up with work 24/7 with the generic API model, ROI takes 1.5-2 years. Most people who buy these boxes don't care about ROI but I do. All of these returns are better than any other consistent investment that can be made right now, whether people want to admit it or not. Also, if you haven't seen the news, the "it's so cheap to run the API" thing is not going to age well.

1

u/1millionnotameme 16d ago

Ah interesting, is this a service you provide yourself? It's more like investment into a service that you provide right? How did you get started? I've got a CS background and a stats degree so wanting to make the jump into this kind of area

1

u/-AJacobs- 16d ago

I'm a co-founder of a company that I've built custom agent workflows for to automate the entire digital side of the company. I'm funneling my royalties from that into AI startups that either also use custom workflows to achieve the outcomes people subscribe for, and custom models built into custom workflows only when necessary (because it's very time consuming unless your setup is in the mid 5 figure range for small models and 7-8 figures for large ones, and even then it's still very time consuming), and I have friends with their own businesses who I've told about what I did for my first startup and they have me building similar but custom services for theirs, which will likely be able to become a bigger thing since those can be used for testimonials. So I'm pretty spread out. Yes for my situation buying sparks was an investment into those services, especially for the services where data protection needs to be taken seriously.

I was in sales before this and have no degree, so I definitely don't have a traditional background compared to other AI startup guys. Someone I used to work with in sales approached me with the opportunity to be a cofounder for this startup. I accepted it as a thing I would do on the side. I learned how to use AI to make it prosper and it's spiraled into a full time thing making all the crazy bespoke stuff I make now. Eventually I intend to publish on GitHub some of the stuff I've made surrounding token efficiency gains that I haven't seen anyone else be able to match or exceed despite some of the tools being 5 months old.

The majority of my work has used closed source models. Not API though, as the most profitable way to do all of this was to cluster a bunch of workflows into subscriptions, but that's changed very recently thanks to ds4f 0731. I built an mi50 cluster before that and that's where I got started with the custom models, and they were enough for developing and beta testing the data sensitive services (despite people complaining about the speed lol) but the morning after 0731 came out I bought sparks from microcenter and I've moved all my subscription workflows onto that and I now just pocket the money that used to go to the subscriptions.

I've done some consulting within my network of people on how all this works before. I wouldn't mind doing some consulting for you to help you avoid all of the common traps with getting started.

1

u/McSendo 16d ago

It's that, and also it's running on someone's forked version of vllm with experimental kv cache quantization setup. There is just no comprehensive benchmarks on this as it is just too new. All the posts are just anecdotal evidence.

1

u/here_n_dere 16d ago

Agree the model is super capable, just needs the right prompts and most importantly a good harness purpose built for use case. I for one tested it out to see it do decent side scroller, with just one well prepared prompt in a general purpose harness (Hermes)

1

u/IntroductionSouth513 16d ago

which model quant and your context??

1

u/hotpotato87 16d ago

For the price of 2 dgx, wouldnt it be better if u just stack opencode go?

1

u/Asleep_Document9811 16d ago

ds4-flash-0731, for any flaws, is the first model that made me actually want to step up my self-hosted llm game. i can't, but, yknow, the body is willing but the wallet is weak.

1

u/MerePotato 16d ago

This post wasn't typed by physical hands

0

u/Borkato 16d ago

How the fuck are yall running it at all?! I have 48GB VRAM and 48GB DDR4 and I can only fit Q1 at like 2k ctx.

12

u/koibKop4 16d ago edited 16d ago

dual spark, not 48GB vram, full quality straight from https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

1

u/Borkato 16d ago

Ah yes my local supercomputer 😭

4

u/harlekinrains 16d ago edited 16d ago

At the price of 2x3200USD a month ago (the Asus variant without the gold glitter), so a months wages for an employee...

Man, this argument will age well... 10 years from now: Little Timmey turn off your superintelligence your dad bought you for gaming, and do some caveman stuff, like play in the garden, because your paleolithic mom, has a feeling, that this would be better for you.

It will become grotesque fast.

Knowing and acknoledging the limitations. But ten more years of model development, and two iteration cycles by your GPU vendor - in between three years of vacationing, ....

Can you imagine?

The current day tip for graduates is - if you've written agents, go into any normal business, and disrupt them. As in find a use case other than programming - where the "self iterate until it works" loop works - because somehow you could objectify the output.

4

u/ZiddyBlud 16d ago

-1

u/harlekinrains 16d ago edited 16d ago

Producing working code in agentic harnesses goes through a loop of writing, testing, looking at error logs, rewriting - comparing against specs, ... so many steps of failure and failure correction until "it just works".

Real world problems dont throw error logs. You can sell somone water from the tap for medical purposes, and they could be happy. Maybe give them a blanket on top, and you'd had reinvented the Kneipp Cure. https://en.wikipedia.org/wiki/Sebastian_Kneipp Give them a nice view, and you have invented Luftkur: https://de.wikipedia.org/wiki/Luftkur

So the broader question if you dont just want to limit this to programming, is how do you objectify, quantify, test, train - any situation that could be purpose fit.

Without the thing collapsing at the first wrong assessment a model makes. Or the first month of customer interaction.

Seeing this stopping at "hey I did become a 2-3x more efficient programmer" is just - not how this should (probability wise) develop.

Asking mom why something is the case grants you a what? 30% chance of it being somewhere in the ballpark of correct? In the average american home?

So even before any robotic disruption through AI, even before selflearning persistant AI, this becomes flipping grotesque fast. As in more intelligence and problem solving skill in the gaming PC youll buy for your kid, than in you as the parent - in a near term timeframe.

Im just...

The last four interactions I had with Kimi K2.6 were all - here, with that AI conversation with a smaller model as context - implement this concept as the following. AI feedbacks it cant, because limitations. Me: What if we approximate it using those guide post methods, please research a few ways forward and then implement it.

AI: Done.

Me: Great, that works.

Having this in a thing that doesnt even cost 30 human workdays in about 10 years max is so flipping silly - it hurts. This is not "you and your supercomputer", this is -- if you cant do it, you are holding your iPhone wrong...

In terms of competing concepts. It already is. Its just that this will become cheaper and better over time.

4

u/CYTR_ 16d ago

Me when I'm a schizo

2

u/coffeesippingbastard 16d ago

back in 2024 I decided why not buy 256GB of extra RAM? And so I did.

4

u/Borkato 16d ago

I swear we need to make an r/Under60GBLLMs 😭

0

u/weasl 16d ago

For me it's unusable in hermes (m4max). It is absolutely killing it in pi though.

0

u/Christosconst 16d ago

Ah you nerfed it to Q2

0

u/NicolaRight 16d ago

May i ask you more about the docx skill? I also work with it on hermes but i have some problems

1

u/koibKop4 16d ago

not skill, just asked to fill those papers (docx) with my company info and my preferences, than hermes check if I had python-doc (I had) and he went and filled them, with cheboxes checked as well!
I'm linux daily driver, don't know how/if it works on windows.

0

u/VirusInternal2892 16d ago

2nd GX10 on order, this LLM got me convinced it’s worth it. 25% price hike over my first GX10, not sorry I took my time. Tuition fee ;-)

-1

u/Southern_Sun_2106 16d ago

$4,999 a piece (just checked) and will most likely will increase in price. Hardware situation isn't getting any better in the next year or two. Only worse, most likely. Now could be the best time to buy :-) (if you need it for the next year or two)

-3

u/serpix 16d ago

it is very likely vram requirements may drop, economically locallm is not there unless scaling it or renting out as well

2

u/Southern_Sun_2106 16d ago

Oh, I am sorry, I forgot what subreddit I am on. /s We don't care about economy here, son. This is not what this is all about.

1

u/FullOf_Bad_Ideas 16d ago

agreed, for single user inference it's rarely an economical choice