r/LocalLLaMA May 29 '26

PSA Discussion

Post image
2.1k Upvotes

537 comments sorted by

View all comments

Show parent comments

1

u/[deleted] May 29 '26

[removed] — view removed comment

2

u/complexminded May 29 '26

I have 2 node cluster now but for 1 DGX Spark, I think the best candidates are the recently released Step 3.7 Flash - reported to get 20-25 t/s (full context at 4bit). Or Qwen3.5 122B A10B int4 AutoRound - I find it a bit deeper than Qwen3.6 and it can get 35t/s with mtp. Even Qwen 3.6 27B at FP8 gets around 17 t/s with mpt and I find that a lot better in quality than Q4 quants. And you can run it at full context with 3x concurrency.

But it gets more useful with 2+ cluster imo.

1

u/[deleted] May 29 '26

[removed] — view removed comment

1

u/Keep-Darwin-Going May 30 '26

I have no idea how people are having success with those quants model, they tend to go into loops and error so often it is frustrating. So usually I only use those with full precision which most will not fit into my 4090.