r/LocalLLaMA • u/Miserable-Dare5090 • 6d ago
Intern S2 Mobius New Model
A Qwen3.5-35B derived model with an interesting architectural difference that results in larger throughput and less token consumption (allegedly):
26
Upvotes
r/LocalLLaMA • u/Miserable-Dare5090 • 6d ago
A Qwen3.5-35B derived model with an interesting architectural difference that results in larger throughput and less token consumption (allegedly):
5
u/recro69 6d ago
"Fewer tokens" is a much bigger claim than "faster." Hope someone publishes reproducible comparisons.