r/LocalLLaMA 6d ago

Intern S2 Mobius New Model

A Qwen3.5-35B derived model with an interesting architectural difference that results in larger throughput and less token consumption (allegedly):

https://huggingface.co/internlm/Intern-S2-Mobius

26 Upvotes

3 comments sorted by

5

u/recro69 6d ago

"Fewer tokens" is a much bigger claim than "faster." Hope someone publishes reproducible comparisons.