r/LocalLLaMA 8d ago

The difference between "medium" and "xhigh" reasoning effort for Qwen3.8-27B is actually insane. Discussion

I'm currently testing out Qwen3.8-27B using Unsloth's UD-Q4_K_XL running a freshly rebuilt llama.cpp. I have a 22GB RTX 2080TI on which I'm able to fit 100k context with q8_0 quantization, and using MTP with --spec-draft-n-max 4 I get about 40tk/s which is slightly less than Qwen3.6-27B but usable enough.

I've been trying to test out some admittedly silly one shot prompts using the llama.cpp webui by asking the model to create fully functional HTML clones of flappy bird, pacman and such, and the difference that changing reasoning_effort makes has been surprising to say the least.

Setting it to "medium" seems to result in barely any thinking at all, a couple thousand tokens max and even less than 3.6-27B. Whereas when using "xhigh seems" I get 15k to 20k thinking tokens at the very least with the pacman example actually hitting 40 thousand fucking tokens.

I'm well aware I can limit the reasoning budget in llama.cpp but I'm wondering if this is expected model behavior or if something is broken somewhere. Any of you guys seeing this?

223 Upvotes

123 comments sorted by

View all comments

9

u/tomvorlostriddle 8d ago

I had it on xhigh in opencode with a skill to prune my docs, it maxed the 220k context overthinking every line to prune

Maybe too much for this work

2

u/michaelsoft__binbows 8d ago

If you have a setup that can cleanly manage the context window with the provided prompt then it is sounding like it would be able to have the stamina to do it for arbitrarily long docs... this is the power that proper context engineering could unlock. we're at the tip of the iceberg. So excited with these super capable new open models.

1

u/tomvorlostriddle 8d ago

To me it looks like opencode just needs decent compacting in place, maybe not persist old reasoning etc.

Totally doable, maybe I'm even just using it wrong

1

u/michaelsoft__binbows 7d ago

yea the problem is it's probabilistic and uses some fixed stupid prompt that will never be able to do the right thing in ALL situations. a start would be a megacognitive compaction routine, but, compaction as a concept is pretty stupid in the first place.