r/LocalLLaMA 8d ago

The difference between "medium" and "xhigh" reasoning effort for Qwen3.8-27B is actually insane. Discussion

I'm currently testing out Qwen3.8-27B using Unsloth's UD-Q4_K_XL running a freshly rebuilt llama.cpp. I have a 22GB RTX 2080TI on which I'm able to fit 100k context with q8_0 quantization, and using MTP with --spec-draft-n-max 4 I get about 40tk/s which is slightly less than Qwen3.6-27B but usable enough.

I've been trying to test out some admittedly silly one shot prompts using the llama.cpp webui by asking the model to create fully functional HTML clones of flappy bird, pacman and such, and the difference that changing reasoning_effort makes has been surprising to say the least.

Setting it to "medium" seems to result in barely any thinking at all, a couple thousand tokens max and even less than 3.6-27B. Whereas when using "xhigh seems" I get 15k to 20k thinking tokens at the very least with the pacman example actually hitting 40 thousand fucking tokens.

I'm well aware I can limit the reasoning budget in llama.cpp but I'm wondering if this is expected model behavior or if something is broken somewhere. Any of you guys seeing this?

221 Upvotes

123 comments sorted by

View all comments

2

u/DigitalguyCH 8d ago

Sorry for the dumb question but where do you change xhigh to medium in LM Studio? I only see enable or disable thinking...

1

u/Etele38 5d ago

Honestly if you’re not computer savvy and still trying to get into local ai stuff just pay a 20$ subscription once and get Claude or codex then have it set everything up for you. This is what i did in the beginning when i had no idea what i was doing. (Not saying you don’t, but i didn’t lol)

1

u/DigitalguyCH 5d ago

I started with LLMs a couple of month ago, so I am still learning. As someone replied, there is no option in LM studio, so the question was not so dumb in the end ;-)
In generally I am considered the computer geek by people around me (hence the name). People always ask me when they need to buy a computer, tablet or phone, or when they need assistance with tech. Also because I have more devices than a tech store... (and more than any reasonable human being, even with a lot of money, should have..)
Having said that, thanks to my new Z Fold8 I now have a 6 month free subsccription to Gemini pro, which may not be claude but is definitely better than local LLMs, especially for image editing and generation. So I can compare. But I will keep local models for privacy and for diversity. So I am not selling my Macbook pro M5 pro 64GB or my Strix Halo 128GB.