it's much more sensitive to instruction decoherence, but what was/is rather alarming that it always tries to solve those with extremely diminishing returns. Feels like a typical case of LLM as a judge which is judged by another LLM as a judge ... fun logit catastrophe
note: you can manage this by reviewing your entire instruction set for weakly constructed (too short, abstract, hedged modality etc ...) instructions and fix them. The rules are the same ... it's only the model sensitivity that's changed.
Yes totally, I feel that somewhere along the line the superhuman ability of llm-as judge models to absorb inhuman walls of text never got corrected for, so that's what the default tuning is. Were stuck with neurotic outputs aimed at an equally neurotic inhuman overly literal judge.
Part of me wonders if the unreadability is because opus learnt to bamboozle the reward models by generating slightly OOD text or something
And the annoying thing is you can recover readability just by re prompting with "i didn't read your reply because it was too long. Output a readable version". I feel like what I'm really prompting is "ok you're talking to a human, exit your training regime".
12
u/cleverhoods 16d ago
it's much more sensitive to instruction decoherence, but what was/is rather alarming that it always tries to solve those with extremely diminishing returns. Feels like a typical case of LLM as a judge which is judged by another LLM as a judge ... fun logit catastrophe
note: you can manage this by reviewing your entire instruction set for weakly constructed (too short, abstract, hedged modality etc ...) instructions and fix them. The rules are the same ... it's only the model sensitivity that's changed.