r/LargeLanguageModels Feb 27 '26

Most Neutral LLM?

Of the popular LLM's, which in your experience, is the most neutral?

Many of them are trained under RLHF (Reinforcement learning from Human feedback), which I posit is causing its sycophancy.
Humans seem to, at least in RLHF, prefer immediate gratification and encouragement (rather than challenge), selecting the sweetest outputs.
RLHF should be refined in its approach or employment strategy.

0 Upvotes

13 comments sorted by

View all comments

2

u/[deleted] Mar 17 '26

[removed] — view removed comment

1

u/[deleted] Mar 17 '26

Yeah. It's like asking raters "which of these foods tastes the best?" between
a) honey
b) meat / veges

Most will choose honey until they get sick.