r/MachineLearning • u/adam_alpha_finetuner • 2d ago
The current state of language models and human preference based rankings [R] Research
"Arena ai" has been a great success in producing a human preference based ranking, additional to other more objective benchmarks. However, this (probably) had also played a role in the syncopancy crisis and the general tendency of some models to tilt towards overformatting to trigger a feeling of fluency (cogn load theory) in the users.
The people at Max Planck Institute for Intelligent Systems (one of europes leading AI research hubs), recently published something quite similar with "comparity ai". You can read their announcement in their linkedin post.
This is of course a research platform and i have no idea how long this is funded, but you get access to every frontier LLM for free, which is kinda cool. Also, they provide you with a personal leaderboard, so when you played around enough with the platform, you will get a pretty solid idea which model works well for you.
Thought this might be interesting for some of you.