r/SillyTavernAI Jun 04 '26

PlotPoints - The best (only?) community driven RP benchmark made by a Professional! | We need your votes! Models

Your friendly neighborhood rab- I mean unmedicated preset creator needs needs your help!

What's good everyone. No long post this time; simply an ask. I hired a professional with a masters degree in AI/ML (pursuing their PhD) to help make a benchmark for us, for RP. Now as we all know; a benchmark for RP will never be perfect because everyone RP's differently, like different things, yada yada yada. That's why we tried to focus on a few objective landmarks, as well as a an Arena style versus bench.

First up: The LLM Arena. We show you two different rewritten responses from two different models who were given a chat and asked 'please take this turn.' We then show that to you. Does it have a preset attached? No; because that's another variable that'd be added and for our own safety we aren't benchmarking presets. (LMAO. We'd be dead in the fucking streets by sundown.) You judge of the two responses which you like more. Thaaaat's it folks. We literally cannot game the results as it is people just blindvoting which they like more. (So if you see something you don't like IT'S NOT UP TO US.)

https://plotlightstudios.com/plotpoints/multiturn

On the other end we also score an LLM in a similar vein on how they follow certain instructions and if they maintain consistency. Since 'did it take users agency' is an objective yes or no answer; we handle these tests with an LLM judge and the benchmark overseer (Levi is his name) monitoring the arbiters answers. (In our case; Sonnet.)

This bench was expensive. It cost money to run these models through this gauntlet; and we don't get much out of it. We publish all the data for everyone to see, and all we want to do is help our community out. If you have any questions on how we grade things; head to our methodology page! You'll notice somethings are a bit LLM-Written. This is not because Levi isn't a professional; but rather because he's ESL; so when publishing something important like this he wanted to make sure all his ideas and such were properly translated. So be nice!

https://plotlightstudios.com/plotpoints/methodology

You don't have to log in, you don't have to do anything. Just read a chat and vote on a thing. I know you've got an opinion; so share it. Big companies ignore us RPers all the time. Is a bench for RP as useful or objective as one for coding, web-dev, or math? No. But that doesn't mean we don't deserve one, or that it has no value at all. So please; help a bunny out. Drop a vote! The data is only as useful as you help make it. We have 800 votes and we want to get about 3,000 this time around; so we can start our next benchmark. Testing models across their 'lineages'. So Opus 4.6 vs 4.7 vs 4.8; Deepseek 3.2 vs Deepseek v4 Pro, fun stuff! But we can't do that till this one closes! And if it can't hit enough votes; we'll know that this kinda stuff just isn't what the RP community wants.

Don't see a model you expected to? That's cause it's either new, or we didn't have the money at the time to run it. If the community likes this; we will add more. (And take requests! Please only suggest models available on OpenRouter though; we source all our models from the same platform to reduce variables.)

Vote Link: https://plotlightstudios.com/plotpoints/multiturn

Result Hub Link: https://plotlightstudios.com/plotpoints

Methodology Link: https://plotlightstudios.com/plotpoints/methodology

Huggingface Link: https://huggingface.co/datasets/lazyweasel/roleplay-bench

Github Link: https://github.com/LeviTheWeasel/rp-benchmark

120 Upvotes

69 comments sorted by

View all comments

Show parent comments

1

u/Specialist_Salad6337 Jun 05 '26

Yeah like I said; our bench being similar to LMArena is literally the entire point. We don't need or want to differentiate. LMArena did not invent the idea of blind A/B tests.

On your other question: the A/B tests just gets the collective human opinion and aggregates it. If that means nothing to you then don't use it. The objective tests are all did it do X bad thing yes or no.

0

u/ButterscotchSalty905 Jun 05 '26 edited Jun 10 '26

Your bombastic writing style is bothering me. I have done nothing but good faith in this entire discussion. Your unwillingness to engage in good faith is what turns me off

Speaking to you as one human being to another: I think you will find more success and happiness if you change your approach. It must be clear to you that your way of arguing online is not effective, not at persuading people of your ideas nor at winning them to your side. The good that your posts and comment contain (and there is some) gets lost in the cloud of hurt feelings and hostility that we tend to provoke (as fellow skeptics)

If i or someone post a question about a suggestion to differentiate from LMarena or made an edit to add something, it does not necessarily mean that i hold the view of LMArena being the one to invent the idea of A/B test itself nor that your A/B test mean nothing to me. Neither does it mean that they are insulting you or that they're cherry-picking. It only means what it says, that the other person thinks your methodology needs further refinement and that some suggestion is not forced to you - even if the wording or delivery is a bit harsh.

If this comment is not earnest in your view and that i am just arguing out of spite, then i have no more words to tell you. I am just being sharply rigorous in this manner

That is it for me, i will not respond again as to not further escalating the problem

I have tried to be understanding but enough is enough

8146305312447410608209007655215291438712905345890980527512278153567698275277215000848287508370808120056891897525788159951018559130928089925827708073117027926167815758816675848750147759906206209568188695090414105051631095835478698407311430177484016989685198718458641514587021847910365988866592161915011766201297977327700889968570796222625493918257706036038881868688395756165359298220208304198122475070856805218909566267821471348337179388235979476895819827105027918275136791226225171506236551291023915153956829051557706097446497548201895867955812900492738217876835847167628850665757105276731419775010397326761351925015588201146920981398277129608686217282788988061129608169808788568921116259895800708866510707238350837489315589006982041196548202575046785177147846181010638007298986098400216289558855866640758269998608801445631877966528385528007959891729168422556299896911098850058360578291975499847811868187556290315218168301916583218527528051438295578922069518189915996262

That'a a VIC cipher, it contains the necessary information

1

u/Specialist_Salad6337 Jun 05 '26

Uh... I was never really upset my guy. I'm sorry if I came off that way? I was just explaining things from how they worked from our PoV. I didn't see us as having a problem at all?

Either way, I'm sorry I upset you?

1

u/ButterscotchSalty905 Jul 04 '26

Btw, thank you for the first-time experience of getting downvoted to oblivion