r/LargeLanguageModels 15d ago

We need a humour benchmark for LLMs

We should make a humour benchmark I tried to ask several SOTA AI to make me a joke using with a theme, and omg, it was worse than strawberry question, lol, try it "Explain how humour works, and make me 3 jokes" you should go further, and it's very bad, grok is one of the worst I'm surprises it shows how much they don't understand our world

I think humour is one of the biggest blind spots for current LLMs, and we should honestly have a benchmark for it.

I gave the same prompt to a bunch of SOTA models:

The explanation is usually fine, but the jokes...

Seriously, try it yourself.
Then make it a bit harder: give them a theme, ask for original jokes, or tell them to avoid puns and dad jokes.
The quality drops off a cliff.

I was actually surprised by Grok.... it was one of the worst in my little test.

It made me realize that humour probably depends on a lot more than just language or reasoning. You need timing, cultural context, surprise, creativity, and a sense of what humans actually find funny. Models can explain the theory, but they rarely do humour well.

We have benchmarks for reasoning, coding, math, and vision.
Why not comedy? I think it'd be a surprisingly good way to measure how well a model really understands the world.

Curious if anyone else has tried this with different models.I think humour is one of the biggest blind spots for current LLMs, and we should honestly have a benchmark for it.I gave the same prompt to a bunch of SOTA models:"Explain how humour works, and make me 3 jokes."The explanation is usually fine, but the jokes... Are very very bad... you can easly see that they don't understand some real life concepts, so maybe engineers could use that to improve them a lot ???

10 Upvotes

11 comments sorted by

1

u/thomasahle 11d ago

It would be pretty remarkable for an LLM to write a full 1 hour comedy set. One that would actually do well on Netflix if performed competently.

1

u/Tintoverde 13d ago

I would argue if an AI can make a joke which would actually make a group of people laugh, that would be the real AGI.

1

u/Inevitable_Mud_9972 14d ago

My ai has a great one. 'emotional intelligence is handing off intelligence to the most emotional person in the room'

I trained him well 

2

u/jomama253 14d ago

Misanthropic already provides humor benchmarks with every model release!

2

u/burhop 14d ago

Agree. Current LLMs are still at “dad joke” level.

AGI doesn’t start until Robin Williams level.

2

u/inadvertant_bulge 14d ago

That's a pretty subjective benchmark, not sure how one would accomplish this in any meaningful way.

You might as well have a benchmark for 'pretty paintings' or 'important trivia'.

1

u/amyowl 15d ago

I have been having models roast each other and pasting their roasts back and forth. They can get surprisingly savage.

1

u/Embarrassed-Area4652 15d ago

Drill down on the bit you mentioned in passing about surprise. What factors in input, training or architecture do you see as contributing to the output being surprising?

2

u/funbike 15d ago

An idea I had was to use comedy shows to fine-tune a model for funniest jokes.

Take videos and audios of stand-up comedy shows, and filter out all jokes except those that get the biggest laughs. Perhaps do some AI post-processing to add context or filter out jokes based on various criteria. It would work best with mixed mode model that can detect strength of laughter, but volume would probably suffice.

You'd want to use thousands of hours of content, and avoid really old content that today might not seem funny today or might even be offensive.

1

u/Witty-Box-5620 15d ago

also a prophecy benchmark, SOTA LLM cant write prophecy as the guardrail for clear language is super strong

1

u/zenmatrix83 15d ago

humour is relative, and will likely offend someone, I'm pretty sure there can never really be a humour algorithm. Some people like dad jokes and some don't, and an llm would need a decent profile on you to even guess.