Humor Arena: Which LLM is the funniest?
We compared 20 model versions on 360 frozen joke prompts, with four jokes requested per prompt and model names hidden from our humor-trained judge.
Fable 5 had the highest estimated score: 66.8 points per 100 comparisons against the rival field. Fable 5.1 scored 58.2. Scores count a win as one point and a tie as half. The models near the top have overlapping uncertainty intervals. The reason we say estimated is that we scale the scores based on the scores of
The scores come from our own automated judge. A separate audit checked the judge with 1,400 ratings from 50 people; that was not a fresh human evaluation of Fable 5.1’s outputs. We specifically fine tuned an OS judge to rate the jokes and it correlates more highly with human preferences than any other model.
The full details here:
https://laugh.so/research/joke-generation/
Would love to hear your feedback!