Does post training make LLMs funnier?
We did a study: does post training actually make LLMs funnier?
We used open models that publish every stage of post training, so we could compare a base model with the future models it became: Tulu 3 (on Llama 3.1 70B), OLMo 3.1 32B and Qwen2.5. We tracked 11 stages, 100 joke prompts, 64 human raters and 2,330 head-to-head judgments.
What we found: post training makes models funnier, but reduces diversity of response.
\- In 5 of 7 training steps, the later model's jokes were judged funnier. Jokes also got 10–20 words shorter after early post training, so they get to the punchline faster.
\- In 6 of 7 steps, the jokes a model wrote for the same prompt got more similar to each other. Ask for eight jokes on one premise and you get eight versions of the same joke. The biggest drop was Qwen2.5 base to instruct.
\- Asking the model to plan a line or two before the joke cut variety in all 4 models we tried, with no reliable gain in funniness.
\- A comedian persona won back a little variety in all 4 models, but only made the jokes funnier in 2 of them.
Humans judged the base versus final. A model judge calibrated on those votes compares the stages in between.
Full report and paper below. Which open models should we run through this next?