← r/LocalLLaMA
▲
62
-4
28👁
r/LocalLLaMA · u/drooolingidiot · 21d ago

We benchmarked 24 LLMs against human writers on 475 creative writing prompts

post image

We just released the first version of our Creative Writing benchmark, comparing 24 LLMs against human writers across 475 writing prompts.

Creative writing is subjective, so the rankings aren't meant to predict what any one person will prefer. Instead, they predict what a large group of readers would prefer, using a custom reward model trained specifically on human preferences for creative writing.

Surprisingly, the strongest frontier models already rank above the talented amateur writer cohort, while professional writers still lead by a wide margin.

You can browse the full benchmark, compare the model outputs side by side, and see how the benchmark works here:

https://vulsar.ai/benchmarks/creative-writing-v1/

Curious what you all think of the results!

70 0 62 10/3 04:49 10/8 21:23 UTC
scorecomments28 sightings
first seen 2026-10-03 04:49 UTClast seen 2026-10-08 21:23 UTCscore then 66score now 62gained -4sightings 28
open on reddit ↗ 💬 77 (+1)