← r/LocalLLaMA
▲
13
+7
16👁
r/LocalLLaMA · u/empirical-sadboy · 3d ago

Can we please have some error bars?

I am sure this gripe has been raised many times before, but every time a new model is released it seems like it's routinely only a few percentage points higher than previous models on benchmarks.

How do we know this is even a "real" difference and not just within the window of measurement error or noise?

Some quick back-of-envelope math: HumanEval has 164 problems, so a model scoring \~70% has a standard error of roughly 3.5 points from question sampling alone. GSM8K (\~1.3k questions) is closer to 1 point. A 2-point "improvement" on either is well inside the noise, and that's before counting anything else that varies: sampling temperature, prompt template, few-shot examples, eval harness version, and possible contamination. There's work showing that trivial formatting changes can swing scores by many points, which is often bigger than the gap between models on the leaderboard.

None of this is hard to fix. Report the number of items, bootstrap confidence intervals, and ideally multiple seeds. Since two models are scored on the same questions, a paired test is much more powerful than eyeballing two accuracies. Miller's "Adding Error Bars to Evals" lays this out well.

Am I missing something, or is a lot of the benchmark chasing just reading tea leaves? Does anyone know of leaderboards or labs that routinely report uncertainty?

16 0 13 10/6 17:47 10/8 19:41 UTC
scorecomments16 sightings
first seen 2026-10-06 17:47 UTClast seen 2026-10-08 19:41 UTCscore then 6score now 13gained +7sightings 16
open on reddit ↗ 💬 4 (+1)