1 of 5 in this categoryNext: 17 The Averaging Trap →
No. 16Resultsv1.0For: Buying · Building

No Error Bars, No Result

“A difference smaller than the noise is not a difference.”

Last reviewed
15 Sep 2026
Sources
5 · Cite
Source types
4 peer-reviewed or classic, 1 preprint
Version
v1.0 5 Oct 2026

The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status

In plain terms

Every score is an estimate. If two systems differ by less than the noise, nobody has shown that they differ.

Takeaways

  • Every eval score is an estimate from a limited set of questions and, often, a random process.
  • Small test sets and small improvements frequently sit inside the margin of error.
  • Report confidence intervals, and use paired comparisons when models answer the same questions.
  • Decide how many test items you need before running the eval.

What it means

If a model scores 80% on 200 questions, the “true” score on similar questions could easily be a few points higher or lower. A rough standard error for accuracy is the square root of p(1 - p) / n. For 80% on 200 items, that’s about 2.8 points, so a 95% confidence interval runs roughly 5.5 points in each direction. A rival model at 83% on the same test hasn’t clearly won anything.

Randomness adds more noise: sampling temperature, random seeds, and data order all shift results between runs.

The evidence

reviewed 15 Sep 2026

Significance testing is often skipped. Dror et al. (2018) surveyed NLP research practice, found statistical significance testing was often missing or misapplied, and wrote a practical guide to choosing the right test.

Many experiments are underpowered. Card et al. (2020) showed that many NLP experiments don’t have enough data to reliably detect the small improvements typically claimed, which means published gains can be exaggerated.

Variance rivals the gains. Bouthillier et al. (2021) found that variance from sources like data sampling, weight initialization, and hyperparameter choices is often as large as the differences papers report. They recommend randomizing as many sources of variation as possible.

A playbook for LLM evals. Miller (2024) treats eval questions as a sample from a larger population and lays out how to compute standard errors, cluster them when questions are related, compare models with paired differences, and plan sample sizes with power analysis.

Benchmarks rarely report it. Reuel et al. (2024) assessed 24 AI benchmarks against 46 best practices and found most don’t report statistical significance.

Use it

  1. Put a 95% confidence interval next to every score.
  2. Compare models on the same items with paired tests.
  3. Cluster standard errors when items come in groups, like several questions about one document.
  4. For stochastic outputs, run multiple samples per item.
  5. Treat any gap inside the interval as a tie.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

Error bars show noise, not whether the metric is the right one or the sample is representative. When a difference is very large next to run-to-run variation, simple reporting can be enough.

Origins

“No Error Bars, No Result” is this site’s name for a basic principle of statistics that AI evaluation has been slow to adopt consistently.

Sources

  1. [1]
    Bouthillier, X., et al. (2021). Accounting for variance in machine learning benchmarks.Proceedings of Machine Learning and Systems (MLSys 2021)
    Open ↗ (opens in a new tab)
  2. [2]
    Card, D., Henderson, P., Khandelwal, U., Jia, R., Mahowald, K., & Jurafsky, D. (2020). With little power comes great responsibility.Proceedings of EMNLP 2020
    Open ↗ (opens in a new tab)
  3. [3]
    Dror, R., et al. (2018). The hitchhiker’s guide to testing statistical significance in natural language processing.Proceedings of ACL 2018
    Open ↗ (opens in a new tab)
  4. [4]
    Miller, E. (2024). Adding error bars to evals: A statistical approach to language model evaluations.arXiv:2411.00640 Preprint
    Open ↗ (opens in a new tab)
  5. [5]
    Reuel, A., Hardy, A., Smith, C., Lamparth, M., Hardy, M., & Kochenderfer, M. J. (2024). BetterBench: Assessing AI benchmarks, uncovering issues, and establishing best practices.NeurIPS 2024 Datasets and Benchmarks Track
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). No Error Bars, No Result (v1.0). https://evalfieldguide.com/patterns/no-error-bars-no-result

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 17 →