No Error Bars, No Result
“A difference smaller than the noise is not a difference.”
- Last reviewed
- 15 Sep 2026
- Source types
- 4 peer-reviewed or classic, 1 preprint
- Version
- v1.0 5 Oct 2026
The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status
In plain terms
Every score is an estimate. If two systems differ by less than the noise, nobody has shown that they differ.
Takeaways #
- Every eval score is an estimate from a limited set of questions and, often, a random process.
- Small test sets and small improvements frequently sit inside the margin of error.
- Report confidence intervals, and use paired comparisons when models answer the same questions.
- Decide how many test items you need before running the eval.
What it means #
If a model scores 80% on 200 questions, the “true” score on similar questions could easily be a few points higher or lower. A rough standard error for accuracy is the square root of p(1 - p) / n. For 80% on 200 items, that’s about 2.8 points, so a 95% confidence interval runs roughly 5.5 points in each direction. A rival model at 83% on the same test hasn’t clearly won anything.
Randomness adds more noise: sampling temperature, random seeds, and data order all shift results between runs.
The evidence #
reviewed 15 Sep 2026Significance testing is often skipped. Dror et al. (2018) surveyed NLP research practice, found statistical significance testing was often missing or misapplied, and wrote a practical guide to choosing the right test.
Many experiments are underpowered. Card et al. (2020) showed that many NLP experiments don’t have enough data to reliably detect the small improvements typically claimed, which means published gains can be exaggerated.
Variance rivals the gains. Bouthillier et al. (2021) found that variance from sources like data sampling, weight initialization, and hyperparameter choices is often as large as the differences papers report. They recommend randomizing as many sources of variation as possible.
A playbook for LLM evals. Miller (2024) treats eval questions as a sample from a larger population and lays out how to compute standard errors, cluster them when questions are related, compare models with paired differences, and plan sample sizes with power analysis.
Benchmarks rarely report it. Reuel et al. (2024) assessed 24 AI benchmarks against 46 best practices and found most don’t report statistical significance.
Use it #
- Put a 95% confidence interval next to every score.
- Compare models on the same items with paired tests.
- Cluster standard errors when items come in groups, like several questions about one document.
- For stochastic outputs, run multiple samples per item.
- Treat any gap inside the interval as a tie.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
Error bars show noise, not whether the metric is the right one or the sample is representative. When a difference is very large next to run-to-run variation, simple reporting can be enough.
Origins #
“No Error Bars, No Result” is this site’s name for a basic principle of statistics that AI evaluation has been slow to adopt consistently.
Sources #
- [1]Bouthillier, X., et al. (2021). Accounting for variance in machine learning benchmarks.Proceedings of Machine Learning and Systems (MLSys 2021)Open ↗ (opens in a new tab)
- [2]Card, D., Henderson, P., Khandelwal, U., Jia, R., Mahowald, K., & Jurafsky, D. (2020). With little power comes great responsibility.Proceedings of EMNLP 2020Open ↗ (opens in a new tab)
- [3]Dror, R., et al. (2018). The hitchhiker’s guide to testing statistical significance in natural language processing.Proceedings of ACL 2018Open ↗ (opens in a new tab)
- [4]Miller, E. (2024). Adding error bars to evals: A statistical approach to language model evaluations.arXiv:2411.00640 PreprintOpen ↗ (opens in a new tab)
- [5]Reuel, A., Hardy, A., Smith, C., Lamparth, M., Hardy, M., & Kochenderfer, M. J. (2024). BetterBench: Assessing AI benchmarks, uncovering issues, and establishing best practices.NeurIPS 2024 Datasets and Benchmarks TrackOpen ↗ (opens in a new tab)
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). No Error Bars, No Result (v1.0). https://evalfieldguide.com/patterns/no-error-bars-no-result@misc{lai-no-error-bars-no-result,
title = {No Error Bars, No Result},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/no-error-bars-no-result}}
}https://evalfieldguide.com/patterns/no-error-bars-no-resultRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.