The Baseline Rule
“A result is only as strong as the baseline it beats.”
- Last reviewed
- 15 Sep 2026
- Source types
- 4 peer-reviewed or classic
- Version
- v1.0 5 Oct 2026
The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status
In plain terms
A win is only as strong as what it beat. Compare against the simplest thing that could work and today’s actual process, with the same tuning effort.
Takeaways #
- Beating a weak or untuned baseline can create the illusion of progress.
- A fair comparison gives the old method the same tuning effort as the new one.
- Always include a simple baseline: a rule, a smaller model, a non-AI process, or today’s human workflow.
- In product work, the real baseline is how people do the task right now.
What it means #
“Our model is 20% better” means nothing until you know better than what. New methods usually get weeks of tuning. Baselines often get copied from an old paper with default settings. The gap between them can be mostly effort, not ideas. In product settings, teams sometimes compare an AI feature to a different AI feature and forget to compare it to the plain workflow people already use.
The evidence #
reviewed 15 Sep 2026Tuned old methods win. Melis, Dyer, and Blunsom (2018) found that standard LSTM language models, carefully tuned, outperformed more recent and more complex architectures.
Neural hype in search. Lin (2019) documented information retrieval papers reporting gains over weak baselines that didn’t hold against well-tuned classic methods.
Recommenders didn’t reproduce. Ferrari Dacrema, Cremonesi, and Jannach (2019) tried to reproduce 18 neural recommendation algorithms from top conferences. Only 7 could be reproduced with reasonable effort, and 6 of those were often outperformed by simple heuristic methods.
Simple methods capture most of the value. Hand (2006) argued that simple classifiers often capture most of the achievable predictive power, and that the gains from sophisticated methods can vanish in real use.
Use it #
- Include “the simplest thing that could work” and “the current process” in every comparison.
- Give baselines an equal tuning budget, and say so.
- Be suspicious when a new method beats baselines by a wide margin on an old benchmark.
- For AI features, compare against the non-AI workflow with real users.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
A baseline is a check on a comparison, not a verdict against new methods. Where a complex method really adds value, a fair comparison will show it.
Origins #
“The Baseline Rule” is this site’s name for a principle that shows up in nearly every field that runs experiments. Hand’s “illusion of progress” (2006) is an early, clear statement in machine learning.
Sources #
- [1]Ferrari Dacrema, M., Cremonesi, P., & Jannach, D. (2019). Are we really making much progress? A worrying analysis of recent neural recommendation approaches.Proceedings of the ACM Conference on Recommender Systems (RecSys 2019)Open ↗ (opens in a new tab)
- [2]Hand, D. J. (2006). Classifier technology and the illusion of progress.Statistical Science, 21(1), 1-14Open ↗ (opens in a new tab)
- [3]Lin, J. (2019). The neural hype and comparisons against weak baselines.ACM SIGIR Forum, 52(2), 40-51Open ↗ (opens in a new tab)
- [4]Melis, G., Dyer, C., & Blunsom, P. (2018). On the state of the art of evaluation in neural language models.International Conference on Learning Representations (ICLR 2018)Open ↗ (opens in a new tab)
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). The Baseline Rule (v1.0). https://evalfieldguide.com/patterns/the-baseline-rule@misc{lai-the-baseline-rule,
title = {The Baseline Rule},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/the-baseline-rule}}
}https://evalfieldguide.com/patterns/the-baseline-ruleRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.