The Metric Mirage
“Change the metric and you can change the conclusion.”
- Last reviewed
- 15 Sep 2026
- Source types
- 4 peer-reviewed or classic
- Version
- v1.0 5 Oct 2026
The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status
In plain terms
The scoring method is part of the result. Score the same outputs a different way and a sudden leap can turn into steady progress, or the reverse.
Takeaways #
- The same model outputs can look like steady progress or a sudden leap, depending on how you score them.
- All-or-nothing metrics like exact match make gradual improvement look abrupt.
- Cheap automatic metrics often disagree with people about quality.
- Before believing a trend, re-score it with another reasonable metric and see if the story holds.
What it means #
The metric is part of the finding. Pick a strict pass/fail score and a model that gets closer and closer to right answers looks like it’s stuck at zero, then suddenly “gets it.” Pick a word-overlap score and a fluent, correct answer phrased differently than the reference looks wrong. Neither number is lying, exactly. Each answers a narrower question than the headline suggests.
The evidence #
reviewed 15 Sep 2026Emergence can be a scoring artifact. Schaeffer, Miranda, and Koyejo (2023) argue that many “emergent abilities” in large language models appear because of the researcher’s choice of metric. Nonlinear or discontinuous metrics produce sharp jumps. Linear or continuous metrics applied to the same models show smooth, predictable change.
Automatic metrics have narrow validity. Reiter (2018) reviewed the evidence on BLEU and found it supports using BLEU for diagnostic evaluation of machine translation systems, not for evaluating individual texts or other language generation tasks.
Overlap isn’t quality. Novikova et al. (2017) found that widely used word-overlap metrics correlate weakly with human judgments of generated text.
One metric hides tradeoffs. Liang et al. (2023) built HELM around seven metrics (accuracy, calibration, robustness, fairness, bias, toxicity, efficiency) because accuracy alone hides the tradeoffs that matter in use.
Use it #
- Report at least one continuous metric alongside any pass/fail metric.
- Check automatic metrics against human judgment on a sample before relying on them.
- When someone shows a dramatic curve, ask what the y-axis measures and what the curve looks like with a different metric.
- Choose metrics before seeing results, and write down why.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
Automatic metrics are not useless. Reiter found BLEU supports diagnostic evaluation of machine translation systems. The problem is using a metric outside the range where it has been shown to work.
Origins #
“The Metric Mirage” is the name this site uses, borrowed from the title of Schaeffer, Miranda, and Koyejo’s 2023 paper. The broader concern about automatic metrics goes back years in natural language generation research.
Sources #
- [1]Liang, P., et al. (2023). Holistic evaluation of language models.Transactions on Machine Learning Research (TMLR)Open ↗ (opens in a new tab)
- [2]Novikova, J., Dušek, O., Cercas Curry, A., & Rieser, V. (2017). Why we need new evaluation metrics for NLG.Proceedings of EMNLP 2017Open ↗ (opens in a new tab)
- [3]Reiter, E. (2018). A structured review of the validity of BLEU.Computational Linguistics, 44(3), 393-401Open ↗ (opens in a new tab)
- [4]Schaeffer, R., Miranda, B., & Koyejo, S. (2023). Are emergent abilities of large language models a mirage?.Advances in Neural Information Processing Systems (NeurIPS 2023)Open ↗ (opens in a new tab)
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). The Metric Mirage (v1.0). https://evalfieldguide.com/patterns/the-metric-mirage@misc{lai-the-metric-mirage,
title = {The Metric Mirage},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/the-metric-mirage}}
}https://evalfieldguide.com/patterns/the-metric-mirageRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.