5 of 5 in this categoryNext: 06 Benchmark Saturation →
No. 05Measuringv1.0For: Building

The Metric Mirage

“Change the metric and you can change the conclusion.”

Last reviewed
15 Sep 2026
Sources
4 · Cite
Source types
4 peer-reviewed or classic
Version
v1.0 5 Oct 2026

The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status

In plain terms

The scoring method is part of the result. Score the same outputs a different way and a sudden leap can turn into steady progress, or the reverse.

Takeaways

  • The same model outputs can look like steady progress or a sudden leap, depending on how you score them.
  • All-or-nothing metrics like exact match make gradual improvement look abrupt.
  • Cheap automatic metrics often disagree with people about quality.
  • Before believing a trend, re-score it with another reasonable metric and see if the story holds.

What it means

The metric is part of the finding. Pick a strict pass/fail score and a model that gets closer and closer to right answers looks like it’s stuck at zero, then suddenly “gets it.” Pick a word-overlap score and a fluent, correct answer phrased differently than the reference looks wrong. Neither number is lying, exactly. Each answers a narrower question than the headline suggests.

The evidence

reviewed 15 Sep 2026

Emergence can be a scoring artifact. Schaeffer, Miranda, and Koyejo (2023) argue that many “emergent abilities” in large language models appear because of the researcher’s choice of metric. Nonlinear or discontinuous metrics produce sharp jumps. Linear or continuous metrics applied to the same models show smooth, predictable change.

Automatic metrics have narrow validity. Reiter (2018) reviewed the evidence on BLEU and found it supports using BLEU for diagnostic evaluation of machine translation systems, not for evaluating individual texts or other language generation tasks.

Overlap isn’t quality. Novikova et al. (2017) found that widely used word-overlap metrics correlate weakly with human judgments of generated text.

One metric hides tradeoffs. Liang et al. (2023) built HELM around seven metrics (accuracy, calibration, robustness, fairness, bias, toxicity, efficiency) because accuracy alone hides the tradeoffs that matter in use.

Use it

  1. Report at least one continuous metric alongside any pass/fail metric.
  2. Check automatic metrics against human judgment on a sample before relying on them.
  3. When someone shows a dramatic curve, ask what the y-axis measures and what the curve looks like with a different metric.
  4. Choose metrics before seeing results, and write down why.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

Automatic metrics are not useless. Reiter found BLEU supports diagnostic evaluation of machine translation systems. The problem is using a metric outside the range where it has been shown to work.

Origins

“The Metric Mirage” is the name this site uses, borrowed from the title of Schaeffer, Miranda, and Koyejo’s 2023 paper. The broader concern about automatic metrics goes back years in natural language generation research.

Sources

  1. [1]
    Liang, P., et al. (2023). Holistic evaluation of language models.Transactions on Machine Learning Research (TMLR)
    Open ↗ (opens in a new tab)
  2. [2]
    Novikova, J., Dušek, O., Cercas Curry, A., & Rieser, V. (2017). Why we need new evaluation metrics for NLG.Proceedings of EMNLP 2017
    Open ↗ (opens in a new tab)
  3. [3]
    Reiter, E. (2018). A structured review of the validity of BLEU.Computational Linguistics, 44(3), 393-401
    Open ↗ (opens in a new tab)
  4. [4]
    Schaeffer, R., Miranda, B., & Koyejo, S. (2023). Are emergent abilities of large language models a mirage?.Advances in Neural Information Processing Systems (NeurIPS 2023)
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). The Metric Mirage (v1.0). https://evalfieldguide.com/patterns/the-metric-mirage

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 06 →