4 of 5 in this categoryNext: 20 The Reproducibility Rule →
No. 19Resultsv1.0For: Buying

The Benchmark Lottery

“Which benchmarks you pick can decide who wins.”

Last reviewed
15 Sep 2026
Sources
4 · Cite
Source types
3 peer-reviewed or classic, 1 preprint
Version
v1.0 5 Oct 2026

Established term: this name is used in the research literature. About this guide’s status

In plain terms

Which benchmarks get reported can decide who looks best. Pick the suite before you see results, and read the whole table.

Takeaways

  • Model rankings often change depending on which benchmarks are included.
  • A method can look like a breakthrough on some tasks and ordinary on others.
  • Choosing benchmarks after seeing results is a quiet form of cherry-picking.
  • Pick benchmarks that match your use case, and pick them before you run anything.

What it means

There are thousands of benchmarks. Any given paper or product announcement reports a handful. If the handful was chosen after seeing results, the table tells you more about the selection than about the system. Even without bad intent, communities settle on benchmarks that happen to favor certain approaches, and newcomers get judged by rules that weren’t written for them.

The evidence

reviewed 15 Sep 2026

Named and measured. Dehghani et al. (2021) showed that the relative ranking of methods can change substantially depending on the benchmark tasks chosen, and called this the benchmark lottery.

Models weren’t even compared on the same tests. Liang et al. (2023) found that before HELM, language models had on average been evaluated on just 17.9% of HELM’s core scenarios, with some prominent models sharing no scenarios at all. HELM raised that to 96%.

Whose utility? Ethayarajh and Jurafsky (2020) show that a leaderboard’s ranking reflects one implicit set of priorities that may not match any particular user’s.

An interdisciplinary warning. Eriksson et al. (2025) reviewed about 100 studies on benchmark shortcomings and found recurring problems, including data contamination, construct validity issues, misaligned incentives, and gaming of results, all shaped by commercial and competitive pressure.

Use it

  1. Choose your benchmark suite before seeing any results, and write it down.
  2. Weight benchmarks by relevance to your use case, not by popularity.
  3. When reading a claim, notice which common benchmarks are missing.
  4. Look at the full results table, not just the highlighted wins.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

A benchmark can still be the right test for a clearly defined task. The lottery shows up when one benchmark’s result is treated as proof of broad superiority.

Origins

The term comes from Mostafa Dehghani and colleagues’ 2021 paper, “The Benchmark Lottery.”

Sources

  1. [1]
    Dehghani, M., Tay, Y., Gritsenko, A. A., et al. (2021). The benchmark lottery.arXiv:2107.07002 Preprint
    Open ↗ (opens in a new tab)
  2. [2]
    Eriksson, M., et al. (2025). Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation.Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2025)
    Open ↗ (opens in a new tab)
  3. [3]
    Ethayarajh, K., & Jurafsky, D. (2020). Utility is in the eye of the user: A critique of NLP leaderboards.Proceedings of EMNLP 2020
    Open ↗ (opens in a new tab)
  4. [4]
    Liang, P., et al. (2023). Holistic evaluation of language models.Transactions on Machine Learning Research (TMLR)
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). The Benchmark Lottery (v1.0). https://evalfieldguide.com/patterns/the-benchmark-lottery

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 20 →