5 of 5 in this categoryNext: 21 Presence, Not Absence →
No. 20Resultsv1.0For: Building · Buying

The Reproducibility Rule

“An eval nobody can rerun is an anecdote.”

Last reviewed
15 Sep 2026
Sources
6 · Cite
Source types
5 peer-reviewed or classic, 1 preprint
Version
v1.0 5 Oct 2026

The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status

In plain terms

If nobody can rerun an evaluation, it is an anecdote. Record the model version, prompts, settings, data, and scoring code, and save the raw outputs.

Takeaways

  • Eval results depend on details: prompts, decoding settings, data versions, scoring code, and the exact model version.
  • Without those details, nobody can check, compare, or build on a result.
  • Documentation standards like model cards and datasheets exist so results carry their context.
  • Save everything needed to rerun the eval, including raw outputs.

What it means

Hosted AI models change under the same name. Prompts get tweaked. Scoring scripts get patched. Six months later, nobody can say why the number was 82% or whether today’s 79% is worse. A result you can’t reproduce can’t be trusted, and it can’t be compared to anything.

The evidence

reviewed 15 Sep 2026

The field built programs for it. Pineau et al. (2021) report on the NeurIPS 2019 reproducibility program, which introduced a code submission policy, a community reproducibility challenge, and a reproducibility checklist.

Details change scores. Biderman et al. (2024) share lessons from maintaining a widely used language model evaluation harness and show how small methodological choices shift results. They recommend sharing code, prompts, and outputs.

Benchmarks fall short. Reuel et al. (2024) found most of the 24 AI benchmarks they assessed can’t be easily replicated.

Agents especially. Kapoor et al. (2025) describe a pervasive lack of reproducibility in AI agent evaluation caused by nonstandard practices.

Standard formats exist. Mitchell et al. (2019) proposed model cards to document intended use and evaluation results. Gebru et al. (2021) proposed datasheets to document how datasets were created and what they’re suited for.

Use it

  1. Record the model name and version (or date), system prompt, decoding settings, dataset version, scoring code version, and run date.
  2. Save raw outputs, not just scores.
  3. Write a one-page eval card: purpose, data, metrics, slices, and known limits.
  4. Keep a small canary set and rerun it whenever a vendor updates a model.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

Full reproducibility is not always possible, for example with proprietary models or APIs that change. In those cases the goal shifts to documenting everything that can be shared.

Origins

“The Reproducibility Rule” is this site’s name. Reproducibility has been a stated norm in science for centuries. Machine learning formalized it through checklists, model cards, and datasheets starting in the late 2010s.

Sources

  1. [1]
    Biderman, S., et al. (2024). Lessons from the trenches on reproducible evaluation of language models.arXiv:2405.14782 Preprint
    Open ↗ (opens in a new tab)
  2. [2]
    Gebru, T., et al. (2021). Datasheets for datasets.Communications of the ACM, 64(12), 86-92
    Open ↗ (opens in a new tab)
  3. [3]
    Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2025). AI agents that matter.Transactions on Machine Learning Research (TMLR)
    Open ↗ (opens in a new tab)
  4. [4]
    Mitchell, M., et al. (2019). Model cards for model reporting.Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2019)
    Open ↗ (opens in a new tab)
  5. [5]
    Pineau, J., et al. (2021). Improving reproducibility in machine learning research (A report from the NeurIPS 2019 Reproducibility Program).Journal of Machine Learning Research, 22(164), 1-20
    Open ↗ (opens in a new tab)
  6. [6]
    Reuel, A., Hardy, A., Smith, C., Lamparth, M., Hardy, M., & Kochenderfer, M. J. (2024). BetterBench: Assessing AI benchmarks, uncovering issues, and establishing best practices.NeurIPS 2024 Datasets and Benchmarks Track
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). The Reproducibility Rule (v1.0). https://evalfieldguide.com/patterns/the-reproducibility-rule

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 21 →