3 of 5 in this categoryNext: 09 The Clever Hans Effect →
No. 08The testv1.0For: Building

Adaptive Overfitting

“Every time you tune against the test set, it’s worth a little less.”

Last reviewed
15 Sep 2026
Sources
4 · Cite
Source types
4 peer-reviewed or classic
Version
v1.0 5 Oct 2026

Established term: this name is used in the research literature. About this guide’s status

In plain terms

Each time you adjust a prompt or a model because of how it did on the test, the test tells you a little less. Keep a locked test that you use only to decide.

Takeaways

  • A test set gives an unbiased read only the first time you use it.
  • Choosing prompts, models, or settings based on test results slowly fits your system to that specific test.
  • The effect is sometimes smaller than people fear, but you can’t know without a fresh test.
  • Separate a dev set you iterate on from a locked test set you decide with.

What it means

Standard statistical guarantees assume you pick your analysis before looking at the data. Real teams don’t work that way. They run the eval, tweak the prompt, run it again, try another model, run it again. Each look leaks a little information about the test set into the choices being made. After enough rounds, a high score partly reflects how well the team learned that particular test.

The evidence

reviewed 15 Sep 2026

Reuse breaks the math. Dwork et al. (2015) showed that adaptive reuse of a holdout set undermines its validity, and proposed a “reusable holdout” method that limits how much each query reveals.

Rebuilt test sets score lower. Recht et al. (2019) rebuilt test sets for CIFAR-10 and ImageNet following the original procedures. Accuracy dropped 3% to 15% on CIFAR-10 and 11% to 14% on ImageNet. The model rankings mostly held, and the authors attribute the drop mainly to subtle differences in the new data rather than years of adaptive overfitting. That nuance matters: the pattern is real, and its size varies.

Competitions held up better than expected. Roelofs et al. (2019) analyzed 120 Kaggle competitions and found little evidence of substantial overfitting from leaderboard reuse.

Agent benchmarks are exposed. Kapoor et al. (2025) found many AI agent benchmarks have inadequate holdout sets, and sometimes none, which lets agents overfit and take shortcuts.

Use it

  1. Keep three splits: iterate on dev, check on validation, decide on a locked test.
  2. Log each time someone looks at locked test results.
  3. Refresh the locked test once it has driven many decisions.
  4. Treat small wins on a heavily reused test with suspicion.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

How large the effect is varies. Recht et al. attribute most of their drop to subtle differences in the new data, and Roelofs et al. found little evidence of substantial overfitting across 120 Kaggle competitions. Treat it as a risk that grows with reuse, not a constant.

Origins

The holdout set is one of the oldest ideas in machine learning. “Adaptive data analysis” became its own research area in the 2010s as researchers formalized what happens when analysts reuse data across many decisions.

Sources

  1. [1]
    Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). The reusable holdout: Preserving validity in adaptive data analysis.Science, 349(6248), 636-638
    Open ↗ (opens in a new tab)
  2. [2]
    Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2025). AI agents that matter.Transactions on Machine Learning Research (TMLR)
    Open ↗ (opens in a new tab)
  3. [3]
    Recht, B., Roelofs, R., Schmidt, L., & Shankar, V. (2019). Do ImageNet classifiers generalize to ImageNet?.Proceedings of the International Conference on Machine Learning (ICML 2019)
    Open ↗ (opens in a new tab)
  4. [4]
    Roelofs, R., Shankar, V., Recht, B., Fridovich-Keil, S., Hardt, M., Miller, J., & Schmidt, L. (2019). A meta-analysis of overfitting in machine learning.Advances in Neural Information Processing Systems (NeurIPS 2019)
    No link

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). Adaptive Overfitting (v1.0). https://evalfieldguide.com/patterns/adaptive-overfitting

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 09 →