Adaptive Overfitting
“Every time you tune against the test set, it’s worth a little less.”
- Last reviewed
- 15 Sep 2026
- Source types
- 4 peer-reviewed or classic
- Version
- v1.0 5 Oct 2026
Established term: this name is used in the research literature. About this guide’s status
In plain terms
Each time you adjust a prompt or a model because of how it did on the test, the test tells you a little less. Keep a locked test that you use only to decide.
Takeaways #
- A test set gives an unbiased read only the first time you use it.
- Choosing prompts, models, or settings based on test results slowly fits your system to that specific test.
- The effect is sometimes smaller than people fear, but you can’t know without a fresh test.
- Separate a dev set you iterate on from a locked test set you decide with.
What it means #
Standard statistical guarantees assume you pick your analysis before looking at the data. Real teams don’t work that way. They run the eval, tweak the prompt, run it again, try another model, run it again. Each look leaks a little information about the test set into the choices being made. After enough rounds, a high score partly reflects how well the team learned that particular test.
The evidence #
reviewed 15 Sep 2026Reuse breaks the math. Dwork et al. (2015) showed that adaptive reuse of a holdout set undermines its validity, and proposed a “reusable holdout” method that limits how much each query reveals.
Rebuilt test sets score lower. Recht et al. (2019) rebuilt test sets for CIFAR-10 and ImageNet following the original procedures. Accuracy dropped 3% to 15% on CIFAR-10 and 11% to 14% on ImageNet. The model rankings mostly held, and the authors attribute the drop mainly to subtle differences in the new data rather than years of adaptive overfitting. That nuance matters: the pattern is real, and its size varies.
Competitions held up better than expected. Roelofs et al. (2019) analyzed 120 Kaggle competitions and found little evidence of substantial overfitting from leaderboard reuse.
Agent benchmarks are exposed. Kapoor et al. (2025) found many AI agent benchmarks have inadequate holdout sets, and sometimes none, which lets agents overfit and take shortcuts.
Use it #
- Keep three splits: iterate on dev, check on validation, decide on a locked test.
- Log each time someone looks at locked test results.
- Refresh the locked test once it has driven many decisions.
- Treat small wins on a heavily reused test with suspicion.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
How large the effect is varies. Recht et al. attribute most of their drop to subtle differences in the new data, and Roelofs et al. found little evidence of substantial overfitting across 120 Kaggle competitions. Treat it as a risk that grows with reuse, not a constant.
Origins #
The holdout set is one of the oldest ideas in machine learning. “Adaptive data analysis” became its own research area in the 2010s as researchers formalized what happens when analysts reuse data across many decisions.
Sources #
- [1]Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). The reusable holdout: Preserving validity in adaptive data analysis.Science, 349(6248), 636-638Open ↗ (opens in a new tab)
- [2]Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2025). AI agents that matter.Transactions on Machine Learning Research (TMLR)Open ↗ (opens in a new tab)
- [3]Recht, B., Roelofs, R., Schmidt, L., & Shankar, V. (2019). Do ImageNet classifiers generalize to ImageNet?.Proceedings of the International Conference on Machine Learning (ICML 2019)Open ↗ (opens in a new tab)
- [4]Roelofs, R., Shankar, V., Recht, B., Fridovich-Keil, S., Hardt, M., Miller, J., & Schmidt, L. (2019). A meta-analysis of overfitting in machine learning.Advances in Neural Information Processing Systems (NeurIPS 2019)No link
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). Adaptive Overfitting (v1.0). https://evalfieldguide.com/patterns/adaptive-overfitting@misc{lai-adaptive-overfitting,
title = {Adaptive Overfitting},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/adaptive-overfitting}}
}https://evalfieldguide.com/patterns/adaptive-overfittingRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.