Data Contamination
“If the test was in the training data, the score measures memory, not skill.”
- Last reviewed
- 15 Sep 2026
- Source types
- 3 peer-reviewed or classic, 1 report
- Version
- v1.0 5 Oct 2026
Established term: this name is used in the research literature. About this guide’s status
In plain terms
If a model saw the test questions while it was training, a high score may only mean it remembers them. Fresh or private questions are the reliable check.
Takeaways #
- Large models train on huge scrapes of the internet, and public benchmarks live on the internet.
- Contaminated scores overstate how well a model handles problems it hasn’t seen.
- You usually can’t inspect a model’s training data, so contamination has to be tested indirectly.
- Fresh, private, or post-cutoff test sets are the most reliable defense.
What it means #
Contamination, also called leakage, happens when information from the test gets into training. In classic machine learning it’s usually a pipeline bug, like normalizing data before splitting it. With large language models it’s often invisible and hard to avoid, because nobody outside the lab knows exactly what went into training. A model that has seen the questions, or close paraphrases, can score well without the skill the test was built to measure.
The evidence #
reviewed 15 Sep 2026Fresh questions, lower scores. Zhang et al. (2024) wrote GSM1k, a new set of grade school math problems matched to the popular GSM8K benchmark. Some model families dropped up to 8% in accuracy. Models that were more likely to generate GSM8K examples verbatim showed bigger gaps (Spearman’s r² = 0.36), which suggests partial memorization. Many frontier models held up well, so this is a check to run, not a blanket verdict.
Measure it per benchmark. Sainz et al. (2023) argue that contamination should be measured for each benchmark and model, because contaminated results lead to wrong conclusions in published research.
Leakage spans fields. Kapoor and Narayanan (2023) documented leakage errors affecting hundreds of papers across 17 scientific fields that use ML, and proposed a taxonomy of leakage types.
Still a current problem. The International AI Safety Report 2026 lists test questions that already appear in training data as one reason benchmark scores fail to reflect real-world use (Bengio et al., 2026).
Use it #
- Build a small private test set from your own data that has never been posted online.
- Compare scores on public items against freshly written look-alike items. A big gap is a red flag.
- Check the benchmark’s release date against the model’s training cutoff.
- Ask vendors what contamination checks they ran and how.
- In your own pipelines, split data by time, user, or source before any preprocessing.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
Many frontier models held up well on fresh questions, so contamination is a check to run, not a verdict. A freshly written test set, or documentation of what training data was excluded, reduces the concern.
Origins #
Leakage has been a known hazard in statistics and machine learning for decades. The scale of web-trained language models turned it from an occasional bug into a standing concern for every public benchmark.
Sources #
- [1]Bengio, Y., et al. (2026). International AI Safety Report 2026.International AI Safety Report; arXiv:2602.21012 ReportOpen ↗ (opens in a new tab)
- [2]Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science.Patterns, 4(9), 100804Open ↗ (opens in a new tab)
- [3]Sainz, O., Campos, J. A., García-Ferrero, I., Etxaniz, J., Lopez de Lacalle, O., & Agirre, E. (2023). NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark.Findings of EMNLP 2023Open ↗ (opens in a new tab)
- [4]Zhang, H., Da, J., Lee, D., Robinson, V., Wu, C., et al. (2024). A careful examination of large language model performance on grade school arithmetic.NeurIPS 2024 Datasets and Benchmarks TrackOpen ↗ (opens in a new tab)
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). Data Contamination (v1.0). https://evalfieldguide.com/patterns/data-contamination@misc{lai-data-contamination,
title = {Data Contamination},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/data-contamination}}
}https://evalfieldguide.com/patterns/data-contaminationRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.