2 of 5 in this categoryNext: 08 Adaptive Overfitting →
No. 07The testv1.0For: Buying · Building

Data Contamination

“If the test was in the training data, the score measures memory, not skill.”

Last reviewed
15 Sep 2026
Sources
4 · Cite
Source types
3 peer-reviewed or classic, 1 report
Version
v1.0 5 Oct 2026

Established term: this name is used in the research literature. About this guide’s status

In plain terms

If a model saw the test questions while it was training, a high score may only mean it remembers them. Fresh or private questions are the reliable check.

Takeaways

  • Large models train on huge scrapes of the internet, and public benchmarks live on the internet.
  • Contaminated scores overstate how well a model handles problems it hasn’t seen.
  • You usually can’t inspect a model’s training data, so contamination has to be tested indirectly.
  • Fresh, private, or post-cutoff test sets are the most reliable defense.

What it means

Contamination, also called leakage, happens when information from the test gets into training. In classic machine learning it’s usually a pipeline bug, like normalizing data before splitting it. With large language models it’s often invisible and hard to avoid, because nobody outside the lab knows exactly what went into training. A model that has seen the questions, or close paraphrases, can score well without the skill the test was built to measure.

The evidence

reviewed 15 Sep 2026

Fresh questions, lower scores. Zhang et al. (2024) wrote GSM1k, a new set of grade school math problems matched to the popular GSM8K benchmark. Some model families dropped up to 8% in accuracy. Models that were more likely to generate GSM8K examples verbatim showed bigger gaps (Spearman’s r² = 0.36), which suggests partial memorization. Many frontier models held up well, so this is a check to run, not a blanket verdict.

Measure it per benchmark. Sainz et al. (2023) argue that contamination should be measured for each benchmark and model, because contaminated results lead to wrong conclusions in published research.

Leakage spans fields. Kapoor and Narayanan (2023) documented leakage errors affecting hundreds of papers across 17 scientific fields that use ML, and proposed a taxonomy of leakage types.

Still a current problem. The International AI Safety Report 2026 lists test questions that already appear in training data as one reason benchmark scores fail to reflect real-world use (Bengio et al., 2026).

Use it

  1. Build a small private test set from your own data that has never been posted online.
  2. Compare scores on public items against freshly written look-alike items. A big gap is a red flag.
  3. Check the benchmark’s release date against the model’s training cutoff.
  4. Ask vendors what contamination checks they ran and how.
  5. In your own pipelines, split data by time, user, or source before any preprocessing.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

Many frontier models held up well on fresh questions, so contamination is a check to run, not a verdict. A freshly written test set, or documentation of what training data was excluded, reduces the concern.

Origins

Leakage has been a known hazard in statistics and machine learning for decades. The scale of web-trained language models turned it from an occasional bug into a standing concern for every public benchmark.

Sources

  1. [1]
    Bengio, Y., et al. (2026). International AI Safety Report 2026.International AI Safety Report; arXiv:2602.21012 Report
    Open ↗ (opens in a new tab)
  2. [2]
    Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science.Patterns, 4(9), 100804
    Open ↗ (opens in a new tab)
  3. [3]
    Sainz, O., Campos, J. A., García-Ferrero, I., Etxaniz, J., Lopez de Lacalle, O., & Agirre, E. (2023). NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark.Findings of EMNLP 2023
    Open ↗ (opens in a new tab)
  4. [4]
    Zhang, H., Da, J., Lee, D., Robinson, V., Wu, C., et al. (2024). A careful examination of large language model performance on grade school arithmetic.NeurIPS 2024 Datasets and Benchmarks Track
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). Data Contamination (v1.0). https://evalfieldguide.com/patterns/data-contamination

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 08 →