Distribution Shift
“Performance measured on one population doesn’t automatically carry over to another.”
- Last reviewed
- 15 Sep 2026
- Source types
- 4 peer-reviewed or classic
- Version
- v1.0 5 Oct 2026
Established term: this name is used in the research literature. About this guide’s status
In plain terms
A score measured on one group of people, places, or times doesn’t automatically carry over to another. Test on data from where the system will actually run.
Takeaways #
- A test set is a snapshot of certain people, places, times, and conditions.
- Change any of those and performance can drop, sometimes sharply.
- Even careful rebuilds of the same benchmark produce lower scores.
- Evaluate on data from the setting where the system will run, and keep checking over time.
What it means #
Models learn patterns from the data they’re trained on and get tested on data that usually looks a lot like it. The real world rarely cooperates. A new hospital has different patients and equipment. A new market uses different language. Next year’s users behave differently than last year’s. Each change is a shift, and a benchmark score doesn’t tell you how the system handles it.
The evidence #
reviewed 15 Sep 2026Same recipe, lower scores. Recht et al. (2019) rebuilt the CIFAR-10 and ImageNet test sets by following the original collection process. Accuracy fell 3% to 15% on CIFAR-10 and 11% to 14% on ImageNet. Small differences in how data was gathered were enough.
Real-world shifts, measured. Koh et al. (2021) built WILDS, a benchmark of ten datasets with real distribution shifts across hospitals, cameras, regions, and time. Models showed large gaps between in-distribution and out-of-distribution performance.
A clinical example. Wong et al. (2021) found a widely deployed sepsis model performed far worse at an independent health system than its developer reported.
Drift erodes gains. Hand (2006) argued that population drift over time is one reason gains from sophisticated models shrink in practice.
Use it #
- Ask where and when the eval data came from, and compare that to your actual users.
- Split test results by site, time period, or source to see how performance moves.
- Run a local validation before rolling out to a new market, language, or customer segment.
- Monitor production for drift and re-evaluate on a schedule.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
Not every change in data hurts. In Recht et al., model rankings mostly held. The question is whether your deployment differs from the test data in ways the model is sensitive to.
Origins #
Distribution shift, also called dataset shift, is a long-standing problem in statistics and machine learning. It’s closely related to external validity in the social and medical sciences.
Sources #
- [1]Hand, D. J. (2006). Classifier technology and the illusion of progress.Statistical Science, 21(1), 1-14Open ↗ (opens in a new tab)
- [2]Koh, P. W., et al. (2021). WILDS: A benchmark of in-the-wild distribution shifts.Proceedings of the International Conference on Machine Learning (ICML 2021)Open ↗ (opens in a new tab)
- [3]Recht, B., Roelofs, R., Schmidt, L., & Shankar, V. (2019). Do ImageNet classifiers generalize to ImageNet?.Proceedings of the International Conference on Machine Learning (ICML 2019)Open ↗ (opens in a new tab)
- [4]Wong, A., Otles, E., Donnelly, J. P., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients.JAMA Internal Medicine, 181(8), 1065-1070Open ↗ (opens in a new tab)
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). Distribution Shift (v1.0). https://evalfieldguide.com/patterns/distribution-shift@misc{lai-distribution-shift,
title = {Distribution Shift},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/distribution-shift}}
}https://evalfieldguide.com/patterns/distribution-shiftRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.