3 of 6 in this categoryNext: 24 The Lab-to-Field Gap →
No. 23Beyondv1.0For: Building · Buying

Distribution Shift

“Performance measured on one population doesn’t automatically carry over to another.”

Last reviewed
15 Sep 2026
Sources
4 · Cite
Source types
4 peer-reviewed or classic
Version
v1.0 5 Oct 2026

Established term: this name is used in the research literature. About this guide’s status

In plain terms

A score measured on one group of people, places, or times doesn’t automatically carry over to another. Test on data from where the system will actually run.

Takeaways

  • A test set is a snapshot of certain people, places, times, and conditions.
  • Change any of those and performance can drop, sometimes sharply.
  • Even careful rebuilds of the same benchmark produce lower scores.
  • Evaluate on data from the setting where the system will run, and keep checking over time.

What it means

Models learn patterns from the data they’re trained on and get tested on data that usually looks a lot like it. The real world rarely cooperates. A new hospital has different patients and equipment. A new market uses different language. Next year’s users behave differently than last year’s. Each change is a shift, and a benchmark score doesn’t tell you how the system handles it.

The evidence

reviewed 15 Sep 2026

Same recipe, lower scores. Recht et al. (2019) rebuilt the CIFAR-10 and ImageNet test sets by following the original collection process. Accuracy fell 3% to 15% on CIFAR-10 and 11% to 14% on ImageNet. Small differences in how data was gathered were enough.

Real-world shifts, measured. Koh et al. (2021) built WILDS, a benchmark of ten datasets with real distribution shifts across hospitals, cameras, regions, and time. Models showed large gaps between in-distribution and out-of-distribution performance.

A clinical example. Wong et al. (2021) found a widely deployed sepsis model performed far worse at an independent health system than its developer reported.

Drift erodes gains. Hand (2006) argued that population drift over time is one reason gains from sophisticated models shrink in practice.

Use it

  1. Ask where and when the eval data came from, and compare that to your actual users.
  2. Split test results by site, time period, or source to see how performance moves.
  3. Run a local validation before rolling out to a new market, language, or customer segment.
  4. Monitor production for drift and re-evaluate on a schedule.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

Not every change in data hurts. In Recht et al., model rankings mostly held. The question is whether your deployment differs from the test data in ways the model is sensitive to.

Origins

Distribution shift, also called dataset shift, is a long-standing problem in statistics and machine learning. It’s closely related to external validity in the social and medical sciences.

Sources

  1. [1]
    Hand, D. J. (2006). Classifier technology and the illusion of progress.Statistical Science, 21(1), 1-14
    Open ↗ (opens in a new tab)
  2. [2]
    Koh, P. W., et al. (2021). WILDS: A benchmark of in-the-wild distribution shifts.Proceedings of the International Conference on Machine Learning (ICML 2021)
    Open ↗ (opens in a new tab)
  3. [3]
    Recht, B., Roelofs, R., Schmidt, L., & Shankar, V. (2019). Do ImageNet classifiers generalize to ImageNet?.Proceedings of the International Conference on Machine Learning (ICML 2019)
    Open ↗ (opens in a new tab)
  4. [4]
    Wong, A., Otles, E., Donnelly, J. P., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients.JAMA Internal Medicine, 181(8), 1065-1070
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). Distribution Shift (v1.0). https://evalfieldguide.com/patterns/distribution-shift

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 24 →