4 of 6 in this categoryNext: 25 The Team Is the System →
No. 24Beyondv1.0For: Designing · Building · Buying

The Lab-to-Field Gap

“A system that works in the lab can fail in the places people actually use it.”

Last reviewed
15 Sep 2026
Sources
6 · Cite
Source types
3 peer-reviewed or classic, 2 preprints, 1 report
Version
v1.0 5 Oct 2026

The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status

In plain terms

A system that works in the lab can fail in real conditions: workflows, time pressure, lighting, bandwidth, and people. Watch it in use before calling it validated.

Takeaways

  • Real use adds lighting, bandwidth, workflows, time pressure, incentives, and people.
  • Model accuracy is one layer. How people interact with the system, and its wider effects, are layers too.
  • Many AI failures are sociotechnical. The model does what it was built to do, and the situation defeats it.
  • Watch the system in use, in context, before calling it validated.

What it means

Lab evaluation holds everything constant except the model. The field holds nothing constant. The value of an AI system comes from the whole setup: the model, the interface, the workflow, the people, and the organization around them. Evaluating only the model is like usability testing only the backend.

The evidence

reviewed 15 Sep 2026

Accurate in validation, rough in clinics. Beede et al. (2020) studied a deep learning system for detecting diabetic retinopathy deployed in 11 clinics in Thailand. The system had performed well in validation. In the clinics, the authors found that socio-environmental factors, like lighting conditions and internet speed, affected model performance, nursing workflows, and the patient experience. Many images were rejected as ungradable.

Three layers of evaluation. Weidinger et al. (2023) propose evaluating generative AI at the capability layer, the human interaction layer, and the systemic impact layer, and note that most safety evaluation so far has focused on the first.

Narrow the sociotechnical gap. Liao and Xiao (2023) frame evaluation as narrowing the gap between what gets measured and what real people need from a system.

Abstraction traps. Selbst et al. (2019) describe traps that come from abstracting away social context, including assuming a solution built for one context will work in another.

The official name for it. The International AI Safety Report 2026 describes an “evaluation gap”: pre-deployment test results are not always strongly predictive of real-world capabilities or risks (Bengio et al., 2026).

Practice lags need. Hutchinson et al. (2022) found ML evaluation practice often fails to reflect the needs of real application contexts.

Use it

  1. Do field observation. Sit with people using the system in their actual environment.
  2. Evaluate the workflow, not just the model: time to complete, error recovery, handoffs.
  3. Pilot in a small number of real sites before scaling.
  4. Define success with the people affected, including people who never touch the interface.
  5. Bring in UX research methods like contextual inquiry and diary studies. They were built for exactly this gap.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

Lab results still carry information. The gap is about what they leave out, such as context, workflows, and people. Findings from one field setting, like clinics in Thailand, may not transfer to yours either.

Origins

“The Lab-to-Field Gap” is this site’s name. Researchers also call it the sociotechnical gap or the evaluation gap.

Sources

  1. [1]
    Beede, E., Baylor, E., Hersch, F., Iurchenko, A., Wilcox, L., Ruamviboonsuk, P., & Vardoulakis, L. M. (2020). A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy.Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2020)
    Open ↗ (opens in a new tab)
  2. [2]
    Bengio, Y., et al. (2026). International AI Safety Report 2026.International AI Safety Report; arXiv:2602.21012 Report
    Open ↗ (opens in a new tab)
  3. [3]
    Hutchinson, B., et al. (2022). Evaluation gaps in machine learning practice.Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT 2022)
    Open ↗ (opens in a new tab)
  4. [4]
    Liao, Q. V., & Xiao, Z. (2023). Rethinking model evaluation as narrowing the socio-technical gap.arXiv:2306.03100 Preprint
    Open ↗ (opens in a new tab)
  5. [5]
    Selbst, A. D., boyd, d., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems.Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2019)
    Open ↗ (opens in a new tab)
  6. [6]
    Weidinger, L., Rauh, M., Marchal, N., Manzini, A., et al. (2023). Sociotechnical safety evaluation of generative AI systems.arXiv:2310.11986 Preprint
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). The Lab-to-Field Gap (v1.0). https://evalfieldguide.com/patterns/the-lab-to-field-gap

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 25 →