2 of 5 in this categoryNext: 18 The Baseline Rule →
No. 17Resultsv1.0For: Designing · Building · Buying

The Averaging Trap

“An average score can hide exactly who the system fails.”

Last reviewed
15 Sep 2026
Sources
4 · Cite
Source types
4 peer-reviewed or classic
Version
v1.0 5 Oct 2026

The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status

In plain terms

A strong average can hide a group the system fails badly. Report results for the groups that matter, and show the worst one.

Takeaways

  • Aggregate accuracy blends easy and hard cases, common and rare groups, into one number.
  • Serious failures often cluster in subgroups that are small in the test set and important in the world.
  • Break results down by user group, input type, language, and difficulty.
  • Share instance-level results when you can, so others can slice them too.

What it means

A 95% overall score can be 99% for most people and 70% for a group that makes up a small slice of the test set. The average looks great. The people in that slice have a very different experience. The more a system is used across different populations, the more likely this is, and the more it matters.

The evidence

reviewed 15 Sep 2026

Gender Shades. Buolamwini and Gebru (2018) evaluated commercial gender classification systems and found error rates up to 34.7% for darker-skinned women, compared to at most 0.8% for lighter-skinned men.

Hidden stratification in medicine. Oakden-Rayner et al. (2020) showed that medical imaging models can have relative performance differences of over 20% on clinically important subsets. A model detecting pneumothorax (collapsed lung) had an AUC of 0.94 on X-rays that showed a chest drain, but only 0.77 on X-rays without one. Chest drains are the treatment, so the model did worst on the untreated, most urgent cases.

Report the breakdown. Burnell et al. (2023), writing in Science, argue that aggregate metrics and lack of access to instance-level results limit understanding of AI systems, and call for granular reporting.

Build it into documentation. Mitchell et al. (2019) made disaggregated evaluation across relevant groups a standard section of model cards.

Use it

  1. Define the slices that matter for your users before you look at the overall number.
  2. Report the worst-performing slice next to the average.
  3. Make sure each important slice has enough examples to support its own error bars.
  4. Read failures by hand. Clusters often reveal slices you didn’t think to define.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

Breakdowns need enough examples in each group to mean anything, and slicing too finely produces noisy numbers. Choose the groups that matter for how the system will be used.

Origins

“The Averaging Trap” is this site’s name. “Hidden stratification” comes from Oakden-Rayner et al. (2020), and “disaggregated evaluation” is standard language in fairness research.

Sources

  1. [1]
    Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification.Proceedings of the Conference on Fairness, Accountability and Transparency (FAT* 2018), PMLR 81, 77-91
    Open ↗ (opens in a new tab)
  2. [2]
    Burnell, R., et al. (2023). Rethink reporting of evaluation results in AI.Science, 380(6641), 136-138
    Open ↗ (opens in a new tab)
  3. [3]
    Mitchell, M., et al. (2019). Model cards for model reporting.Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2019)
    Open ↗ (opens in a new tab)
  4. [4]
    Oakden-Rayner, L., Dunnmon, J., Carneiro, G., & Ré, C. (2020). Hidden stratification causes clinically meaningful failures in machine learning for medical imaging.Proceedings of the ACM Conference on Health, Inference, and Learning (CHIL 2020)
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). The Averaging Trap (v1.0). https://evalfieldguide.com/patterns/the-averaging-trap

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 18 →