5 of 5 in this categoryNext: 11 Prompt Sensitivity →
No. 10The testv1.0For: Building · Designing

The Gold Standard Myth

“Labels and human ratings are measurements with error, not ground truth.”

Last reviewed
15 Sep 2026
Sources
5 · Cite
Source types
5 peer-reviewed or classic
Version
v1.0 5 Oct 2026

The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status

In plain terms

The answer key and the human ratings are measurements too, and they contain errors and disagreement. Treat them as evidence, not as ground truth.

Takeaways

  • Test sets contain wrong labels. Near the top of a leaderboard, those errors can decide who “wins.”
  • People disagree. For many tasks that disagreement is real information, not noise to average away.
  • Untrained or rushed human raters can be close to random on hard judgments.
  • Report agreement between raters, and audit labels on the items that matter most.

What it means

Every eval compares a system’s output to something treated as correct: a label, a reference answer, a human rating. That “correct” answer was produced by people working under time pressure, with instructions that couldn’t cover every case. Some answers are wrong. Some questions have more than one defensible answer. Calling it ground truth hides all of that.

The evidence

reviewed 15 Sep 2026

Answer keys have errors. Northcutt, Athalye, and Mueller (2021) found an average of at least 3.3% label errors across 10 widely used test sets, and at least 6% in the ImageNet validation set. With corrected labels, rankings between models could flip. On ImageNet, ResNet-18 outperformed ResNet-50 if the share of originally mislabeled examples rose by just 6%.

Disagreement is signal. Aroyo and Welty (2015) describe seven myths of human annotation, including that every item has one right answer and that disagreement is bad. They argue disagreement tells you something about the item, the instructions, or the task.

Raters can be at chance. Clark et al. (2021) found that, without training, evaluators distinguished GPT-3 text from human-written text at chance level. Quick training raised accuracy only to about 55%.

Crowdsourced ratings can mislead. Karpinska, Akoury, and Krishna (2021) found that Mechanical Turk workers, unlike English teachers, failed to tell model-generated stories apart from human-written references. Most papers they surveyed also left out key details about how their crowdsourced ratings were collected.

Practice is inconsistent. van der Lee et al. (2021) reviewed human evaluation in natural language generation, found wide variation in how it’s done and reported, and proposed best-practice guidelines.

Use it

  1. Have two people label a sample independently and measure their agreement before trusting the labels.
  2. Audit labels on items where strong models disagree with the “answer.”
  3. Use raters with real expertise for expert tasks, and train and calibrate them.
  4. Keep the disagreement data. Report distributions, not just majority votes.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

Human judgment is still the best reference for many tasks. The pattern says to treat labels as evidence that can contain errors, and disagreement as information. It does not say to drop human evaluation.

Origins

“The Gold Standard Myth” is this site’s name for a theme that runs through annotation research, crowdsourcing studies, and human evaluation guidelines. Aroyo and Welty’s “Truth Is a Lie” (2015) is the clearest early statement of it.

Sources

  1. [1]
    Aroyo, L., & Welty, C. (2015). Truth is a lie: Crowd truth and the seven myths of human annotation.AI Magazine, 36(1), 15-24
    Open ↗ (opens in a new tab)
  2. [2]
    Clark, E., August, T., Serrano, S., Haduong, N., Gururangan, S., & Smith, N. A. (2021). All that’s ‘human’ is not gold: Evaluating human evaluation of generated text.Proceedings of ACL-IJCNLP 2021
    Open ↗ (opens in a new tab)
  3. [3]
    Karpinska, M., Akoury, N., & Krishna, K. (2021). The perils of using Mechanical Turk to evaluate open-ended text generation.Proceedings of EMNLP 2021
    Open ↗ (opens in a new tab)
  4. [4]
    Northcutt, C. G., Athalye, A., & Mueller, J. (2021). Pervasive label errors in test sets destabilize machine learning benchmarks.NeurIPS 2021 Datasets and Benchmarks Track
    Open ↗ (opens in a new tab)
  5. [5]
    van der Lee, C., Gatt, A., van Miltenburg, E., & Krahmer, E. (2021). Human evaluation of automatically generated text: Current trends and best practice guidelines.Computer Speech & Language, 67, 101151
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). The Gold Standard Myth (v1.0). https://evalfieldguide.com/patterns/the-gold-standard-myth

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 11 →