3 of 5 in this categoryNext: 14 Once Is Not Reliable →
No. 13Runningv1.0For: Designing · Building

Criteria Drift

“You figure out what ‘good’ means by grading outputs, so your criteria will change.”

Last reviewed
15 Sep 2026
Sources
2 · Cite
Source types
1 peer-reviewed or classic, 1 preprint
Version
v1.0 5 Oct 2026

Established term: this name is used in the research literature. About this guide’s status

In plain terms

You work out what “good” means by reading real outputs, so your rubric will change as you grade. Version it, and re-grade a fixed set when it does.

Takeaways

  • Nobody can write a complete rubric before seeing real outputs.
  • Grading surfaces new failure modes and new preferences, which changes the rubric.
  • Treat evaluation criteria as a living design artifact, versioned like code.
  • When criteria change, re-grade a fixed reference set or your trend lines stop being comparable.

What it means

If you’ve done qualitative research, this will feel familiar. You don’t know the themes until you’ve read the transcripts. Evaluation works the same way. Teams sit down to define “a good summary,” write a few criteria, then start reading outputs and realize they care about things they never wrote down, like tone, or whether the summary admits uncertainty. The rubric grows. Some criteria turn out to matter only for certain kinds of outputs.

That’s healthy. The risk is pretending the criteria were fixed all along and comparing numbers graded under different rubrics.

The evidence

reviewed 15 Sep 2026

Named in a user study. Shankar et al. (2024) observed people building LLM-based evaluators with their EvalGen tool and described criteria drift: “users need criteria to grade outputs, but grading outputs helps users define criteria.” Some criteria depended on the specific outputs people had seen.

Start from real needs. Liao and Xiao (2023) argue that evaluation should be grounded in the real-world needs of the people using a system, and that methods from human-computer interaction can help close the gap between what gets measured and what matters.

Use it

  1. Start with open coding. Read 30 to 50 real outputs and note what’s good and bad before writing any rubric.
  2. Version the rubric, with a date and a reason for each change.
  3. Keep a fixed reference set and re-grade it when the rubric changes.
  4. Involve the people who will actually use the output in defining the criteria.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

This pattern rests on fewer studies than most, mainly a user study by Shankar et al. and related work. It matters most while you are still working out what good looks like. With settled criteria it matters less.

Origins

The term comes from Shreya Shankar, J.D. Zamfirescu-Pereira, Björn Hartmann, Aditya Parameswaran, and Ian Arawjo’s 2024 UIST paper, “Who Validates the Validators?”

Sources

  1. [1]
    Liao, Q. V., & Xiao, Z. (2023). Rethinking model evaluation as narrowing the socio-technical gap.arXiv:2306.03100 Preprint
    Open ↗ (opens in a new tab)
  2. [2]
    Shankar, S., Zamfirescu-Pereira, J. D., Hartmann, B., Parameswaran, A. G., & Arawjo, I. (2024). Who validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences.Proceedings of the ACM Symposium on User Interface Software and Technology (UIST 2024)
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). Criteria Drift (v1.0). https://evalfieldguide.com/patterns/criteria-drift

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 14 →