Criteria Drift
“You figure out what ‘good’ means by grading outputs, so your criteria will change.”
- Last reviewed
- 15 Sep 2026
- Source types
- 1 peer-reviewed or classic, 1 preprint
- Version
- v1.0 5 Oct 2026
Established term: this name is used in the research literature. About this guide’s status
In plain terms
You work out what “good” means by reading real outputs, so your rubric will change as you grade. Version it, and re-grade a fixed set when it does.
Takeaways #
- Nobody can write a complete rubric before seeing real outputs.
- Grading surfaces new failure modes and new preferences, which changes the rubric.
- Treat evaluation criteria as a living design artifact, versioned like code.
- When criteria change, re-grade a fixed reference set or your trend lines stop being comparable.
What it means #
If you’ve done qualitative research, this will feel familiar. You don’t know the themes until you’ve read the transcripts. Evaluation works the same way. Teams sit down to define “a good summary,” write a few criteria, then start reading outputs and realize they care about things they never wrote down, like tone, or whether the summary admits uncertainty. The rubric grows. Some criteria turn out to matter only for certain kinds of outputs.
That’s healthy. The risk is pretending the criteria were fixed all along and comparing numbers graded under different rubrics.
The evidence #
reviewed 15 Sep 2026Named in a user study. Shankar et al. (2024) observed people building LLM-based evaluators with their EvalGen tool and described criteria drift: “users need criteria to grade outputs, but grading outputs helps users define criteria.” Some criteria depended on the specific outputs people had seen.
Start from real needs. Liao and Xiao (2023) argue that evaluation should be grounded in the real-world needs of the people using a system, and that methods from human-computer interaction can help close the gap between what gets measured and what matters.
Use it #
- Start with open coding. Read 30 to 50 real outputs and note what’s good and bad before writing any rubric.
- Version the rubric, with a date and a reason for each change.
- Keep a fixed reference set and re-grade it when the rubric changes.
- Involve the people who will actually use the output in defining the criteria.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
This pattern rests on fewer studies than most, mainly a user study by Shankar et al. and related work. It matters most while you are still working out what good looks like. With settled criteria it matters less.
Origins #
The term comes from Shreya Shankar, J.D. Zamfirescu-Pereira, Björn Hartmann, Aditya Parameswaran, and Ian Arawjo’s 2024 UIST paper, “Who Validates the Validators?”
Sources #
- [1]Liao, Q. V., & Xiao, Z. (2023). Rethinking model evaluation as narrowing the socio-technical gap.arXiv:2306.03100 PreprintOpen ↗ (opens in a new tab)
- [2]Shankar, S., Zamfirescu-Pereira, J. D., Hartmann, B., Parameswaran, A. G., & Arawjo, I. (2024). Who validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences.Proceedings of the ACM Symposium on User Interface Software and Technology (UIST 2024)Open ↗ (opens in a new tab)
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). Criteria Drift (v1.0). https://evalfieldguide.com/patterns/criteria-drift@misc{lai-criteria-drift,
title = {Criteria Drift},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/criteria-drift}}
}https://evalfieldguide.com/patterns/criteria-driftRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.