How to use this guide
AI Evaluation, an Overview
What AI evaluation is
AI evaluation is how you gather evidence about what an AI system can do, how it behaves, and what happens when people use it, so that someone can make a decision. Ship it or hold it. Buy it or pass. Trust it here, check it there.
The decision is the point. An eval produces evidence for it, and the quality of that evidence should match the stakes.
At its core, evaluation is measurement. Wallach et al. (2025) argue that evaluating generative AI is a social science measurement challenge: the things people care about, like helpfulness, safety, or reasoning, are abstract concepts that have to be carefully defined before they can be measured at all. That framing runs through this whole site.
Why it’s hard
- One system, countless tasks. General-purpose models get used for things no benchmark anticipated, and a finite test can’t cover “everything in the whole wide world” (Raji et al., 2021).
- No single right answer. Open-ended outputs like summaries, code, and advice can be good in many ways and bad in many ways. Human raters struggle with this too (Clark et al., 2021), and AI judges bring their own biases (Zheng et al., 2023).
- Unknown training data. You usually can’t check whether a model has already seen the test (Sainz et al., 2023).
- Moving targets. Benchmarks saturate quickly (Ott et al., 2022), and hosted models change under the same name.
- Systems that respond to the test. Models can often tell when they’re being evaluated (Needham et al., 2025), and they increasingly find loopholes in evaluations (Bengio et al., 2026).
- Context decides value. The same model can succeed in a lab and struggle in a clinic (Beede et al., 2020).
Three layers of evaluation
Weidinger et al. (2023) describe three layers. Most published evaluation lives in the first one. Most real-world consequences show up in the other two.
| Layer | The question | Typical methods |
|---|---|---|
| Capability | What can the system do, and how well? | Benchmarks, behavioral tests, red teaming, capability evals |
| Human interaction | What happens when people use it? | User studies, field studies, human-AI team experiments |
| Systemic impact | What happens to organizations, markets, and communities over time? | Deployment monitoring, audits, longitudinal and policy research |
The measurement chain
Every score sits at the end of a chain. Wallach et al. (2025), adapting Adcock and Collier (2001), describe four levels:
- Background concept. The broad idea. Example: “helpful.”
- Systematized concept. An explicit definition. Example: “The response resolves the user’s stated request without needing a follow-up.”
- Measurement instrument. The actual test. Example: 300 real support tickets, a written rubric, two trained raters.
- Measurements. The numbers. Example: 71% resolved, with a 95% confidence interval of about 66% to 76%.
When two people argue about whether a model is “good at reasoning,” they’re usually disagreeing at level 1 or 2 while pointing at numbers from level 4. Validity is about the whole chain holding together (Messick, 1995).
Methods at a glance
| Method | What it tells you | Watch out for | Related patterns |
|---|---|---|---|
| Static benchmarks | Performance on a fixed, shared set of tasks | Saturation, contamination, construct gaps | Benchmark Saturation, Data Contamination, The Construct Gap |
| Dynamic or adversarial benchmarks | Performance on fresh, hard examples, often written to beat current models (Kiela et al., 2021) | Can overweight adversarial cases that rarely occur in real use | The Clever Hans Effect |
| Behavioral testing | Whether specific behaviors hold up under controlled changes (Ribeiro et al., 2020) | Only tests the behaviors you think to write down | The Clever Hans Effect, Presence, Not Absence |
| Human evaluation | How people judge output quality (van der Lee et al., 2021) | Rater expertise, agreement, and fatigue | The Gold Standard Myth, Criteria Drift |
| Pairwise preference arenas | Which outputs people prefer at scale (Chiang et al., 2024) | Who votes, what they ask, and who gets to test privately | Goodhart’s Law, Campbell’s Law |
| LLM-as-a-judge | Cheap, fast grading at scale (Zheng et al., 2023) | Position, length, and self-preference bias | Judge Bias |
| Red teaming | Failure modes found by people trying to break the system (Ganguli et al., 2022) | Coverage depends on who’s on the team | Presence, Not Absence |
| Dangerous capability evals | Whether a model has capabilities that could cause severe harm (Shevlane et al., 2023) | Elicitation effort, sandbagging, eval awareness | Presence, Not Absence, Evaluation Awareness |
| User and field studies | What happens when real people use the system in context (Beede et al., 2020; Bansal et al., 2021) | Cost, time, and small samples | The Lab-to-Field Gap, The Team Is the System |
| Production monitoring | How the system performs on live traffic over time | Privacy, consent, and slow feedback | Distribution Shift, Evaluation Awareness |
| Audits and documentation | Whether claims hold up and are recorded in a checkable way (Mitchell et al., 2019; Raji et al., 2022) | Access to models, data, and results | The Functionality Fallacy, The Reproducibility Rule |
Who evaluates, and why
Different people need different evidence from the same system.
- Model developers want to know if a new model is better than the last one.
- Product teams want to know if a model works for their users, in their workflow, at a cost they can afford. Kapoor et al. (2025) point out that benchmarks often mix up what model developers need with what downstream developers need.
- Researchers want to understand what systems can and can’t do, and why.
- Buyers want to know if a vendor’s claims hold up on their own data.
- Auditors and regulators want evidence that a system is valid, reliable, and safe for its intended use. The NIST AI Risk Management Framework (2023) makes “Measure” one of its four core functions.
Hutchinson et al. (2022) found that ML evaluation practice often serves the first group and underserves the rest. A lot of the patterns on this site are about closing that gap.
How to use the patterns
To learn, read them by category:
- What you’re measuring: Goodhart’s Law, Campbell’s Law, The Construct Gap, The Functionality Fallacy, The Metric Mirage
- The test itself: Benchmark Saturation, Data Contamination, Adaptive Overfitting, The Clever Hans Effect, The Gold Standard Myth
- Running the eval: Prompt Sensitivity, Judge Bias, Criteria Drift, Once Is Not Reliable, The Cost Frontier
- Reading the results: No Error Bars, No Result, The Averaging Trap, The Baseline Rule, The Benchmark Lottery, The Reproducibility Rule
- Beyond the benchmark: Presence, Not Absence, Evaluation Awareness, Distribution Shift, The Lab-to-Field Gap, The Team Is the System, The Jagged Frontier
As a tool, use the Playbook when you’re reading an AI claim, designing an eval, or reviewing a launch.
In a team, start with Being Pragmatic. It covers how to match rigor to stakes, fit evaluation into existing rituals, and bring people along without becoming the eval police.
A note on the word “pattern”
A pattern here means a reliable way evaluation goes wrong, with research behind it. It is not a law of physics. Some have established names, like Goodhart’s Law. Others are named on this site to make a well-documented idea easier to remember, and each page says which. Every page links to its sources, and preprints are labeled as preprints.
Good starting reads
- Chang et al. (2024), a broad survey of how large language models are evaluated.
- Burnell et al. (2023), a short Science piece on how evaluation results should be reported.
- Eriksson et al. (2025), an interdisciplinary review of what’s wrong with AI benchmarks.
- Wallach et al. (2025), the case for treating AI evaluation as measurement.
References
- Adcock, R., & Collier, D. (2001). Measurement validity: A shared standard for qualitative and quantitative research. American Political Science Review, 95(3), 529-546.
- Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2021).
- Beede, E., Baylor, E., Hersch, F., Iurchenko, A., Wilcox, L., Ruamviboonsuk, P., & Vardoulakis, L. M. (2020). A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2020).
- Bengio, Y., et al. (2026). International AI Safety Report 2026. International AI Safety Report; arXiv:2602.21012. (Report)
- Burnell, R., et al. (2023). Rethink reporting of evaluation results in AI. Science, 380(6641), 136-138.
- Chang, Y., et al. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3).
- Chiang, W.-L., et al. (2024). Chatbot Arena: An open platform for evaluating LLMs by human preference. Proceedings of the International Conference on Machine Learning (ICML 2024).
- Clark, E., August, T., Serrano, S., Haduong, N., Gururangan, S., & Smith, N. A. (2021). All that’s ‘human’ is not gold: Evaluating human evaluation of generated text. Proceedings of ACL-IJCNLP 2021.
- Eriksson, M., et al. (2025). Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2025).
- Ganguli, D., et al. (2022). Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv:2209.07858. (Preprint)
- Hutchinson, B., et al. (2022). Evaluation gaps in machine learning practice. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT 2022).
- Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2025). AI agents that matter. Transactions on Machine Learning Research (TMLR).
- Kiela, D., et al. (2021). Dynabench: Rethinking benchmarking in NLP. Proceedings of NAACL-HLT 2021.
- Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749.
- Mitchell, M., et al. (2019). Model cards for model reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2019).
- National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. U.S. Department of Commerce. (Report)
- Needham, J., Edkins, G., Pimpale, G., Bartsch, H., & Hobbhahn, M. (2025). Large language models often know when they are being evaluated. arXiv:2505.23836. (Preprint)
- Ott, S., Barbosa-Silva, A., Blagec, K., Brauner, J., & Samwald, M. (2022). Mapping global dynamics of benchmark creation and saturation in artificial intelligence. Nature Communications, 13, 6793.
- Raji, I. D., Bender, E. M., Paullada, A., Denton, E., & Hanna, A. (2021). AI and the everything in the whole wide world benchmark. NeurIPS 2021 Datasets and Benchmarks Track.
- Raji, I. D., Kumar, I. E., Horowitz, A., & Selbst, A. (2022). The fallacy of AI functionality. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT 2022).
- Ribeiro, M. T., Wu, T., Guestrin, C., & Singh, S. (2020). Beyond accuracy: Behavioral testing of NLP models with CheckList. Proceedings of ACL 2020.
- Sainz, O., Campos, J. A., García-Ferrero, I., Etxaniz, J., Lopez de Lacalle, O., & Agirre, E. (2023). NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark. Findings of EMNLP 2023.
- Shevlane, T., et al. (2023). Model evaluation for extreme risks. arXiv:2305.15324. (Preprint)
- van der Lee, C., Gatt, A., van Miltenburg, E., & Krahmer, E. (2021). Human evaluation of automatically generated text: Current trends and best practice guidelines. Computer Speech & Language, 67, 101151.
- Wallach, H., Desai, M., Cooper, A. F., Wang, A., Atalla, C., et al. (2025). Position: Evaluating generative AI systems is a social science measurement challenge. Proceedings of the International Conference on Machine Learning (ICML 2025), PMLR 267.
- Weidinger, L., Rauh, M., Marchal, N., Manzini, A., et al. (2023). Sociotechnical safety evaluation of generative AI systems. arXiv:2310.11986. (Preprint)
- Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., et al. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks Track.