How to use this guide

AI Evaluation, an Overview

What AI evaluation is

AI evaluation is how you gather evidence about what an AI system can do, how it behaves, and what happens when people use it, so that someone can make a decision. Ship it or hold it. Buy it or pass. Trust it here, check it there.

The decision is the point. An eval produces evidence for it, and the quality of that evidence should match the stakes.

At its core, evaluation is measurement. Wallach et al. (2025) argue that evaluating generative AI is a social science measurement challenge: the things people care about, like helpfulness, safety, or reasoning, are abstract concepts that have to be carefully defined before they can be measured at all. That framing runs through this whole site.

Why it’s hard

Three layers of evaluation

Weidinger et al. (2023) describe three layers. Most published evaluation lives in the first one. Most real-world consequences show up in the other two.

LayerThe questionTypical methods
CapabilityWhat can the system do, and how well?Benchmarks, behavioral tests, red teaming, capability evals
Human interactionWhat happens when people use it?User studies, field studies, human-AI team experiments
Systemic impactWhat happens to organizations, markets, and communities over time?Deployment monitoring, audits, longitudinal and policy research

The measurement chain

Every score sits at the end of a chain. Wallach et al. (2025), adapting Adcock and Collier (2001), describe four levels:

  1. Background concept. The broad idea. Example: “helpful.”
  2. Systematized concept. An explicit definition. Example: “The response resolves the user’s stated request without needing a follow-up.”
  3. Measurement instrument. The actual test. Example: 300 real support tickets, a written rubric, two trained raters.
  4. Measurements. The numbers. Example: 71% resolved, with a 95% confidence interval of about 66% to 76%.

When two people argue about whether a model is “good at reasoning,” they’re usually disagreeing at level 1 or 2 while pointing at numbers from level 4. Validity is about the whole chain holding together (Messick, 1995).

Methods at a glance

MethodWhat it tells youWatch out forRelated patterns
Static benchmarksPerformance on a fixed, shared set of tasksSaturation, contamination, construct gapsBenchmark Saturation, Data Contamination, The Construct Gap
Dynamic or adversarial benchmarksPerformance on fresh, hard examples, often written to beat current models (Kiela et al., 2021)Can overweight adversarial cases that rarely occur in real useThe Clever Hans Effect
Behavioral testingWhether specific behaviors hold up under controlled changes (Ribeiro et al., 2020)Only tests the behaviors you think to write downThe Clever Hans Effect, Presence, Not Absence
Human evaluationHow people judge output quality (van der Lee et al., 2021)Rater expertise, agreement, and fatigueThe Gold Standard Myth, Criteria Drift
Pairwise preference arenasWhich outputs people prefer at scale (Chiang et al., 2024)Who votes, what they ask, and who gets to test privatelyGoodhart’s Law, Campbell’s Law
LLM-as-a-judgeCheap, fast grading at scale (Zheng et al., 2023)Position, length, and self-preference biasJudge Bias
Red teamingFailure modes found by people trying to break the system (Ganguli et al., 2022)Coverage depends on who’s on the teamPresence, Not Absence
Dangerous capability evalsWhether a model has capabilities that could cause severe harm (Shevlane et al., 2023)Elicitation effort, sandbagging, eval awarenessPresence, Not Absence, Evaluation Awareness
User and field studiesWhat happens when real people use the system in context (Beede et al., 2020; Bansal et al., 2021)Cost, time, and small samplesThe Lab-to-Field Gap, The Team Is the System
Production monitoringHow the system performs on live traffic over timePrivacy, consent, and slow feedbackDistribution Shift, Evaluation Awareness
Audits and documentationWhether claims hold up and are recorded in a checkable way (Mitchell et al., 2019; Raji et al., 2022)Access to models, data, and resultsThe Functionality Fallacy, The Reproducibility Rule

Who evaluates, and why

Different people need different evidence from the same system.

Hutchinson et al. (2022) found that ML evaluation practice often serves the first group and underserves the rest. A lot of the patterns on this site are about closing that gap.

How to use the patterns

To learn, read them by category:

As a tool, use the Playbook when you’re reading an AI claim, designing an eval, or reviewing a launch.

In a team, start with Being Pragmatic. It covers how to match rigor to stakes, fit evaluation into existing rituals, and bring people along without becoming the eval police.

A note on the word “pattern”

A pattern here means a reliable way evaluation goes wrong, with research behind it. It is not a law of physics. Some have established names, like Goodhart’s Law. Others are named on this site to make a well-documented idea easier to remember, and each page says which. Every page links to its sources, and preprints are labeled as preprints.

Good starting reads

References