Guide

Glossary

Adaptive overfitting
Fitting a system to a specific test set by repeatedly making choices based on its results. See Adaptive Overfitting.
Aggregate metric
A single number that summarizes performance across all test cases, like overall accuracy. See The Averaging Trap.
AUC (area under the ROC curve)
A measure of how well a classifier separates positive from negative cases across all thresholds. 0.5 is chance, 1.0 is perfect.
Baseline
The comparison point for a result: a simpler method, an older model, or the current human workflow. See The Baseline Rule.
Behavioral testing
Checking specific behaviors with targeted test cases, like whether a model’s answer stays the same when a name is changed (Ribeiro et al., 2020).
Benchmark
A fixed dataset and scoring method used to compare systems on a task.
Calibration
How well a system’s confidence matches how often it’s actually right. A calibrated model that says “90% sure” is right about 90% of the time. Modern neural networks are often overconfident (Guo et al., 2017).
Capability evaluation
Testing what a system can do, as opposed to how people use it or what effects it has.
Complementary performance
When a human-AI team does better than either the human or the AI alone (Bansal et al., 2021).
Confidence interval
A range that likely contains the true value of a measurement, given sampling noise. See No Error Bars, No Result.
Construct
An abstract quality you want to measure but can’t observe directly, like reasoning or helpfulness (Cronbach & Meehl, 1955).
Construct validity
How well a measurement actually captures the construct it claims to measure (Messick, 1995). See The Construct Gap.
Contamination
When test data, or close copies of it, appear in a model’s training data. See Data Contamination.
Criteria drift
The way evaluation criteria change as people grade real outputs. See Criteria Drift.
Datasheet
A standard document describing how a dataset was created, what’s in it, and what it’s suited for (Gebru et al., 2021).
Disaggregated evaluation
Reporting results separately for different groups or input types instead of only as an average.
Distribution shift
A difference between the data a system was built and tested on and the data it meets in use. See Distribution Shift.
Dynamic benchmark
A benchmark that keeps adding new examples, often written to defeat current models.
Elicitation
The effort put into drawing out a model’s full capability, through prompting, tools, fine-tuning, or scaffolding. See Presence, Not Absence.
Error analysis
Grouping and studying a system’s mistakes to understand their causes. Works best with precisely defined error groups, enough examples, and tested hypotheses (Wu et al., 2019). See Being Pragmatic.
Error budget
An agreed-on acceptable failure rate, with an agreed-on response when it’s exceeded. Borrowed from site reliability engineering (Beyer et al., 2016). See Being Pragmatic.
Evaluation awareness
A model’s ability to tell that it’s being evaluated rather than used for real (Needham et al., 2025). See Evaluation Awareness.
External validation
Testing a system on data from a different setting than the one it was developed in, ideally by an independent team.
Ground truth
The answer treated as correct in an evaluation. Often a human judgment with its own error. See The Gold Standard Myth.
Holdout set
Data kept separate from training and tuning so it can give an unbiased estimate of performance (Dwork et al., 2015).
Inter-rater agreement
How often independent human raters give the same judgment. Low agreement means the task, the instructions, or the construct needs work.
Leaderboard
A public ranking of systems on one or more benchmarks. See Campbell’s Law.
LLM-as-a-judge
Using a language model to grade or compare outputs from AI systems (Zheng et al., 2023). See Judge Bias.
Model card
A standard document describing a model’s intended use, evaluation results, and limitations (Mitchell et al., 2019).
Online controlled experiment
Randomly exposing some users to a change and comparing outcomes against users who didn’t get it, often called an A/B test (Kohavi et al., 2020).
Overreliance
Accepting AI output when it’s wrong. Its opposite, underreliance, is ignoring AI output when it’s right (Parasuraman & Riley, 1997). See The Team Is the System.
Pairwise comparison
Asking raters which of two outputs is better, instead of scoring each on its own. Used at scale by preference arenas (Chiang et al., 2024).
Pareto frontier
The set of options where you can’t improve one metric, like accuracy, without making another, like cost, worse. See The Cost Frontier.
pass@k and pass^k
pass@k is the chance of at least one success in k attempts. pass^k is the chance of success on all k attempts, a measure of consistency (Yao et al., 2025). See Once Is Not Reliable.
Prompt sensitivity
How much a model’s results change with small changes to how a prompt is worded or formatted. See Prompt Sensitivity.
Red teaming
Deliberately trying to make a system fail or cause harm, to find problems before others do.
Reliability (measurement)
How consistent a measurement is when repeated. A measurement can be reliable and still not valid.
Reward hacking
When an optimized system finds a way to score well on its objective without doing what the objective was meant to encourage. See Goodhart’s Law.
Risk tier
A category that sets how much evaluation a system needs, based on its potential impact. Risk-based prioritization is central to the NIST AI Risk Management Framework (2023). See Being Pragmatic.
Sandbagging
Strategic underperformance on an evaluation (van der Weij et al., 2025).
Saturation
When top systems score near the maximum on a benchmark, so it no longer separates them. See Benchmark Saturation.
Shortcut learning
Learning a pattern that predicts the right answer on the test without the intended skill (Geirhos et al., 2020). See The Clever Hans Effect.
Sociotechnical evaluation
Evaluating a system together with the people, workflows, and institutions around it. See The Lab-to-Field Gap.
Staged rollout
Releasing a system to progressively larger groups of users while monitoring results at each step.
Statistical power
The chance an experiment detects a real difference of a given size. Small test sets have low power. See No Error Bars, No Result.
Validity
How well the evidence supports the interpretations and uses of a score (Messick, 1995).
Vibe check
Informal, exploratory trying-out of an AI system. Practitioners describe it as the first line of evaluation (van der Maden et al., 2026). Useful as a start, risky as the whole process.