Guide
Glossary
- Adaptive overfitting
- Fitting a system to a specific test set by repeatedly making choices based on its results. See Adaptive Overfitting.
- Aggregate metric
- A single number that summarizes performance across all test cases, like overall accuracy. See The Averaging Trap.
- AUC (area under the ROC curve)
- A measure of how well a classifier separates positive from negative cases across all thresholds. 0.5 is chance, 1.0 is perfect.
- Baseline
- The comparison point for a result: a simpler method, an older model, or the current human workflow. See The Baseline Rule.
- Behavioral testing
- Checking specific behaviors with targeted test cases, like whether a model’s answer stays the same when a name is changed (Ribeiro et al., 2020).
- Benchmark
- A fixed dataset and scoring method used to compare systems on a task.
- Calibration
- How well a system’s confidence matches how often it’s actually right. A calibrated model that says “90% sure” is right about 90% of the time. Modern neural networks are often overconfident (Guo et al., 2017).
- Capability evaluation
- Testing what a system can do, as opposed to how people use it or what effects it has.
- Complementary performance
- When a human-AI team does better than either the human or the AI alone (Bansal et al., 2021).
- Confidence interval
- A range that likely contains the true value of a measurement, given sampling noise. See No Error Bars, No Result.
- Construct
- An abstract quality you want to measure but can’t observe directly, like reasoning or helpfulness (Cronbach & Meehl, 1955).
- Construct validity
- How well a measurement actually captures the construct it claims to measure (Messick, 1995). See The Construct Gap.
- Contamination
- When test data, or close copies of it, appear in a model’s training data. See Data Contamination.
- Criteria drift
- The way evaluation criteria change as people grade real outputs. See Criteria Drift.
- Datasheet
- A standard document describing how a dataset was created, what’s in it, and what it’s suited for (Gebru et al., 2021).
- Disaggregated evaluation
- Reporting results separately for different groups or input types instead of only as an average.
- Distribution shift
- A difference between the data a system was built and tested on and the data it meets in use. See Distribution Shift.
- Dynamic benchmark
- A benchmark that keeps adding new examples, often written to defeat current models.
- Elicitation
- The effort put into drawing out a model’s full capability, through prompting, tools, fine-tuning, or scaffolding. See Presence, Not Absence.
- Error analysis
- Grouping and studying a system’s mistakes to understand their causes. Works best with precisely defined error groups, enough examples, and tested hypotheses (Wu et al., 2019). See Being Pragmatic.
- Error budget
- An agreed-on acceptable failure rate, with an agreed-on response when it’s exceeded. Borrowed from site reliability engineering (Beyer et al., 2016). See Being Pragmatic.
- Evaluation awareness
- A model’s ability to tell that it’s being evaluated rather than used for real (Needham et al., 2025). See Evaluation Awareness.
- External validation
- Testing a system on data from a different setting than the one it was developed in, ideally by an independent team.
- Ground truth
- The answer treated as correct in an evaluation. Often a human judgment with its own error. See The Gold Standard Myth.
- Holdout set
- Data kept separate from training and tuning so it can give an unbiased estimate of performance (Dwork et al., 2015).
- Inter-rater agreement
- How often independent human raters give the same judgment. Low agreement means the task, the instructions, or the construct needs work.
- Leaderboard
- A public ranking of systems on one or more benchmarks. See Campbell’s Law.
- LLM-as-a-judge
- Using a language model to grade or compare outputs from AI systems (Zheng et al., 2023). See Judge Bias.
- Model card
- A standard document describing a model’s intended use, evaluation results, and limitations (Mitchell et al., 2019).
- Online controlled experiment
- Randomly exposing some users to a change and comparing outcomes against users who didn’t get it, often called an A/B test (Kohavi et al., 2020).
- Overreliance
- Accepting AI output when it’s wrong. Its opposite, underreliance, is ignoring AI output when it’s right (Parasuraman & Riley, 1997). See The Team Is the System.
- Pairwise comparison
- Asking raters which of two outputs is better, instead of scoring each on its own. Used at scale by preference arenas (Chiang et al., 2024).
- Pareto frontier
- The set of options where you can’t improve one metric, like accuracy, without making another, like cost, worse. See The Cost Frontier.
- pass@k and pass^k
- pass@k is the chance of at least one success in k attempts. pass^k is the chance of success on all k attempts, a measure of consistency (Yao et al., 2025). See Once Is Not Reliable.
- Prompt sensitivity
- How much a model’s results change with small changes to how a prompt is worded or formatted. See Prompt Sensitivity.
- Red teaming
- Deliberately trying to make a system fail or cause harm, to find problems before others do.
- Reliability (measurement)
- How consistent a measurement is when repeated. A measurement can be reliable and still not valid.
- Reward hacking
- When an optimized system finds a way to score well on its objective without doing what the objective was meant to encourage. See Goodhart’s Law.
- Risk tier
- A category that sets how much evaluation a system needs, based on its potential impact. Risk-based prioritization is central to the NIST AI Risk Management Framework (2023). See Being Pragmatic.
- Sandbagging
- Strategic underperformance on an evaluation (van der Weij et al., 2025).
- Saturation
- When top systems score near the maximum on a benchmark, so it no longer separates them. See Benchmark Saturation.
- Shortcut learning
- Learning a pattern that predicts the right answer on the test without the intended skill (Geirhos et al., 2020). See The Clever Hans Effect.
- Sociotechnical evaluation
- Evaluating a system together with the people, workflows, and institutions around it. See The Lab-to-Field Gap.
- Staged rollout
- Releasing a system to progressively larger groups of users while monitoring results at each step.
- Statistical power
- The chance an experiment detects a real difference of a given size. Small test sets have low power. See No Error Bars, No Result.
- Validity
- How well the evidence supports the interpretations and uses of a score (Messick, 1995).
- Vibe check
- Informal, exploratory trying-out of an AI system. Practitioners describe it as the first line of evaluation (van der Maden et al., 2026). Useful as a start, risky as the whole process.