Guide

Playbook

For how to introduce these practices to a team without slowing it down, see Being Pragmatic.

1. Reading an AI claim

Ten questions for any benchmark result, vendor deck, or launch post.

  1. What decision is this number supposed to support? If nobody can say, the number is marketing. (The Functionality Fallacy)
  2. What exactly was measured? Read some test items. Does the benchmark’s name match its contents? (The Construct Gap)
  3. Could the model have seen the test? Check the benchmark’s release date against the training cutoff. (Data Contamination)
  4. How big is the test, and where are the error bars? A 2-point gap on 200 items is probably noise. (No Error Bars, No Result)
  5. Better than what? Was there a simple baseline, and was it tuned fairly? (The Baseline Rule)
  6. Who’s hidden in the average? Is there a breakdown by group, language, or input type? (The Averaging Trap)
  7. How many tries did this take? Prompts, variants, runs, and seeds all count. (Prompt Sensitivity, Once Is Not Reliable, Campbell’s Law)
  8. Who did the grading, and was the grader checked? (Judge Bias, The Gold Standard Myth)
  9. What did it cost to get this score? (The Cost Frontier)
  10. Has anyone independent reproduced it in a setting like yours? (The Reproducibility Rule, Distribution Shift, The Lab-to-Field Gap)

2. Building your own eval

A practical sequence for product teams. Scale each step to the stakes.

  1. Start with the decision. Write down what you’ll do differently depending on the result.
  2. Define the construct in one sentence. “Good” is not a definition. “Resolves the request without a follow-up” is.
  3. Collect test cases from real use. Include edge cases and known past failures. Keep a private, locked test set that never touches prompt tuning.
  4. Read outputs before you write the rubric. Expect the rubric to change, and version it (Shankar et al., 2024).
  5. Pick metrics before running. Include at least one continuous metric, plus cost and latency.
  6. Set baselines. The current workflow, a simple method, and the model you use today.
  7. Vary prompts and repeat runs. Three to five prompt variants and several samples per case is a reasonable floor (Sclar et al., 2024; Yao et al., 2025).
  8. Validate your graders. Measure agreement between human raters. If you use an LLM judge, check it against human labels first (Zheng et al., 2023).
  9. Compute intervals and slice the results. Report confidence intervals, paired comparisons, and the worst slice (Miller, 2024).
  10. Probe for what the score can’t show. Try shortcut tests, red teaming, and serious elicitation effort.
  11. Test people and AI together, in context. When you can, compare human alone, AI alone, and human with AI (Vaccaro et al., 2024).
  12. Document it, rerun it, monitor it. Write an eval card, save raw outputs, rerun a canary set on every model update, and watch production after launch.

3. Red flags

You seeIt might meanPattern
One headline number, no intervalThe difference could be noiseNo Error Bars, No Result
“Superhuman” on a benchmark that’s several years oldSaturation, contamination, or bothBenchmark Saturation, Data Contamination
Big win on public benchmarks, nothing privateMemorization or overfitting to the public testData Contamination, Adaptive Overfitting
Results shown only on benchmarks where the system wonSelective reportingThe Benchmark Lottery
A model grading its own outputsSelf-preference biasJudge Bias
Scores jumped after lots of prompt tweaking on the test setFitting to the testAdaptive Overfitting, Goodhart’s Law
“We found no evidence of dangerous capability X”Limited elicitation effortPresence, Not Absence
A great demo and no field pilotUntested in real conditionsThe Lab-to-Field Gap
Accuracy compared across systems with very different costsAn unfair comparisonThe Cost Frontier
Explanations added mainly to raise user trustPossible overrelianceThe Team Is the System
High average accuracy, no subgroup breakdownHidden failuresThe Averaging Trap
A single run per test case for an agentUnknown reliabilityOnce Is Not Reliable

4. Eval card template

Copy this into any eval you run. It borrows from model cards (Mitchell et al., 2019).

Eval card

Name / version:

Owner:

Date run:


Decision this informs:

Construct (one sentence):

Intended use and users:


System under test (model + version, system prompt, settings):

Baselines:


Test data (source, size, date range, how it was protected from training and tuning):

Slices reported:


Metrics (and why):

Grading (human raters, agreement, LLM judge validation):

Prompt variants and runs per case:


Results (with 95% intervals and worst slice):

Cost and latency:


Elicitation and red teaming effort:

Known limits and open questions:

Where raw outputs and code live:

Next rerun trigger:

References