Guide
Playbook
For how to introduce these practices to a team without slowing it down, see Being Pragmatic.
1. Reading an AI claim
Ten questions for any benchmark result, vendor deck, or launch post.
- What decision is this number supposed to support? If nobody can say, the number is marketing. (The Functionality Fallacy)
- What exactly was measured? Read some test items. Does the benchmark’s name match its contents? (The Construct Gap)
- Could the model have seen the test? Check the benchmark’s release date against the training cutoff. (Data Contamination)
- How big is the test, and where are the error bars? A 2-point gap on 200 items is probably noise. (No Error Bars, No Result)
- Better than what? Was there a simple baseline, and was it tuned fairly? (The Baseline Rule)
- Who’s hidden in the average? Is there a breakdown by group, language, or input type? (The Averaging Trap)
- How many tries did this take? Prompts, variants, runs, and seeds all count. (Prompt Sensitivity, Once Is Not Reliable, Campbell’s Law)
- Who did the grading, and was the grader checked? (Judge Bias, The Gold Standard Myth)
- What did it cost to get this score? (The Cost Frontier)
- Has anyone independent reproduced it in a setting like yours? (The Reproducibility Rule, Distribution Shift, The Lab-to-Field Gap)
2. Building your own eval
A practical sequence for product teams. Scale each step to the stakes.
- Start with the decision. Write down what you’ll do differently depending on the result.
- Define the construct in one sentence. “Good” is not a definition. “Resolves the request without a follow-up” is.
- Collect test cases from real use. Include edge cases and known past failures. Keep a private, locked test set that never touches prompt tuning.
- Read outputs before you write the rubric. Expect the rubric to change, and version it (Shankar et al., 2024).
- Pick metrics before running. Include at least one continuous metric, plus cost and latency.
- Set baselines. The current workflow, a simple method, and the model you use today.
- Vary prompts and repeat runs. Three to five prompt variants and several samples per case is a reasonable floor (Sclar et al., 2024; Yao et al., 2025).
- Validate your graders. Measure agreement between human raters. If you use an LLM judge, check it against human labels first (Zheng et al., 2023).
- Compute intervals and slice the results. Report confidence intervals, paired comparisons, and the worst slice (Miller, 2024).
- Probe for what the score can’t show. Try shortcut tests, red teaming, and serious elicitation effort.
- Test people and AI together, in context. When you can, compare human alone, AI alone, and human with AI (Vaccaro et al., 2024).
- Document it, rerun it, monitor it. Write an eval card, save raw outputs, rerun a canary set on every model update, and watch production after launch.
3. Red flags
| You see | It might mean | Pattern |
|---|---|---|
| One headline number, no interval | The difference could be noise | No Error Bars, No Result |
| “Superhuman” on a benchmark that’s several years old | Saturation, contamination, or both | Benchmark Saturation, Data Contamination |
| Big win on public benchmarks, nothing private | Memorization or overfitting to the public test | Data Contamination, Adaptive Overfitting |
| Results shown only on benchmarks where the system won | Selective reporting | The Benchmark Lottery |
| A model grading its own outputs | Self-preference bias | Judge Bias |
| Scores jumped after lots of prompt tweaking on the test set | Fitting to the test | Adaptive Overfitting, Goodhart’s Law |
| “We found no evidence of dangerous capability X” | Limited elicitation effort | Presence, Not Absence |
| A great demo and no field pilot | Untested in real conditions | The Lab-to-Field Gap |
| Accuracy compared across systems with very different costs | An unfair comparison | The Cost Frontier |
| Explanations added mainly to raise user trust | Possible overreliance | The Team Is the System |
| High average accuracy, no subgroup breakdown | Hidden failures | The Averaging Trap |
| A single run per test case for an agent | Unknown reliability | Once Is Not Reliable |
4. Eval card template
Copy this into any eval you run. It borrows from model cards (Mitchell et al., 2019).
Eval card
Name / version:
Owner:
Date run:
Decision this informs:
Construct (one sentence):
Intended use and users:
System under test (model + version, system prompt, settings):
Baselines:
Test data (source, size, date range, how it was protected from training and tuning):
Slices reported:
Metrics (and why):
Grading (human raters, agreement, LLM judge validation):
Prompt variants and runs per case:
Results (with 95% intervals and worst slice):
Cost and latency:
Elicitation and red teaming effort:
Known limits and open questions:
Where raw outputs and code live:
Next rerun trigger:
References
- Miller, E. (2024). Adding error bars to evals: A statistical approach to language model evaluations. arXiv:2411.00640. (Preprint)
- Mitchell, M., et al. (2019). Model cards for model reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2019).
- Sclar, M., Choi, Y., Tsvetkov, Y., & Suhr, A. (2024). Quantifying language models’ sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting. International Conference on Learning Representations (ICLR 2024).
- Shankar, S., Zamfirescu-Pereira, J. D., Hartmann, B., Parameswaran, A. G., & Arawjo, I. (2024). Who validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences. Proceedings of the ACM Symposium on User Interface Software and Technology (UIST 2024).
- Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8, 2293-2303.
- Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2025). τ-bench: A benchmark for tool-agent-user interaction in real-world domains. International Conference on Learning Representations (ICLR 2025).
- Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., et al. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks Track.