Guide
Questions to ask when building an AI evaluation
Most evaluation mistakes are made before the first score is computed. These questions, drawn from the patterns, help you check the design before you trust the result.
Goodhart’s Law
If you reward people or models for hitting a number, they get good at the number. Whether they got better at the actual job is a separate question that needs its own check.
- Which metric did you optimize during training and model selection, and is the reported score on that same metric?
- Was the reported test set ever used for tuning, prompt iteration, or choosing between checkpoints?
- Show us five outputs that score highly but are still wrong. What do they have in common?
- What will you measure after launch, and who decides when that metric has stopped being useful?
The Construct Gap
A test called “reasoning” only measures how a system does on that test’s items, in that format. Whether that says anything about reasoning in general needs its own evidence.
- Which of our real tasks does this benchmark resemble, and how was that checked?
- What would a high score here fail to tell us?
The Metric Mirage
The scoring method is part of the result. Score the same outputs a different way and a sudden leap can turn into steady progress, or the reverse.
- Does the result hold when it is re-scored with a second, continuous metric?
- Were the metrics chosen, with the reasons written down, before the results were seen?
Data Contamination
If a model saw the test questions while it was training, a high score may only mean it remembers them. Fresh or private questions are the reliable check.
- What overlap analysis was run between the test set and the training data?
- Were results confirmed on data created after the training cutoff?
Adaptive Overfitting
Each time you adjust a prompt or a model because of how it did on the test, the test tells you a little less. Keep a locked test that you use only to decide.
- Is there a locked test set, kept separate from the dev set used for iteration?
- How many times has the locked test set been consulted to make decisions?
The Clever Hans Effect
A model can pass a test by picking up an accidental pattern, such as a watermark or a telltale word, instead of the skill you meant to test. Build cases where the shortcut fails.
- Was a partial-input baseline run to see whether accuracy stays high without the part of the input that should matter?
- Do contrast pairs or behavioral tests show the model gets the right answers for the right reasons?
The Gold Standard Myth
The answer key and the human ratings are measurements too, and they contain errors and disagreement. Treat them as evidence, not as ground truth.
- How often do independent raters agree on the labels, and was that measured before the labels were trusted?
- Who audited the labels on the items where strong models disagree with the answer key?
Prompt Sensitivity
Small changes in wording or formatting can swing a model’s score. A result from one prompt is a result for that prompt.
- How much does the score move across reasonable rephrasings?
- Was the prompt tuned on the test set?
Judge Bias
Using an AI to grade AI is fast and cheap, but graders have favorites: the first answer, the longer answer, and answers that sound like their own. Check the grader against people.
- Was the AI grader checked against human judgments on a sample?
- Were answer order and length controlled for, and is the grader from the same model family?
Criteria Drift
You work out what “good” means by reading real outputs, so your rubric will change as you grade. Version it, and re-grade a fixed set when it does.
- Was the rubric written after reading real outputs, and is it versioned with a date and a reason for each change?
- When the criteria changed, was a fixed reference set re-graded so trends stay comparable?
Once Is Not Reliable
A system that works once may fail the next time. Customers live with the all-runs-pass rate, not the best-run rate.
- How many runs per item, and what is the pass rate across all of them?
- Is the result reported as pass@1, pass@k, or passing on every run?
The Cost Frontier
You can buy accuracy with more compute, so a score without its cost is half a result. Compare systems at matched cost, or show the whole tradeoff.
- What did each task cost in tokens or dollars, and what was the latency, next to every accuracy number?
- Was a cheap baseline, such as a single call, a smaller model, or a simple retry loop, included at a matched budget?
No Error Bars, No Result
Every score is an estimate. If two systems differ by less than the noise, nobody has shown that they differ.
- What is the 95% confidence interval on each score, and how many items and runs does it rest on?
- Is the gap between systems larger than the interval, using a paired comparison on the same items?
The Averaging Trap
A strong average can hide a group the system fails badly. Report results for the groups that matter, and show the worst one.
- Which user groups, languages, or input types matter here, and were results reported for each?
- What is the worst-performing slice, and does it have enough examples to support its own error bars?
The Baseline Rule
A win is only as strong as what it beat. Compare against the simplest thing that could work and today’s actual process, with the same tuning effort.
- Was the current process, or the simplest thing that could work, included as a baseline?
- Did the baselines get the same tuning budget as the new method?
The Reproducibility Rule
If nobody can rerun an evaluation, it is an anecdote. Record the model version, prompts, settings, data, and scoring code, and save the raw outputs.
- Are the model version, system prompt, decoding settings, dataset version, and scoring code recorded so the eval can be rerun?
- Are raw outputs saved, and is there a canary set to rerun whenever the vendor updates the model?
Presence, Not Absence
An evaluation can show that a capability or a risk exists. It can’t show that none does, only that it wasn’t found with this much effort.
- How hard did anyone try to elicit this capability or failure: which prompts, tools, fine-tuning, and how many attempts?
- Who red-teamed this, and do they have real domain expertise?
The Team Is the System
When people work with AI, the person and the AI together produce the outcome. Test them together, including how often people accept advice that is wrong.
- Was the person plus the AI tested together, and compared with the human alone and the AI alone?
- How accurate are people on the cases where the AI is wrong?
Looking for something shorter? Use it now has friendly questions for common moments. To see which patterns fit your situation, try Find your patterns.