Guide

Questions to ask when building an AI evaluation

Most evaluation mistakes are made before the first score is computed. These questions, drawn from the patterns, help you check the design before you trust the result.

Print this as a one-page checklist

Goodhart’s Law

If you reward people or models for hitting a number, they get good at the number. Whether they got better at the actual job is a separate question that needs its own check.

Read the evidence: Goodhart’s Law

The Construct Gap

A test called “reasoning” only measures how a system does on that test’s items, in that format. Whether that says anything about reasoning in general needs its own evidence.

Read the evidence: The Construct Gap

The Metric Mirage

The scoring method is part of the result. Score the same outputs a different way and a sudden leap can turn into steady progress, or the reverse.

Read the evidence: The Metric Mirage

Data Contamination

If a model saw the test questions while it was training, a high score may only mean it remembers them. Fresh or private questions are the reliable check.

Read the evidence: Data Contamination

Adaptive Overfitting

Each time you adjust a prompt or a model because of how it did on the test, the test tells you a little less. Keep a locked test that you use only to decide.

Read the evidence: Adaptive Overfitting

The Clever Hans Effect

A model can pass a test by picking up an accidental pattern, such as a watermark or a telltale word, instead of the skill you meant to test. Build cases where the shortcut fails.

Read the evidence: The Clever Hans Effect

The Gold Standard Myth

The answer key and the human ratings are measurements too, and they contain errors and disagreement. Treat them as evidence, not as ground truth.

Read the evidence: The Gold Standard Myth

Prompt Sensitivity

Small changes in wording or formatting can swing a model’s score. A result from one prompt is a result for that prompt.

Read the evidence: Prompt Sensitivity

Judge Bias

Using an AI to grade AI is fast and cheap, but graders have favorites: the first answer, the longer answer, and answers that sound like their own. Check the grader against people.

Read the evidence: Judge Bias

Criteria Drift

You work out what “good” means by reading real outputs, so your rubric will change as you grade. Version it, and re-grade a fixed set when it does.

Read the evidence: Criteria Drift

Once Is Not Reliable

A system that works once may fail the next time. Customers live with the all-runs-pass rate, not the best-run rate.

Read the evidence: Once Is Not Reliable

The Cost Frontier

You can buy accuracy with more compute, so a score without its cost is half a result. Compare systems at matched cost, or show the whole tradeoff.

Read the evidence: The Cost Frontier

No Error Bars, No Result

Every score is an estimate. If two systems differ by less than the noise, nobody has shown that they differ.

Read the evidence: No Error Bars, No Result

The Averaging Trap

A strong average can hide a group the system fails badly. Report results for the groups that matter, and show the worst one.

Read the evidence: The Averaging Trap

The Baseline Rule

A win is only as strong as what it beat. Compare against the simplest thing that could work and today’s actual process, with the same tuning effort.

Read the evidence: The Baseline Rule

The Reproducibility Rule

If nobody can rerun an evaluation, it is an anecdote. Record the model version, prompts, settings, data, and scoring code, and save the raw outputs.

Read the evidence: The Reproducibility Rule

Presence, Not Absence

An evaluation can show that a capability or a risk exists. It can’t show that none does, only that it wasn’t found with this much effort.

Read the evidence: Presence, Not Absence

The Team Is the System

When people work with AI, the person and the AI together produce the outcome. Test them together, including how often people accept advice that is wrong.

Read the evidence: The Team Is the System

Looking for something shorter? Use it now has friendly questions for common moments. To see which patterns fit your situation, try Find your patterns.