Guide
Questions to ask when buying an AI model or vendor
A vendor’s numbers answer the questions the vendor chose to ask. These questions are drawn from the patterns on this site. Each one comes from a documented pattern in the research, with the evidence one click away.
Goodhart’s Law
If you reward people or models for hitting a number, they get good at the number. Whether they got better at the actual job is a separate question that needs its own check.
- Which metric did you optimize during training and model selection, and is the reported score on that same metric?
- Was the reported test set ever used for tuning, prompt iteration, or choosing between checkpoints?
- Show us five outputs that score highly but are still wrong. What do they have in common?
- What will you measure after launch, and who decides when that metric has stopped being useful?
Campbell’s Law
The more a score decides who gets funded, hired, or bought, the more people bend their work around the score. The damage reaches the work itself, not just the number.
- Does a real decision rest on a single ranking, or on at least two independent measures plus a qualitative review?
- How many variants, seeds, or prompts were tried before this result was reported?
The Construct Gap
A test called “reasoning” only measures how a system does on that test’s items, in that format. Whether that says anything about reasoning in general needs its own evidence.
- Which of our real tasks does this benchmark resemble, and how was that checked?
- What would a high score here fail to tell us?
The Functionality Fallacy
Before asking whether an AI system is fair or safe, ask whether it works at all. Plenty of deployed systems fail at their stated job, and a vendor’s claim is not evidence.
- What evidence shows the system works on data like ours, gathered by someone other than the vendor?
- What does the system do when it’s wrong, and how often does that happen?
Benchmark Saturation
Tests stop being useful once everyone scores near the top. The remaining differences are mostly noise, and the real gaps are somewhere else.
- How close are top systems to this benchmark’s ceiling?
- Is there a newer or harder test that still separates systems?
Data Contamination
If a model saw the test questions while it was training, a high score may only mean it remembers them. Fresh or private questions are the reliable check.
- What overlap analysis was run between the test set and the training data?
- Were results confirmed on data created after the training cutoff?
The Gold Standard Myth
The answer key and the human ratings are measurements too, and they contain errors and disagreement. Treat them as evidence, not as ground truth.
- How often do independent raters agree on the labels, and was that measured before the labels were trusted?
- Who audited the labels on the items where strong models disagree with the answer key?
Prompt Sensitivity
Small changes in wording or formatting can swing a model’s score. A result from one prompt is a result for that prompt.
- How much does the score move across reasonable rephrasings?
- Was the prompt tuned on the test set?
Judge Bias
Using an AI to grade AI is fast and cheap, but graders have favorites: the first answer, the longer answer, and answers that sound like their own. Check the grader against people.
- Was the AI grader checked against human judgments on a sample?
- Were answer order and length controlled for, and is the grader from the same model family?
Once Is Not Reliable
A system that works once may fail the next time. Customers live with the all-runs-pass rate, not the best-run rate.
- How many runs per item, and what is the pass rate across all of them?
- Is the result reported as pass@1, pass@k, or passing on every run?
The Cost Frontier
You can buy accuracy with more compute, so a score without its cost is half a result. Compare systems at matched cost, or show the whole tradeoff.
- What did each task cost in tokens or dollars, and what was the latency, next to every accuracy number?
- Was a cheap baseline, such as a single call, a smaller model, or a simple retry loop, included at a matched budget?
No Error Bars, No Result
Every score is an estimate. If two systems differ by less than the noise, nobody has shown that they differ.
- What is the 95% confidence interval on each score, and how many items and runs does it rest on?
- Is the gap between systems larger than the interval, using a paired comparison on the same items?
The Averaging Trap
A strong average can hide a group the system fails badly. Report results for the groups that matter, and show the worst one.
- Which user groups, languages, or input types matter here, and were results reported for each?
- What is the worst-performing slice, and does it have enough examples to support its own error bars?
The Baseline Rule
A win is only as strong as what it beat. Compare against the simplest thing that could work and today’s actual process, with the same tuning effort.
- Was the current process, or the simplest thing that could work, included as a baseline?
- Did the baselines get the same tuning budget as the new method?
The Benchmark Lottery
Which benchmarks get reported can decide who looks best. Pick the suite before you see results, and read the whole table.
- Was the benchmark suite chosen and written down before any results were seen?
- Which common benchmarks are missing from the results table, and why?
The Reproducibility Rule
If nobody can rerun an evaluation, it is an anecdote. Record the model version, prompts, settings, data, and scoring code, and save the raw outputs.
- Are the model version, system prompt, decoding settings, dataset version, and scoring code recorded so the eval can be rerun?
- Are raw outputs saved, and is there a canary set to rerun whenever the vendor updates the model?
Distribution Shift
A score measured on one group of people, places, or times doesn’t automatically carry over to another. Test on data from where the system will actually run.
- When was the test data collected, and from which users?
- How does it compare with current production inputs?
The Lab-to-Field Gap
A system that works in the lab can fail in real conditions: workflows, time pressure, lighting, bandwidth, and people. Watch it in use before calling it validated.
- What changed for real users in a pilot, measured how?
- Which lab results did not hold up in the pilot?
Looking for something shorter? Use it now has friendly questions for common moments. To see which patterns fit your situation, try Find your patterns.