Guide

Questions to ask when buying an AI model or vendor

A vendor’s numbers answer the questions the vendor chose to ask. These questions are drawn from the patterns on this site. Each one comes from a documented pattern in the research, with the evidence one click away.

Print this as a one-page checklist

Goodhart’s Law

If you reward people or models for hitting a number, they get good at the number. Whether they got better at the actual job is a separate question that needs its own check.

Read the evidence: Goodhart’s Law

Campbell’s Law

The more a score decides who gets funded, hired, or bought, the more people bend their work around the score. The damage reaches the work itself, not just the number.

Read the evidence: Campbell’s Law

The Construct Gap

A test called “reasoning” only measures how a system does on that test’s items, in that format. Whether that says anything about reasoning in general needs its own evidence.

Read the evidence: The Construct Gap

The Functionality Fallacy

Before asking whether an AI system is fair or safe, ask whether it works at all. Plenty of deployed systems fail at their stated job, and a vendor’s claim is not evidence.

Read the evidence: The Functionality Fallacy

Benchmark Saturation

Tests stop being useful once everyone scores near the top. The remaining differences are mostly noise, and the real gaps are somewhere else.

Read the evidence: Benchmark Saturation

Data Contamination

If a model saw the test questions while it was training, a high score may only mean it remembers them. Fresh or private questions are the reliable check.

Read the evidence: Data Contamination

The Gold Standard Myth

The answer key and the human ratings are measurements too, and they contain errors and disagreement. Treat them as evidence, not as ground truth.

Read the evidence: The Gold Standard Myth

Prompt Sensitivity

Small changes in wording or formatting can swing a model’s score. A result from one prompt is a result for that prompt.

Read the evidence: Prompt Sensitivity

Judge Bias

Using an AI to grade AI is fast and cheap, but graders have favorites: the first answer, the longer answer, and answers that sound like their own. Check the grader against people.

Read the evidence: Judge Bias

Once Is Not Reliable

A system that works once may fail the next time. Customers live with the all-runs-pass rate, not the best-run rate.

Read the evidence: Once Is Not Reliable

The Cost Frontier

You can buy accuracy with more compute, so a score without its cost is half a result. Compare systems at matched cost, or show the whole tradeoff.

Read the evidence: The Cost Frontier

No Error Bars, No Result

Every score is an estimate. If two systems differ by less than the noise, nobody has shown that they differ.

Read the evidence: No Error Bars, No Result

The Averaging Trap

A strong average can hide a group the system fails badly. Report results for the groups that matter, and show the worst one.

Read the evidence: The Averaging Trap

The Baseline Rule

A win is only as strong as what it beat. Compare against the simplest thing that could work and today’s actual process, with the same tuning effort.

Read the evidence: The Baseline Rule

The Benchmark Lottery

Which benchmarks get reported can decide who looks best. Pick the suite before you see results, and read the whole table.

Read the evidence: The Benchmark Lottery

The Reproducibility Rule

If nobody can rerun an evaluation, it is an anecdote. Record the model version, prompts, settings, data, and scoring code, and save the raw outputs.

Read the evidence: The Reproducibility Rule

Distribution Shift

A score measured on one group of people, places, or times doesn’t automatically carry over to another. Test on data from where the system will actually run.

Read the evidence: Distribution Shift

The Lab-to-Field Gap

A system that works in the lab can fail in real conditions: workflows, time pressure, lighting, bandwidth, and people. Watch it in use before calling it validated.

Read the evidence: The Lab-to-Field Gap

Looking for something shorter? Use it now has friendly questions for common moments. To see which patterns fit your situation, try Find your patterns.