Guide
Questions for an AI launch review
A launch review is the last cheap moment to ask what the evaluation did not cover. These questions come from the patterns most relevant to shipping an AI feature, each with its evidence.
The Functionality Fallacy
Before asking whether an AI system is fair or safe, ask whether it works at all. Plenty of deployed systems fail at their stated job, and a vendor’s claim is not evidence.
- What evidence shows the system works on data like ours, gathered by someone other than the vendor?
- What does the system do when it’s wrong, and how often does that happen?
Judge Bias
Using an AI to grade AI is fast and cheap, but graders have favorites: the first answer, the longer answer, and answers that sound like their own. Check the grader against people.
- Was the AI grader checked against human judgments on a sample?
- Were answer order and length controlled for, and is the grader from the same model family?
Once Is Not Reliable
A system that works once may fail the next time. Customers live with the all-runs-pass rate, not the best-run rate.
- How many runs per item, and what is the pass rate across all of them?
- Is the result reported as pass@1, pass@k, or passing on every run?
No Error Bars, No Result
Every score is an estimate. If two systems differ by less than the noise, nobody has shown that they differ.
- What is the 95% confidence interval on each score, and how many items and runs does it rest on?
- Is the gap between systems larger than the interval, using a paired comparison on the same items?
The Averaging Trap
A strong average can hide a group the system fails badly. Report results for the groups that matter, and show the worst one.
- Which user groups, languages, or input types matter here, and were results reported for each?
- What is the worst-performing slice, and does it have enough examples to support its own error bars?
Presence, Not Absence
An evaluation can show that a capability or a risk exists. It can’t show that none does, only that it wasn’t found with this much effort.
- How hard did anyone try to elicit this capability or failure: which prompts, tools, fine-tuning, and how many attempts?
- Who red-teamed this, and do they have real domain expertise?
Evaluation Awareness
Models can often tell a test from real use, and may behave differently under test. Pair pre-launch tests with monitoring of real use.
- Are the eval scenarios built from real, messy usage, or do they look obviously synthetic?
- How does behavior in the eval compare with behavior in monitored real use?
Distribution Shift
A score measured on one group of people, places, or times doesn’t automatically carry over to another. Test on data from where the system will actually run.
- When was the test data collected, and from which users?
- How does it compare with current production inputs?
The Lab-to-Field Gap
A system that works in the lab can fail in real conditions: workflows, time pressure, lighting, bandwidth, and people. Watch it in use before calling it validated.
- What changed for real users in a pilot, measured how?
- Which lab results did not hold up in the pilot?
The Team Is the System
When people work with AI, the person and the AI together produce the outcome. Test them together, including how often people accept advice that is wrong.
- Was the person plus the AI tested together, and compared with the human alone and the AI alone?
- How accurate are people on the cases where the AI is wrong?
The Jagged Frontier
AI can excel at a task that seems hard and fail at one that seems easy, and people can’t see where the edge is. Map it for your own tasks with direct tests.
- Were tasks that are easy for people tested, not only hard ones?
- Where are the known failure pockets, and are users told about them?
Looking for something shorter? Use it now has friendly questions for common moments. To see which patterns fit your situation, try Find your patterns.