Guide

Questions for an AI launch review

A launch review is the last cheap moment to ask what the evaluation did not cover. These questions come from the patterns most relevant to shipping an AI feature, each with its evidence.

Print this as a one-page checklist

The Functionality Fallacy

Before asking whether an AI system is fair or safe, ask whether it works at all. Plenty of deployed systems fail at their stated job, and a vendor’s claim is not evidence.

Read the evidence: The Functionality Fallacy

Judge Bias

Using an AI to grade AI is fast and cheap, but graders have favorites: the first answer, the longer answer, and answers that sound like their own. Check the grader against people.

Read the evidence: Judge Bias

Once Is Not Reliable

A system that works once may fail the next time. Customers live with the all-runs-pass rate, not the best-run rate.

Read the evidence: Once Is Not Reliable

No Error Bars, No Result

Every score is an estimate. If two systems differ by less than the noise, nobody has shown that they differ.

Read the evidence: No Error Bars, No Result

The Averaging Trap

A strong average can hide a group the system fails badly. Report results for the groups that matter, and show the worst one.

Read the evidence: The Averaging Trap

Presence, Not Absence

An evaluation can show that a capability or a risk exists. It can’t show that none does, only that it wasn’t found with this much effort.

Read the evidence: Presence, Not Absence

Evaluation Awareness

Models can often tell a test from real use, and may behave differently under test. Pair pre-launch tests with monitoring of real use.

Read the evidence: Evaluation Awareness

Distribution Shift

A score measured on one group of people, places, or times doesn’t automatically carry over to another. Test on data from where the system will actually run.

Read the evidence: Distribution Shift

The Lab-to-Field Gap

A system that works in the lab can fail in real conditions: workflows, time pressure, lighting, bandwidth, and people. Watch it in use before calling it validated.

Read the evidence: The Lab-to-Field Gap

The Team Is the System

When people work with AI, the person and the AI together produce the outcome. Test them together, including how often people accept advice that is wrong.

Read the evidence: The Team Is the System

The Jagged Frontier

AI can excel at a task that seems hard and fail at one that seems easy, and people can’t see where the edge is. Map it for your own tasks with direct tests.

Read the evidence: The Jagged Frontier

Looking for something shorter? Use it now has friendly questions for common moments. To see which patterns fit your situation, try Find your patterns.