Tool

Use it now

You do not need anyone’s buy-in to ask a good question. These are phrased as curiosity, not challenge, so you can use them in a meeting, a Slack thread, or a vendor call tomorrow. Tap a question to copy it.

When you hear

“It scores 94% on the benchmark.”

A good answer: Names the task, gives a range or number of runs, and says how contamination was checked, or admits it wasn’t.

Why: The Construct Gap, Once Is Not Reliable, Data Contamination

When you hear

“Model A beats Model B.”

A good answer: Same setup for both, a gap larger than the noise, and a cost number next to the quality number.

Why: Prompt Sensitivity, No Error Bars, No Result, The Cost Frontier

When you hear

A great demo, or “it worked on my examples.”

A good answer: Examples were chosen before seeing results, our own data gets tried, and failures are shown without being asked twice.

Why: The Jagged Frontier, Distribution Shift, Presence, Not Absence

When you hear

“We had an AI grade the answers.”

A good answer: A human spot-check exists, the grader is separate, and there is a known list of biases it was checked for.

Why: Judge Bias, The Gold Standard Myth

When you hear

“Users love it” or “the win rate is 70%.”

A good answer: A described group, blind comparison, and a separate check for correctness.

Why: The Metric Mirage, Criteria Drift, The Construct Gap

When you hear

“It’s ready to ship.”

A good answer: A named owner, a monitoring plan, and a stated list of known gaps.

Why: The Lab-to-Field Gap, The Team Is the System, Presence, Not Absence

When you hear

“It’s 95% accurate.”

A good answer: A defined population, a baseline for comparison, and error types broken out instead of one average.

Why: The Averaging Trap, The Baseline Rule

When you hear

“We’re improving the metric every week.”

A good answer: A fresh test set is held back, and someone has looked at what the metric leaves out.

Why: Goodhart’s Law, Adaptive Overfitting, Campbell’s Law

Want the full reasoning? Each pattern page has the evidence and sources. Reviewing a specific claim? Try the claim checker. Longer lists: buying an AI model or vendor, a launch review, building an evaluation.