Tool
Use it now
You do not need anyone’s buy-in to ask a good question. These are phrased as curiosity, not challenge, so you can use them in a meeting, a Slack thread, or a vendor call tomorrow. Tap a question to copy it.
When you hear
“It scores 94% on the benchmark.”
A good answer: Names the task, gives a range or number of runs, and says how contamination was checked, or admits it wasn’t.
Why: The Construct Gap, Once Is Not Reliable, Data Contamination
When you hear
“Model A beats Model B.”
A good answer: Same setup for both, a gap larger than the noise, and a cost number next to the quality number.
Why: Prompt Sensitivity, No Error Bars, No Result, The Cost Frontier
When you hear
A great demo, or “it worked on my examples.”
A good answer: Examples were chosen before seeing results, our own data gets tried, and failures are shown without being asked twice.
Why: The Jagged Frontier, Distribution Shift, Presence, Not Absence
When you hear
“We had an AI grade the answers.”
A good answer: A human spot-check exists, the grader is separate, and there is a known list of biases it was checked for.
When you hear
“Users love it” or “the win rate is 70%.”
A good answer: A described group, blind comparison, and a separate check for correctness.
When you hear
“It’s ready to ship.”
A good answer: A named owner, a monitoring plan, and a stated list of known gaps.
Why: The Lab-to-Field Gap, The Team Is the System, Presence, Not Absence
When you hear
“It’s 95% accurate.”
A good answer: A defined population, a baseline for comparison, and error types broken out instead of one average.
When you hear
“We’re improving the metric every week.”
A good answer: A fresh test set is held back, and someone has looked at what the metric leaves out.
Want the full reasoning? Each pattern page has the evidence and sources. Reviewing a specific claim? Try the claim checker. Longer lists: buying an AI model or vendor, a launch review, building an evaluation.