Tool
Check a claim
Paste a claim about an AI system. You get the patterns that apply, why, and a question to ask for each. Not checking a claim? Describe your situation instead.
What kind of claim? Detected from the text, edit if wrong
Matching is rule-based against the patterns’ published criteria. How matching works
Results
Results appear here. Paste a claim and choose Find patterns that apply.
How matching works
Each claim type maps to the patterns a careful reader would check first, using the manuscript’s Playbook and red-flag table. A few phrases (such as “internal benchmark” or “superhuman”) add matches of their own. A pattern that matches more than one type moves up a step. This is a reading aid, not an assessment of the claim.
Show the rules
Benchmark score
- Data Contamination strong match The score is on a public benchmark. Check whether the test items overlap with the training data.
- No Error Bars, No Result strong match A headline number is reported without an interval or a count of items and runs.
- The Construct Gap likely A benchmark measures what its items test, not what its name says.
- Benchmark Saturation likely Near the ceiling of a benchmark, small gaps carry little information.
- Goodhart’s Law likely If anyone iterated against this benchmark, it is a target rather than a measure.
Comparison to other models
- The Baseline Rule strong match A comparison is only as strong as the baseline it beats.
- No Error Bars, No Result strong match A win is claimed with no interval or number of runs.
- The Benchmark Lottery likely Which benchmarks are reported can decide who wins.
- The Cost Frontier likely Systems that cost very different amounts to run are not a fair comparison.
- The Averaging Trap worth checking An average gap can hide the groups where the system loses.
- The Reproducibility Rule worth checking Competitor numbers are only comparable if they came from the same prompts, settings, and scoring code.
Human-preference win rate
- Campbell’s Law strong match Rankings that drive buying and press attention invite gaming, including selective reporting of variants.
- Goodhart’s Law likely Optimizing for votes can raise the number without raising quality.
- The Averaging Trap likely A win rate averages over who voted and what they asked.
- The Gold Standard Myth likely Human ratings are measurements with error, and raters disagree.
- No Error Bars, No Result likely A win rate is an estimate and needs an interval.
- The Benchmark Lottery worth checking Check which comparisons were left out.
AI-graded result
- Judge Bias strong match An AI grader has known biases: answer order, answer length, and a preference for its own writing.
- The Gold Standard Myth likely A grader is only as good as the human labels it was checked against.
- No Error Bars, No Result likely AI-graded scores are still estimates and need intervals.
- Criteria Drift worth checking Rubrics change as people read real outputs. Check whether the criteria were versioned.
- Prompt Sensitivity worth checking The grader’s own prompt affects the scores it gives.
Production accuracy
- The Functionality Fallacy strong match Vendor-reported performance is a claim. Independent validation on your own population is evidence.
- Distribution Shift likely Performance on one population does not automatically carry over to another.
- The Lab-to-Field Gap likely Real use adds workflows, time pressure, and people that the lab left out.
- The Averaging Trap likely Production accuracy is an average. Check the worst-performing groups.
- Once Is Not Reliable likely Customer-facing results depend on how often it works every time, not once.
- Evaluation Awareness worth checking Behavior under test may differ from behavior in real use, so check monitoring.