Tool

Check a claim

Paste a claim about an AI system. You get the patterns that apply, why, and a question to ask for each. Not checking a claim? Describe your situation instead.

What kind of claim? Detected from the text, edit if wrong

Matching is rule-based against the patterns’ published criteria. How matching works

Results

Results appear here. Paste a claim and choose Find patterns that apply.

How matching works

Each claim type maps to the patterns a careful reader would check first, using the manuscript’s Playbook and red-flag table. A few phrases (such as “internal benchmark” or “superhuman”) add matches of their own. A pattern that matches more than one type moves up a step. This is a reading aid, not an assessment of the claim.

Show the rules

Benchmark score

  • Data Contamination strong match The score is on a public benchmark. Check whether the test items overlap with the training data.
  • No Error Bars, No Result strong match A headline number is reported without an interval or a count of items and runs.
  • The Construct Gap likely A benchmark measures what its items test, not what its name says.
  • Benchmark Saturation likely Near the ceiling of a benchmark, small gaps carry little information.
  • Goodhart’s Law likely If anyone iterated against this benchmark, it is a target rather than a measure.

Comparison to other models

Human-preference win rate

AI-graded result

  • Judge Bias strong match An AI grader has known biases: answer order, answer length, and a preference for its own writing.
  • The Gold Standard Myth likely A grader is only as good as the human labels it was checked against.
  • No Error Bars, No Result likely AI-graded scores are still estimates and need intervals.
  • Criteria Drift worth checking Rubrics change as people read real outputs. Check whether the criteria were versioned.
  • Prompt Sensitivity worth checking The grader’s own prompt affects the scores it gives.

Production accuracy

  • The Functionality Fallacy strong match Vendor-reported performance is a claim. Independent validation on your own population is evidence.
  • Distribution Shift likely Performance on one population does not automatically carry over to another.
  • The Lab-to-Field Gap likely Real use adds workflows, time pressure, and people that the lab left out.
  • The Averaging Trap likely Production accuracy is an average. Check the worst-performing groups.
  • Once Is Not Reliable likely Customer-facing results depend on how often it works every time, not once.
  • Evaluation Awareness worth checking Behavior under test may differ from behavior in real use, so check monitoring.