Principles for judging whether an AI system actually holds up.
26 short, sourced laws for people who build, buy, or design with AI. How to use this guide
All laws, by where they bite
I.
What you’re measuring
5 of 5 published- No. 01MeasuringGoodhart’s Law“When a measure becomes a target, it ceases to be a good measure.”Read the law →
- No. 02MeasuringCampbell’s Law“The more a number drives decisions, the more it gets corrupted, and the more it distorts the work it was meant to track.”Read the law →
- No. 03MeasuringThe Construct Gap“A benchmark measures what it measures, not what its name says.”Read the law →
- No. 04MeasuringThe Functionality Fallacy“Don’t assume an AI system works. Whether it works is the first question, not a settled one.”Read the law →
- No. 05MeasuringThe Metric Mirage“Change the metric and you can change the conclusion.”Read the law →
II.
The test itself
5 of 5 published- No. 06The testBenchmark Saturation“Every benchmark has a shelf life.”Read the law →
- No. 07The testData Contamination“If the test was in the training data, the score measures memory, not skill.”Read the law →
- No. 08The testAdaptive Overfitting“Every time you tune against the test set, it’s worth a little less.”Read the law →
- No. 09The testThe Clever Hans Effect“A model can get the right answer for the wrong reason.”Read the law →
- No. 10The testThe Gold Standard Myth“Labels and human ratings are measurements with error, not ground truth.”Read the law →
III.
Running the eval
5 of 5 published- No. 11RunningPrompt Sensitivity“A score from one prompt is a score for that prompt.”Read the law →
- No. 12RunningJudge Bias“An AI grader has preferences, including a preference for itself.”Read the law →
- No. 13RunningCriteria Drift“You figure out what ‘good’ means by grading outputs, so your criteria will change.”Read the law →
- No. 14RunningOnce Is Not Reliable“Succeeding once is not the same as succeeding every time.”Read the law →
- No. 15RunningThe Cost Frontier“Accuracy without cost is half a result.”Read the law →
IV.
Reading the results
5 of 5 published- No. 16ResultsNo Error Bars, No Result“A difference smaller than the noise is not a difference.”Read the law →
- No. 17ResultsThe Averaging Trap“An average score can hide exactly who the system fails.”Read the law →
- No. 18ResultsThe Baseline Rule“A result is only as strong as the baseline it beats.”Read the law →
- No. 19ResultsThe Benchmark Lottery“Which benchmarks you pick can decide who wins.”Read the law →
- No. 20ResultsThe Reproducibility Rule“An eval nobody can rerun is an anecdote.”Read the law →
V.
Beyond the benchmark
6 of 6 published- No. 21BeyondPresence, Not Absence“An evaluation can show a capability or risk exists. It can’t prove one doesn’t.”Read the law →
- No. 22BeyondEvaluation Awareness“A model that can tell it’s being tested may not act the way it will in the real world.”Read the law →
- No. 23BeyondDistribution Shift“Performance measured on one population doesn’t automatically carry over to another.”Read the law →
- No. 24BeyondThe Lab-to-Field Gap“A system that works in the lab can fail in the places people actually use it.”Read the law →
- No. 25BeyondThe Team Is the System“When people work with AI, evaluate the person and the AI together, not the AI alone.”Read the law →
- No. 26BeyondThe Jagged Frontier“AI capability is uneven, and its edges don’t line up with what people find hard.”Read the law →
What you’re measuring
- No. 01Goodhart’s Law“When a measure becomes a target, it ceases to be a good measure.”
- No. 02Campbell’s Law“The more a number drives decisions, the more it gets corrupted, and the more it distorts the work it was meant to track.”
- No. 03The Construct Gap“A benchmark measures what it measures, not what its name says.”
- No. 04The Functionality Fallacy“Don’t assume an AI system works. Whether it works is the first question, not a settled one.”
- No. 05The Metric Mirage“Change the metric and you can change the conclusion.”
The test itself
- No. 06Benchmark Saturation“Every benchmark has a shelf life.”
- No. 07Data Contamination“If the test was in the training data, the score measures memory, not skill.”
- No. 08Adaptive Overfitting“Every time you tune against the test set, it’s worth a little less.”
- No. 09The Clever Hans Effect“A model can get the right answer for the wrong reason.”
- No. 10The Gold Standard Myth“Labels and human ratings are measurements with error, not ground truth.”
Running the eval
- No. 11Prompt Sensitivity“A score from one prompt is a score for that prompt.”
- No. 12Judge Bias“An AI grader has preferences, including a preference for itself.”
- No. 13Criteria Drift“You figure out what ‘good’ means by grading outputs, so your criteria will change.”
- No. 14Once Is Not Reliable“Succeeding once is not the same as succeeding every time.”
- No. 15The Cost Frontier“Accuracy without cost is half a result.”
Reading the results
- No. 16No Error Bars, No Result“A difference smaller than the noise is not a difference.”
- No. 17The Averaging Trap“An average score can hide exactly who the system fails.”
- No. 18The Baseline Rule“A result is only as strong as the baseline it beats.”
- No. 19The Benchmark Lottery“Which benchmarks you pick can decide who wins.”
- No. 20The Reproducibility Rule“An eval nobody can rerun is an anecdote.”
Beyond the benchmark
- No. 21Presence, Not Absence“An evaluation can show a capability or risk exists. It can’t prove one doesn’t.”
- No. 22Evaluation Awareness“A model that can tell it’s being tested may not act the way it will in the real world.”
- No. 23Distribution Shift“Performance measured on one population doesn’t automatically carry over to another.”
- No. 24The Lab-to-Field Gap“A system that works in the lab can fail in the places people actually use it.”
- No. 25The Team Is the System“When people work with AI, evaluate the person and the AI together, not the AI alone.”
- No. 26The Jagged Frontier“AI capability is uneven, and its edges don’t line up with what people find hard.”