Campbell’s Law
“The more a number drives decisions, the more it gets corrupted, and the more it distorts the work it was meant to track.”
- Last reviewed
- 15 Sep 2026
- Source types
- 4 peer-reviewed or classic, 1 preprint
- Version
- v1.0 5 Oct 2026
Established term: this name is used in the research literature. About this guide’s status
In plain terms
The more a score decides who gets funded, hired, or bought, the more people bend their work around the score. The damage reaches the work itself, not just the number.
Takeaways #
- Goodhart’s Law is about the measure breaking. Campbell’s Law is about people and institutions bending around it.
- Leaderboards influence funding, press, hiring, and buying decisions. That gives everyone a reason to game them, often without anyone deciding to cheat.
- The damage spreads past the number. It shapes which research gets done and which products get built.
- Big decisions deserve several indicators plus qualitative evidence, never a single rank.
What it means #
When a benchmark becomes the currency of a field, it changes behavior. Teams pick problems that score well, report the best of many runs, quietly tune on the test, and skip work that doesn’t move the rank. None of this has to be fraud. Small, reasonable-looking choices add up to a distorted picture.
The same thing happens inside companies. If a launch review only asks for a win rate, teams learn to deliver a win rate.
The evidence #
reviewed 15 Sep 2026Scholarship bends. Lipton and Steinhardt (2019) documented troubling trends in ML papers, including failing to identify where empirical gains came from (for example, crediting a new architecture when the gains came from tuning) and presenting speculation as explanation.
Access is uneven. Singh et al. (2025) estimated that Google and OpenAI received about 19.2% and 20.4% of all Chatbot Arena data, while 83 open-weight models together received about 29.7%. More data from the arena means more ability to fit the arena.
Leaderboards encode someone’s values. Ethayarajh and Jurafsky (2020) argue that leaderboards reward accuracy while ignoring things real users care about, like model size, energy use, and robustness. A leaderboard’s idea of “best” is not yours.
Selection shapes the story. Dehghani et al. (2021) show that which benchmarks a community adopts can favor certain methods over others.
Use it #
- When a number drives a real decision (a vendor, a launch, a promotion), require at least two independent measures and a qualitative review.
- Ask how many variants, seeds, or prompts were tried before the reported result.
- Watch for teams narrowing scope to whatever the metric rewards.
- Reward people for finding flaws in the eval, not only for moving it.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
The pattern describes what happens under pressure and incentives, so a measure that nobody competes on or can easily game is less exposed. The leaderboard figures here describe one leaderboard, not every ranking.
Origins #
Social scientist Donald T. Campbell wrote in 1979: “The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it was intended to monitor.” He was writing about program evaluation in areas like education and policing. AI leaderboards fit the pattern closely.
Sources #
- [1]Campbell, D. T. (1979). Assessing the impact of planned social change.Evaluation and Program Planning, 2(1), 67-90Open ↗ (opens in a new tab)
- [2]Dehghani, M., Tay, Y., Gritsenko, A. A., et al. (2021). The benchmark lottery.arXiv:2107.07002 PreprintOpen ↗ (opens in a new tab)
- [3]Ethayarajh, K., & Jurafsky, D. (2020). Utility is in the eye of the user: A critique of NLP leaderboards.Proceedings of EMNLP 2020Open ↗ (opens in a new tab)
- [4]Lipton, Z. C., & Steinhardt, J. (2019). Troubling trends in machine learning scholarship.ACM Queue, 17(1)Open ↗ (opens in a new tab)
- [5]Singh, S., Nan, Y., Wang, A., D’Souza, D., Kapoor, S., et al. (2025). The leaderboard illusion.Advances in Neural Information Processing Systems (NeurIPS 2025)Open ↗ (opens in a new tab)
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). Campbell's Law (v1.0). https://evalfieldguide.com/patterns/campbells-law@misc{lai-campbells-law,
title = {Campbell's Law},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/campbells-law}}
}https://evalfieldguide.com/patterns/campbells-lawRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.