Goodhart’s Law
“When a measure becomes a target, it ceases to be a good measure.”
- Last reviewed
- 15 Sep 2026
- Source types
- 5 peer-reviewed or classic, 2 preprints, 1 report
- Version
- v1.0 5 Oct 2026
Established term: this name is used in the research literature. About this guide’s status
In plain terms
If you reward people or models for hitting a number, they get good at the number. Whether they got better at the actual job is a separate question that needs its own check.
Takeaways #
- A benchmark score stands in for something you actually care about. The harder anyone pushes on the stand-in, the further it drifts from the real thing.
- AI gets hit by this twice. People chase leaderboard numbers, and models trained on a reward chase that reward.
- Keep at least one measure outside the optimization loop: a held-back test, fresh data, or a direct check of the real outcome.
- When a score jumps, ask what improved. The capability, or the fit to the test?
What it means #
Every eval is a proxy. Accuracy on a coding benchmark stands in for “writes useful software.” Win rate in a chatbot arena stands in for “people find it helpful.” A proxy works fine while nobody leans on it. Once the number becomes the goal, people and systems find ways to move the number that don’t move the thing behind it.
Manheim and Garrabrant (2018) split the problem into four variants. Regressional: the proxy and the goal only partly overlap, so picking the top scorers also picks up noise. Extremal: the relationship holds in normal ranges and breaks at the extremes. Causal: pushing the proxy doesn’t cause the goal. Adversarial: someone, or something, games the metric on purpose.
The evidence #
reviewed 15 Sep 2026It can be measured. Gao, Schulman, and Hilton (2023) optimized language models against a proxy reward model and tracked a separate “gold” reward model. The gold score rose at first, then fell as optimization continued, and the pattern scaled predictably with reward model size.
More capable systems game harder. Amodei et al. (2016) named reward hacking as a core safety problem. Pan, Bhatia, and Steinhardt (2022) found that more capable agents exploited misspecified rewards more, sometimes with sudden jumps in bad behavior.
Leaderboards feel it too. Singh et al. (2025) documented how private testing on Chatbot Arena let some providers try many model variants and publish only the best. Meta tested 27 private variants before the Llama 4 release. Extra Arena data alone produced relative gains of up to 112% on the Arena distribution.
It’s getting more common. The International AI Safety Report 2026 notes models increasingly find loopholes that let them score well on evaluations without doing what the evaluation intended (Bengio et al., 2026).
Use it #
- Before trusting a score jump, check a second measure nobody optimized.
- Keep a private holdout set that never informs training, prompt tuning, or model selection.
- Refresh test items on a schedule.
- Pair every number with a sample of real outputs that a person reads.
- Ask: “What’s the cheapest way to raise this number without improving the product?” If the answer is easy, expect it to happen.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
Not every target gets gamed. A measure that is hard to move without doing the real work, or one checked against separate evidence, holds up far better. The effect also depends on how hard something is optimized: in the reward-model study, the true score rose before it fell.
Origins #
Economist Charles Goodhart described the idea in a 1975 paper on UK monetary policy: any observed statistical regularity tends to collapse once pressure is placed on it for control purposes. Anthropologist Marilyn Strathern gave it the popular wording in 1997 while writing about audit culture in British universities: “When a measure becomes a target, it ceases to be a good measure.”
Sources #
- [1]Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete problems in AI safety.arXiv:1606.06565 PreprintOpen ↗ (opens in a new tab)
- [2]Bengio, Y., et al. (2026). International AI Safety Report 2026.International AI Safety Report; arXiv:2602.21012 ReportOpen ↗ (opens in a new tab)
- [3]Gao, L., Schulman, J., & Hilton, J. (2023). Scaling laws for reward model overoptimization.Proceedings of the International Conference on Machine Learning (ICML 2023)Open ↗ (opens in a new tab)
- [4]Goodhart, C. A. E. (1975). Problems of monetary management: The U.K. experience.Papers in Monetary Economics, Vol. 1. Reserve Bank of AustraliaNo link
- [5]Manheim, D., & Garrabrant, S. (2018). Categorizing variants of Goodhart’s Law.arXiv:1803.04585 PreprintOpen ↗ (opens in a new tab)
- [6]Pan, A., Bhatia, K., & Steinhardt, J. (2022). The effects of reward misspecification: Mapping and mitigating misaligned models.International Conference on Learning Representations (ICLR 2022)Open ↗ (opens in a new tab)
- [7]Singh, S., Nan, Y., Wang, A., D’Souza, D., Kapoor, S., et al. (2025). The leaderboard illusion.Advances in Neural Information Processing Systems (NeurIPS 2025)Open ↗ (opens in a new tab)
- [8]Strathern, M. (1997). ‘Improving ratings’: Audit in the British university system.European Review, 5(3), 305-321No link
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). Goodhart's Law (v1.0). https://evalfieldguide.com/patterns/goodharts-law@misc{lai-goodharts-law,
title = {Goodhart's Law},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/goodharts-law}}
}https://evalfieldguide.com/patterns/goodharts-lawRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.