The Construct Gap
“A benchmark measures what it measures, not what its name says.”
- Last reviewed
- 15 Sep 2026
- Source types
- 7 peer-reviewed or classic, 1 preprint
- Version
- v1.0 5 Oct 2026
The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status
In plain terms
A test called “reasoning” only measures how a system does on that test’s items, in that format. Whether that says anything about reasoning in general needs its own evidence.
Takeaways #
- Reasoning, safety, understanding, and helpfulness are constructs. Nothing measures them directly.
- A benchmark is one way of turning a construct into a test. Treat its name as a claim to check.
- Trace the chain before trusting a score: concept, definition, test items, scoring, number.
- Most LLM benchmarks don’t spell that chain out.
What it means #
Measurement researchers separate the thing you care about (the construct) from the thing you can observe (the measurement). Validity is how well the evidence supports the jump from one to the other. A “reasoning” benchmark made of multiple-choice logic puzzles measures performance on those puzzles, in that format. Whether that says anything about reasoning in general is a separate question that needs its own evidence.
Wallach et al. (2025), building on Adcock and Collier (2001), lay out four levels: a background concept (the broad, fuzzy idea), a systematized concept (an explicit definition), a measurement instrument (the actual test), and the measurements it produces. Moving between levels takes systematization, operationalization, and application, and then the whole chain has to be interrogated for validity. A lot of AI eval arguments happen because people skip the first two levels and fight about the numbers.
The evidence #
reviewed 15 Sep 2026Big names, narrow tests. Raji et al. (2021) show how benchmarks like ImageNet and GLUE get treated as measures of general capability, the “everything in the whole wide world,” which no finite dataset can support.
The problem is widespread. Bean et al. (2025) had 29 expert reviewers systematically review 445 LLM benchmarks from leading NLP and ML venues and found recurring patterns that undermine the validity of the claims made from them. They offer eight recommendations.
Validity lives in the interpretation. Messick (1995) argued that validity isn’t a property of a test. It’s a property of the inferences and uses drawn from its scores, including their consequences.
Mismatches cause harm. Jacobs and Wallach (2021) show how gaps between constructs like “risk” or “fairness” and their operational definitions lead to real-world harm.
Claims need matching evidence. Salaudeen et al. (2025) propose a validity-centered framework: the same score can support a narrow claim and fail to support a broad one.
Use it #
- Finish this sentence: “This score is evidence that the system can ___ under ___ conditions.” If you can’t fill in the blanks, you don’t know what the score means.
- Read 20 test items yourself. Would a person who aced them actually have the skill in the benchmark’s name?
- Match the size of the claim to the evidence. “Scores 85% on X” is safe. “Can reason” is not.
- Check whether the test format (multiple choice, single turn, short answer) matches how the system will actually be used.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
A narrow benchmark is a fine basis for a narrow claim. The gap opens when a score is used to claim something broader, such as general ability. If your claim is only that a system does these particular tasks, the score may support it.
Origins #
Lee Cronbach and Paul Meehl introduced construct validity in psychology in 1955. Samuel Messick’s 1995 framework made it the center of modern validity theory. ML researchers brought these ideas into AI evaluation over the last several years. “The Construct Gap” is the name this site uses for the idea.
Sources #
- [1]Adcock, R., & Collier, D. (2001). Measurement validity: A shared standard for qualitative and quantitative research.American Political Science Review, 95(3), 529-546Open ↗ (opens in a new tab)
- [2]Bean, A. M., Kearns, R. O., Romanou, A., Hafner, F. S., Mayne, H., et al. (2025). Measuring what matters: Construct validity in large language model benchmarks.Advances in Neural Information Processing Systems (NeurIPS 2025), Datasets and Benchmarks TrackOpen ↗ (opens in a new tab)
- [3]Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests.Psychological Bulletin, 52(4), 281-302Open ↗ (opens in a new tab)
- [4]Jacobs, A. Z., & Wallach, H. (2021). Measurement and fairness.Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT 2021)Open ↗ (opens in a new tab)
- [5]Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning.American Psychologist, 50(9), 741-749Open ↗ (opens in a new tab)
- [6]Raji, I. D., Bender, E. M., Paullada, A., Denton, E., & Hanna, A. (2021). AI and the everything in the whole wide world benchmark.NeurIPS 2021 Datasets and Benchmarks TrackOpen ↗ (opens in a new tab)
- [7]Salaudeen, O., et al. (2025). Measurement to meaning: A validity-centered framework for AI evaluation.arXiv:2505.10573 PreprintOpen ↗ (opens in a new tab)
- [8]Wallach, H., Desai, M., Cooper, A. F., Wang, A., Atalla, C., et al. (2025). Position: Evaluating generative AI systems is a social science measurement challenge.Proceedings of the International Conference on Machine Learning (ICML 2025), PMLR 267Open ↗ (opens in a new tab)
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). The Construct Gap (v1.0). https://evalfieldguide.com/patterns/the-construct-gap@misc{lai-the-construct-gap,
title = {The Construct Gap},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/the-construct-gap}}
}https://evalfieldguide.com/patterns/the-construct-gapRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.