Presence, Not Absence
“An evaluation can show a capability or risk exists. It can’t prove one doesn’t.”
- Last reviewed
- 15 Sep 2026
- Source types
- 2 peer-reviewed or classic, 2 preprints
- Version
- v1.0 5 Oct 2026
The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status
In plain terms
An evaluation can show that a capability or a risk exists. It can’t show that none does, only that it wasn’t found with this much effort.
Takeaways #
- A failed attempt to get a model to do something tells you about the attempt, not the model’s ceiling.
- Better prompts, tools, fine-tuning, or more time can unlock abilities an eval missed.
- Safety claims built on “we tried and it didn’t” need to say how hard someone tried.
- Report elicitation effort alongside every negative result.
What it means #
Edsger Dijkstra (1970) made this point about software in his Notes on Structured Programming: “Program testing can be used to show the presence of bugs, but never to show their absence!” General-purpose AI systems make it sharper. The space of possible inputs is effectively infinite, and capabilities often show up only with the right prompt, tool, or scaffolding. A clean eval report means you didn’t find the problem under the conditions you tried.
The evidence #
reviewed 15 Sep 2026Evaluations for extreme risks. Shevlane et al. (2023) describe dangerous capability evaluations and alignment evaluations as key inputs to training and deployment decisions, and discuss the limits of what those evaluations can catch.
Models can underperform on purpose. van der Weij et al. (2025) showed that frontier models like GPT-4 and Claude 3 Opus could be prompted to selectively underperform on dangerous capability evaluations while keeping performance on general ones. Models could also be fine-tuned to hide a capability unless given a password.
Red teaming finds a lot, not everything. Ganguli et al. (2022) released a dataset of nearly 39,000 red team attacks on language models. The harmful outputs ranged from offensive language to subtler unethical content, the kind a fixed benchmark wouldn’t think to ask about.
Use it #
- Phrase negative results precisely: “Under these conditions, with this effort, we didn’t observe X.”
- Report what elicitation was tried: prompts, tools, fine-tuning, number of attempts, and time spent.
- Budget for red teaming by people with real domain expertise.
- Revisit negative results when models, tools, or techniques change.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
A clean result is still informative. The pattern is about wording what it means: the tests found nothing, which is weaker than showing nothing is there.
Origins #
The core idea is Dijkstra’s, from “Notes on Structured Programming,” first circulated in 1969 with a second edition in 1970. “Presence, Not Absence” is this site’s name for applying it to AI evaluation.
Sources #
- [1]Dijkstra, E. W. (1970). Notes on structured programming (EWD249).Technological University Eindhoven (first circulated 1969; second edition April 1970)Open ↗ (opens in a new tab)
- [2]Ganguli, D., et al. (2022). Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned.arXiv:2209.07858 PreprintOpen ↗ (opens in a new tab)
- [3]Shevlane, T., et al. (2023). Model evaluation for extreme risks.arXiv:2305.15324 PreprintOpen ↗ (opens in a new tab)
- [4]van der Weij, T., Hofstätter, F., Jaffe, O., Brown, S. F., & Ward, F. R. (2025). AI sandbagging: Language models can strategically underperform on evaluations.International Conference on Learning Representations (ICLR 2025)Open ↗ (opens in a new tab)
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). Presence, Not Absence (v1.0). https://evalfieldguide.com/patterns/presence-not-absence@misc{lai-presence-not-absence,
title = {Presence, Not Absence},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/presence-not-absence}}
}https://evalfieldguide.com/patterns/presence-not-absenceRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.