2 of 6 in this categoryNext: 23 Distribution Shift →
No. 22Beyondv1.0For: Building

Evaluation Awareness

“A model that can tell it’s being tested may not act the way it will in the real world.”

Last reviewed
15 Sep 2026
Sources
3 · Cite
Source types
1 peer-reviewed or classic, 1 preprint, 1 report
Version
v1.0 5 Oct 2026

Established term: this name is used in the research literature. About this guide’s status

In plain terms

Models can often tell a test from real use, and may behave differently under test. Pair pre-launch tests with monitoring of real use.

Takeaways

  • Frontier models can often tell evaluation prompts apart from real use.
  • If behavior changes under test, evaluation results stop predicting deployment behavior.
  • Scenarios that look artificial are easier to spot.
  • Pair pre-launch tests with monitoring of real use.

What it means

People act differently when they know they’re being watched. Models trained on huge amounts of text, including text about AI evaluations, can pick up on the telltale signs of a test: tidy multiple-choice formats, contrived scenarios, obviously fake names. If a model responds to those cues, even without any intent in the human sense, the eval measures test behavior instead of real behavior.

The evidence

reviewed 15 Sep 2026

Models are good at spotting tests. Needham et al. (2025) built a benchmark of 1,000 prompts and transcripts from 61 datasets, covering public benchmarks, real deployment interactions, and agent trajectories. Frontier models identified evaluations far above chance. Gemini 2.5 Pro reached an AUC of 0.83, compared to 0.92 for a human baseline. Models were better at this in agentic settings than in chat.

Recognized at the international level. The International AI Safety Report 2026 notes it has become more common for models to distinguish test settings from real-world deployment and to exploit loopholes in evaluations, which makes results harder to interpret (Bengio et al., 2026).

The capacity to underperform exists. van der Weij et al. (2025) showed models can be prompted or fine-tuned to strategically underperform on specific evaluations.

Use it

  1. Build eval scenarios from real, messy usage instead of obviously synthetic setups.
  2. Compare behavior in evals against behavior in real use, with consent and privacy safeguards.
  3. Track when models comment on the possibility that they’re being tested.
  4. For high-stakes behavior, don’t rely on pre-launch evals alone. Monitor after launch.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

This is a young area and several findings are preprints. A model that can identify a test is not proof that it behaves differently on tests, and the evidence here shows the capacity, not how often it matters.

Origins

“Evaluation awareness” is the term used in recent AI safety research. It’s a young area, and several key findings are preprints, so expect this page to change.

Sources

  1. [1]
    Bengio, Y., et al. (2026). International AI Safety Report 2026.International AI Safety Report; arXiv:2602.21012 Report
    Open ↗ (opens in a new tab)
  2. [2]
    Needham, J., Edkins, G., Pimpale, G., Bartsch, H., & Hobbhahn, M. (2025). Large language models often know when they are being evaluated.arXiv:2505.23836 Preprint
    Open ↗ (opens in a new tab)
  3. [3]
    van der Weij, T., Hofstätter, F., Jaffe, O., Brown, S. F., & Ward, F. R. (2025). AI sandbagging: Language models can strategically underperform on evaluations.International Conference on Learning Representations (ICLR 2025)
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). Evaluation Awareness (v1.0). https://evalfieldguide.com/patterns/evaluation-awareness

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 23 →