1 of 5 in this categoryNext: 12 Judge Bias →
No. 11Runningv1.0For: Building

Prompt Sensitivity

“A score from one prompt is a score for that prompt.”

Last reviewed
15 Sep 2026
Sources
3 · Cite
Source types
2 peer-reviewed or classic, 1 preprint
Version
v1.0 5 Oct 2026

Established term: this name is used in the research literature. About this guide’s status

In plain terms

Small changes in wording or formatting can swing a model’s score. A result from one prompt is a result for that prompt.

Takeaways

  • Small formatting changes (spacing, separators, option labels, wording) can swing LLM results by a lot.
  • The best format for one model often isn’t the best for another, so single-prompt comparisons can be unfair.
  • Report results across several reasonable prompts, including the spread.
  • Test with the prompts real users actually write.

What it means

Traditional software gives the same result for the same input. Language models respond to surface details a person would ignore. If you evaluate two models with one fixed prompt template, you’re partly measuring how well each model happens to like that template. The ranking you get might not survive a rewrite.

The evidence

reviewed 15 Sep 2026

Formatting alone moves scores. Sclar et al. (2024) found that several widely used open-source LLMs were extremely sensitive to subtle formatting changes in few-shot prompts, with differences of up to 76 accuracy points for LLaMA-2-13B. The best format for one model correlated only weakly with the best format for another. They propose FormatSpread to report a range instead of one number.

Single-prompt results are brittle. Mizrahi et al. (2024) analyzed 6.5 million instances across 20 models and 39 tasks and found results from single-instruction evaluation brittle. They call for evaluating with a diverse set of prompts, and for choosing metrics that fit the needs of different users, like model developers versus teams building products.

Implementation details matter. Biderman et al. (2024), drawing on experience maintaining a widely used LLM evaluation harness, show that small choices in prompts, answer extraction, and normalization change scores, and argue for sharing exact prompts and code.

Use it

  1. Test at least three to five prompt variants and report the mean and range.
  2. Use the same set of variants for every model you compare.
  3. Include messy, realistic user phrasings, not just clean templates.
  4. Save the exact prompts and scoring code alongside the results.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

Sensitivity varies by model and task, and the largest swings in the evidence come from specific models and formats. If you run one fixed, well-tested prompt, the question becomes how that prompt holds up on your real inputs.

Origins

Prompt sensitivity became a named research topic with the rise of few-shot prompting in large language models. The term describes a family of findings rather than one paper.

Sources

  1. [1]
    Biderman, S., et al. (2024). Lessons from the trenches on reproducible evaluation of language models.arXiv:2405.14782 Preprint
    Open ↗ (opens in a new tab)
  2. [2]
    Mizrahi, M., et al. (2024). State of what art? A call for multi-prompt LLM evaluation.Transactions of the Association for Computational Linguistics, 12
    Open ↗ (opens in a new tab)
  3. [3]
    Sclar, M., Choi, Y., Tsvetkov, Y., & Suhr, A. (2024). Quantifying language models’ sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting.International Conference on Learning Representations (ICLR 2024)
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). Prompt Sensitivity (v1.0). https://evalfieldguide.com/patterns/prompt-sensitivity

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 12 →