Once Is Not Reliable
“Succeeding once is not the same as succeeding every time.”
- Last reviewed
- 15 Sep 2026
- Source types
- 1 peer-reviewed or classic, 1 preprint
- Version
- v1.0 5 Oct 2026
The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status
In plain terms
A system that works once may fail the next time. Customers live with the all-runs-pass rate, not the best-run rate.
Takeaways #
- AI outputs vary from run to run. A single pass hides that.
- pass@k asks, “Did it work at least once in k tries?” pass^k asks, “Did it work all k times?” Users live with the second one.
- Capability gains don’t automatically bring reliability gains.
- For anything customer-facing or agentic, measure consistency across repeated runs.
What it means #
Many evals run each test case once, or report the best of several attempts. That’s fine for measuring whether a model can do something. It’s the wrong number for deciding whether to put it in front of customers. A support agent that resolves a refund correctly 60% of the time will get it wrong for a lot of real people, even if it “passes” the test on a good run.
The evidence #
reviewed 15 Sep 2026Reliability drops fast with repetition. Yao et al. (2025) built τ-bench, which simulates conversations between users and agents that have tools and policies to follow. State-of-the-art function-calling agents like GPT-4o succeeded on fewer than 50% of tasks. Their pass^k metric showed consistency was worse: pass^8 fell below 25% in the retail domain.
Accuracy and reliability are different axes. Rabanser et al. (2026) proposed twelve reliability metrics across consistency, robustness, predictability, and safety. Across 15 models, recent capability gains produced only small improvements in reliability.
Use it #
- Run each test case several times (at least three to five for LLM outputs) and report both the average success rate and the all-runs-pass rate.
- Separate random, occasional failures from consistent failures on specific cases. They need different fixes.
- Include paraphrased versions of the same task in your test set.
- Set reliability thresholds based on the cost of a failure, not on what the model can currently hit.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
Consistency matters most for agents and multi-step tasks. If one good answer is enough, or a person reviews every output, a strict repeat-success metric such as pass^k may ask for more than you need.
Origins #
“Once Is Not Reliable” is this site’s name. The pass^k metric comes from Yao et al.’s τ-bench, and the broader framing draws on reliability engineering in safety-critical fields.
Sources #
- [1]Rabanser, S., Kapoor, S., Kirgis, P., Liu, K., Utpala, S., & Narayanan, A. (2026). Towards a science of AI agent reliability.arXiv:2602.16666 PreprintOpen ↗ (opens in a new tab)
- [2]Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2025). τ-bench: A benchmark for tool-agent-user interaction in real-world domains.International Conference on Learning Representations (ICLR 2025)Open ↗ (opens in a new tab)
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). Once Is Not Reliable (v1.0). https://evalfieldguide.com/patterns/once-is-not-reliable@misc{lai-once-is-not-reliable,
title = {Once Is Not Reliable},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/once-is-not-reliable}}
}https://evalfieldguide.com/patterns/once-is-not-reliableRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.