4 of 5 in this categoryNext: 15 The Cost Frontier →
No. 14Runningv1.0For: Building · Buying

Once Is Not Reliable

“Succeeding once is not the same as succeeding every time.”

Last reviewed
15 Sep 2026
Sources
2 · Cite
Source types
1 peer-reviewed or classic, 1 preprint
Version
v1.0 5 Oct 2026

The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status

In plain terms

A system that works once may fail the next time. Customers live with the all-runs-pass rate, not the best-run rate.

Takeaways

  • AI outputs vary from run to run. A single pass hides that.
  • pass@k asks, “Did it work at least once in k tries?” pass^k asks, “Did it work all k times?” Users live with the second one.
  • Capability gains don’t automatically bring reliability gains.
  • For anything customer-facing or agentic, measure consistency across repeated runs.

What it means

Many evals run each test case once, or report the best of several attempts. That’s fine for measuring whether a model can do something. It’s the wrong number for deciding whether to put it in front of customers. A support agent that resolves a refund correctly 60% of the time will get it wrong for a lot of real people, even if it “passes” the test on a good run.

The evidence

reviewed 15 Sep 2026

Reliability drops fast with repetition. Yao et al. (2025) built τ-bench, which simulates conversations between users and agents that have tools and policies to follow. State-of-the-art function-calling agents like GPT-4o succeeded on fewer than 50% of tasks. Their pass^k metric showed consistency was worse: pass^8 fell below 25% in the retail domain.

Accuracy and reliability are different axes. Rabanser et al. (2026) proposed twelve reliability metrics across consistency, robustness, predictability, and safety. Across 15 models, recent capability gains produced only small improvements in reliability.

Use it

  1. Run each test case several times (at least three to five for LLM outputs) and report both the average success rate and the all-runs-pass rate.
  2. Separate random, occasional failures from consistent failures on specific cases. They need different fixes.
  3. Include paraphrased versions of the same task in your test set.
  4. Set reliability thresholds based on the cost of a failure, not on what the model can currently hit.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

Consistency matters most for agents and multi-step tasks. If one good answer is enough, or a person reviews every output, a strict repeat-success metric such as pass^k may ask for more than you need.

Origins

“Once Is Not Reliable” is this site’s name. The pass^k metric comes from Yao et al.’s τ-bench, and the broader framing draws on reliability engineering in safety-critical fields.

Sources

  1. [1]
    Rabanser, S., Kapoor, S., Kirgis, P., Liu, K., Utpala, S., & Narayanan, A. (2026). Towards a science of AI agent reliability.arXiv:2602.16666 Preprint
    Open ↗ (opens in a new tab)
  2. [2]
    Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2025). τ-bench: A benchmark for tool-agent-user interaction in real-world domains.International Conference on Learning Representations (ICLR 2025)
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). Once Is Not Reliable (v1.0). https://evalfieldguide.com/patterns/once-is-not-reliable

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 15 →