The Cost Frontier
“Accuracy without cost is half a result.”
- Last reviewed
- 15 Sep 2026
- Source types
- 3 peer-reviewed or classic
- Version
- v1.0 5 Oct 2026
The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status
In plain terms
You can buy accuracy with more compute, so a score without its cost is half a result. Compare systems at matched cost, or show the whole tradeoff.
Takeaways #
- You can buy accuracy with more compute: retries, voting, longer reasoning, bigger models.
- Comparing systems that cost very different amounts to run isn’t a fair comparison.
- Plot accuracy against cost and latency, and look at the frontier of best tradeoffs.
- Simple, cheap approaches sometimes match elaborate ones.
What it means #
A leaderboard that only shows accuracy invites a quiet arms race. Call the model five times and take a vote, and the score goes up. So does the bill. For a product team, a system that’s two points better and ten times more expensive might be the wrong choice. A fair comparison shows both numbers, so each reader can decide which tradeoff fits their situation.
The evidence #
reviewed 15 Sep 2026Accuracy-only evaluation produces bloated agents. Kapoor et al. (2025) found that a narrow focus on accuracy leads to needlessly complex and costly AI agents. They argue for jointly optimizing cost and accuracy, and show that doing so can cut cost substantially while keeping accuracy. They also found simple baseline strategies competitive with more complex agent designs on coding tasks.
Efficiency belongs in the scorecard. Liang et al. (2023) include efficiency as one of HELM’s seven core metrics.
Users care about more than accuracy. Ethayarajh and Jurafsky (2020) note that leaderboards ignore costs like model size and energy use that matter to real users.
Use it #
- Record cost per task (tokens or dollars) and latency next to every accuracy number.
- Compare systems at matched budgets, or show the full cost-accuracy curve.
- Always include a cheap baseline: a single call, a smaller model, a simple retry loop.
- Ask whether the extra points are worth the extra cost for your specific use case.
Questions to ask #
For vendor reviews, model cards, and launch reviews.
Where this doesn’t apply #
If cost truly does not matter, as with a rare, high-value task, putting accuracy first can be reasonable. The evidence on cost savings comes from agent benchmarks and may not carry over to every setting.
Origins #
“The Cost Frontier” is this site’s name. The idea of comparing options along a Pareto frontier comes from economics and engineering. Kapoor et al. (2025) made the case for it in AI agent evaluation.
Sources #
- [1]Ethayarajh, K., & Jurafsky, D. (2020). Utility is in the eye of the user: A critique of NLP leaderboards.Proceedings of EMNLP 2020Open ↗ (opens in a new tab)
- [2]Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2025). AI agents that matter.Transactions on Machine Learning Research (TMLR)Open ↗ (opens in a new tab)
- [3]Liang, P., et al. (2023). Holistic evaluation of language models.Transactions on Machine Learning Research (TMLR)Open ↗ (opens in a new tab)
Cite this pattern #
AI Evaluation Field Guide. (2026, October 5). The Cost Frontier (v1.0). https://evalfieldguide.com/patterns/the-cost-frontier@misc{lai-the-cost-frontier,
title = {The Cost Frontier},
author = {{AI Evaluation Field Guide}},
year = {2026},
month = oct,
note = {Version 1.0},
howpublished = {\url{https://evalfieldguide.com/patterns/the-cost-frontier}}
}https://evalfieldguide.com/patterns/the-cost-frontierRevision history #
- v1.05 Oct 2026Published.
Full changelog · Spot a mistake or a better source? Report it. Corrections are logged in the changelog.