4 of 5 in this categoryNext: 05 The Metric Mirage →
No. 04Measuringv1.0For: Buying · Designing

The Functionality Fallacy

“Don’t assume an AI system works. Whether it works is the first question, not a settled one.”

Last reviewed
15 Sep 2026
Sources
4 · Cite
Source types
3 peer-reviewed or classic, 1 report
Version
v1.0 5 Oct 2026

Established term: this name is used in the research literature. About this guide’s status

In plain terms

Before asking whether an AI system is fair or safe, ask whether it works at all. Plenty of deployed systems fail at their stated job, and a vendor’s claim is not evidence.

Takeaways

  • Debates about fairness, ethics, and safety often skip a more basic question: does the system do what it claims?
  • Plenty of deployed AI systems fail at their stated task. That failure is a harm on its own.
  • Vendor-reported performance is a claim. Independent validation on your own population is evidence.
  • Put “does it work, for whom, and how do we know?” at the top of every review.

What it means

Raji et al. (2022) point out that critics and policymakers often accept a vendor’s performance claims at face value and jump straight to downstream concerns. Meanwhile, some systems don’t work at all, some work only in narrow conditions, and some have capabilities that were overstated from the start. If nobody checks functionality, broken systems ship and nobody measures the damage.

The evidence

reviewed 15 Sep 2026

Failure is common and varied. Raji et al. (2022) catalog deployed AI failures, from tasks that can’t be done at all, to engineering mistakes, to failures that appear after deployment, to capabilities that were misrepresented.

A widely used model underperformed badly. Wong et al. (2021) externally validated the Epic Sepsis Model, which was in use at hundreds of US hospitals. Its area under the curve was 0.63, well below the 0.76 to 0.83 the developer reported. It missed 67% of patients who had sepsis while generating alerts for 18% of all hospitalized patients.

Practice lags need. Hutchinson et al. (2022) found that ML evaluation practice focuses on a narrow set of metrics and datasets that often don’t match what deployment contexts actually require.

Standards bodies agree. The NIST AI Risk Management Framework (2023) makes measuring whether a system is valid and reliable a core function, not an afterthought.

Use it

  1. Ask for evidence the system works on data like yours, gathered by someone other than the vendor.
  2. Ask what the system does when it’s wrong, and how often that happens.
  3. Define “works” in terms of the outcome you care about (patients treated, tickets resolved, time saved), not only model metrics.
  4. If nobody has checked, run a small pilot with a real comparison group before scaling.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

The pattern is about assuming a system works, not about every system failing. Many deployed systems do what they claim. The evidence shows that failure is common enough to check for, not that checking will always find a problem.

Origins

The term comes from Inioluwa Deborah Raji, I. Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst’s 2022 FAccT paper, “The Fallacy of AI Functionality.”

Sources

  1. [1]
    Hutchinson, B., et al. (2022). Evaluation gaps in machine learning practice.Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT 2022)
    Open ↗ (opens in a new tab)
  2. [2]
    National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1.U.S. Department of Commerce Report
    Open ↗ (opens in a new tab)
  3. [3]
    Raji, I. D., Kumar, I. E., Horowitz, A., & Selbst, A. (2022). The fallacy of AI functionality.Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT 2022)
    Open ↗ (opens in a new tab)
  4. [4]
    Wong, A., Otles, E., Donnelly, J. P., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients.JAMA Internal Medicine, 181(8), 1065-1070
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). The Functionality Fallacy (v1.0). https://evalfieldguide.com/patterns/the-functionality-fallacy

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 05 →