Article
FDA clearance ≠ clinical benefit: what the outcomes-evidence gap means for procurement
Most FDA-cleared clinical AI tools were never tested on patient outcomes. What that means for procurement teams and how to evaluate post-market evidence.
- FDA
- AI regulation
- procurement
- clinical evidence
- health systems
There is a version of due diligence that health system procurement teams perform on AI tools that goes roughly like this: Does the vendor have FDA clearance? Yes? Add it to the approved list. This approach is understandable — FDA clearance signals that a tool has cleared a meaningful regulatory bar, and procurement teams are not clinical researchers. But the approach has a critical flaw, and the evidence for that flaw is now well-documented.
Research published in the last two years, including a widely circulated Healio analysis of FDA-cleared AI diagnostic tools, found that a substantial majority of cleared devices had never been evaluated for patient outcome impact. They were cleared based on technical performance — sensitivity, specificity, AUC — measured against reference standards, often on curated datasets. Whether using the tool actually improved what happened to patients was, in most cases, simply not studied.
What FDA clearance actually certifies
To understand the gap, it helps to understand what the three FDA pathways for AI/ML-based Software as a Medical Device (SaMD) actually require.
The 510(k) pathway — the most common route for AI diagnostic tools — requires a manufacturer to demonstrate substantial equivalence to a predicate device. The evidentiary bar is analytical and technical performance, not clinical outcomes. A tool that detects a finding on a chest X-ray with high sensitivity compared to radiologist reads can be cleared through 510(k) without a single patient-level outcome being measured.
The De Novo pathway, used for novel device types without predicates, requires more rigorous performance evaluation but still typically stops at diagnostic accuracy rather than outcomes. Only the PMA (Premarket Approval) pathway — required for Class III, high-risk devices — mandates clinical studies with sufficient rigor to demonstrate clinical benefit, and relatively few AI tools have gone through PMA.
The FDA has been thoughtful about not imposing PMA-level requirements on every AI diagnostic aid, recognizing that doing so would suppress beneficial innovation. But the practical consequence is that “FDA-cleared” encompasses an enormous range of evidentiary standards, and conflating clearance with proven clinical benefit is a meaningful error for procurement teams.
The real question: does it change what happens to patients?
The alternative to checking the clearance box is harder but more meaningful: evaluating the post-market evidence trail for actual patient outcomes.
The first question is whether the vendor has conducted or sponsored any prospective or pragmatic clinical studies assessing outcome impact. Not accuracy studies — those tell you the tool can detect things. Outcome studies: does detection with this tool, integrated into this workflow, lead to faster treatment, fewer errors, shorter hospital stays, lower mortality? A vendor that has invested in outcomes research is signaling something important about their confidence in the product’s actual clinical value.
The second question is whether real-world performance data exists from health systems comparable to yours. Tools trained and validated on academic medical center data may perform very differently in community hospital settings, on patients with different demographic profiles, or with different EHR workflows. Asking vendors for post-market surveillance data from installed sites — not cherry-picked case studies but systematic performance monitoring data — is a reasonable demand.
The third question concerns how performance was measured during validation. Many AI tools report aggregate performance metrics that obscure differential performance across patient subgroups. A tool with 92% sensitivity overall may have 78% sensitivity in patients over 75, or in patients with comorbidities that shift image characteristics. Procurement teams should request subgroup performance breakdowns for populations that will actually use the tool in their system.
How to build an evidence-based procurement framework
Health systems building serious AI procurement capability are moving toward structured evidence review frameworks that go beyond regulatory status. The model looks more like a pharmacy and therapeutics committee than an IT procurement review — a standing group with clinical, informatics, and operational expertise that evaluates the evidence base for AI tools against defined standards before deployment.
Key elements of a functional framework include: a tiered evidence requirement matched to clinical risk (a documentation AI tool needs a lighter evidence review than a sepsis-prediction algorithm); a vendor evidence dossier requirement that goes beyond the FDA clearance certificate; and a post-deployment performance monitoring obligation built into the contract.
On the contract side, health systems with leverage are beginning to require performance warranties — provisions that entitle the health system to remedies if real-world performance falls below the thresholds specified in the vendor’s evidence dossier. This is not yet standard practice, but it is gaining traction among larger IDNs as a way to align vendor incentives with actual deployment performance.
The FDA’s evolving posture
It is worth noting that the FDA recognizes the outcomes-evidence gap and is working to address it through the Digital Health Center of Excellence and updated guidance on AI/ML-based SaMD. The agency’s post-market surveillance framework for AI tools — which requires manufacturers to monitor performance in real-world conditions and report significant performance changes — is a step toward better post-market evidence generation.
But this will take time to produce a meaningful body of outcomes evidence for cleared tools. In the interim, procurement teams cannot wait for the regulatory framework to close the gap. The due diligence responsibility sits with health systems, and the tools to do it — structured evidence review, vendor performance dossiers, post-deployment monitoring — are available now.
FDA clearance is a necessary condition for deploying a clinical AI tool. It is not a sufficient condition for believing the tool will benefit your patients.