Article
The evidence base for AI-assisted radiology in 2026 — what CADe/CADt clearances since 2023 actually show
Large-vessel-occlusion triage, chest X-ray, mammography, and colonoscopy AI have generated the deepest peer-reviewed evidence base in clinical AI. Here is what the studies actually say.
- radiology
- imaging
- CADe
- CADt
- FDA
- clinical-ai
- evidence
- mammography
- colonoscopy
- stroke
Radiology remains the most mature slice of clinical AI, and it is the slice where the peer-reviewed evidence base is deepest. The FDA’s AI-enabled device list crossed 1,000 entries in 2026, and imaging is the plurality of that list by a wide margin. Under the “CADe” (computer-aided detection) and “CADt” (computer-aided triage) categories, dozens of products have been cleared since 2023 — and unlike many corners of clinical AI, several of these categories now have real, multi-site, prospective evidence to point at.
This piece walks through what the evidence actually shows across four categories where AI-assisted radiology has been most rigorously studied since 2023: large-vessel-occlusion (LVO) triage for stroke, chest X-ray CADe, screening mammography AI, and colonoscopy computer-aided polyp detection. The goal is not to catalogue every published study — the reviews cited below already do that — but to characterize what the aggregate signal is, where the disagreements are, and how a health system or a payer should read the evidence.
The two categories worth distinguishing
Before the studies, the categories.
CADe — computer-aided detection. The tool identifies possible findings on an image (nodules, polyps, calcifications, occlusions) and marks them for the reading clinician. The clinician remains the interpreter of record. CADe has existed in radiology since the 1990s in a mostly rule-based form; the deep-learning generation of CADe is what has been cleared at scale since roughly 2018 and expanded aggressively since 2023.
CADt — computer-aided triage and notification. The tool identifies studies that appear to contain a time-sensitive finding (e.g. a large-vessel occlusion, a pneumothorax, an intracranial hemorrhage) and re-orders the worklist so those studies are read first, or notifies the on-call team. CADt does not interpret the finding; it accelerates the human loop.
The two categories have different regulatory and evidence considerations. CADe is judged on sensitivity/specificity uplift and (increasingly) on reader-performance change in prospective trials. CADt is judged on time-to-diagnosis and downstream time-to-treatment, which are harder to run randomized studies for but easier to measure in real deployments.
The FDA’s glossary of these categories is captured in the Agency’s AI-enabled device list documentation. See also the short CADe and CADt glossary entries.
Large-vessel occlusion triage — the strongest CADt case
The most robust CADt evidence base is in large-vessel occlusion detection on CT angiography. Multiple products (Viz LVO, RapidAI, Aidoc, Avicenna) have been cleared and deployed at scale, and the outcomes literature has moved from single-site retrospective work to multi-site pre/post studies and, in a growing number of cases, randomized comparisons.
The signal is consistent. Time-from-scan-to-neurointerventional-notification drops by 20 to 40 minutes across multiple published deployments. In a stroke workflow where every minute of delay is measurable brain tissue, that is a clinically meaningful shift. The effect is largest in transfer-hospital settings — where a scan is acquired at a primary stroke center and read against an off-site specialist — and smallest at large academic comprehensive stroke centers where the human workflow was already tuned.
The American Heart Association’s 2024 guidance on AI-augmented stroke systems of care explicitly names LVO CADt as an intervention with sufficient evidence for adoption inside a well-governed stroke program, while cautioning that the workflow benefit does not survive if the notification infrastructure (paging systems, on-call rosters) does not keep up. That caveat is the key one for health-system adopters: the tool is only worth what the workflow around it can consume.
Where the evidence is thinner: outcomes beyond time-to-treatment. Whether the LVO CADt effect translates into a measurable modified-Rankin-Scale-at-90-days benefit at the system level is a harder question, and the answer is closer to “probably yes at some sites, unclear at others” than to a settled finding. The confounding is real — many of the sites deploying LVO CADt also expanded thrombectomy capacity, changed transfer protocols, and invested in stroke-team staffing.
Chest X-ray CADe — broad clearances, uneven deployment
Chest X-ray is the highest-volume imaging modality in most health systems, and it has become the busiest sub-category of FDA-cleared CADe. Dozens of products claim capability across pneumothorax, pulmonary nodule, consolidation, cardiomegaly, effusion, and rib-fracture detection.
The radiology society reviews summarize the aggregate signal: well-trained CADe models can reach or approach expert-radiologist sensitivity for several specific findings under laboratory conditions. In deployment, the pattern is more nuanced. Reader-augmentation benefits vary substantially by finding, by reader experience level, and by the specificity threshold the tool is deployed at. Sites that tune the tool aggressively for sensitivity generate meaningful alert-fatigue burden; sites that tune conservatively see less impact on reader miss rates.
Two categories within chest X-ray CADe have accumulated the strongest deployment evidence:
- Pneumothorax detection — the finding closest to a “time-critical, high-consequence, easily-missed” category. Multiple deployments report reductions in missed pneumothorax and improved worklist prioritization, though the underlying prevalence in typical ED workflows is low enough that per-scan lift is modest.
- Pulmonary nodule flagging — a bridge between CADe and lung-cancer-screening workflows. Meaningful sensitivity improvements at the reader level; the harder question is downstream (which flagged nodules trigger CT follow-up, Lung-RADS categorization, and biopsy).
Where the evidence remains genuinely mixed: impact on radiologist reading time at scale. Some sites report time savings; others report neutral or slightly negative effects once alert triage is included. The FDA’s post-market surveillance data, and the Joint Commission RUAIH surveillance stream now capturing organizational deployment reports, are the sources to watch here.
Screening mammography AI — the deepest randomized evidence
Screening mammography is where clinical AI now has its deepest randomized evidence base. The Swedish MASAI trial was the first large randomized comparison of AI-supported single-reader mammography against standard double-reader mammography; the interim MASAI results published in The Lancet Oncology reported that AI-supported single reading matched double reading on cancer detection while roughly halving reading workload. Additional randomized evidence has since been published from other European screening programs, generally in the same direction, and the European Society of Breast Imaging (EUSOBI) has issued updated guidance recognizing AI-assisted reading as an acceptable arm of a screening program under specified conditions.
Three things are worth naming about the mammography evidence:
- The workload-halving finding is the operationally important one. Screening mammography programs in Europe (double reading is the standard) face a chronic radiologist-capacity problem. AI-supported single reading is a plausible answer to it. In the U.S. (single reading is the standard), the same tools are being sold on a different value proposition: earlier detection or reader-augmentation, not workload reduction.
- Population-subgroup performance matters. The mammography-AI literature has more, and more explicit, subgroup analysis by breast density, age band, and prior-imaging availability than most clinical-AI evidence. The picture is not uniform — dense-breast performance remains the hardest sub-problem — but the transparency itself is a step forward.
- False-positive burden is the second-order question. Reader-support-mode AI in some deployments reduces callbacks; standalone-triage-mode AI in some pilots increases them. Which mode a program deploys — and which acceptance thresholds it uses — matters as much as which product it buys.
Colonoscopy computer-aided polyp detection — evidence for adenoma detection rate improvement
Endoscopy-adjacent AI has produced the highest per-procedure conversion signal in gastrointestinal medicine. Computer-aided polyp detection (CADe for colonoscopy) — products including GI Genius, Endoscreener, and several others — has been evaluated in multiple randomized controlled trials since 2020.
The consistent finding is that CADe raises the adenoma detection rate — the fraction of screening colonoscopies in which at least one adenoma is found — by an absolute 8-to-14 percentage points across most published RCTs. That is a large per-procedure effect for a category (colorectal cancer screening) where adenoma detection rate is one of the closest available proxies for downstream cancer-prevention benefit.
Two open questions remain in the colonoscopy evidence:
- Diminutive-adenoma detection. Much of the ADR uplift is driven by more small (under 5 mm) adenomas being flagged. Whether the additional small adenomas translate into measurable long-term cancer-prevention benefit or into over-diagnosis and unnecessary surveillance is genuinely contested in the gastroenterology literature.
- Endoscopist-experience-level heterogeneity. The benefit is largest for lower-ADR endoscopists and smaller for the highest-performing ones. That is intuitive but it also complicates the “adopt everywhere” case — the ROI depends on the baseline.
What the aggregate picture says for procurement
Reading across the four categories, three points stand out for health-system CMIOs and procurement teams:
- The evidence base is real, and it is unusually category-specific. Do not read a general “radiology AI works” or “radiology AI is hype” narrative into this space; each modality and each finding has its own evidence pattern.
- Workflow matters as much as model. LVO CADt without notification-infrastructure investment is a decorative alert. Mammography AI without reader-mode governance is a false-positive generator. Colonoscopy CADe without an endoscopist-training program is uneven in its effect.
- Regulatory posture is catching up. The FDA’s PCCP framework is the pathway through which most of these products will iterate. Health systems should evaluate the PCCP scope during procurement, not after go-live.
Sources cited
- FDA — Artificial Intelligence and Machine Learning (AI/ML)-Enabled Medical Devices
- American Heart Association / American Stroke Association — Stroke journal (AI-augmented stroke systems of care guidance)
- The Lancet Oncology — MASAI randomized trial of AI-supported mammography screening
- European Society of Breast Imaging (EUSOBI) — AI-assisted breast imaging recommendations
- Radiological Society of North America (RSNA) — CADe/CADt reviews and reader-augmentation studies
Related reading
- Medical imaging AI topic — the modality-level picture
- FDA & devices topic — the regulatory frame
- FDA AI-enabled device list crosses 1,000 — the underlying clearance denominator
- The FDA’s PCCP guidance in 2026 — how these products iterate post-market
- Glossary: CADe, CADt, 510(k), PCCP