Skip to content
AI in Healthcare

Article

AI safety incidents in healthcare — what RUAIH, MAUDE, and MedWatch are starting to show

The first year of RUAIH certification plus the growing MAUDE and MedWatch signal are starting to show how AI-enabled devices fail in production. What the data actually says.

By AI in Healthcare Editorial Updated
  • safety
  • RUAIH
  • MAUDE
  • MedWatch
  • governance
  • FDA
  • joint-commission
  • clinical-ai

For most of the last decade, the AI-safety-in-healthcare conversation was theoretical. We had case reports of specific failures — the sepsis-prediction tool with lower external-validation performance than the vendor claimed, the diabetic-retinopathy screening model that failed on a specific camera vendor, the ambient scribe that dropped a medication reconciliation — but we did not have a systematic reporting signal telling us how these systems fail in production, at what frequency, or with what pattern.

That is changing. Three separate reporting channels — the Joint Commission’s Responsible Use of AI in Healthcare (RUAIH) certification safety-and-effectiveness monitoring domain, the FDA’s MAUDE Medical Device Reports database, and the FDA’s MedWatch signal — are starting to accumulate enough AI-related reports that patterns are visible. This piece walks through what the data is starting to show, what the reporting infrastructure still misses, and what health systems and vendors should be doing about it.

The three channels

The three signals are complementary. Each captures a different slice of the failure spectrum.

  • MAUDE (Manufacturer and User Facility Device Experience) is the FDA’s public database of medical-device incident reports. Manufacturers, hospitals, and other reporters submit reports on device malfunctions, injuries, and deaths. For AI-enabled devices, MAUDE captures the tightest, most-documented reports — the ones tied to a specific cleared product with a specific 510(k), De Novo, or PMA number.
  • MedWatch is the FDA’s broader safety-reporting program covering drugs, biologics, and devices. It includes both formal MDR reports and voluntary signals from clinicians and patients. MedWatch captures reports on AI-enabled products, but also captures reports on AI-adjacent software (CDS tools, clinical LLMs operating under the Cures Act CDS exclusion) that MAUDE does not systematically cover.
  • RUAIH safety-and-effectiveness monitoring is the newest channel. As part of the Joint Commission’s RUAIH certification framework, certified health systems monitor their deployed AI for pre-specified performance and safety metrics and file incident reports through the Joint Commission’s Sentinel Event / Sentinel Event Alert infrastructure. This captures organizational-level signals — how a health system’s own AI deployment is performing — that FDA-focused reporting does not.

Sources:

What the data is starting to show

Two years into the systematic AI-in-healthcare surveillance signal, the patterns are still preliminary. But enough signal has accumulated that we can describe what the shape looks like.

Failure modes cluster in a small number of categories

  • Distribution shift. Model deployed on a patient population, scanner-vendor mix, or care-setting distribution meaningfully different from the training set. Most-cited MAUDE reports of imaging-AI performance degradation trace back here.
  • Workflow drift. The deployment context changes in a way the model was not built for. A CADt triage tool built for a 24/7 emergency-department workflow, deployed as an overnight-only tool at a smaller hospital, misses cases that arrive during the day. The model didn’t change; the deployment did.
  • Silent data-source changes. The upstream EHR reformats a data field, the scanner firmware updates, the lab changes reference ranges. The model keeps producing outputs, but on inputs that have subtly changed. Detection often lags weeks or months.
  • Automation bias. Clinicians defer to the model in cases where they should not have. This is the class of failure most likely to be captured in MedWatch (via a treating-clinician report) rather than in MAUDE (which requires a device-side incident narrative).
  • Prompt-injection / adversarial input. Newer, and particularly relevant to clinical LLMs and patient-facing chat systems. Documented in security-research reports; not yet showing up materially in MAUDE, but appearing in emerging Joint Commission safety alerts.

The signal is heavily biased toward imaging

Imaging AI is the largest slice of MAUDE-reported AI incidents, because imaging AI is the largest slice of cleared AI devices. This is a base-rate story, not a “imaging is riskier” story. Non-imaging AI-enabled devices are underrepresented in MAUDE partly because they are fewer in number and partly because their failure modes are less amenable to the MDR-report template.

Ambient scribes and LLM-based products underreport

A meaningful fraction of clinical LLM deployments — including many ambient scribes — operate under the Cures Act CDS exclusion and therefore do not have a device number, do not enter MAUDE, and do not have a formal MDR channel. Some vendors have voluntary reporting mechanisms; most do not. This is a real gap in the surveillance picture, and one the Joint Commission’s RUAIH framework is starting to fill in for RUAIH-certified deployments.

The RUAIH-certified cohort is starting to publish

Health systems in the initial RUAIH-certified cohort — many of them overlapping with the mature governance programs discussed in the health-system governance article — are starting to publish aggregated incident-signal reports through the Joint Commission’s Sentinel Event Alert stream. Early alerts have focused on drift-related performance degradation in sepsis-prediction tools and specific failure patterns in ambient-scribe medication reconciliation.

What health systems should be doing

For any health-system CMIO, safety officer, or AI-governance-committee member reading this, the surveillance-infrastructure checklist is:

  1. File MAUDE reports on cleared devices. Under 21 CFR Part 803, user facilities are required to report device-related deaths and serious injuries. AI-enabled device incidents are covered. If your incident-reporting process treats “the model made a mistake” as a workflow issue rather than a device event, you are underreporting.
  2. Use MedWatch for AI-adjacent-CDS reports. If a CDS tool operating under the Cures Act CDS exclusion contributed to an incident, MedWatch is the right channel. Filing is voluntary but the aggregate signal is what drives future FDA guidance.
  3. Connect internal monitoring to the Joint Commission Sentinel Event pathway. If your governance program has drift-detection instrumentation, connect the alert thresholds to your incident-reporting workflow so material signals get escalated.
  4. Publish, where possible. Aggregated, de-identified incident reports at the health-system level are one of the strongest ways to accelerate the field’s collective learning.

What vendors should be doing

For AI-enabled device vendors, the equivalent checklist is:

  1. Own the MAUDE-reporting duty for your cleared devices. 21 CFR Part 803 requires manufacturers to report events; underreporting is a compliance risk that has caught up with vendors in other device categories and will catch up here.
  2. Ship real drift-monitoring instrumentation. Not a dashboard the health system builds on top of your API — instrumentation your product ships with, that customers can inspect.
  3. Publish your PCCP transparency posture. Under the FDA’s PCCP framework, you have flexibility to update the model. Your customers need to know what you have changed and when. Silent updates are a governance red flag.
  4. Support RUAIH-aligned procurement conversations. Model cards, evaluation reports, monitoring plans, and incident-reporting playbooks are the artifacts your customers’ governance programs will ask for.

What to expect through 2027

Three predictions worth stating explicitly:

  • A small number of high-profile FDA safety communications on cleared AI devices. These will feel like a shock when they arrive; they will actually be a sign the system is working. Post-market surveillance is supposed to catch things.
  • RUAIH-certified health systems will publish an aggregated safety signal report — probably in mid-to-late 2027, likely in a peer-reviewed venue — that changes how the field talks about ambient-scribe and sepsis-prediction incident rates.
  • The Cures-Act-CDS-exclusion boundary will keep narrowing. As more incident data accumulates for LLM-based CDS operating outside the device framework, expect the FDA to tighten the interpretation of “independent review” further.

The AI-safety-in-healthcare conversation is finally becoming an empirical conversation. Health systems, vendors, and regulators that engage with it substantively will be much better positioned than those that treat “AI safety” as a compliance-check exercise.