Article
What mature health-system AI governance actually looks like — Mayo, Kaiser, Mass General Brigham
Three of the most-cited AI-governance programs in U.S. healthcare have published enough of their processes that patterns are visible. What mature looks like, and what most systems miss.
- governance
- health-system
- mayo
- kaiser
- mass-general-brigham
- clinical-ai
- RUAIH
- CMIO
“AI governance” is one of the most-invoked and least-specified concepts in healthcare AI in 2026. Every vendor deck mentions it. Every board-of-trustees meeting worries about it. Every RFP asks about it. But if you ask ten health-system CMIOs to describe their governance program, you get ten different answers — and if you compare those answers to what the leading academic centers have actually published, a pattern emerges.
This piece walks through what mature health-system AI governance actually looks like, using three of the most-documented programs — Mayo Clinic, Kaiser Permanente, and Mass General Brigham — and identifies the specific pieces most health systems are still missing.
The three reference programs
Mayo Clinic, Kaiser Permanente, and Mass General Brigham have each published enough of their AI-governance process — through peer-reviewed papers, conference talks, and NAM working-group participation — that outsiders can reconstruct the overall shape of what they do. The specifics differ; the common structure is remarkably consistent.
- Mayo Clinic’s AI Assurance Lab operates as an internal review body for any AI product Mayo intends to deploy. Its structure and process are documented in several published pieces including NEJM AI perspective pieces and Mayo’s own operational publications.
- Kaiser Permanente’s Center for Advanced Analytics functions as both an AI R&D center and a governance-and-deployment gate. Kaiser has published on its ambient-scribe evaluation process and its algorithmic-bias review methodology.
- Mass General Brigham’s AI Governance Committee was one of the first formal cross-hospital AI review bodies in the U.S. and has published its intake-and-review process for anyone wanting to submit a use case.
The common structure
Across all three programs, the intake-through-sunset flow follows roughly the same six stages:
1. Intake and use-case scoring
Anyone in the health system — a clinician, a department chief, a data-science team, a vendor pitching a pilot — can propose an AI use case. A small committee scores it against pre-published criteria: clinical risk (potential harm if the model is wrong), workflow disruption (how much the deployment changes clinician behavior), evidence quality (what supports the manufacturer’s claims), and strategic fit (does the system want to be doing this at all).
Not every proposal advances. All three programs publicly discuss saying “no” to use cases that look interesting but do not clear the intake criteria — often because the evidence base is too thin, or the risk-benefit calculation does not close.
2. Vendor and model due-diligence
For proposals that advance, the due-diligence phase covers:
- Security and privacy review (BAA, data-flow diagrams, PHI handling).
- Model-card review (what is the training data, how was it evaluated, what are the reported performance stratifications).
- PCCP review (what post-market modifications is the vendor committed to under the FDA’s PCCP framework, and what are the notification requirements).
- Dataset-representativeness assessment — critical, and where many vendor pitches fall down. If the model was trained on a population that does not resemble the health system’s patient mix, that gap needs to be understood before deployment.
- Independent performance replication where possible. Mayo’s AI Assurance Lab, in particular, has emphasized running its own retrospective evaluation on institutional data rather than accepting the vendor’s evaluation at face value.
3. Pilot with defined success criteria
Pilots are scoped with real KPIs — not “clinicians liked it” — and with a pre-specified stop condition. All three programs have publicly stopped pilots that failed to meet their pilot-phase success criteria. This is a genuinely differentiating practice: most healthcare-AI pilots run indefinitely and quietly transition to “we deployed it,” regardless of whether the pilot data supported the deployment.
4. Deployment with post-deployment surveillance
Deployment includes a monitoring plan with pre-specified performance metrics, thresholds, and drift-detection cadence. When the model’s performance drifts outside the envelope, a designated owner is paged. Reporting into MAUDE, MedWatch, and the Joint Commission RUAIH surveillance stream is part of the plan, not an afterthought.
5. Ongoing review and update handling
Any material model change — even a PCCP-covered update — is reviewed against the deployment’s monitoring data before it takes effect. Kaiser and Mass General Brigham have both publicly declined vendor updates that did not meet the health-system’s internal validation criteria, even when the update was FDA-cleared.
6. Sunset triggers
Explicit criteria for pulling the tool: sustained drift beyond threshold, safety-signal accumulation, evidence-base changes that undermine the deployment case, contract termination. All three programs have retired tools that no longer met the sunset criteria.
Sources cited
- Mayo Clinic — AI in Practice / AI Assurance Lab overview
- Kaiser Permanente — Advanced Analytics AI Governance
- Joint Commission — Responsible Use of AI in Healthcare (RUAIH) certification
- Coalition for Health AI (CHAI) — Assurance Framework
What most health systems are still missing
The Mayo/Kaiser/Mass-General model is not exotic. It is well-documented, publicly discussed, and widely acknowledged as best practice. And yet the majority of U.S. health systems in 2026 are missing pieces of it.
The most common gaps we see when talking to CMIOs and vendors:
- No formal intake. The health system has a “we should have an AI committee” story but no actual, publicly-known process by which someone submits a use case for review. Result: vendors bypass the committee by selling directly to service-line leaders.
- No independent evaluation. The health system accepts the vendor’s model-card and performance-evaluation without running its own retrospective on institutional data. Result: performance in production is often materially worse than the vendor claim.
- No monitoring plan at deployment. The tool is deployed; the monitoring plan is a future work item. Result: drift is not detected until a clinician escalates or an incident report surfaces.
- No sunset criteria. The tool is deployed without any pre-specified criteria for when it should be pulled. Result: underperforming tools accumulate; the governance program becomes a “yes” gate, not a lifecycle owner.
- No connection to safety-reporting infrastructure. Incidents are handled ad-hoc rather than being routed through MAUDE, MedWatch, and internal safety-reporting systems.
Any one of these gaps, on its own, is common. Two or three of them together is where the meaningful risk lives.
The RUAIH standardization signal
The Joint Commission’s RUAIH certification program, launched in 2026, is quickly becoming the de-facto framework health systems use to close these gaps — even if they do not formally certify. The five domains RUAIH assesses — governance, data management, risk-and-bias reduction, safety-and-effectiveness monitoring, and transparency-and-education — map cleanly onto the Mayo/Kaiser/Mass-General structure and give health systems a shared vocabulary and audit surface.
For AI vendors, expect procurement conversations in 2026–2027 to increasingly reference RUAIH: “how does your product plug into a RUAIH-certified governance program?” is a legitimate question that a vendor should be able to answer concretely.
What to do if you are starting from behind
For health systems that recognize themselves in the “missing pieces” description above, the path forward is not exotic:
- Publish the intake process. Anyone in the system should be able to find how to propose an AI use case.
- Staff a small standing committee with real authority to say no. Two clinicians, one data-science lead, one privacy officer, one general counsel is a functioning core.
- Adopt RUAIH’s five domains as the internal-review structure, even if you do not certify.
- Run at least one retrospective replication of a vendor evaluation on your own data before your next deployment. It will surface issues.
- Require monitoring plans and sunset criteria as a condition of deployment, not as future work.
The Mayo/Kaiser/Mass-General programs have decade-long head starts and orders-of-magnitude larger data-science staffs. That is not a reason to defer governance; the underlying process does not require a hundred data scientists. It requires a clear intake, an honest evaluation loop, and the willingness to say no.
Related reading
- Health-system copilots topic — the topic hub
- Joint Commission RUAIH certification — governance-framework news
- Glossary: RUAIH, Model Card, Drift
- FDA PCCP guidance in 2026 — the regulatory update layer