Article
How health systems can scale AI beyond the pilot: lessons from the leaders
Why AI pilots succeed but enterprise rollouts fail — and what leading health systems do differently to move from proof-of-concept to production at scale.
- AI strategy
- health systems
- enterprise AI
- AI governance
- digital transformation
The graveyard of healthcare AI is littered with successful pilots. A radiology department runs a six-month proof-of-concept with a computer vision tool that flags critical findings. Sensitivity is excellent, radiologists are enthusiastic, and the vendor’s case study goes into the sales deck. Then the health system tries to expand to five more sites, and the thing quietly dies — integration complexity, inconsistent workflows, physician resistance in new contexts, a data governance question that nobody resolved.
This is the pattern. And it is not primarily a technology problem.
Why pilots succeed and rollouts fail
Pilots succeed because they are designed to succeed. A motivated champion, a controlled environment, hand-holding from the vendor, and a narrow success metric chosen before the pilot started. The organizational friction that will kill enterprise deployment — credentialing, IT security review, EHR workflow integration, change management across a diverse medical staff — is either absent or temporarily suspended.
Rollout fails because the organization discovers, at scale, that it never built the operating model to run AI as a system. Governance decisions get made by exception rather than policy. There is no clear answer to who owns model performance after go-live. Physician training is handled differently in every department. The data infrastructure that worked at one site has a different schema at another.
The health systems that have cracked this are not necessarily the ones with the biggest AI budgets. They are the ones that stopped treating AI deployment as a series of point solutions and started treating it as an operational function.
What the leaders do differently
HCA Healthcare’s approach is instructive. Rather than running an enterprise AI team that approves vendor pilots and then hands off to IT, HCA built what amounts to an AI operations function — a group responsible not just for evaluating tools but for the entire deployment lifecycle: integration, monitoring, retraining triggers, and deprecation. The function sits between the clinical informatics team and operations, reporting to a Chief Digital Officer with explicit authority to set standards that cross service lines.
Mass General Brigham took a different path but arrived at a similar structural answer. Their AI governance committee operates as a standing clinical body with the same authority as a pharmacy and therapeutics committee — meaning AI tools used at the clinical point of care require review before deployment, not just legal and IT sign-off. That shift in framing, from IT procurement to clinical governance, changes everything about how tools get evaluated and how performance gets monitored post-deployment.
The pattern across leading systems: governance is not a committee that approves pilots. It is a function that owns the lifecycle.
Governance structures that work
Effective AI governance in health systems has a few consistent features. First, a single accountable executive — typically a CDIO or Chief AI Officer — with cross-functional authority. Second, a tiered review process that matches scrutiny to clinical risk: administrative and revenue cycle tools follow a lighter path than tools that directly inform diagnostic or treatment decisions. Third, mandatory post-deployment surveillance with defined performance thresholds and a clear process for pulling a tool if those thresholds are breached.
What does not work is a committee of stakeholders who provide input, after which nobody is actually accountable for outcomes. This structure is extremely common and produces exactly the diffusion of responsibility that allows underperforming tools to linger in production.
AI operations as an organizational function
The concept of “AI Ops” in healthcare is still nascent, but its outline is becoming clear. Health systems that have scaled successfully have teams responsible for: maintaining a registry of deployed AI tools and their performance status; managing the vendor relationships and data-sharing agreements that underpin those tools; running or overseeing model performance monitoring; and coordinating with clinical informatics on workflow integration.
This is not a small team. Geisinger, which has emerged as one of the more operationally mature AI health systems, has built a function that spans data engineering, clinical informatics, and a dedicated group of AI implementation specialists who work site by site on change management. The headcount investment is real, and it precedes the value realization — which is one reason most health systems underinvest in it.
The data-infrastructure prerequisite
No governance structure can fix a broken data foundation. The most consistent predictor of AI deployment failure is the absence of a normalized, accessible data layer across the enterprise. This means not just a data warehouse but clinical data that has been harmonized across EHR instances, coded consistently, and made accessible in near-real time for inference pipelines.
Health systems running multiple EHR versions, or those that have grown through acquisition without data integration investment, face a structural disadvantage that AI vendor promises cannot paper over. The FHIR R4 mandate for payer-provider interoperability has helped at the edges, but internal enterprise data unification remains a health system responsibility.
The practical implication: before the next pilot, ask whether your data infrastructure can support production inference across your enterprise. If the answer is no, that problem deserves more investment than the next vendor contract.
What separates the next tier
The gap between leading health systems and the middle of the market is widening. As systems like HCA, Mass General Brigham, and Geisinger accumulate operational AI expertise and infrastructure, the cost of building those capabilities from scratch rises for everyone else. The organizations that will close the gap fastest are not necessarily the largest — they are the ones willing to treat AI deployment as an organizational transformation problem rather than an IT procurement problem.
That reframe, more than any specific tool or vendor relationship, is what distinguishes the leaders.