Article
AI and health equity in 2026: where algorithmic bias is showing up and what to do about it
Documented bias in clinical AI — from training data demographics to differential performance by race and language — requires systematic procurement responses.
- health equity
- algorithmic bias
- AI ethics
- procurement
- health disparities
- NIH
Algorithmic bias in clinical AI is not a theoretical concern. The documented cases are specific, the mechanisms are understood, and the consequences — differential diagnostic accuracy, resource misallocation, systematically worse outcomes for already-disadvantaged populations — are real. What is less developed is the organizational response: the procurement frameworks, monitoring systems, and institutional accountability structures that health systems need to identify and address bias in the tools they deploy.
The field has moved past the stage of awareness. In August 2026, the question is not whether algorithmic bias in clinical AI is a problem — it is whether health systems are building the organizational capability to detect it and respond to it.
Where bias is showing up: the documented patterns
Training data demographics are the primary driver of AI performance disparities. AI systems trained predominantly on data from academic medical center populations — typically whiter, more insured, and more English-speaking than the U.S. patient population — systematically underperform on patients who differ from that training distribution. This is not a hypothesis; it is a pattern documented across radiology AI, dermatology algorithms, sepsis prediction models, and clinical NLP tools.
The mechanisms vary. Dermatology AI trained predominantly on lighter-skinned patients shows significantly lower sensitivity for skin findings on darker skin tones — a performance disparity with direct diagnostic consequences. Sepsis prediction models calibrated on insured hospital populations have miscalibrated risk scores for uninsured patients who have different baseline care-seeking patterns and comorbidity profiles. NLP tools that extract structured data from clinical notes perform worse on notes written about non-English-speaking patients, where the documentation is often sparser, less complete, and contains translated phrasing that NLP models do not handle well.
Proxy variable bias is a subtler and in some ways more pernicious problem. Several widely deployed resource-allocation algorithms — the most documented being a commercial algorithm used by major health systems to identify high-risk patients for care management — use healthcare utilization and cost as proxies for health need. Because Black patients historically have received less healthcare for the same health needs, cost-based proxies systematically underestimate their health needs and route fewer resources toward them. The algorithm was not explicitly using race — it was using cost — but the effect was a significant racial disparity in care management enrollment.
Insurance status bias operates similarly. Clinical prediction tools trained on commercially insured populations may have different calibration in Medicaid or uninsured populations, because the diagnostic and treatment patterns in the training data reflect care access patterns as much as clinical reality.
The HHS equity review framework
The Department of Health and Human Services has responded to documented bias with a structured equity review framework for AI tools used in CMS programs. The framework requires that AI tools deployed in Medicare and Medicaid programs undergo subgroup performance evaluation across race, ethnicity, language, disability status, and geographic categories before deployment, and that post-deployment monitoring include ongoing subgroup performance tracking.
This framework applies directly to vendors selling AI tools into Medicare Advantage plans, accountable care organizations, and other CMS program contexts. It creates a compliance floor — vendors that cannot demonstrate subgroup performance evaluation will face procurement barriers in government programs. For health systems procuring for CMS program contexts, the framework also provides a template for procurement standards that can be extended to all clinical AI procurement, not just CMS-covered tools.
The specifics of the HHS framework include requiring vendors to provide disaggregated performance metrics by demographic categories, documenting the demographic composition of training and validation datasets, and specifying what monitoring data will be collected in post-deployment surveillance. These are reasonable requirements that any serious clinical AI vendor should be able to meet.
What health systems can do in procurement
Procurement is the highest-leverage point for health systems that want to reduce algorithmic bias in their AI portfolio. The decisions made at contract execution — what evidence is required, what representations vendors must make, what monitoring obligations are built into the agreement — determine the equity profile of the deployed AI environment.
The specific requests that procurement teams should make: training dataset demographic composition disclosure (geographic origin, race/ethnicity distribution, insurance mix, language distribution); subgroup performance validation results for the patient populations the health system serves; documentation of what bias testing was performed and what mitigation steps were applied; and a commitment to post-market subgroup performance monitoring with data sharing to the health system.
Health systems with significant Medicaid, uninsured, or limited-English-proficiency patient populations should require validation performance on patient populations with those characteristics — not just aggregate performance metrics. A vendor that cannot provide subgroup performance data is implicitly telling you that they have not tested it, which is information about their equity commitment.
Contract terms matter. Building equity monitoring obligations into vendor agreements — requiring vendors to flag performance disparities in their deployed tools and notify the health system if post-market surveillance reveals differential performance — creates accountability that a one-time procurement review cannot.
Post-deployment monitoring for bias
Procurement standards are necessary but not sufficient. AI tools can exhibit bias at deployment that was not detected in pre-deployment validation, for several reasons: the health system’s patient population may differ from the validation population in ways that matter; algorithmic drift over time may affect subgroups differentially; and local workflow factors — documentation patterns, EHR configuration, ordering behavior — may interact with the algorithm in ways that produce disparate outcomes.
Health systems with mature AI governance have begun incorporating equity monitoring into their AI operations frameworks alongside standard performance monitoring. This means tracking disaggregated performance metrics by race, ethnicity, insurance status, and language for deployed AI tools, with defined thresholds that trigger review. A sepsis model that performs significantly worse for Spanish-speaking patients than for English-speaking patients at your institution is a performance problem and an equity problem, and the monitoring system should surface it.
The practical challenge is that health systems often lack the data infrastructure to do this well. Reliable race and ethnicity data, language preference data, and insurance status data need to be consistently structured and accessible for monitoring analysis. Many health systems have invested in equity dashboards for clinical outcomes; extending that infrastructure to AI performance monitoring is the next step.
The NIH training data diversity investment
The National Institutes of Health has responded to the training data diversity problem with a multi-year investment in building diverse, representative biomedical datasets that can serve as training and validation resources for clinical AI development. This includes the All of Us Research Program, which has enrolled over a million participants with demographic diversity that exceeds typical clinical research populations, and targeted investments in imaging datasets that include representative samples of underrepresented skin tones, body habitus, and imaging acquisition conditions.
The time horizon for these investments to produce meaningfully better-calibrated AI tools is years, not months. In the near term, health systems cannot wait for the training data problem to be solved by NIH — they need to use procurement standards and post-deployment monitoring to manage the bias risks that exist in the tools available today.
The equity implications of clinical AI are too significant to treat as a future problem. The tools are deployed now, they affect patient care now, and the organizational capability to manage their equity dimensions needs to be built now.