Article
How to run a pragmatic clinical trial for ambient AI: the NEJM AI playbook
Traditional RCT design was not built for AI tools that change clinician behavior. NEJM AI's pragmatic trial framework offers a better evaluation approach.
- clinical trials
- ambient AI
- evidence generation
- research methods
- NEJM AI
- pragmatic trials
- study design
When a pharmaceutical company wants to know if a drug works, it runs a randomized controlled trial. The logic is elegant: randomize patients, control the exposure, measure the outcome. The design has known limitations — generalizability, blinding, selection — but its strengths are well understood and its outputs are interpretable.
Ambient AI tools do not fit this design. Not because the underlying question is less rigorous, but because the intervention behaves differently from a drug in ways that matter for evaluation. Ambient AI changes as it’s used. It affects the behavior of clinicians who know they’re using it. Its effects propagate through documentation into downstream clinical decisions. And the outcomes that matter most — documentation quality, physician burnout, care efficiency, patient outcomes — are heterogeneous, temporally dispersed, and confounded by the same operational variables that made the deployment complex to begin with.
The pragmatic trial framework that NEJM AI published in the first half of 2026 is the most useful published guidance on how to evaluate ambient AI tools in clinical deployment conditions. It doesn’t solve every problem, but it establishes a set of methodological standards that health systems should be demanding from vendors — and that vendors should be demonstrating before enterprise rollout claims are credible.
Why traditional RCT design breaks down
The Hawthorne effect — the well-documented tendency of people to perform differently when they know they’re being observed — is particularly pronounced in ambient AI studies. Clinicians randomized to the intervention arm know they’re using the AI tool; they may document more carefully, speak more deliberately, or engage differently with the AI’s outputs. The control arm, knowing they don’t have access, may experience resentment or compensatory changes in their own behavior. Neither group is behaving as they would in routine deployment.
Blinding is effectively impossible. You cannot blind a clinician to the presence of an ambient microphone producing real-time transcription. Pre-post designs are a common workaround, but they are deeply confounded: the pre-period and post-period differ not just in the presence of the AI tool but in season, staffing, patient acuity, whatever else was happening in the health system, and the learning curve effects of the tool itself.
Crossover designs — in which the same clinicians experience both intervention and control periods — are the most promising traditional design for this setting, but they carry strong carryover effects. Documentation habits formed during the AI-assisted period don’t fully reverse. Clinicians who have experienced AI-assisted note completion have different expectations of documentation effort than they did before, and those expectations shape their behavior during the control period in ways that compress the measured effect size.
Outcome selection is the deepest problem. The outcomes that ambient AI vendors are most eager to claim — time savings, note quality, clinician satisfaction — are also the easiest to manipulate through study design. Time savings measured from login to sign-off include time the tool was not running. Note quality assessed by blinded clinician review is subject to rater variability and the specific rubric used. Satisfaction measured immediately post-deployment captures novelty effect, not steady-state experience. The outcomes that are hardest to manipulate — downstream patient outcomes, care process adherence, medication errors, unplanned readmissions — require longer follow-up and larger sample sizes than most vendor-sponsored studies are designed to deliver.
The NEJM AI pragmatic trial framework
The framework’s core insight is that ambient AI evaluation should be designed as a deployment study, not an efficacy study. The question is not “does this tool work under ideal conditions” — that question is easier to answer and less interesting. The question is “does this tool improve care delivery when implemented in a real health system with real operational complexity.”
The framework specifies several design requirements that distinguish pragmatic ambient AI trials from the typical vendor white paper.
Cluster-level randomization is preferred over individual-level randomization. Clinicians within a unit, clinic, or service line are randomized together, reducing spillover between arms and matching the natural unit of implementation (ambient AI is typically deployed to teams, not individuals). This requires larger total sample sizes but produces estimates that are more credible for enterprise deployment decisions.
Pre-registration of primary outcomes is required. The framework insists that the primary outcome — and the analysis plan for that outcome — be specified before data collection begins and registered in a public repository. This is standard for pharmaceutical trials and almost entirely absent from AI medical device evaluation. The absence of pre-registration is itself a quality signal: when a vendor cannot specify in advance what success looks like, the study is designed to find significance rather than test a hypothesis.
Instrumentation requirements specify what data must be captured during the deployment. At minimum: time-stamped usage logs (when was the ambient tool active, when were notes signed, how often were AI-generated sections accepted versus edited versus deleted), structured quality metrics on documentation (completeness, accuracy against structured data sources), and provider experience measures using validated instruments rather than ad hoc satisfaction surveys. The instrumentation burden is real but not unreasonable, and vendors who cannot provide usage telemetry to this specification are making credible evaluation impossible.
Active control is preferred over usual care. Where possible, the comparison condition should be a structured usual-care documentation workflow rather than unstructured clinical practice. This reduces variance in the control arm and makes the treatment contrast more interpretable.
Common design pitfalls
Beyond the framework’s explicit requirements, several pitfalls recur in the published ambient AI literature.
Selection of eager adopters for pilot studies is the most common source of inflated effect sizes. Clinicians who volunteer for an ambient AI pilot are systematically different from the full deployment population — they are more tech-comfortable, more likely to invest in learning the tool, and less likely to represent the physician who will be the marginal adopter in a mandatory rollout. Pilots that enroll volunteers produce effect sizes that do not generalize.
Outcome timeframe misalignment is underappreciated. Documentation time savings are measurable within weeks. Burnout reduction is measurable over months. Patient outcome effects may require years. Studies that report time savings and satisfaction after four weeks while claiming burnout reduction benefits are mismatching the timeframe of the outcome to the timeframe of the evaluation.
The note quality confound deserves specific attention. AI-generated notes tend to be longer than human-generated notes in early deployment. Length is not quality. Evaluations that use length as a proxy for completeness will systematically overrate AI-assisted documentation, particularly in specialties where over-documentation is already a problem.
What health systems should require
Health systems deploying ambient AI at enterprise scale — hundreds of clinicians, significant IT investment, contractual commitments — should require vendors to provide or support a pragmatic trial design before rollout. This does not mean delaying deployment indefinitely for research. It means instrumenting the deployment to generate comparative evidence prospectively, committing to pre-registered evaluation against meaningful outcomes, and sharing the results regardless of direction.
Vendors who agree to this are confident in their tool’s performance under scrutiny. Vendors who cannot agree to this are asking health systems to take the risk of an enterprise deployment based on evidence that cannot be independently assessed.
The pragmatic trial framework does not guarantee certainty. It provides a structure for generating evidence that is credible enough to act on. In a market where the gap between ambient AI marketing claims and verifiable performance remains wide, that standard is overdue.