Skip to content
AI in Healthcare

News

Harvard study in Science: OpenAI's o1 outdiagnoses emergency physicians on real ER cases

Researchers at Beth Israel Deaconess Medical Center and Harvard Medical School published results in Science showing that OpenAI's o1 model generated the correct diagnosis in 67.1% of real ER cases versus 55.3% and 50% for two attending physicians, and scored 89% on clinical vignette diagnosis versus 34% for human doctors. The study used 76 de-identified cases from Beth Israel's ED, with the AI working from the same messy EHR data available to clinicians, without access to imaging, physical examination findings, or vocal cues.

NPR By AI in Healthcare Editorial Source dated
  • research
  • clinical-AI
  • diagnostics
  • LLM
  • emergency-medicine

The most important methodological detail in this study is what the AI did not have access to. Most clinician-versus-AI comparisons that show AI performing well use curated, clean cases from standardized question banks. The Beth Israel study used 76 de-identified real ER cases — messy EHR notes, inconsistent data entry, incomplete documentation — and the AI worked from exactly the same data the physicians had, with no access to imaging, physical exam findings, or the vocal and behavioral cues that inform real clinical judgment. The performance gap (67.1% vs ~52% average across the two physicians) under those conditions is the part that matters.

The 89% versus 34% gap on clinical vignettes — structured written cases — is a larger magnitude than most prior studies, and reflects o1’s step-by-step reasoning architecture. Chain-of-thought reasoning on medical cases produces different outputs than next-token prediction on medical text: the model is working through differential diagnoses sequentially rather than pattern-matching to the most likely answer in the training distribution. That said, clinical vignettes are an easier target than real ER cases, and the 67.1% real-world number is the more meaningful figure.

Senior author Adam Rodman (Harvard/Beth Israel) was explicit that the results do not imply AI replaces emergency physicians — a position that is both accurate and slightly evasive of the harder downstream question. If o1 (a model now two generations behind o3, per the Reason magazine analysis) outperforms attending physicians on real ER diagnostic tasks, the question for health systems is not “should AI replace doctors?” but “what does the right human-AI division of diagnostic labor look like, and what liability framework governs it?” That question has no good answer yet.

The automation-bias risk runs in both directions in this context. Physicians handed an AI-generated differential may anchor on it in ways that reduce independent diagnostic reasoning. But physicians working without any AI assistance on high-volume ER shifts are already subject to cognitive overload, time pressure, and anchoring on the presenting complaint. The question is which set of cognitive biases is more dangerous in which clinical scenario — and that is an empirical question, not a rhetorical one.

Related coverage: medical imaging AI topic, FDA & devices topic.

Primary source: Read the full original on NPR ↗