OpenAI’s latest reasoning model, o1-preview, just aced a high-stakes medical exam. Researchers at Harvard Medical School and Beth Israel Deaconess Medical Center pitted it against experienced emergency physicians using raw electronic health records from 76 real patients. The result? AI nailed exact or near-exact diagnoses 67% of the time at triage—beating two attending doctors’ 50-55% scores. By admission, AI hit 82%, matching humans more closely. This isn’t some polished benchmark. It’s messy ER data: vital signs, nurse notes, fragmented histories.
The study, published Thursday in Science, spanned five experiments. o1-preview crushed NEJM clinicopathological conferences, including the correct diagnosis in 78% of cases and topping physician baselines by wide margins. On management plans from Grey Matters cases, it scored a median 89%, outpacing doctors using conventional tools by 48 points. “The model outperformed our very large physician baseline,” said Arjun Manrai, assistant professor of biomedical informatics at Harvard and a co-author, during a press briefing covered by NPR.
But hold on. No one’s handing stethoscopes to robots yet. The tests fed AI text only—no images, no patient grimaces, no bedside whispers. Real ER chaos? Absent. “This is valuable… but it skips a central part of the job of ‘being a doctor,'” noted Dr. Ashwin Ramaswamy, a urology instructor at Mount Sinai, in Mashable. Patients don’t arrive pre-filtered. They arrive scared, slurring, deteriorating.
Authors hammered the point. “I don’t think our findings mean that AI replaces doctors,” Manrai told The Guardian, adding we’re seeing “a really profound change in technology that will reshape medicine.” Co-author Dr. Adam Rodman, a Beth Israel clinician, warned of hype. “I get a little bit queasy about how some of these results might be used,” he said in Vox, fearing AI-doctor startups might spin it to sideline physicians. Medicine demands clinical trials. Accountability. Oversight.
Diagnostic errors kill. They injure 795,000 Americans yearly, per society estimates. AI could slash that—if integrated right. In one case, o1 spotted lupus lung inflammation behind a clot that stumped humans. Yet errors lurk. o1’s slip-ups? Often omissions, missing the unseen. Real-world pilots lag. A Mass General Brigham study last month found AI chatbots botch initial differentials over 80% of the time, as reported by Euronews.
Adoption creeps in. Nearly 20% of U.S. doctors tap AI for diagnoses, says AMA data cited in The Guardian. In the UK, 16% use it daily. Patients? A West Health-Gallup poll shows 66 million Americans have quizzed AI on health—mostly to supplement, not supplant, visits (West Health). But 14% skipped a doctor based on bot advice. Risky.
And business eyes the prize. Hospital CEOs salivate over costs. Radiologists? On the chopping block, claims one public system head, fueling X debates. Yet liability gaps yawn wide. Who pays when AI errs? No framework exists. Rodman again: “There is not a formal framework right now for accountability.” Patients crave human hands for life-or-death calls.
Earlier wins hint at promise. Google’s mammography AI boosted sensitivity to 54%, detecting 25% more interval cancers in a 2026 Nature Cancer trial (Nature Cancer). Med-PaLM 2 hit 86% on MedQA benchmarks. But benchmarks ≠ bedsides. Stanford’s 2026 clinical AI review flags the gap: models shine on exams, falter amid uncertainty (Stanford Medicine).
So where next? Collaboration. Triadic care: doctor, patient, AI. o1 thinks cyclically—weighing, checking logic. Doctors bring context, empathy, escalation. “AI as an extension of a physician, not a replacement,” praises one commentator in Science News. Trials underway. Infrastructure investments pour in—NIH funds, hospital pilots.
Change accelerates. OpenAI’s o1-preview? Already outdated by o3 whispers. Exponential gains double capabilities quarterly, per benchmarks. Professions adapt or atrophy. Radiology routine reads speed up. Complex cases? Humans lead.
Regulation trails. FDA clears narrow tools; broad LLMs roam free. Equity worries mount—does AI falter on diverse data? Elderly prompts? Non-English charts? Unproven.
The verdict. AI outperforms on paper diagnostics. Boom. It augments, doesn’t erase, skilled hands. But deploy rashly? Disaster. Forward trials decide. Medicine evolves. Humans steer.


WebProNews is an iEntry Publication