Briev
Live
Technology
UNDERREPORTED

AI Model Outperforms Doctors in Multiple Diagnostic Benchmarks

A study by researchers at Beth Israel Deaconess Medical Center and Harvard Medical School found that the OpenAI LLM o1 surpassed physicians in several diagnostic tests.

Researchers at Beth Israel Deaconess Medical Center and Harvard Medical School evaluated the OpenAI large language model o1 against practicing physicians in a series of diagnostic challenges. When asked to generate differential diagnoses from case studies, the AI included the correct answer in 78% of instances, compared with about 30% for doctors. In a set of five real-world clinical vignettes, o1 received an average score of 89%, while human clinicians scored 34% on the same material.

Additional testing on emergency-room scenarios showed the model identifying the exact or near-exact diagnosis in 67.1% of cases, outperforming two attending physicians who achieved 55.3% and 50% respectively. The results, appearing in Science, align with prior research indicating AI’s superiority in mammogram interpretation and early pancreatic cancer detection. While the authors stress that human clinicians remain the ultimate standard, they argue that AI could substantially improve diagnostic accuracy and patient outcomes. Meanwhile, states such as Nevada and Illinois have enacted laws restricting AI from providing mental-health services, potentially limiting its broader clinical deployment.

Why it matters

The study shows AI can diagnose more accurately than doctors, hinting at future improvements in patient care and medical efficiency.

In this story

AI diagnostic toolslarge language modelclinical reasoningmedical benchmarkingregulatory restrictionsearly disease detectionphysician performance