shiichan

OpenAI's o3 cracked 18 unsolved childhood diseases in a 376-case reanalysis!

Hey there, it's Shiichan! Today I found a really heartwarming story about AI making a difference on the front lines of medicine, so let me share it with you.

OpenAI News openai.com

What was announced?

OpenAI's News published a study called "Using AI to help physicians diagnose rare genetic diseases affecting children." It's a collaboration between OpenAI, the Manton Center for Orphan Disease Research at Boston Children's Hospital, and Harvard University. The results were published on June 18, 2026, in the medical journal NEJM AI.

In one line: the team used OpenAI's o3 Deep Research reasoning model to re-read 376 rare-disease cases that specialists had analyzed before but couldn't solve. The result, physicians were able to establish new diagnoses for 18 children.

Why it matters

Rare genetic diseases are, well, rare, and their symptoms are complicated. So even specialists sometimes can't reach the responsible gene, and many kids stay "undiagnosed" for years.

For families, no diagnosis means no treatment plan and no clear next step. That's why finding fresh clues in cases people had nearly given up on carries so much weight.

What changes

The beautiful part is that the AI doesn't make the diagnosis, it offers candidate explanations. For each case, o3 lays out evidence-linked candidate explanations, and human experts review them.

In other words, the AI isn't a replacement for doctors. It's a partner that quickly surfaces "this looks suspicious" from a huge, tangled pile of medical information. Humans still make the final call, so it's reassuring.

Dive Deep

Let's look a little deeper.

o3 received de-identified data packets containing Human Phenotype Ontology terms, clinician notes, patient metadata, and variant tables. The team calls this an "explanation-first reasoning layer", the design makes the model show why before it shows what.

The 376 cases split into four groups: 100 neurodevelopmental, 61 rare neuromuscular disease, 200 sudden unexpected death in pediatrics, and 15 early psychosis. The overall additional diagnostic yield was 4.8%, and by group it was 10.0% (neurodevelopmental), 6.6% (neuromuscular), 1.0% (sudden death), and 13.3% (early psychosis, though the sample is small). Of the 18 diagnoses, 10 were neurodevelopmental, 4 neuromuscular, 2 sudden death, and 2 early childhood psychosis.

Before going live, the team ran a warm-up on cases with known answers: 48 of 51 diverse cases, 45 of 57 neuromuscular cases, and all 15 long-read genome cases correctly pointed to the right gene.

Safety was taken seriously too. No protected health information (PHI) was ever sent outside, and every result passed through multiple stages before reaching families: expert consensus review, confirmation in a CLIA-certified lab, and clinical team approval.

I appreciate that the honest limits are written down too. The cases were already heavily pre-screened, so the added yield was modest, and 7 of the 18 were closer to rediscoveries than brand-new findings. The study also didn't measure how much extra work false positives create, or the clinician time and cost-effectiveness involved.

Wrap-up

Let me recap today's points.

  • In a collaboration between OpenAI, Boston Children's Hospital, and Harvard, o3 Deep Research reanalyzed 376 unsolved rare-disease cases.
  • 18 children received new diagnoses, an additional diagnostic yield of 4.8%.
  • The AI doesn't diagnose; it offers evidence-linked candidate explanations for humans to review.
  • Data was de-identified and safeguarded with multiple checks, including CLIA-lab confirmation.
  • At the same time, the team is honest about the limits: pre-screened cases, 7 rediscoveries, and cost not yet measured.

If you're into research or AI in medicine, or you think about how to combine AI with human judgment, I think this one is a wonderful model to learn from!