AI Chemist GPT-5.4 Cracks a Stubborn Drug Reaction with TEMPO!
Hi everyone, it's Shii-chan! Today I'm excited to share a story where an AI actually made a real chemical reaction in a flask better.
OpenAI News
What was announced?
This one comes from OpenAI News. OpenAI teamed up with Molecule.one, a drug-discovery chemistry company, and gave their model GPT-5.4 a wonderfully open-ended homework: pick one of several important reactions and make it better. Working with Molecule.one's Maria — an agentic chemistry AI wired to a high-throughput lab (Maria Lab) — GPT-5.4 generated research proposals, designed and ran experiments, analyzed the data, and even proposed follow-up experiments.
The target was Chan-Lam coupling, a reaction that forms carbon-nitrogen bonds and shows up a lot in medicines. GPT-5.4 decided on its own to focus on primary sulfonamides, a hard but high-value substrate class, and suggested that a mild oxidant like TEMPO might help. That hunch turned out to be spot on.
Why it matters
When you're hunting for small drug-like molecules, there's a wall: you can only test the molecules you can actually make. When a reaction gives low yields or too many byproducts, chemists may have to drop a promising molecule or design a whole new route. That makes synthesis a major bottleneck in drug discovery.
The sulfonamide group shows up across many kinds of medicine — anticancer drugs, antimicrobials, diuretics — yet the Chan-Lam coupling of primary sulfonamides with boronic acids has historically given low yields. Making it more reliable could give medicinal chemists a much broader, more practical set of molecules to explore.
OpenAI has already shown AI contributing to new results in math (the unit distance problem), theoretical physics (gluon amplitudes), and biology (lowering the cost of cell-free protein synthesis). This project extends that trajectory into medicinal chemistry — the same direction as GPT-Rosalind for life sciences.
What changes
The big point is that the AI read the literature, proposed a surprising hypothesis, helped design and analyze experiments, and reached a finding that human chemists could evaluate. And it wasn't just on paper — it worked with real molecules and instruments.
Under the optimized conditions, yields improved for 88% of the boronic acids and 83% of the sulfonamides tested. The mean yield rose from 16.6% to 25.2%, and the share of reactions clearing 30% yield climbed from 15.6% to 37.5%. Beyond the microliter scale, human chemists reproduced representative reactions by hand at bench scale and saw higher yields for 11 of 14 substrate pairs, with a more than twofold increase in eight of them. That replication matters, because very small-scale experiments can sometimes produce artifacts that vanish at larger scale.
Dive Deep
Here's how the system worked. Prompts written for Maria AI were used with GPT-5.4 inside a harness to generate and rank thousands of research proposals. Human chemists reviewed the top-ranked subset, picked four for the lab, and Maria AI translated the chosen plans into detailed lab instructions, ran the high-throughput experiments, analyzed the raw data, and returned structured results to GPT-5.4.
One of those four proposals, OAI-M1-03, was the "use a mild oxidant like TEMPO" idea. Across two cycles, Maria ran a total of 10,080 reactions — more than a chemist running three reactions a day would do in a decade. That scale let the system pick TEMPO out of ten oxidants tested, see the effect repeat across diverse combinations, and map its limits. In the second cycle it also found a nice bonus: TEMPO could be swapped for a much cheaper analog, 4-hydroxy-TEMPO, with little loss in performance.
The largest human correction was avoiding DMSO as a solvent, since chemists worried it could react with the stronger oxidants used as comparisons. The whole process took three months, from the first prompt on March 4th to sharing the OAI-M1-03 results with independent experts on June 4th. That's why OpenAI calls this near-autonomous, not fully autonomous — human chemists handled proposal selection, experimental corrections, prep of consumables and reagents, and hands-on replication throughout. Four external experts, including Tim Cernak, Associate Professor of Medicinal Chemistry at the University of Michigan, reviewed the preprint and judged the result novel and worth sharing.
The write-up is honest about the limits, too. This does not show that AI can run a chemistry research program end to end, and it doesn't guarantee the method generalizes to other coupling reactions, substrate classes, or manufacturing conditions. The yields came from a high-throughput platform, and bench validation covered just 14 representative substrate pairs. Mechanism work and independent reproduction are still homework.
There's serious attention to Preparedness. The work was deliberately scoped to a legitimate medicinal-chemistry problem — no toxins, chemical weapons, or harmful compound design. The model was evaluated under the Preparedness Framework and by the UK AI Security Institute, and was built to refuse requests aimed at harmful uses. Humans choosing which proposals enter the lab and holding the physical infrastructure add another layer of safety.
If you want the details, OpenAI has published the paper and a rewritten model chain-of-thought for OAI-M1-03.
Wrap-up
- OpenAI and Molecule.one used GPT-5.4 plus the autonomous Maria Lab to improve a tricky Chan-Lam coupling (primary sulfonamides)
- The key was a surprising suggestion to add TEMPO, found by running 10,080 reactions in total
- Mean yield 16.6% -> 25.2%, share above 30% yield 15.6% -> 37.5%, and bench replication improved 11 of 14 pairs
- Near-autonomous, not fully autonomous: the AI proposed hypotheses while humans handled selection, corrections, replication, and safety
- If you get excited about running the organic-chemistry research loop faster with AI, this one's for drug-discovery and synthesis folks