shiichan

Overnight and Unstoppable: How Cognition Saw Claude Fable 5's Real Power

Hey there, it's Shiichan! Today I found a pretty surprising story about an agent working through the entire night. Just imagining waking up to finished work gets me excited.

Claude Blog claude.com

What was announced?

In a "Working at the frontier" post on the Claude Blog, Cognition, the team behind the autonomous software engineering AI Devin, shared how they're using Claude Fable 5. The post features Silas Alberti, Cognition's SVP of Research.

According to Alberti, Fable 5 stays clear-headed in complex contexts, properly reaches for debugging tools, and can keep working continuously. He describes leaving a task overnight: "I wake up, and it's been working for eight hours straight and actually making real progress."

The story so far

Before Fable, you could only delegate agents that stayed on-task for "a couple of minutes, maybe an hour." Past that point, sessions would start to drift away from the goal. Trusting an agent with long, unsupervised autonomous work just wasn't realistic yet.

What changes

On Cognition's own "Frontier Code" benchmark, their internal anti-slop standard, the hardest subset told a clear story: prior Opus scored around 10%, while Claude Fable 5 reached about 30%.

That means handing Devin a big task and stepping away for a while is becoming much more practical. Cognition expects that within the next one to two years, 90% of agent sessions will be proactive: automatically spotting problems, analyzing the codebase, and arriving at solutions before anyone even asks.

Wrap-up

  • Devin, Cognition's coding agent, handles long autonomous work more reliably now that it runs on Claude Fable 5
  • Earlier models tended to drift after minutes to an hour; Fable 5 pushed forward through 8 straight hours overnight
  • On the hardest subset of Cognition's "Frontier Code" benchmark, scores jumped from about 10% to about 30%
  • Cognition expects 90% of agent sessions to be proactive within 1-2 years

Worth a close read if you're curious how far long-running AI agent reliability has actually come.