shiichan

GPT-6 Astra Reviewed 41 Financial Statements in Minutes—and Caught All 4 Planted Errors!

Hey there, it's Shii! Today I came across a pretty cool case study from the legal-tech world: an AI wrapping up a financial-document check at an incredible speed. Let me walk you through it!

OpenAI News openai.com

What was announced?

OpenAI News shared a case study of Legora, an agentic platform for legal work, using OpenAI's latest model, GPT-6 Astra, to handle a financial-statement review task.

Legora is an agentic operating system covering legal work end to end, from contract and agreement review to legal research. It's used by more than 100,000 professionals across more than 1,800 in-house legal departments and law firms in over 50 markets.

The workflow highlighted here is called "financial-statement tie-out": checking every figure in draft accounts against trial balances, a consolidation schedule, and the previous year's accounts until every item agrees.

The story so far

According to Legora Legal Engineer Percevale Perks, this reconciliation work "can take an entire evening, sometimes days." It's the kind of tedious, manual task where someone has to check every single figure by hand.

The more documents involved, the longer it takes, and the more room there is for something to slip through unnoticed.

What changes

Using GPT-6 Astra, Legora's Agent completed the tie-out across 41 documents in a single run, and it only took minutes.

The Agent checked every balance against its supporting schedule, surfaced breaks in the amounts, and recorded each check along the way. That leaves the legal professional with a granular record covering every line item and figure, which makes the follow-up review much easier.

Here's how Percevale Perks put it:

"I think what changed before and after is the processing power, the ability to ingest such a large number of documents, digest really complex information, and get all of those different line items and figures."

That said, the final call still belongs to a human expert. The Agent handles the exhaustive comparison, and the expert makes the judgment call on each result — that human-in-the-loop approach hasn't changed. It's also central to how Legora is extending its platform beyond legal work into audit, tax, compliance, and risk.

Dive Deep

Legora evaluated GPT-6 Astra using its own benchmark, the Legora Benchmark for Agentic Reasoning (BAR), which measures performance on end-to-end legal tasks drawn from real-world use cases.

On this specific financial-statement workflow, GPT-6 Astra improved performance by nearly 40% over the previous model. Across all tasks in the BAR, though, the average improvement was only about 3%, so this financial-review task seems to be a particularly good fit for the new model.

The accuracy numbers are striking too: GPT-6 Astra found all four errors Legora had planted in the accounts, including a £500,000 gap hidden in the revenue note. On top of retaining every check the previous model got right, it completed around 50 more checks as well.

Wrap-up

  • Legora used GPT-6 Astra to complete a financial-statement tie-out across 41 documents in a single run, taking only minutes
  • It found all four planted errors, including a £500,000 gap in the revenue note
  • On Legora's BAR benchmark, this workflow saw nearly 40% improvement over the previous model, versus about 3% averaged across all tasks
  • The final decision still rests with a human expert, under a human-in-the-loop approach
  • Legora is expanding beyond legal work into audit, tax, and compliance

If you're buried in financial-statement checks in legal or audit work, this is a case study worth paying attention to!