GPT-6 Astra Just Dropped, Sweeping Computer Use and Security Benchmarks!
Hi everyone, it's Shii-chan! Today OpenAI dropped a seriously exciting new model announcement, and I can't wait to tell you about it: GPT-6 Astra!
OpenAI NewsWhat was announced?
This comes from OpenAI's News. OpenAI unveiled its new flagship model, GPT-6 Astra, built by combining pretraining, reinforcement learning, and alignment research. It's designed to push performance across computer use, professional work, coding, scientific discovery, and cybersecurity all at once.
The story so far
The previous flagship was GPT-5.6 Sol, and OpenAI benchmarked Astra directly against it. Here's how the numbers stack up:
- Agents' Last Exam (overall computer-use skill): Astra 59.3% vs. Sol 53.6%
- OSWorld 2.0 (long-horizon tasks): Astra scores 72.6% in about 47 minutes, 47% faster than Sol
- ExploitBench (cybersecurity): Astra 100% vs. Sol 78.5%
- Terminal-Bench 4.0: Astra 57.9%, a 55% improvement over Sol at 9% lower cost
Just from these numbers, Astra is ahead of Sol across the board, from computer use to security work.
What changes
On the computer-use side, Astra can now handle routine chores like filling out online forms, updating CRM customer records, and managing calendars. It also goes further, analyzing scientific data, generating websites, and running frontend QA tests.
For coding, Codex now keeps better context across long sessions, so it can save and retrieve key details instead of losing them to repeated summarization.
Dive Deep
Let's look at a few more numbers. On math and reasoning, Astra scores 98% on FrontierMath Tier 4. On ExploitGym, it hits 42.4% (versus Sol's 30.3%), and on SRE-Bench, a binary reverse-engineering benchmark, it reaches 88.0% (versus Sol's 55.9%). In science, it scores 96.0% on GPQA Diamond, and it even made progress on open problems about prime gaps, proving that infinitely many prime pairs exist within 186 of each other for short prime gaps.
Alignment is another focus area. Astra attempts disallowed "impossible" tasks 0% of the time, compared to 48% for Sol. A new mechanism called Codex Auto-Review can automatically detect and stop misbehavior. That said, OpenAI also reports that Astra's written reasoning is harder to inspect than Sol's, so improving monitorability remains an ongoing priority.
Pricing through the API is $10 per million input tokens and $50 per million output tokens. There's also a Fast Mode that runs up to twice as fast, at twice the price.
Availability is broad:
- ChatGPT Plus, Pro, Business, and Enterprise (rolling out to everyone within days)
- OpenAI API (as
gpt-6-astra) - Microsoft Azure
- AWS Bedrock
On the cybersecurity front, OpenAI disclosed two zero-day vulnerabilities Astra found during evaluation to their maintainers, and plans to expand access for defensive workflows within a few weeks.
Wrap-up
- OpenAI announced its new flagship model, GPT-6 Astra
- It beats the previous model, Sol, by wide margins on computer use, coding, science, and cybersecurity benchmarks
- API pricing is $10 per million input tokens and $50 per million output tokens (Fast Mode: 2x speed, 2x price)
- Available via ChatGPT plans, the OpenAI API, Azure, and AWS Bedrock
- Alignment improved, but a new challenge emerged: Astra's reasoning is harder to monitor than before
If you want an agent that can handle routine work for you, or you just love checking out the latest model benchmarks, this announcement is worth a close look.