# GPT-5.6 Sol Hits Up to 14x the Speed! OpenAI Previews New 'Ultrafast' Mode

Hi, it's me, Shiichan! Today's news is all about speed, and I can't help getting excited when a frontier model gets this fast.

## What was announced?

OpenAI News shared an early look at a new service tier called Ultrafast. It runs GPT-5.6 Sol up to 14x faster than Standard processing, launching first in the OpenAI API.

Under the hood, it's powered by Cerebras. Ultrafast generates up to 750 output tokens per second, bringing OpenAI's most intelligent model to products and workflows where every second counts.

## Why it matters

With GPT-5.6, OpenAI has been pushing the frontier of what its models can do while also making the whole stack more efficient. Those improvements have made advanced intelligence more affordable and more useful to more people.

Until now, though, getting real-time speed usually meant picking a smaller or more specialized model. Ultrafast opens up a different path: doing more useful work per second, without giving up intelligence.

When speed no longer trades off against intelligence, AI can move into the most time-sensitive parts of a business. OpenAI points to scenarios like:

- Incident response and reliability: analyze application logs, recent code changes, and engineer reports to find the likely cause and help prepare a fix while the outage is still unfolding
- Financial research and security: analyze market signals, assess transactions, and catch suspicious activity while conditions are still changing
- Customer support and voice: resolve complex, multi-step customer issues in real time without interrupting the conversation
- Commerce: answer product questions, check inventory, personalize recommendations, and resolve checkout issues while the shopper is still deciding, before hesitation becomes an abandoned cart
- Live research and experimentation: turn overnight research runs into an interactive session where you can test an idea, check the results, adjust, and run another experiment without breaking your flow

## What changes

During the preview, OpenAI is working with an initial group of companies to learn where this speed matters most. Inside OpenAI itself, a group of developers has been testing GPT-5.6 Sol on Ultrafast mode to see which workflows benefit from real-time, frontier-level intelligence.

For incident response, teams use it to read logs, analyze traces, synthesize conversations, and identify next checks — helping prepare or validate a fix in a fraction of the time, while engineers stay responsible for judgment and deployment.

For research, teams use it to rapidly search knowledge sources, query data, and gather and summarize information across tools. A common pattern used to be launching a batch of experiments overnight and reviewing results the next morning; with Ultrafast, that loop can tighten into several iterations during a single workday.

## Dive Deep

Ultrafast is the next step in OpenAI's partnership with Cerebras to bring ultra-low-latency inference to its platform. Running GPT-5.6 Sol on Ultrafast mode, Cerebras supports OpenAI's most intelligent model at up to 750 output tokens per second.

On availability:

- It's in limited preview today, for a select group of customers
- Early access starts with companies in coding, commerce, financial research, support, and other interactive applications
- OpenAI will expand access as capacity grows
- If your business needs frontier intelligence at the highest speed, you can [sign up to get notified](https://openai.com/form/ultrafast/) when access expands

## Wrap-up

- OpenAI announced a new service tier, Ultrafast, that runs GPT-5.6 Sol up to 14x faster than Standard.
- It's powered by Cerebras and can generate up to 750 output tokens per second.
- Target use cases include incident response, financial research, customer support, commerce, and research — anywhere speed is critical.
- OpenAI itself is already using it to speed up incident response and research workflows.
- It's in limited preview now, rolling out to select customers first — worth watching if you need frontier-level speed.
