OpenAI Fires Up Its In-House Chip Jalapeño — Inside the Full Compute Strategy
Hey everyone, it's Shii-chan! Today's story is about how OpenAI thinks about its computing stack from the ground up, and there's a shiny new chip in the mix.
OpenAI NewsWhat was announced?
This is from OpenAI's News. OpenAI's CFO, Sarah Friar, lays out the company's compute strategy: one integrated system spanning data centers and chips, frontier models, the developer platform, consumer and enterprise products, and AI-native devices, where each layer strengthens the next.
Alongside the piece, OpenAI shared the first measured performance results for Jalapeño, its first custom-designed inference chip. On InferenceX, a public benchmark using GPT-OSS 120B, Jalapeño delivered more peak throughput per kilowatt and lower token latency than the commercial systems it was compared against. It also performed well on DeepSeek R1 and Kimi K2, showing the gains hold across different model families.
Why it matters
Up to now, OpenAI has run its models mostly on compute from outside partners, like Microsoft's cloud and NVIDIA's chips, and that foundation still matters a lot for OpenAI's growth. But when you can co-design the model, serving software, chip, memory, and network together, you get to optimize throughput, latency, energy efficiency, and cost as a single system. Jalapeño matters because it's a real, measured proof point that OpenAI can now do that in-house, not just rely on partners.
What changes
OpenAI is actively mixing and matching partners depending on what each workload needs. Its current portfolio includes:
- Microsoft and NVIDIA — the foundation that has driven OpenAI's growth so far
- AWS, AMD, Broadcom, Cerebras, CoreWeave, and Oracle — partners with different strengths in cloud infrastructure and low-latency inference
- SB Energy and SoftBank — partners for data-center development and energy supply
The approach is to use premium systems where raw capability matters most, and efficiency-optimized systems where scale and cost matter more. An in-house chip like Jalapeño adds one more option to that mix.
On the data-center side, the piece points to Project Camellia in Georgia as an example of designing facilities around customer workloads while also creating local jobs, supporting local businesses, using a closed-loop system to conserve water, and subjecting its commitments to an annual independent public audit.
Dive Deep
The piece also connects efficiency gains to real economic value. Better models reach correct answers in fewer attempts, smarter routing and context management cut wasted computation, and optimized software paired with purpose-built hardware improves speed and energy efficiency.
As a concrete example, on the Artificial Analysis Coding Agent Index, GPT-5.6 Sol running with max reasoning hit a new high score while using 54% fewer output tokens than another leading model. Fewer tokens means faster results, fewer retries, agents that can complete longer workflows, and lower cost per successful task.
The piece closes on Jevons paradox: as efficiency improves, more use cases become economically worthwhile, which expands overall usage rather than shrinking it. Examples given include reviewing every contract, running live financial scenarios, and letting engineers test more ideas than before.
Wrap-up
- OpenAI's CFO frames its compute strategy as one system spanning chips through products
- Jalapeño, OpenAI's first custom inference chip, posted its first measured results, beating compared commercial systems on throughput per kilowatt and latency on GPT-OSS 120B
- The partner portfolio spans Microsoft, NVIDIA, AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy, and SoftBank
- GPT-5.6 Sol hit a new high on a coding agent benchmark while using 54% fewer output tokens
- Jevons paradox ties efficiency gains to expanding, not shrinking, overall AI usage
If you're curious about the business side of OpenAI's AI infrastructure strategy, this one's worth a read.