shiichan

Is Your AI Spend Actually Paying Off? OpenAI Proposes "Useful Intelligence per Dollar"

Hi everyone, it's me, Shiichan! Today isn't about a shiny new feature — it's about a question that matters just as much: is your AI spend actually paying off? It might sound dry, but it's a big deal.

OpenAI News openai.com

What was announced?

OpenAI's News published a piece written by CFO Sarah Friar about how companies can measure the value they're getting from their AI investment. She says the question she hears from CFOs everywhere is simple: how do we get more value from our AI spend?

Why it matters

For years, software success was measured by adoption: seats purchased, active users, licenses renewed. But Friar argues that understanding AI's value demands a more powerful measure, one that looks at the work actually accomplished.

Her proposed scorecard is called "Useful Intelligence per Dollar." It answers four key questions:

  • Is AI completing work that matters?
  • What does each successful task actually cost?
  • Can people depend on the result?
  • Does each AI dollar produce more value as usage grows?

What changes

The first question starts with the work itself: how many customer issues got resolved, how many code changes shipped, how many contracts got reviewed, how much time came back to people.

One example given is a finance team preparing for a forecast review — finding the latest forecast, moving data into Excel or Sheets, identifying changes, reconciling tabs, and rebuilding slides. Letting ChatGPT Work take on much of that process frees the team to focus on the real judgment calls: what changed, why, and what to do next.

The second question is about cost. A lower price per token doesn't automatically mean a lower cost per outcome — a cheaper model that needs more retries or more human review can end up costing more overall. What actually matters is the full cost of completing the work, divided by the number of tasks that met the quality bar.

Dive Deep

That cost question gets a concrete example in GPT‑5.6, the model family OpenAI released last week. It comes in three tiers:

  • Sol — the flagship, for the strongest reasoning
  • Terra — balances performance and cost
  • Luna — the fastest and most affordable

The idea is to pick Luna for fast, high-volume work, Terra for deeper tasks, and Sol when getting it right in fewer attempts matters most.

On the benchmark side, on the Artificial Analysis Coding Agent Index's DeepSWE v1.1 test for long-horizon engineering tasks, GPT‑5.6 Sol (with max reasoning) reached 72.7%, ahead of Claude Fable 5's 69.9%, at an estimated 36.2% lower API cost. Separately, GPT‑5.6 also set a new state of the art on that index while using 54% fewer output tokens than another leading model.

The third question is dependability. AI adoption tends to deepen in stages — first drafting, then reasoning across tools and data, and eventually taking action while people retain judgment and control. Friar suggests tracking three concrete outcomes:

  • Ready to use — met the quality bar as delivered
  • Needs correction — required another attempt or human edits
  • Needs escalation — a person had to step in and finish the work

Alongside that, organizations should define what data AI can access, what systems it can use or change, and when a person should review or approve an action. ChatGPT Work is built on the security, privacy, compliance, and workspace-management foundation of ChatGPT Enterprise to support exactly that.

The fourth question is about economics at scale: tracking the same workflow over time to see whether completed work grows faster than total cost while quality holds. Compute sits at the center of that — training compute builds future capability, while inference compute delivers useful work today.

Wrap-up

  • OpenAI CFO Sarah Friar proposes "Useful Intelligence per Dollar" as a new way to measure AI investment value.
  • The four angles to track are: work completed, real cost per task, dependability, and economics at scale.
  • GPT‑5.6, released last week, comes in three tiers (Sol, Terra, Luna); on DeepSWE v1.1 it hit 72.7% versus Claude Fable 5's 69.9%, at 36.2% lower estimated cost.
  • The core idea: judge AI spend by the full cost per successful task, not just the price per token.
  • A good read if you need to explain AI ROI to leadership or you're wrestling with how to justify AI spend.