The Information Machine

OpenAI Pushes 'Useful Work Per Dollar' Framework for Enterprise AI Measurement · history

Version 2

2026-07-19 08:11 UTC · 30 items

What

OpenAI has published two advisory pieces arguing that standard business metrics miss the value of AI deployments, and CFO Sarah Friar formalized the argument into a four-dimension scorecard — useful work, cost per successful task, dependability, and return on compute — targeted at enterprise finance audiences. [1][2] A companion case study on Cars24 offers the kind of concrete output numbers the scorecard is designed to track: 1 million monthly conversation minutes, a 50% improvement in customer support resolution rates, and 80% reduction in workflow turnaround time. [3] A counter-narrative is now on record: Simon Willison, endorsing Nik Suresh's account, argues that financial incentive alignment between vendor and customer executives sustains unrealistic AI productivity claims and actively suppresses honest assessment. [4]

Why it matters

OpenAI's framework reframes AI evaluation around output value — a framing that favors agentic, higher-cost deployments where OpenAI competes, and one whose definitions OpenAI as the vendor controls. The Willison/Suresh counter points to a structural problem: the same financial incentives that drive enterprise AI adoption also make honest measurement institutionally difficult to sustain.

Open questions

  • Will independent analysts (Gartner, Forrester) adopt, contest, or offer alternatives to OpenAI's four-dimension scorecard? [2]

  • How does OpenAI define 'useful work' and 'successful task' operationally, and who controls those definitions in enterprise deployments? [1][2]

  • Does the incentive structure Willison and Suresh describe — where contradicting AI productivity claims risks contract cancellations — make honest adoption of any measurement framework institutionally impossible? [4]

  • Are rival AI vendors (Anthropic, Google, Microsoft) proposing their own ROI measurement frameworks, or accepting OpenAI's framing by default?

Narrative

OpenAI published two pieces within three days in mid-July 2026, making the case that standard business metrics — cost per seat, token prices, headcount saved — do not capture what enterprises receive from AI deployments. The first, a general advisory, frames the right lens as 'useful work per dollar': measuring the value of outputs rather than the price of inputs. [1] The follow-up, authored by CFO Sarah Friar, formalizes this into a named scorecard with four dimensions: useful work (volume and quality of tasks completed), cost per successful task (spend normalized to completed work), dependability (consistency and error rate), and return on compute (value extracted per unit of compute invested). [2] Friar's byline is deliberate — the content is aimed at finance and executive audiences rather than technical buyers.

A companion case study on Cars24, an Indian used-car marketplace, provides the concrete numbers OpenAI is asking enterprises to track: over 1 million monthly conversation minutes handled by AI agents, a 50% increase in customer support resolution rates, 12% recovery of previously dropped seller leads, and an 80% reduction in turnaround time across key workflows. [3] ChatGPT Enterprise and Codex reached 85–90% daily active usage across approximately 600 employees, spreading from engineering into finance, legal, marketing, and operations. [3] OpenAI presents this as evidence that AI can move from pilots to production when tied to specific business workflows — exactly the kind of output-value story the scorecard is designed to justify.

Against this, Simon Willison endorsed a sharply critical account by Nik Suresh of what corporate AI measurement actually looks like in practice. [4] Suresh describes executives producing AI-centered technical strategies for multi-billion-dollar organizations without ever having used an AI tool, engineers running AI rewrites of codebases solely to inflate token usage leaderboards, and a culture of silence around AI productivity claims. The silence, Suresh argues, is not accidental: when customer executives publicly claim 100x productivity gains, any vendor representative who disputes those claims risks being seen as undermining the customer's credibility and inviting contract cancellation. The dominant force sustaining unrealistic numbers is financial incentive alignment, not genuine belief.

Both OpenAI pieces are openly promotional, and neither references independent measurement research or third-party audits. Willison and Suresh's account, if accurate, describes an institutional environment where even a well-designed measurement framework faces a headwind: the organizations most eager to adopt output-value metrics are also the ones with the strongest incentives to report favorable results.

Timeline

  • 2026-07-14: OpenAI publishes advisory arguing enterprises should measure AI investment by 'useful work per dollar' rather than traditional cost metrics. [1]
  • 2026-07-16: OpenAI publishes Cars24 case study citing 1M monthly conversation minutes, 50% resolution rate improvement, and 80% workflow turnaround reduction as evidence of output-value measurement in practice. [3]
  • 2026-07-17: OpenAI CFO Sarah Friar publishes a four-dimension AI ROI scorecard (useful work, cost per successful task, dependability, return on compute) aimed at enterprise CFOs. [2]
  • 2026-07-19: Simon Willison endorses Nik Suresh's account arguing that financial incentive alignment between vendor and customer executives sustains unrealistic AI productivity claims and suppresses honest assessment. [4]

Perspectives

OpenAI / Sarah Friar

Traditional business metrics are inadequate for AI; proposes a vendor-defined four-dimension scorecard centered on output value and task completion to justify agentic AI spend.

Evolution: Consistent across both pieces; the July 17 scorecard formalizes the July 14 advisory framing. The Cars24 case study provides empirical support for the output-value argument.

Simon Willison / Nik Suresh

Corporate AI measurement is distorted by financial incentives — customer executives make inflated productivity claims, and anyone who disputes them risks contract cancellations; the result is institutional silence, not honest evaluation.

Evolution: New voice this pass; directly contests the premise that enterprise AI ROI can be measured honestly under current incentive structures.

Tensions

  • OpenAI argues enterprises should measure AI by output value using its scorecard; Willison and Suresh argue the incentive structure that sustains enterprise AI adoption makes honest measurement institutionally impossible. [2][4]
  • OpenAI's framework measures ROI by output value, but no independent body has validated the definitions — OpenAI, as the vendor, controls what 'useful work' and 'successful task' mean in its own scorecard. [2][1]

Sources

  1. [1] How to manage AI investments in the agentic era — OpenAI Blog (2026-07-14)
  2. [2] A scorecard for the AI age — OpenAI Blog (2026-07-17)
  3. [3] How Cars24 scales conversations and builds faster with OpenAI — OpenAI Blog (2026-07-16)
  4. [4] AI Mania Is Eviscerating Global Decision-Making — Simon Willison (2026-07-19)