OpenAI Pushes 'Useful Work Per Dollar' Framework for Enterprise AI Measurement
What's new in v3
New items this pass are all empty shells — titles and URLs with no extracted claims, stances, or quotes. They represent broad media amplification of the scorecard story (enterprise finance outlets, LinkedIn reposts) but introduce no substantive new voices or counter-arguments. IBM published a piece on measuring GenAI productivity that could represent an independent vendor perspective, but its content is unavailable. The core story — OpenAI's scorecard versus the Willison/Suresh incentive-alignment counter — is unchanged from the previous pass.
What
OpenAI published a four-dimension AI ROI scorecard — useful work, cost per successful task, dependability, and return on compute — promoted by CFO Sarah Friar and targeted at enterprise finance audiences. [1][2] A Cars24 case study supplies the concrete output numbers the scorecard is designed to track: 1 million monthly conversation minutes, 50% improvement in customer support resolution rates, and 80% reduction in workflow turnaround time. [3] Coverage has spread broadly across enterprise and finance outlets. [5][6][7][8] Against this, Simon Willison endorsing Nik Suresh argues that financial incentive alignment between vendor and customer executives structurally sustains unrealistic AI productivity claims and suppresses honest assessment. [4]
Why it matters
OpenAI's framework reframes AI evaluation around output value — a framing that favors agentic, higher-cost deployments where OpenAI competes, and one whose definitions OpenAI as the vendor controls. The Willison/Suresh counter identifies a structural problem: the same financial incentives driving enterprise AI adoption also make honest measurement institutionally difficult to sustain.
Open questions
Will independent analysts (Gartner, Forrester) or other vendors (IBM, Anthropic, Google) adopt, contest, or offer alternatives to OpenAI's four-dimension scorecard? [2][9]
How does OpenAI define 'useful work' and 'successful task' operationally, and who controls those definitions in enterprise deployments? [1][2]
Does the incentive structure Willison and Suresh describe — where contradicting AI productivity claims risks contract cancellations — make honest adoption of any measurement framework institutionally impossible? [4]
Are rival AI vendors (Anthropic, Google, Microsoft) proposing their own ROI measurement frameworks, or accepting OpenAI's framing by default?
Narrative
OpenAI published two pieces within three days in mid-July 2026, arguing that standard business metrics — cost per seat, token prices, headcount saved — do not capture what enterprises receive from AI deployments. The first, a general advisory, frames the right lens as 'useful work per dollar': measuring the value of outputs rather than the price of inputs. [1] The follow-up, authored by CFO Sarah Friar, formalizes this into a named scorecard with four dimensions: useful work (volume and quality of tasks completed), cost per successful task (spend normalized to completed work), dependability (consistency and error rate), and return on compute (value extracted per unit of compute invested). [2] Friar's byline is deliberate — the content is aimed at finance and executive audiences rather than technical buyers.
A companion case study on Cars24, an Indian used-car marketplace, provides the concrete numbers OpenAI is asking enterprises to track: over 1 million monthly conversation minutes handled by AI agents, a 50% increase in customer support resolution rates, 12% recovery of previously dropped seller leads, and an 80% reduction in turnaround time across key workflows. [3] ChatGPT Enterprise and Codex reached 85–90% daily active usage across approximately 600 employees, spreading from engineering into finance, legal, marketing, and operations. OpenAI presents this as evidence that AI can move from pilots to production when tied to specific business workflows — exactly the output-value story the scorecard is designed to justify.
Against this, Simon Willison endorsed a sharply critical account by Nik Suresh of what corporate AI measurement actually looks like in practice. [4] Suresh describes executives producing AI-centered technical strategies without ever having used an AI tool, engineers running AI rewrites of codebases solely to inflate token usage leaderboards, and a culture of silence around AI productivity claims. The silence, Suresh argues, is not accidental: when customer executives publicly claim 100x productivity gains, any vendor representative who disputes those claims risks being seen as undermining the customer's credibility and inviting contract cancellation. The dominant force sustaining unrealistic numbers is financial incentive alignment, not genuine belief.
Both OpenAI pieces are openly promotional and neither references independent measurement research or third-party audits. The scorecard has drawn broad media coverage across enterprise and finance outlets, but the coverage so far consists of amplification rather than independent analysis or counter-argument. [5][6][7][8] IBM has published its own piece on measuring GenAI productivity in the enterprise, which could represent an independent vendor perspective, but its substantive content has not been extracted. [9] Willison and Suresh's account, if accurate, describes an institutional environment where even a well-designed measurement framework faces a headwind: the organizations most eager to adopt output-value metrics are also the ones with the strongest incentives to report favorable results.
Timeline
- 2026-07-14: OpenAI publishes advisory arguing enterprises should measure AI investment by 'useful work per dollar' rather than traditional cost metrics. [1]
- 2026-07-16: OpenAI publishes Cars24 case study citing 1M monthly conversation minutes, 50% resolution rate improvement, and 80% workflow turnaround reduction. [3]
- 2026-07-17: OpenAI CFO Sarah Friar publishes a four-dimension AI ROI scorecard (useful work, cost per successful task, dependability, return on compute) aimed at enterprise CFOs. [2]
- 2026-07-19: Simon Willison endorses Nik Suresh's account arguing financial incentive alignment sustains unrealistic AI productivity claims and suppresses honest assessment. [4]
- 2026-07-22: Coverage of OpenAI's scorecard spreads across enterprise and finance outlets including AI Business, Yahoo Finance, and Pulse2; no outlet introduces independent counter-analysis. [5][6][7][8]
Perspectives
OpenAI / Sarah Friar
Traditional business metrics are inadequate for AI; the four-dimension output-value scorecard gives enterprise CFOs a vendor-defined framework to justify agentic AI spend.
Evolution: Consistent across both pieces; the July 17 scorecard formalizes the July 14 advisory framing, with Cars24 data as empirical support.
Simon Willison / Nik Suresh
Corporate AI measurement is distorted by financial incentives — inflated productivity claims persist because disputing them risks contract cancellations; the result is institutional silence, not honest evaluation.
Evolution: Consistent since introduced; directly contests the premise that enterprise AI ROI can be measured honestly under current incentive structures.
Tensions
- OpenAI argues enterprises should measure AI by output value using its scorecard; Willison and Suresh argue the incentive structure sustaining enterprise AI adoption makes honest measurement institutionally impossible. [2][4]
- OpenAI's framework measures ROI by output value, but no independent body has validated its definitions — OpenAI as the vendor controls what 'useful work' and 'successful task' mean in its own scorecard. [2][1]
Status: active but slowing
Sources
- [1] How to manage AI investments in the agentic era — OpenAI Blog (2026-07-14)
- [2] A scorecard for the AI age — OpenAI Blog (2026-07-17)
- [3] How Cars24 scales conversations and builds faster with OpenAI — OpenAI Blog (2026-07-16)
- [4] AI Mania Is Eviscerating Global Decision-Making — Simon Willison (2026-07-19)
- [5] OpenAI Proposes 'Useful Intelligence Per Dollar' Scorecard For Enterprise AI Spending — reactive:openai-enterprise-ai-roi
- [6] OpenAI Urges Enterprises to Use Scorecard to Measure Worth of AI — reactive:openai-enterprise-ai-roi
- [7] OpenAI CFO details ‘scorecard’ for measuring AI’s worth — reactive:openai-enterprise-ai-roi
- [8] OpenAI sets AI value scorecard for corporate spend — reactive:openai-enterprise-ai-roi
- [9] Top 5 Tips for Measuring Productivity of Gen AI in the Enterprise | IBM — reactive:openai-enterprise-ai-roi