The Information Machine

AI Agents: 24x Token Growth Projections, Enterprise Cost Pressure, and the Agentic Business Thesis · history

Version 8

2026-07-04 18:26 UTC · 175 items

What

Goldman Sachs projects 24x AI agent token demand growth by 2030 [1], backed by production data showing Meta employees consuming 60+ trillion tokens in 30 days [3], OpenAI's Codex at 99.8% of internal output tokens [2], and Ethan Mollick documenting Opus 4.7 completing 2–17 weeks of engineering work in 14 hours at $251 [4]. Cost dynamics now pull in opposite directions: software optimization and Chinese model pricing (as low as $0.18/million tokens vs. a $4 frontier average [8]) compress costs downward, while Amazon raised GPU compute prices 20% in late June 2026 with the broader infrastructure stack repricing upward [9]. Enterprise adoption splits between a small population of high-intensity power users driving most lab revenue and roughly 60% of large companies slowing AI spending [10][11]. Palantir CEO Alex Karp adds a new named voice to the skeptic camp, arguing that token-based pricing is implicit evidence that AI vendors don't believe their own value claims [13].

Why it matters

The Amazon GPU repricing, if sustained across the infrastructure stack, puts direct pressure on the cost-compression case that underpins the bull thesis for AI agent economics. Karp's pricing-model argument — that a vendor confident in billion-dollar value creation would demand equity, not token fees — gives the skeptic position a specific structural form that the standard enterprise-adoption data does not.

Open questions

  • SemiAnalysis reports Anthropic gross margins above 70% [16] while separate reporting places the figure at approximately 44% [12] — which measurement reflects the current period, and what accounts for the discrepancy?

  • Amazon raised GPU compute prices 20% in late June 2026, with the infrastructure stack following [9] — does this represent a structural shift toward rising compute costs, or a one-time adjustment that will be absorbed by efficiency gains?

  • Open-source model processing on OpenRouter rose from 34% to 65% in five months [8] — is this rate of share gain still accelerating, and at what point does it materially constrain frontier lab pricing power?

  • Over 70% of OpenAI and Anthropic ARR is concentrated in coding [10] — is there evidence that adjacent domains (cybersecurity, tax, legal) are capturing share at scale, or is the pattern still largely confined to the original base?

Narrative

Goldman Sachs Research projects 24x AI agent token demand growth by 2030 [1], a forecast grounded in production-scale data. OpenAI's internal numbers show Codex rising from below 10% to approximately 99.8% of the company's internal output tokens in under a year, with non-developer individual agent use growing 137x and organizational use growing 189x since August 2025 [2]. Meta employees consumed over 60 trillion tokens in a single 30-day period, with one employee alone consuming 280 billion tokens at an implied cost of approximately $50,000 per year [3]. Ethan Mollick documents Opus 4.7 autonomously completing work representing 2–17 weeks of human engineering in 14 hours at $251 [4]. McKinsey projects AI agents could mediate $3–5 trillion of global retail commerce by 2030, requiring brands to expose machine-readable APIs to remain agent-discoverable [5].

Cost dynamics are moving in multiple directions simultaneously. SemiAnalysis reports effective Opus 4.7 costs near $0.99 per million tokens for optimized agentic workloads, because those workloads run at roughly 300:1 input-to-output ratios and achieve cache hit rates above 90% [6]. NVIDIA's Blackwell inference stack reduced DeepSeek V4 token costs by up to 5x within one month, with combined optimizations yielding up to 20x token throughput [7]. Against these efficiency gains, Citibank Research places Chinese AI model prices as low as $0.18 per million tokens versus a $4 average for frontier models, and open-source processing on OpenRouter grew from 34% in January 2026 to 65% in June 2026, driven significantly by Chinese model adoption [8]. Amazon raised GPU compute prices 20% in late June 2026, with the broader infrastructure stack repricing upward in the days following [9]. Gartner estimates AI coding costs will surpass the average developer salary by 2028, making cost control a growing strategic concern [8].

Enterprise adoption is split between a small number of high-intensity users and a broader population cutting back or spending modestly. SemiAnalysis argues that 90th-percentile customers account for most API revenue and are not cutting spend, that the median Fortune 500 employee spends under $100 per year on AI, and that pullbacks at Meta and Uber reflect poor internal incentive structures rather than a general trend [10]. UBS separately reports roughly 60% of large companies slowing AI spending [11]. The gap between individual Meta engineers running agentic loops implying compute costs equivalent to $10 million per year [3] and the median enterprise worker is the structural feature that makes both statistics simultaneously plausible. Over 70% of OpenAI and Anthropic ARR is concentrated in coding [10], and the broader adoption thesis depends on that pattern replicating in adjacent domains. OpenAI's adjusted gross margin fell to 33% in 2025 after inference costs quadrupled, recovering to 39% in Q1 2026 with a 52% year-end target [12].

Skepticism about AI value claims has a new named voice. Palantir CEO Alex Karp argues that token-based pricing is implicit evidence that AI vendors don't believe their own transformational claims: if a vendor genuinely believed it could generate a billion dollars in value for a customer, it would demand equity or revenue share rather than token fees [13]. Tesla has set an AI tool spending limit of approximately $800 per month per software engineer, with one observer arguing that spending above $200 per week is likely waste for most engineers [14]. These data points contrast with SemiAnalysis's argument that 90th-percentile enterprise users account for most lab revenue and are not pulling back [10]. Domain expansion beyond coding is underway but early: tax preparation, built on OpenAI's Codex, offers a concrete adjacent example, with the key architectural finding that the harness and evaluation system around the model — not the model itself — is the value-creating layer [15].

Timeline

  • 2026-05-30: Goldman Sachs 24x AI agent token forecast publicized; first reports that Microsoft and Uber find specific agent deployments more expensive than equivalent human workers. [1]
  • 2026-06-24: DeepMind integrates computer use into Gemini 3.5 Flash with adversarial prompt-injection training and confirmation gates for irreversible actions. [24]
  • 2026-06-25: OpenAI publishes internal agent usage data: Codex at 99.8% of output tokens; non-developer individual use up 137x and organizational use up 189x since August 2025. [2]
  • 2026-06-25: McKinsey projects AI agents will mediate $3–5 trillion of global retail commerce by 2030, requiring brands to expose machine-readable APIs. [5]
  • 2026-06-25: Notion announces Notion Mail shutdown on September 22, 2026, citing users' shift to AI agents for email. [25]
  • 2026-06-25: Multiple reports confirm Microsoft and Uber find AI agent deployments more expensive than equivalent human workers, leading to pullbacks. [20][21][22]
  • 2026-06-27: SemiAnalysis: Anthropic ARR grew from $9B to $44B+, gross margins above 70%; effective Opus 4.7 cost near $0.99/million tokens due to 300:1 input/output ratios and 90%+ cache hit rates. [16][6]
  • 2026-06-29: Scout launches: users specify a business KPI in plain English; system autonomously builds and tests agents, routing money- or integration-related changes to human approval. [26]
  • 2026-06-30: NVIDIA Blackwell inference stack cuts DeepSeek V4 token costs by 5x within one month; combined optimizations yield up to 20x token throughput on Blackwell. [7][19]
  • 2026-06-30: Mollick: Opus 4.7 autonomously builds software representing 2–17 weeks of human engineering work in 14 hours at $251. [4]
  • 2026-06-30: OpenAI adjusted gross margin fell to 33% in 2025 after inference costs quadrupled, recovering to 39% in Q1 2026 with 52% year-end target; Anthropic gross margin reported separately at approximately 44%. [12]
  • 2026-06-30: Citibank Research: Chinese AI models priced as low as $0.18/million tokens vs. $4 average for frontier models; open-source processing on OpenRouter grew from 34% in January to 65% in June 2026. [8]
  • 2026-06-30: SemiAnalysis: Meta/Uber AI pullback reports overstated; over 70% of OpenAI and Anthropic ARR from coding; median Fortune 500 employee spends under $100/year on AI. [10]
  • 2026-07-01: Meta employees consumed 60+ trillion tokens in 30 days; one employee alone consumed 280 billion tokens, implying approximately $50,000/year at average rates. [3]
  • 2026-07-01: UBS: roughly 60% of large companies slowing AI spending; Chinese firms account for 45%+ of OpenRouter traffic as of April 2026. [11]
  • 2026-07-01: Tax AI on OpenAI's Codex prepares returns for accountant review; the harness and evaluation system around the model — not the model itself — framed as the core product. [15]
  • 2026-07-03: Amazon reprices GPU compute +20%; broader infrastructure stack reprices upward in the days following. [9][23]
  • 2026-07-03: Palantir CEO Karp argues token-based pricing is implicit evidence that AI vendors don't believe their own value claims: genuine billion-dollar value creation would be priced as equity or revenue share. [13]
  • 2026-07-03: Tesla sets an AI tool spending cap of approximately $800/month per software engineer; a quoted observer argues spending above $200/week is likely waste for most engineers. [14]

Perspectives

Goldman Sachs Research

Projects 24x AI agent token demand growth by 2030; expects token cost declines to outpace price reductions, positioning cloud providers near a gross-margin turning point.

Evolution: Consistent — the originating analytical source for the thread's central forecast.

SemiAnalysis

AI labs now capture most value in the stack; effective Opus 4.7 costs near $0.99/million tokens for optimized workloads; the enterprise pullback narrative is overstated; over 70% of OpenAI and Anthropic ARR comes from coding, with cybersecurity as the projected next domain.

Evolution: Consistent; the primary counter-voice to enterprise cost-skeptic sources.

OpenAI

Internal data shows agents enabling longer and more complex tasks; adjusted gross margin fell to 33% in 2025 after inference costs quadrupled, recovering to 39% in Q1 2026 with a 52% year-end target.

Evolution: Consistent; the margin trajectory confirms cost pressure existed before the current recovery.

McKinsey

Projects AI agents will mediate $3–5 trillion of global retail commerce by 2030; argues brands must adopt machine-readable API infrastructure or be bypassed by AI purchasing agents.

Evolution: Consistent.

NVIDIA

Makes a technical and commercial case for its vertically integrated inference stack; Blackwell hardware plus software optimizations achieved 5x cost reduction on DeepSeek V4 within one month, with combined layers yielding up to 20x throughput.

Evolution: Consistent.

Enterprise cost-skeptics (Microsoft, Uber, UBS, Citibank, Palantir/Karp, Tesla)

Microsoft and Uber found specific deployments more expensive than equivalent human workers; UBS reports 60% of large companies slowing spend; Citibank documents a ~22x sticker-price gap between Chinese and frontier models; Karp argues token pricing reveals vendor doubt about their own value claims; Tesla caps per-engineer AI spend at $800/month.

Evolution: Karp's pricing-model argument and Tesla's spending cap add a philosophical and benchmarking dimension to what was previously an empirical cost argument.

Ethan Mollick (One Useful Thing)

Argues AI has crossed from chatbot co-intelligence to agentic autonomy; documents Opus 4.7 producing 2–17 weeks of human engineering work in 14 hours at $251; argues institutional planning horizons cannot track current capability curves.

Evolution: Consistent.

Enterprise power users (Meta consumption data)

A small number of high-intensity users — individual Meta engineers running agentic loops implying compute costs equivalent to $10 million per year — drive disproportionate AI token consumption; SemiAnalysis argues these users account for most lab ARR and are not cutting back.

Evolution: Consistent; Meta's 60+ trillion token figure provides the concrete scale behind the power-user concentration pattern.

Tensions

  • SemiAnalysis directly calls the Meta and Uber pullback reports overstated, attributing them to poor internal incentive structures; UBS reports 60% of large companies slowing spend — the two findings may describe different customer segments but neither directly addresses the other's data. [10][11]
  • SemiAnalysis reports Anthropic gross margins above 70%; separate reporting places the figure at approximately 44% — the discrepancy may reflect different metrics, time periods, or definitions, but is unresolved across available sources. [16][12]
  • Citibank Research puts Chinese model prices at $0.18/million versus $4 for frontier models, with open-source share on OpenRouter at 65% in June 2026; SemiAnalysis argues effective Opus 4.7 costs are near $0.99/million for optimized workloads — whether software optimization closes the gap enough to retain enterprise customers is unresolved. [8][6][11]
  • Over 70% of OpenAI and Anthropic ARR is concentrated in coding, while OpenAI, McKinsey, and others project broad agent adoption across legal, finance, and consumer retail — the broader thesis depends on pattern replication not yet observed at scale. [10][2][5]
  • NVIDIA and SemiAnalysis argue hardware and software optimization are driving sustained cost compression; Amazon's +20% GPU compute repricing, with broader infrastructure following, runs counter to that trajectory. [7][6][9][23]
  • Palantir CEO Karp argues token-based pricing is implicit evidence that vendors doubt their own transformational value claims; SemiAnalysis argues 90th-percentile enterprise customers account for most lab revenue and are not cutting back, implying the pricing model reflects market-clearing economics rather than vendor lack of confidence. [13][10]

Sources

  1. [1] Goldman Sachs: "Token use by AI agents is expected to multiply 24 times by 2030" — Rohan Paul Twitter (2026-05-30)
  2. [2] OpenAI just released a paper showing how they are now seeing the first version of office work where agents do most of th… — Rohan Paul Twitter (2026-06-25)
  3. [3] Meta employees used over 60 trillion tokens in 30 days, with one user alone consumed 280 billion. — Rohan Paul Twitter (2026-07-01)
  4. [4] The twilight of the chatbots — One Useful Thing (2026-06-30)
  5. [5] Mckinsey report - AI agents are quietly taking over the retail shopping cart and could mediate $3 Tn to $5 tn of global … — Rohan Paul Twitter (2026-06-25)
  6. [6] The substitution math is the part to internalize. Tasks that used to need a junior analyst for several hours, converting… — SemiAnalysis Twitter (2026-06-27)
  7. [7] How NVIDIA’s Inference Software Stack Powers the Lowest Token Cost — NVIDIA Blog (2026-06-30)
  8. [8] Reuters: Chinese models charge as little as 18 cents per million tokens versus $4 average for top models, says CitiBank … — Rohan Paul Twitter (2026-06-30)
  9. [9] Last week Amazon repriced GPU compute +20%. This week the rest of the infrastructure stack caught up. — reactive:ai-agent-economics-enterprise (2026-07-03)
  10. [10] TokenBudgeting: — SemiAnalysis Twitter (2026-06-30)
  11. [11] UBS says about 60% of big companies are slowing AI spending. — Rohan Paul Twitter (2026-07-01)
  12. [12] The Information reports that OpenAI has cut inference costs by more than half on some existing models, while logged-out … — Rohan Paul Twitter (2026-06-30)
  13. [13] "Are the prompts secure? Is this being transferred to you? If it was so valuable, if I can make you a bn dollars, wouldn… — Rohan Paul Twitter (2026-07-03)
  14. [14] The new AI budget benchmark for software engineers may have just landed at $800/month. — Rohan Paul Twitter (2026-07-03)
  15. [15] 😺 Watch: AI can do your taxes now — The Neuron (2026-07-01)
  16. [16] If you are an operator trying to write down what tokens will cost in 2027, the answer is materially lower than today, an… — SemiAnalysis Twitter (2026-06-27)
  17. [17] The throughput math has gotten the most pushback in our reader notes, so its worth being precise. On the same B300 runni… — SemiAnalysis Twitter (2026-06-27)
  18. [18] One of the more uncomfortable observations in our AI Value Capture piece is internal: our token spend at SemiAnalysis no… — SemiAnalysis Twitter (2026-06-27)
  19. [19] NVIDIA's newly published report says its Blackwell inference stack cut DeepSeek V4 token costs by up to 5x in one month.… — Rohan Paul Twitter (2026-06-30)
  20. [20] AI promised cost savings, but Microsoft and Uber say it’s costing more than human workers | Company Business News — reactive:ai-agent-economics-enterprise
  21. [21] Microsoft and Uber Pull Back on AI Subscriptions Due to Cost and ... — reactive:ai-agent-economics-enterprise
  22. [22] Uber, Microsoft, and Others Burning Through AI Budgets. Now What? — reactive:ai-agent-economics-enterprise
  23. [23] Last week Amazon repriced GPU compute +20%. This week the rest of the infrastructure stack caught up. — reactive:ai-agent-economics-enterprise (2026-07-03)
  24. [24] Introducing computer use in Gemini 3.5 Flash — DeepMind Blog (2026-06-24)
  25. [25] Notion killing Skiff-influenced email app since most users use AI agents instead — Ars Technica AI (2026-06-25)
  26. [26] AI agents to automatically improve business-critical KPIs. — Rohan Paul Twitter (2026-06-29)