The Information Machine

AI Agents: 24x Token Growth Projections, Enterprise Cost Pressure, and the Agentic Business Thesis · history

Version 6

2026-07-02 02:49 UTC · 128 items

What

Goldman Sachs projects 24x AI agent token growth by 2030 [1], a forecast now backed by concrete production data: Meta employees consumed over 60 trillion tokens in 30 days with one user alone consuming 280 billion [3], and Ethan Mollick documents Opus 4.7 autonomously producing 2–17 weeks of human engineering work in 14 hours at $251 in token costs [4]. Cost compression is accelerating from both directions: NVIDIA's Blackwell inference stack cut DeepSeek V4 token costs by 5x within one month [7], and SemiAnalysis reports effective Opus 4.7 costs near $0.99 per million tokens for optimized agentic workloads [6]. Enterprise adoption is split: UBS reports roughly 60% of large companies slowing AI spending [11] while SemiAnalysis directly calls the widely-reported Meta and Uber pullbacks overstated — driven by poor internal incentive structures rather than a structural ceiling — and notes the median Fortune 500 employee spends under $100 per year on AI, suggesting the enterprise adoption curve has barely begun [10].

Why it matters

Over 70% of OpenAI and Anthropic ARR is currently concentrated in coding [10], so the broader agent transition across legal, finance, and consumer commerce remains ahead — the business thesis depends on whether that concentration pattern replicates. Simultaneously, Chinese AI models priced up to 50x cheaper than U.S. counterparts now account for over 45% of OpenRouter traffic [11], introducing competitive pressure that cost-optimization arguments about U.S. frontier labs have not yet fully addressed.

Open questions

  • SemiAnalysis reported Anthropic gross margins above 70% [9] while separate reporting places the figure at approximately 44% [8] — which measurement reflects the current period, and what accounts for the discrepancy between these two sources?

  • SemiAnalysis predicts cybersecurity will be the next domain to replicate what coding has done to AI lab ARR [10] — is there observable evidence that transition has begun at scale?

  • UBS reports 60% of large companies slowing AI spending [11] while SemiAnalysis says 90th-percentile customers account for most API revenue and are not cutting back [10] — do these findings describe non-overlapping market segments, or is one of them wrong about the same population?

  • Chinese models are priced up to 50x cheaper per token and have grown to over 45% of OpenRouter traffic [11] — at what price point does the open-source Chinese ecosystem become the dominant cost-optimized choice for enterprise agentic workloads that do not require frontier reasoning?

Narrative

Goldman Sachs Research's 24x AI agent token growth forecast by 2030 [1] is grounded in a pattern already visible in production deployments. OpenAI's internal data shows Codex went from below 10% to approximately 99.8% of the company's internal output tokens in under a year, with non-developer agent use growing 137x for individuals and 189x for organizations since August 2025 [2]. Meta employees collectively consumed over 60 trillion tokens in a single 30-day period, with one employee alone consuming 280 billion tokens at an implied cost of approximately $50,000 per year [3]. Ethan Mollick reports Opus 4.7 autonomously built a software package representing 2–17 weeks of human engineering work in 14 hours at $251 in token costs, and notes that one quarter of OpenAI employees run four or more AI agents simultaneously at least once per week [4]. McKinsey projects a separate demand ceiling: AI agents could mediate between $3 trillion and $5 trillion of global retail commerce by 2030, requiring brands to expose machine-readable APIs to remain agent-discoverable [5].

The cost structure enabling this growth is compressing from multiple directions. SemiAnalysis reports effective Opus 4.7 costs near $0.99 per million tokens for real agentic workloads, because those workloads run at approximately 300:1 input-to-output ratios and achieve cache hit rates above 90% [6]. NVIDIA's Blackwell inference stack reduced DeepSeek V4 token costs by up to 5x within approximately one month of release; combined optimizations — disaggregated serving, large expert parallelism over NVLink, NVFP4 precision, and multi-token prediction — yield up to 20x token throughput on Blackwell [7]. OpenAI cut inference costs by more than half on some existing models; the company's adjusted gross margin fell to 33% in 2025 from 40% in 2024 after inference costs quadrupled, recovering to 39% in Q1 2026 with a 52% year-end target [8]. SemiAnalysis had separately reported Anthropic gross margins above 70% [9], while other reporting places Anthropic's gross margin at approximately 44% [8] — the discrepancy is unresolved across available sources.

The enterprise adoption picture is actively contested. SemiAnalysis directly argues that widely-reported budget pullbacks at Meta and Uber stem from poor internal incentive structures not found at most organizations, that 90th-percentile customers who account for most AI API revenue are not cutting spend, and that over 70% of OpenAI and Anthropic ARR is attributable to coding, with cybersecurity projected as the next domain to follow [10]. Against this, UBS reports that approximately 60% of large companies are slowing AI spending as executives add guardrails [11]. A separate cost pressure comes from Chinese AI models priced up to 50x cheaper than American counterparts on a per-token basis; Chinese firms accounted for over 45% of all traffic on the aggregation platform OpenRouter by April 2026, up from under 2% in late 2024, with U.S. export controls on chips having unintentionally accelerated the open Chinese AI ecosystem [11]. The gap between a small number of high-intensity power users — some individual Meta engineers running agentic loops that imply compute costs equivalent to $10 million per year — and the median Fortune 500 employee spending under $100 per year on AI [10][3] is the structural feature that reconciles growth projections with survey data showing broad enterprise caution.

On the product side, DeepMind integrated computer use into Gemini 3.5 Flash with adversarial training against prompt injection and confirmation gates for irreversible actions [12]. Notion announced it will shut down Notion Mail on September 22, 2026, citing that most users rely on AI agents for email rather than a dedicated client [13]. Scout, launched June 29, lets users specify a business KPI in plain English; the system autonomously builds and tests agents to achieve it, routing changes involving money or external integrations to human approval before deployment [14]. Mollick argues that organizations still operating from AI plans written before winter 2025 are describing a system far less capable than what currently exists [4].

Timeline

  • 2026-05-30: Goldman Sachs 24x AI agent token forecast publicized; first reports that Microsoft and Uber find agent deployments more expensive than equivalent human workers. [1]
  • 2026-06-24: DeepMind integrates computer use into Gemini 3.5 Flash with adversarial prompt-injection training and confirmation gates for irreversible actions. [12]
  • 2026-06-25: OpenAI publishes internal agent usage data: Codex at 99.8% of output tokens; non-developer individual use up 137x and organizational use up 189x since August 2025. [2]
  • 2026-06-25: Microsoft upgrades Copilot in Excel with external data connectors to FactSet, Morningstar, and others; SKILL.md-defined workflows and Plan mode with pre-edit audit trail. [21]
  • 2026-06-25: McKinsey projects AI agents will mediate $3–5 trillion of global retail commerce by 2030, requiring brands to expose machine-readable APIs. [5]
  • 2026-06-25: Notion announces Notion Mail shutdown on September 22, 2026, citing users' shift to AI agents for email. [13]
  • 2026-06-25: Multiple reports confirm Microsoft and Uber find AI agent deployments more expensive than equivalent human workers, leading to pullbacks. [18][19][20]
  • 2026-06-27: SemiAnalysis: Anthropic ARR grew from $9B to $44B+, gross margins from 38% to 70%+; effective Opus 4.7 cost near $0.99/million tokens due to 300:1 input/output ratios and 90%+ cache hit rates. [9][6]
  • 2026-06-27: SemiAnalysis: B300 GPU software optimization achieves 14x throughput on DeepSeek R1; internal token spend equals roughly 30% of employee compensation, over 5x Meta's per-employee average. [15][16]
  • 2026-06-29: Scout launches: users specify a business KPI in plain English; system autonomously builds and tests agents, routing money- or integration-related changes to human approval. [14]
  • 2026-06-30: NVIDIA Blackwell inference stack cuts DeepSeek V4 token costs by 5x within one month; combined optimizations yield up to 20x token throughput on Blackwell. [7][17]
  • 2026-06-30: Mollick: Opus 4.7 autonomously builds software representing 2–17 weeks of human engineering work in 14 hours at $251; one-quarter of OpenAI employees run 4+ agents simultaneously at least weekly. [4]
  • 2026-06-30: OpenAI cut inference costs by more than half on some models; adjusted gross margin fell to 33% in 2025 after inference costs quadrupled, recovering to 39% in Q1 2026 with 52% year-end target; Anthropic gross margin reported at approximately 44%. [8]
  • 2026-06-30: SemiAnalysis: Meta/Uber AI pullback reports overstated; over 70% of OpenAI and Anthropic ARR from coding; median Fortune 500 employee spends under $100/year on AI. [10]
  • 2026-07-01: Meta employees consumed 60+ trillion tokens in 30 days; one employee alone consumed 280 billion tokens, implying approximately $50,000/year at average rates. [3]
  • 2026-07-01: UBS: roughly 60% of large companies slowing AI spending; Chinese models priced up to 50x cheaper than U.S. counterparts; Chinese firms account for 45%+ of OpenRouter traffic as of April 2026, up from under 2% in late 2024. [11]

Perspectives

Goldman Sachs Research

Projects 24x AI agent token demand growth by 2030; expects token cost declines to outpace price reductions, positioning cloud providers near a gross-margin turning point.

Evolution: Consistent — the originating analytical source for the thread's central forecast.

SemiAnalysis

AI labs now capture most value in the stack; effective Opus 4.7 costs near $0.99/million tokens for optimized workloads; the enterprise pullback narrative at Meta and Uber is overstated, driven by poor internal incentives; over 70% of OpenAI and Anthropic ARR comes from coding, with cybersecurity projected as the next domain.

Evolution: Adds explicit rebuttal of enterprise pullback reports this pass; previously described cost structures without directly contesting the pullback narrative.

McKinsey

Projects AI agents will mediate $3–5 trillion of global retail commerce by 2030; argues brands must adopt machine-readable API infrastructure or be bypassed by AI purchasing agents.

Evolution: Consistent.

OpenAI

Internal research documents agents enabling longer and more complex tasks; Codex at 99.8% of internal output tokens; adjusted gross margin fell to 33% in 2025 after inference costs quadrupled, now recovering with a 52% year-end target.

Evolution: The June 30 margin disclosure adds a cost-pressure dimension: OpenAI's own economics were under strain before this quarter's recovery, complicating earlier advocacy framing.

NVIDIA

Makes a technical and commercial case for its vertically integrated inference software stack; Blackwell hardware plus software optimizations achieved 5x cost reduction on DeepSeek V4 within one month, with combined layers yielding up to 20x throughput.

Evolution: New voice in this thread; quantifies hardware-side contributions to cost compression that prior sources had attributed mainly to software.

Ethan Mollick (One Useful Thing)

Argues AI has crossed from chatbot co-intelligence to agentic autonomy; documents Opus 4.7 producing 2–17 weeks of human engineering work in 14 hours at $251; argues institutional planning horizons cannot track current capability curves.

Evolution: New voice in this thread; provides concrete per-task cost benchmarks and an institutional-adaptation framing absent from prior sources.

Enterprise cost-skeptics (Microsoft, Uber, UBS survey)

Microsoft and Uber found specific agent deployments more expensive than equivalent human workers; UBS separately reports roughly 60% of large companies are slowing AI spending as executives add guardrails.

Evolution: UBS survey data this pass broadens the cost-resistance story from two named companies to a majority of large enterprises.

Enterprise power users (Meta consumption data)

A small number of high-intensity users — individual Meta engineers running agentic loops implying compute costs equivalent to $10 million per year — drive disproportionate AI token consumption; SemiAnalysis argues these users are not cutting back and account for most lab ARR.

Evolution: Meta's 60+ trillion token consumption figure this pass adds concrete scale to the power-user concentration pattern that SemiAnalysis had described theoretically.

Tensions

  • Goldman Sachs, McKinsey, and SemiAnalysis project large-scale agent adoption and value capture by 2030; UBS reports roughly 60% of large companies are currently slowing AI spending, and Microsoft and Uber found specific deployments more expensive than equivalent human workers. [1][5][9][11][18][19]
  • SemiAnalysis directly calls the Meta and Uber pullback reports overstated, attributing them to poor internal incentive structures; UBS reports 60% of large companies slowing spend — the two may describe different customer segments (power users vs. median enterprise) but neither directly addresses the other's data. [10][11]
  • SemiAnalysis reported Anthropic gross margins above 70%; separate reporting places the figure at approximately 44% — the discrepancy may reflect different metrics, time periods, or definitions, but it is unresolved across available sources. [9][8]
  • Over 70% of OpenAI and Anthropic ARR is concentrated in coding, while the narrative from OpenAI, McKinsey, and others projects broad agent adoption across legal, finance, and consumer retail — the broader thesis depends on pattern replication that has not yet been observed at scale. [10][2][5]
  • Chinese AI models are priced up to 50x cheaper than U.S. counterparts and account for 45%+ of OpenRouter traffic; U.S. frontier labs argue software-layer and hardware optimization make effective costs far below sticker prices — whether the price gap narrows enough to retain enterprise market share is unresolved. [11][6][7]

Sources

  1. [1] Goldman Sachs: "Token use by AI agents is expected to multiply 24 times by 2030" — Rohan Paul Twitter (2026-05-30)
  2. [2] OpenAI just released a paper showing how they are now seeing the first version of office work where agents do most of th… — Rohan Paul Twitter (2026-06-25)
  3. [3] Meta employees used over 60 trillion tokens in 30 days, with one user alone consumed 280 billion. — Rohan Paul Twitter (2026-07-01)
  4. [4] The twilight of the chatbots — One Useful Thing (2026-06-30)
  5. [5] Mckinsey report - AI agents are quietly taking over the retail shopping cart and could mediate $3 Tn to $5 tn of global … — Rohan Paul Twitter (2026-06-25)
  6. [6] The substitution math is the part to internalize. Tasks that used to need a junior analyst for several hours, converting… — SemiAnalysis Twitter (2026-06-27)
  7. [7] How NVIDIA’s Inference Software Stack Powers the Lowest Token Cost — NVIDIA Blog (2026-06-30)
  8. [8] The Information reports that OpenAI has cut inference costs by more than half on some existing models, while logged-out … — Rohan Paul Twitter (2026-06-30)
  9. [9] If you are an operator trying to write down what tokens will cost in 2027, the answer is materially lower than today, an… — SemiAnalysis Twitter (2026-06-27)
  10. [10] TokenBudgeting: — SemiAnalysis Twitter (2026-06-30)
  11. [11] UBS says about 60% of big companies are slowing AI spending. — Rohan Paul Twitter (2026-07-01)
  12. [12] Introducing computer use in Gemini 3.5 Flash — DeepMind Blog (2026-06-24)
  13. [13] Notion killing Skiff-influenced email app since most users use AI agents instead — Ars Technica AI (2026-06-25)
  14. [14] AI agents to automatically improve business-critical KPIs. — Rohan Paul Twitter (2026-06-29)
  15. [15] The throughput math has gotten the most pushback in our reader notes, so its worth being precise. On the same B300 runni… — SemiAnalysis Twitter (2026-06-27)
  16. [16] One of the more uncomfortable observations in our AI Value Capture piece is internal: our token spend at SemiAnalysis no… — SemiAnalysis Twitter (2026-06-27)
  17. [17] NVIDIA's newly published report says its Blackwell inference stack cut DeepSeek V4 token costs by up to 5x in one month.… — Rohan Paul Twitter (2026-06-30)
  18. [18] AI promised cost savings, but Microsoft and Uber say it’s costing more than human workers | Company Business News — reactive:ai-agent-economics-enterprise
  19. [19] Microsoft and Uber Pull Back on AI Subscriptions Due to Cost and ... — reactive:ai-agent-economics-enterprise
  20. [20] Uber, Microsoft, and Others Burning Through AI Budgets. Now What? — reactive:ai-agent-economics-enterprise
  21. [21] Microsoft just turned Copilot in Excel into a finance workflow system — Rohan Paul Twitter (2026-06-25)