AI Agents: 24x Token Growth Projections, Enterprise Cost Pressure, and the Agentic Business Thesis · history
Version 7
2026-07-03 08:27 UTC · 152 items
What
Goldman Sachs projects 24x AI agent token growth by 2030 [1], grounded in production data showing Meta employees consuming over 60 trillion tokens in 30 days [3] and Ethan Mollick documenting Opus 4.7 completing 2–17 weeks of engineering work in 14 hours at $251 [4]. Citibank Research puts Chinese AI model prices as low as $0.18 per million tokens versus a $4 average for frontier models, and open-source processing on OpenRouter grew from 34% in January 2026 to 65% in June 2026, driven significantly by Chinese model adoption [8]. Enterprise adoption splits between high-intensity power users driving most AI lab revenue and roughly 60% of large companies currently slowing AI spending [9][11]. Domain expansion beyond coding — which accounts for over 70% of OpenAI and Anthropic ARR — is beginning, with tax preparation offering a concrete example of an agentic expert workflow [12][11].
Why it matters
The Chinese model pricing gap, now anchored by Citibank data rather than inference, is roughly 22x on sticker price and open-source share on OpenRouter nearly doubled in five months [8]. The business thesis for U.S. frontier labs depends on whether software optimization arguments — SemiAnalysis puts effective Opus 4.7 costs near $0.99/million for agentic workloads [6] — can retain enterprise customers when Chinese alternatives are this cheap and gaining share this fast.
Open questions
SemiAnalysis reported Anthropic gross margins above 70% [16] while separate reporting places the figure at approximately 44% [10] — which measurement reflects the current period, and what accounts for the discrepancy?
Open-source model processing on OpenRouter rose from 34% to 65% in five months [8] — is this rate of share gain still accelerating, and at what point does it materially constrain frontier lab pricing power?
SemiAnalysis predicts cybersecurity will replicate coding's share of AI lab ARR [11], while tax preparation now offers a concrete adjacent example [12] — is there evidence either domain is capturing share at scale yet?
UBS reports 60% of large companies slowing AI spending [9] while SemiAnalysis argues 90th-percentile customers account for most API revenue and are not cutting back [11] — do these findings describe non-overlapping market segments, or is one wrong about the same population?
Narrative
Goldman Sachs Research's 24x AI agent token growth forecast by 2030 [1] is supported by production numbers visible across multiple deployments. OpenAI's internal data shows Codex went from below 10% to approximately 99.8% of the company's internal output tokens in under a year, with non-developer agent use growing 137x for individuals and 189x for organizations since August 2025 [2]. Meta employees collectively consumed over 60 trillion tokens in a single 30-day period, with one employee alone consuming 280 billion tokens at an implied cost of approximately $50,000 per year [3]. Ethan Mollick documents Opus 4.7 autonomously building a software package representing 2–17 weeks of human engineering work in 14 hours at $251 in token costs [4]. McKinsey projects AI agents could mediate between $3 trillion and $5 trillion of global retail commerce by 2030, requiring brands to expose machine-readable APIs to remain agent-discoverable [5].
Cost compression is accelerating from hardware and software, while Chinese model competition introduces structural pricing pressure. SemiAnalysis reports effective Opus 4.7 costs near $0.99 per million tokens for real agentic workloads, because those workloads run at approximately 300:1 input-to-output ratios and achieve cache hit rates above 90% [6]. NVIDIA's Blackwell inference stack reduced DeepSeek V4 token costs by up to 5x within approximately one month of release, with combined optimizations yielding up to 20x token throughput [7]. Against these efficiency gains, Citibank Research puts Chinese AI model prices as low as $0.18 per million tokens versus a $4 average for frontier models — roughly a 22x sticker-price gap — and open-source model processing on OpenRouter grew from 34% in January 2026 to 65% in June 2026, driven significantly by Chinese model adoption [8]. UBS separately reported Chinese firms accounting for over 45% of all OpenRouter traffic as of April 2026 [9]. Gartner estimates AI coding costs will surpass the average developer salary by 2028, making cost control a growing strategic concern at the enterprise level [8]. OpenAI's own adjusted gross margin fell to 33% in 2025 after inference costs quadrupled, recovering to 39% in Q1 2026 with a 52% year-end target [10].
Enterprise adoption is contested in both direction and measurement. SemiAnalysis argues that pullbacks at Meta and Uber stem from poor internal incentive structures and are not representative: 90th-percentile customers account for most API revenue and are not cutting spend, and the median Fortune 500 employee spends under $100 per year on AI, suggesting the adoption curve has barely begun [11]. UBS separately reports roughly 60% of large companies slowing AI spending as executives add guardrails [9]. The gap between a small number of high-intensity power users — some individual Meta engineers running agentic loops implying compute costs equivalent to $10 million per year [3] — and the median enterprise worker is the structural feature that makes both statistics simultaneously plausible. Over 70% of OpenAI and Anthropic ARR is concentrated in coding [11], and the broader adoption thesis depends on that pattern replicating in adjacent domains.
Domain expansion beyond coding is beginning, with tax preparation as a concrete example. Tax AI, built on OpenAI's Codex, parses heterogeneous document types — PDFs, spreadsheets, images, and handwritten notes — to prepare returns for accountant review; the system converts expert corrections into structured signals that surface measurable edge cases for engineers to fix before shipping [12]. The central design insight from this deployment: the instruction and workflow layer around the model — the harness, evaluation system, and review interface — is the product, not the underlying model itself [12]. DeepMind integrated computer use into Gemini 3.5 Flash with adversarial training against prompt injection and confirmation gates for irreversible actions [13], and Scout launched with a similar human-approval gate for changes involving money or external integrations [14]. Notion shut down Notion Mail on September 22, 2026, citing users' shift to AI agents for email [15].
Timeline
- 2026-05-30: Goldman Sachs 24x AI agent token forecast publicized; first reports that Microsoft and Uber find agent deployments more expensive than equivalent human workers. [1]
- 2026-06-24: DeepMind integrates computer use into Gemini 3.5 Flash with adversarial prompt-injection training and confirmation gates for irreversible actions. [13]
- 2026-06-25: OpenAI publishes internal agent usage data: Codex at 99.8% of output tokens; non-developer individual use up 137x and organizational use up 189x since August 2025. [2]
- 2026-06-25: McKinsey projects AI agents will mediate $3–5 trillion of global retail commerce by 2030, requiring brands to expose machine-readable APIs. [5]
- 2026-06-25: Notion announces Notion Mail shutdown on September 22, 2026, citing users' shift to AI agents for email. [15]
- 2026-06-25: Multiple reports confirm Microsoft and Uber find AI agent deployments more expensive than equivalent human workers, leading to pullbacks. [20][21][22]
- 2026-06-27: SemiAnalysis: Anthropic ARR grew from $9B to $44B+, gross margins from 38% to 70%+; effective Opus 4.7 cost near $0.99/million tokens due to 300:1 input/output ratios and 90%+ cache hit rates. [16][6]
- 2026-06-29: Scout launches: users specify a business KPI in plain English; system autonomously builds and tests agents, routing money- or integration-related changes to human approval. [14]
- 2026-06-30: NVIDIA Blackwell inference stack cuts DeepSeek V4 token costs by 5x within one month; combined optimizations yield up to 20x token throughput on Blackwell. [7][19]
- 2026-06-30: Mollick: Opus 4.7 autonomously builds software representing 2–17 weeks of human engineering work in 14 hours at $251; one-quarter of OpenAI employees run 4+ agents simultaneously at least weekly. [4]
- 2026-06-30: OpenAI adjusted gross margin fell to 33% in 2025 after inference costs quadrupled, recovering to 39% in Q1 2026 with 52% year-end target; Anthropic gross margin reported separately at approximately 44%. [10]
- 2026-06-30: SemiAnalysis: Meta/Uber AI pullback reports overstated; over 70% of OpenAI and Anthropic ARR from coding; median Fortune 500 employee spends under $100/year on AI. [11]
- 2026-06-30: Citibank Research: Chinese AI models priced as low as $0.18/million tokens vs. $4 average for frontier models; open-source processing on OpenRouter grew from 34% in January to 65% in June 2026; Gartner projects AI coding costs will exceed average developer salary by 2028. [8]
- 2026-07-01: Meta employees consumed 60+ trillion tokens in 30 days; one employee alone consumed 280 billion tokens, implying approximately $50,000/year at average rates. [3]
- 2026-07-01: UBS: roughly 60% of large companies slowing AI spending; Chinese firms account for 45%+ of OpenRouter traffic as of April 2026, up from under 2% in late 2024. [9]
- 2026-07-01: Tax AI on OpenAI's Codex prepares tax returns for accountant review; feedback loop converts expert corrections into structured improvement signals; harness architecture — not the underlying model — framed as the core product. [12]
Perspectives
Goldman Sachs Research
Projects 24x AI agent token demand growth by 2030; expects token cost declines to outpace price reductions, positioning cloud providers near a gross-margin turning point.
Evolution: Consistent — the originating analytical source for the thread's central forecast.
SemiAnalysis
AI labs now capture most value in the stack; effective Opus 4.7 costs near $0.99/million tokens for optimized workloads; the enterprise pullback narrative at Meta and Uber is overstated, driven by poor internal incentives; over 70% of OpenAI and Anthropic ARR comes from coding, with cybersecurity projected as the next domain.
Evolution: Consistent; the primary counter-voice to UBS and enterprise cost-skeptic sources.
McKinsey
Projects AI agents will mediate $3–5 trillion of global retail commerce by 2030; argues brands must adopt machine-readable API infrastructure or be bypassed by AI purchasing agents.
Evolution: Consistent.
OpenAI
Internal research documents agents enabling longer and more complex tasks; Codex at 99.8% of internal output tokens; adjusted gross margin fell to 33% in 2025 after inference costs quadrupled, recovering to 39% in Q1 2026 with a 52% year-end target.
Evolution: Consistent; the margin data confirms cost pressure existed before the current recovery.
NVIDIA
Makes a technical and commercial case for its vertically integrated inference software stack; Blackwell hardware plus software optimizations achieved 5x cost reduction on DeepSeek V4 within one month, with combined layers yielding up to 20x throughput.
Evolution: Consistent.
Enterprise cost-skeptics (Microsoft, Uber, UBS, Citibank)
Microsoft and Uber found specific agent deployments more expensive than equivalent human workers; UBS reports roughly 60% of large companies slowing AI spending; Citibank Research documents Chinese models at $0.18/million versus $4 for frontier models, framing model selection as a cost-control problem rather than a capability contest.
Evolution: Citibank data this pass sharpens the pricing gap with specific figures and adds OpenRouter share data showing open-source at 65% in June 2026, up from 34% in January — more precise and faster-moving than the April 2026 UBS figures.
Ethan Mollick (One Useful Thing)
Argues AI has crossed from chatbot co-intelligence to agentic autonomy; documents Opus 4.7 producing 2–17 weeks of human engineering work in 14 hours at $251; argues institutional planning horizons cannot track current capability curves.
Evolution: Consistent.
Enterprise power users (Meta consumption data)
A small number of high-intensity users — individual Meta engineers running agentic loops implying compute costs equivalent to $10 million per year — drive disproportionate AI token consumption; SemiAnalysis argues these users account for most lab ARR and are not cutting back.
Evolution: Consistent; Meta's 60+ trillion token figure provides the concrete scale behind the power-user concentration pattern.
Tensions
- Goldman Sachs, McKinsey, and SemiAnalysis project large-scale agent adoption and value capture by 2030; UBS reports roughly 60% of large companies are currently slowing AI spending, and Microsoft and Uber found specific deployments more expensive than equivalent human workers. [1][5][16][9][20][21]
- SemiAnalysis directly calls the Meta and Uber pullback reports overstated, attributing them to poor internal incentive structures; UBS reports 60% of large companies slowing spend — the two findings may describe different customer segments but neither directly addresses the other's data. [11][9]
- SemiAnalysis reports Anthropic gross margins above 70%; separate reporting places the figure at approximately 44% — the discrepancy may reflect different metrics, time periods, or definitions, but is unresolved across available sources. [16][10]
- Over 70% of OpenAI and Anthropic ARR is concentrated in coding, while OpenAI, McKinsey, and others project broad agent adoption across legal, finance, and consumer retail — the broader thesis depends on pattern replication not yet observed at scale. [11][2][5]
- Citibank Research puts Chinese model prices at $0.18/million versus $4 for frontier models, with open-source share on OpenRouter at 65% in June 2026; SemiAnalysis argues effective Opus 4.7 costs are near $0.99/million for optimized workloads — whether software optimization closes the gap enough to retain enterprise customers is unresolved. [8][6][9]
Sources
- [1] Goldman Sachs: "Token use by AI agents is expected to multiply 24 times by 2030" — Rohan Paul Twitter (2026-05-30)
- [2] OpenAI just released a paper showing how they are now seeing the first version of office work where agents do most of th… — Rohan Paul Twitter (2026-06-25)
- [3] Meta employees used over 60 trillion tokens in 30 days, with one user alone consumed 280 billion. — Rohan Paul Twitter (2026-07-01)
- [4] The twilight of the chatbots — One Useful Thing (2026-06-30)
- [5] Mckinsey report - AI agents are quietly taking over the retail shopping cart and could mediate $3 Tn to $5 tn of global … — Rohan Paul Twitter (2026-06-25)
- [6] The substitution math is the part to internalize. Tasks that used to need a junior analyst for several hours, converting… — SemiAnalysis Twitter (2026-06-27)
- [7] How NVIDIA’s Inference Software Stack Powers the Lowest Token Cost — NVIDIA Blog (2026-06-30)
- [8] Reuters: Chinese models charge as little as 18 cents per million tokens versus $4 average for top models, says CitiBank … — Rohan Paul Twitter (2026-06-30)
- [9] UBS says about 60% of big companies are slowing AI spending. — Rohan Paul Twitter (2026-07-01)
- [10] The Information reports that OpenAI has cut inference costs by more than half on some existing models, while logged-out … — Rohan Paul Twitter (2026-06-30)
- [11] TokenBudgeting: — SemiAnalysis Twitter (2026-06-30)
- [12] 😺 Watch: AI can do your taxes now — The Neuron (2026-07-01)
- [13] Introducing computer use in Gemini 3.5 Flash — DeepMind Blog (2026-06-24)
- [14] AI agents to automatically improve business-critical KPIs. — Rohan Paul Twitter (2026-06-29)
- [15] Notion killing Skiff-influenced email app since most users use AI agents instead — Ars Technica AI (2026-06-25)
- [16] If you are an operator trying to write down what tokens will cost in 2027, the answer is materially lower than today, an… — SemiAnalysis Twitter (2026-06-27)
- [17] The throughput math has gotten the most pushback in our reader notes, so its worth being precise. On the same B300 runni… — SemiAnalysis Twitter (2026-06-27)
- [18] One of the more uncomfortable observations in our AI Value Capture piece is internal: our token spend at SemiAnalysis no… — SemiAnalysis Twitter (2026-06-27)
- [19] NVIDIA's newly published report says its Blackwell inference stack cut DeepSeek V4 token costs by up to 5x in one month.… — Rohan Paul Twitter (2026-06-30)
- [20] AI promised cost savings, but Microsoft and Uber say it’s costing more than human workers | Company Business News — reactive:ai-agent-economics-enterprise
- [21] Microsoft and Uber Pull Back on AI Subscriptions Due to Cost and ... — reactive:ai-agent-economics-enterprise
- [22] Uber, Microsoft, and Others Burning Through AI Budgets. Now What? — reactive:ai-agent-economics-enterprise