The Information Machine

AI Labs Simultaneously Acknowledge Recursive Self-Improvement Threshold · history

Version 9

2026-06-14 08:15 UTC · 180 items

What

In June 2026, Anthropic and OpenAI separately disclosed evidence of recursive self-improvement in deployed systems — Claude authoring more than 80% of Anthropic's production code [1], OpenAI calling RSI 'potentially the most consequential frontier safety issue of the coming decade' [3] — and joined Google DeepMind in calling for international mechanisms to slow frontier AI development [7]. Sam Altman has reportedly warned OpenAI staff internally that a major RSI breakthrough could justify delaying the company's IPO, since some work is strategically easier to pursue while the company remains private [11]. External challenges remain: Jeremy Howard argues Anthropic's slowdown call is logically negated by its own continued use of Claude Mythos for frontier research [14], Geoffrey Irving founded Sequent on the premise that lab alignment is too reactive for ASI [15], and the US government directed CAISI — the oversight body OpenAI proposed — to stop publishing public model evaluations [8].

Why it matters

Altman's reported internal warning that a major RSI breakthrough might lead OpenAI to delay its IPO to preserve strategic flexibility suggests the company's actual posture on RSI timing is more conditional than its public endorsement of coordinated slowdowns implies. Three labs endorsing international coordination while each continues frontier development concentrates agenda-setting power among incumbents who also control when and how that coordination gets defined.

Open questions

  • Does Anthropic plan to stop using Claude Mythos for its own frontier research? Anthropic has not publicly responded to Howard's argument that this practice negates its slowdown call [14].

  • If Altman's reported internal position is that a major RSI breakthrough might lead OpenAI to delay its IPO to preserve strategic flexibility while private [11], does the company's public endorsement of coordinated slowdowns reflect its actual institutional incentives?

  • The US government directed CAISI to stop publishing public AI model evaluations [8] — without that transparency function, what enforcement mechanism does OpenAI's governance blueprint actually provide?

  • Can any international coordination mechanism for slowing frontier AI be verified across jurisdictions? All three labs call for one, but none has specified how compliance would be established or enforced [7].

Narrative

In early June 2026, Anthropic and OpenAI independently published documents acknowledging that recursive self-improvement may already be underway in deployed systems. Anthropic's 'When AI Builds Itself' disclosed that Claude authored more than 80% of Anthropic's production code merged in May 2026, that per-engineer code output reached 8x the 2024 baseline, and that Claude Mythos Preview accelerated model-training code by approximately 52x [1][2]. Claude's success rate on open-ended coding tasks reached 76%, up 50 points in six months [2]. OpenAI's concurrent policy blueprint called RSI 'potentially the most consequential frontier safety issue of the coming decade,' stating the company sees 'early signs of recursive self-improvement in today's systems' [3][4]. Technical research published in this period adds context: a paper on SIA describes a loop in which one AI agent rewrites its own configuration and updates its model without human intervention, framing this as a solution to the bottleneck of human-in-the-loop development [5], and Google DeepMind researchers identified four distinct technical pathways through which AGI could transition to ASI [6].

On governance, Anthropic, OpenAI, and Google DeepMind have all called for international coordination mechanisms to slow frontier AI development [7]. OpenAI's blueprint proposed a federal oversight body (CAISI) with mandatory evaluation authority; the US government subsequently directed CAISI to stop publishing public AI model evaluations, removing the transparency function at the center of OpenAI's proposal [8]. Analyst Zvi Mowshowitz praised OpenAI's blueprint but warned its proposed federal preemption of state frontier safety laws is its most dangerous element: once states surrender regulatory leverage, Congress may never enact the promised federal replacement [4]. NSPM-11 effectively bans Anthropic from federal contracts while permitting unrestricted government use of any AI model once deployed, and the NSA is reported to be using Claude Mythos for offensive cyber operations with approximately half a dozen embedded Anthropic engineers — a contradiction now receiving mainstream tech press coverage [7][9][10].

A separate internal dimension shapes OpenAI's calculus: Altman reportedly warned staff that a major RSI breakthrough could justify delaying the company's IPO, arguing that some work is strategically easier to pursue while the company remains private and free of public-market pressure on revenue and profit [11]. This sits alongside his March 2028 prediction that AI will conduct a significant fraction of OpenAI's own research [12] and OpenAI's public blog endorsement of coordinated slowdown mechanisms [13]. The three positions — accelerated research targets, public slowdown endorsements, and private acknowledgment that an RSI event might push the company to stay private longer — are not obviously consistent.

Two independent challenges to the labs' safety framing have emerged. Jeremy Howard argues the only internally consistent way to slow RSI is to prohibit the lab with the top-ranked model from using it for frontier AI research: Anthropic fails this test by using Claude Mythos for its own frontier research while calling for others to restrict theirs, and has not publicly responded [14]. Geoffrey Irving separately announced Sequent, founded on the argument that existing lab alignment work is 'predominantly empirical and reactive' and cannot yield principled insight into whether safety generalizes to ASI [15]. Claude Fable 5 launched with benchmark performance similar to GPT-5.5 and Composer 2.5 but at 4-12x higher cost per completed task [8]. Critics have also challenged Anthropic's slowdown call as competitive strategy timed to the company's confidential S-1 filing at a valuation of approximately $965 billion, arguing a global slowdown disproportionately advantages incumbents already positioned at the frontier [16][17][18].

Timeline

  • 2025-05: Claude Opus 4 accelerates Anthropic's model-training code by approximately 3x. [2]
  • 2026-05: Claude authors more than 80% of Anthropic's production code; per-engineer output reaches 8x the 2024 baseline; Claude Mythos Preview accelerates model-training code by approximately 52x; Claude's success rate on open-ended coding tasks reaches 76%, up 50 points in six months. [2][26][1]
  • 2026-06-03: OpenAI publishes 'Democratic Governance of Frontier AI' policy blueprint, acknowledging early RSI signs and proposing CAISI as a federal oversight body with mandatory evaluation authority. [22][3][4]
  • 2026-06-04: Anthropic publishes 'When AI Builds Itself,' disclosing Claude's 80%+ code authorship and calling for a global coordinated slowdown in frontier AI development. [19][17][20]
  • 2026-06-05: Zvi Mowshowitz praises OpenAI blueprint but warns its federal preemption clause could strip states of safety oversight permanently; critics note the timing of Anthropic's slowdown call relative to its confidential S-1 filing at a ~$965B valuation. [4][16][25][17]
  • 2026-06-07: Scientific American covers Anthropic's RSI warning; the 'perfect pre-IPO narrative' critique spreads broadly on social media. [27][18][28]
  • 2026-06-08: Jack Clark treats Anthropic's 8x code increase as preliminary RSI evidence; Sam Altman publishes blog predicting AI will conduct a significant fraction of OpenAI's own research by March 2028. [24][12]
  • 2026-06-09: OpenAI's official blog endorses coordinated mechanisms to slow frontier development when needed; Zvi Mowshowitz reports three-lab convergence (Anthropic, OpenAI, DeepMind) on international slowdown coordination. [13][7]
  • 2026-06-09: NSPM-11 bans Anthropic from government contracts while NSA uses Claude Mythos for offensive cyber with approximately half a dozen embedded Anthropic engineers reported. [7][9]
  • 2026-06-10: Geoffrey Irving announces Sequent, founded on the argument that lab alignment approaches are too reactive to provide principled safety confidence before ASI; plans to raise $100-150M. [15]
  • 2026-06-10: Jeremy Howard argues Anthropic's slowdown call is internally contradictory because Anthropic continues using Claude Mythos for frontier research, advancing the frontier it claims to want to slow. [14]
  • 2026-06-10: Sam Altman reportedly warns OpenAI staff that a major RSI breakthrough could justify delaying the company's IPO, since some work is strategically easier while the company remains private. [11]
  • 2026-06-11: Tom's Hardware and mainstream tech press pick up the NSA/Claude Mythos offensive cyber story; Claude Fable 5 launches at similar benchmark performance to GPT-5.5 and Composer 2.5 but at 4-12x higher cost per task; US government directs CAISI to stop publishing public AI model evaluations. [10][8]
  • 2026-06-11: Research on SIA — a loop in which one AI agent rewrites its own configuration and updates its model without human intervention — published, illustrating the self-improvement architecture the labs are describing. [5]
  • 2026-06-12: Google DeepMind researchers identify four distinct technical pathways through which AGI could transition to ASI, including continued scaling of compute, model size, data, and test-time compute. [6]

Perspectives

Anthropic

Believes current models may be approaching the RSI threshold and calls for a global coordinated slowdown before autonomous self-improvement becomes difficult to control; has not publicly responded to Howard's argument that continued use of Claude Mythos for frontier research contradicts the slowdown call, nor to the NSA offensive cyber reporting.

Evolution: RSI disclosure and global-slowdown call are the most explicit external acknowledgment of near-term risk Anthropic has made; Howard's consistency challenge and the NSA embedded-engineers story remain publicly unanswered.

OpenAI / Sam Altman

Acknowledges 'early signs' of RSI and calls it the top frontier safety issue of the decade; official blog endorses coordinated slowdown mechanisms; Altman targets AI-conducted research at OpenAI by March 2028; and reportedly warns staff internally that a major RSI breakthrough could justify delaying the IPO to preserve strategic flexibility while private.

Evolution: The IPO delay warning adds a conditional, business-strategy dimension to Altman's public slowdown endorsement — the positions are not clearly consistent; CAISI's suppression of public evaluations also undercuts the transparency argument in OpenAI's own blueprint.

Google DeepMind

Supports an international organization to enable coordinated slowdowns of frontier AI development; researchers are simultaneously publishing technical work identifying four pathways from AGI to ASI.

Evolution: Governance position is consistent with prior reporting; the AGI-to-ASI paper adds technical depth but does not change the public stance.

Jeremy Howard

Argues the only internally consistent way to slow RSI is to prohibit the lab with the top-ranked model from using it for frontier AI research; Anthropic fails this test and has not responded.

Evolution: No response from Anthropic; argument stands uncontested.

Geoffrey Irving / Sequent

Founded Sequent on the argument that current lab alignment approaches are empirical and reactive, cannot yield principled confidence that safety generalizes to ASI, and that an independent theory-focused organization is necessary.

Evolution: New entrant; no subsequent statements on record.

Zvi Mowshowitz

Guardedly positive on OpenAI's blueprint but identifies federal preemption of state safety laws as its most dangerous element; reports three-lab convergence on international coordination; treats CAISI evaluation suppression as a significant transparency setback.

Evolution: Coverage extended from blueprint analysis to Fable 5 launch and CAISI suppression, consistently concerned about governance mechanisms losing their teeth.

Jack Clark (Import AI)

Treats Anthropic's 8x code increase as 'preliminary evidence of prosaic recursive self-improvement' and connects RL reward hacking — AI rediscovering regulatory loopholes without instructions — to the institutional rule systems governing societies.

Evolution: Consistent with long-standing editorial urgency on AI risk; no shift.

Critics (NY Post, social media commentators)

Argue Anthropic's slowdown call is competitive strategy timed to the S-1 filing at a ~$965B valuation; the 'perfect pre-IPO narrative' framing is the dominant skeptical frame and has spread widely across platforms.

Evolution: Spreading more broadly; no new analytical positions introduced.

Tensions

  • Howard argues Anthropic's continued use of Claude Mythos for its own frontier research logically negates its slowdown call — the lab with the top model advances the frontier by using it; Anthropic has not responded. [21][14]
  • Altman publicly endorses coordinated slowdown mechanisms while reportedly telling staff internally that a major RSI breakthrough could justify delaying OpenAI's IPO to preserve strategic flexibility while private — the two positions point in different directions on what OpenAI actually wants from RSI timing. [13][11]
  • OpenAI's blueprint proposed CAISI as a federal oversight body with public evaluation authority; the US government directed CAISI to stop publishing public AI model evaluations, removing the transparency function the proposal depended on. [22][8]
  • Anthropic's public safety messaging centers on a global slowdown; Anthropic engineers are simultaneously reported embedded at the NSA using Claude Mythos for offensive cyber operations. [21][7][9][10]
  • Anthropic frames its slowdown call as a genuine safety response to near-term RSI risk; critics argue it is regulatory strategy timed to entrench Anthropic's market position after its IPO filing at a ~$965B valuation. [19][16][17][18]
  • Jack Clark treats Anthropic's 8x code increase as preliminary RSI evidence; other analysts read the same data as a productivity milestone where humans still set goals and review outputs, making 'autonomous self-improvement' a matter of degree rather than a categorical threshold. [2][26][1][24]

Sources

  1. [1] Today’s edition of my newsletter just went out. — Rohan Paul Twitter (2026-06-05)
  2. [2] 😺 Anthropic: AI Is Building AI now — The Neuron (2026-06-05)
  3. [3] Peter Wildeford🇺🇸🚀 on X: "OPENAI: "We also see early signs of recursive self-improvement in today's systems". RSI is "potentially the most consequential frontier safety issue of the coming decade."" / X — reactive:rsi-governance-moment
  4. [4] OpenAI Offers A New Policy Blueprint — Zvi's AI Roundups (2026-06-05)
  5. [5] This paper shows an AI improving itself better when it rewrites its setup and updates its model. — Rohan Paul Twitter (2026-06-11)
  6. [6] Beautiful paper from Google DeepMind. — Rohan Paul Twitter (2026-06-12)
  7. [7] Three Labs With a Plan and A Memorandum — Zvi's AI Roundups (2026-06-09)
  8. [8] AI #172: The First Fable — Zvi's AI Roundups (2026-06-11)
  9. [9] National Security Presidential Memorandum/NSPM-11 — reactive:rsi-governance-moment
  10. [10] NSA using Claude Mythos for 'offensive cyber operations,' report claims — says 'half-a-dozen' Anthropic engineers embedded inside the agency — reactive:rsi-governance-moment
  11. [11] Sam Altman is reportedly warning staff that recursive self-improvement (RSI) could delay its IPO. — Rohan Paul Twitter (2026-06-10)
  12. [12] Sam Altman's new blog about OpenAI's future path says by March-2028 a significant fraction of its own research will be d… — Rohan Paul Twitter (2026-06-08)
  13. [13] OpenAI's latest official blog says the world may need a way to coordinate "slowing frontier development when needed." ht… — Rohan Paul Twitter (2026-06-09)
  14. [14] Quoting Jeremy Howard — Simon Willison (2026-06-10)
  15. [15] Sequent: scale and automation for higher confidence in alignment — Alignment Forum (2026-06-10)
  16. [16] The company that just confidentially filed its S-1 for a trillion-dollar IPO published a blog post four days later askin... — reactive:rsi-governance-moment (2026-06-05)
  17. [17] Anthropic calls for global AI slowdown after $965B valuation. Critics claim it's just to hobble competition. — reactive:rsi-governance-moment
  18. [18] Anthropic has discovered the perfect pre-IPO narrative: its product is so powerful that it justifies a trillion-dollar v... — reactive:rsi-governance-moment (2026-06-07)
  19. [19] Anthropic just called for a global way to slow frontier AI because its own models may be approaching recursive self-impr… — Rohan Paul Twitter (2026-06-05)
  20. [20] Anthropic calls for pause of global AI development — reactive:rsi-governance-moment
  21. [21] Anthropic calls for global pause in AI development before humans ... — reactive:rsi-governance-moment
  22. [22] [PDF] Democratic Governance of Frontier AI - OpenAI — reactive:rsi-governance-moment
  23. [23] https://t.co/ux9SUEPVWG — reactive:rsi-governance-moment (2026-06-05)
  24. [24] Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing — Import AI (2026-06-08)
  25. [25] @unusual_whales Anthropic floated the idea of a global pause on frontier AI development right after securing $65 billion... — reactive:rsi-governance-moment (2026-06-05)
  26. [26] Anthropic just disclosed that Claude now writes more than 80% of the production code it merges. — Rohan Paul Twitter (2026-06-05)
  27. [27] Anthropic warns AI may soon begin recursive self-improvement — reactive:rsi-governance-moment
  28. [28] Anthropic Calls for Global Slowdown in AI Development - Reddit — reactive:rsi-governance-moment