The Information Machine

AI Labs Simultaneously Acknowledge Recursive Self-Improvement Threshold · history

Version 5

2026-06-10 02:25 UTC · 140 items

What

In early June 2026, Anthropic and OpenAI separately acknowledged early signs of recursive self-improvement in deployed systems — Anthropic disclosing that Claude authored more than 80% of its production code in May 2026 [1], OpenAI calling RSI 'potentially the most consequential frontier safety issue of the coming decade' [3]. A governance convergence has followed: Anthropic, OpenAI, and Google DeepMind have all now publicly called for international mechanisms to coordinate slowing frontier AI development when needed [6], with OpenAI's official blog explicitly endorsing the possibility [8]. The picture is complicated by a government dimension: Trump's NSPM-11 reportedly bans Anthropic from federal contracts while permitting unrestricted government use of deployed AI, and the NSA is reported to be using Claude Mythos for offensive cyber operations with Anthropic engineers embedded at the agency [6].

Why it matters

Three major labs publicly agreeing on international slowdown coordination makes that position harder to dismiss as one company's self-interest, but it also concentrates agenda-setting power among the incumbents best positioned to shape what any such mechanism looks like. The NSA/Anthropic arrangement — if accurate — creates a direct tension between Anthropic's public safety messaging and its operational government relationships that neither party has addressed.

Open questions

  • OpenAI's blog now endorses coordinated slowdown mechanisms [8] while Altman's March 2028 roadmap targets AI-conducted research at OpenAI [5] — are these positions reconcilable, or does the gap between OpenAI's safety rhetoric and its operational planning reflect genuine internal disagreement?

  • Does NSPM-11 — banning Anthropic from government contracts while permitting unrestricted use of deployed AI — give the U.S. government effective leverage to override Anthropic's safety commitments after a model is handed over [6]?

  • Can any international coordination mechanism for slowing frontier AI be verified or enforced across jurisdictions? All three labs now call for one, but none has specified a verification mechanism [15][16][6].

  • Does the NSA's reported use of Claude Mythos for offensive cyber operations [6] undermine Anthropic's safety narrative, or does Anthropic employees' public rejection of the 'machine God' framing signal meaningful internal constraints on what they will build?

Narrative

In early June 2026, Anthropic and OpenAI independently published documents acknowledging that recursive self-improvement may already be underway in deployed systems. Anthropic's 'When AI Builds Itself' disclosed that Claude authored more than 80% of Anthropic's production code merged in May 2026, that per-engineer code output reached 8x the 2024 baseline, and that Claude Mythos Preview accelerated model-training code by approximately 52x [1][2]. Claude's success rate on open-ended coding tasks reached 76%, up 50 points in six months, and reliable task length is doubling approximately every four months [2]. OpenAI's concurrent policy blueprint called RSI 'potentially the most consequential frontier safety issue of the coming decade' and stated the company sees 'early signs of recursive self-improvement in today's systems' [3][4]. Sam Altman separately published a blog post predicting AI will conduct a significant fraction of OpenAI's own research by March 2028 [5].

On governance, Anthropic, OpenAI, and Google DeepMind have all publicly called for international coordination mechanisms to slow frontier AI development when needed [6]. Anthropic's June 4 call was the most direct — a request for a global coordinated slowdown before autonomous self-improvement becomes difficult to control [7]. OpenAI's policy blueprint proposed a federal oversight body (CAISI) with mandatory evaluation authority but no deployment veto, and OpenAI's official blog has since stated the world may need mechanisms to coordinate 'slowing frontier development when needed' [8][6]. Analyst Zvi Mowshowitz praised the substance of OpenAI's blueprint but warned that its proposed federal preemption of state frontier safety laws is the most dangerous element: once states surrender regulatory leverage, Congress may never enact the federal replacement [4].

The credibility of both labs' safety positions is complicated by financial and analytical context. Anthropic's slowdown call came shortly after the company confidentially filed an S-1 at a valuation of approximately $965 billion [9][10], and the formulation 'the perfect pre-IPO narrative' circulated widely on social media, with critics arguing a global slowdown would disproportionately disadvantage competitors that have not yet secured Anthropic's position [11]. Jack Clark (Import AI) offered a different angle: he treated Anthropic's 8x code-merge increase as 'preliminary evidence of prosaic recursive self-improvement' while noting that RL systems have separately demonstrated the ability to rediscover regulatory loopholes with 61.25% recall without explicit instructions, connecting reward hacking to the institutional rule systems that govern societies [12]. TechCrunch noted that RSI 'is just as hard to pin down' as AGI once was [13], and multiple analysts observe that the RSI loop is not yet closed — humans still direct goals and review outputs — making 'autonomous self-improvement' a matter of degree rather than a categorical threshold [2][14].

A government dimension has entered the story. Per Zvi Mowshowitz's reporting, Trump's NSPM-11 effectively bans Anthropic and its subcontractors from federal contracts while implementing a legal framework that allows unrestricted government use of any AI model once deployed — removing any contractual enforcement mechanism Anthropic might try to assert [6]. The NSA is reported to be using Claude Mythos for offensive cyber operations, with Anthropic engineers embedded at the agency [6]. Separately, Anthropic employees have publicly disputed a framing by Joshua Achiam characterizing Anthropic as seeking a 'machine God' outcome; multiple staff, including Sarah Chen, described a 'The Culture'-type scenario as a 'disastrous disempowerment' rather than a goal [6]. Zvi frames the two labs' philosophical postures as distinct: Anthropic taking superintelligence and its consequences seriously, OpenAI trying to deny those consequences exist [6].

Timeline

  • 2025-02: Claude Code reaches research preview; Claude's share of Anthropic merged code is in the low single digits. [14]
  • 2025-05: Claude Opus 4 accelerates Anthropic's model-training code by approximately 3x. [2]
  • 2026-05: Claude authors more than 80% of Anthropic's production code; per-engineer output reaches 8x the 2024 baseline; reliable task length doubles approximately every four months; Claude sustains useful work on tasks over 16 hours. [2][14][1]
  • 2026-05: Claude Mythos Preview accelerates model-training code by approximately 52x; Claude's success rate on open-ended coding tasks reaches 76%, up 50 points in six months. [2]
  • 2026-05-28: TechCrunch reports RSI 'is the new AGI' and is just as difficult to define, raising questions about what governance proposals built on the term can bind. [13]
  • 2026-06-03: OpenAI publishes 'Democratic Governance of Frontier AI' policy blueprint, acknowledging early RSI signs in current systems and proposing CAISI as a federal oversight body. [19][3][4]
  • 2026-06-04: Anthropic publishes 'When AI Builds Itself,' disclosing Claude's 80%+ code authorship and calling for a global coordinated slowdown in frontier AI development. [17][10][18]
  • 2026-06-05: Zvi Mowshowitz praises OpenAI blueprint but warns its federal preemption clause could strip states of safety oversight permanently. [4]
  • 2026-06-05: Seth Herd (Alignment Forum) argues augmented LLMs are the most likely first takeover-capable AI and that current alignment research neglects this path. [22]
  • 2026-06-05: Critics note timing of Anthropic's slowdown call relative to its confidential S-1 filing at a ~$965B valuation. [9][21][10]
  • 2026-06-06: WSJ, Business Insider, France24, and SiliconAngle publish mainstream coverage of Anthropic's RSI disclosure and slowdown call. [23][24][25][7]
  • 2026-06-07: Scientific American covers Anthropic's RSI warning; the 'perfect pre-IPO narrative' critique spreads broadly on social media. [26][11][27]
  • 2026-06-08: Jack Clark (Import AI) treats Anthropic's 8x code increase as preliminary RSI evidence; SocioHack benchmark shows RL AI rediscovering regulatory loopholes at 61.25% recall without instructions. [12]
  • 2026-06-08: Sam Altman publishes blog predicting AI will conduct a significant fraction of OpenAI's own research by March 2028, with a three-goal roadmap toward personal AGI for every person. [5]
  • 2026-06-09: OpenAI's official blog endorses coordinated mechanisms to slow frontier development when needed, partially converging with Anthropic's position. [8]
  • 2026-06-09: Zvi Mowshowitz reports three labs (Anthropic, OpenAI, DeepMind) all call for international slowdown coordination; NSPM-11 bans Anthropic from government contracts while NSA uses Claude Mythos for offensive cyber with embedded Anthropic engineers; Anthropic employees publicly reject the 'machine God' framing. [6]

Perspectives

Anthropic

Believes current models may be approaching the RSI threshold and calls for a global coordinated slowdown before autonomous self-improvement becomes difficult to control; internally, employees publicly reject characterizations of Anthropic's goal as a 'machine God' or Culture-type outcome, calling such scenarios 'disastrous disempowerment.'

Evolution: The RSI disclosure and global-slowdown call are the most explicit external acknowledgment of near-term risk Anthropic has made; the internal framing dispute over end-goals is new this pass.

OpenAI / Sam Altman

Acknowledges 'early signs' of RSI in current systems and calls it the top frontier safety issue of the coming decade; official blog now also endorses coordinated slowdown mechanisms when needed; separately targets AI-conducted research at OpenAI by March 2028.

Evolution: The blog's endorsement of coordinated slowdowns moves OpenAI closer to Anthropic's position, though Altman's March 2028 research target points in the opposite direction.

Google DeepMind

Supports an international organization to enable coordinated slowdowns of frontier AI development, per Zvi Mowshowitz's reporting — joining Anthropic and OpenAI on the international coordination position.

Evolution: First appearance in this thread; no prior stance on record.

Jack Clark (Import AI)

Treats Anthropic's 8x code increase as 'preliminary evidence of prosaic recursive self-improvement' and connects RSI to RL findings showing AI rediscovering regulatory loopholes without instructions; describes current AI capabilities and stable societal structures as difficult to reconcile.

Evolution: Consistent with long-standing editorial urgency on AI risk; no shift.

Zvi Mowshowitz

Guardedly positive on OpenAI's blueprint but identifies federal preemption of state safety laws as its most dangerous element; his June 9 piece reports the three-lab convergence on international coordination, the NSPM-11/NSA angle, and frames Anthropic as taking superintelligence consequences seriously versus OpenAI denying those consequences exist.

Evolution: Extended from blueprint analysis to a broader comparative assessment of the three labs' philosophical postures and the government dimension.

Critics (NY Post, social media commentators)

Argue Anthropic's slowdown call is competitive strategy rather than safety response: the S-1 filing at a ~$965B valuation timing and 'perfect pre-IPO narrative' framing have spread broadly.

Evolution: Spreading more widely across platforms; no new analytical positions.

Seth Herd (Alignment Forum)

Argues the first takeover-capable AI will be an augmented LLM rather than a novel architecture; warns that existing alignment research neglects mechanistic prediction of likely failure modes and that motivated reasoning is splitting the safety community.

Evolution: Consistent across appearances; no shift.

Rohan Paul / The Neuron (analytical commentators)

Treat Anthropic's code-authorship disclosure as a significant capability milestone while noting the RSI loop is not yet closed — humans still direct goals and review outputs — making 'autonomous self-improvement' a matter of degree rather than a categorical threshold.

Evolution: Consistent across appearances; no shift.

Tensions

  • Anthropic calls for a global coordinated slowdown in frontier AI development; OpenAI's Altman simultaneously publishes a roadmap predicting AI will conduct a significant fraction of OpenAI's own research by March 2028 — the two positions point in opposite directions on development pacing. [7][5]
  • Anthropic's public safety messaging centers on the need for a global slowdown; Anthropic engineers are simultaneously reported to be embedded at the NSA using Claude Mythos for offensive cyber operations — a contradiction neither party has publicly addressed. [7][6]
  • Anthropic frames its slowdown call as a genuine safety response to near-term RSI risk; critics argue it is regulatory strategy timed to entrench Anthropic's market position after its IPO filing at a ~$965B valuation. [17][9][10][11]
  • OpenAI's blueprint proposes federal preemption of state frontier safety laws; Mowshowitz argues this is the most dangerous provision because Congress may never enact the promised federal replacement once states surrender leverage. [4]
  • Both labs treat RSI as a near-term observable phenomenon; TechCrunch and other analysts note the term remains as difficult to define as AGI once was, leaving the scope of any RSI-based governance proposal ambiguous. [13][3][17]
  • Anthropic's disclosure that Claude writes 80%+ of its production code is read by Jack Clark and Anthropic as preliminary RSI evidence; The Neuron and others read it as a productivity milestone where humans still set goals and review outputs. [2][14][1][12]

Sources

  1. [1] Today’s edition of my newsletter just went out. — Rohan Paul Twitter (2026-06-05)
  2. [2] 😺 Anthropic: AI Is Building AI now — The Neuron (2026-06-05)
  3. [3] Peter Wildeford🇺🇸🚀 on X: "OPENAI: "We also see early signs of recursive self-improvement in today's systems". RSI is "potentially the most consequential frontier safety issue of the coming decade."" / X — reactive:rsi-governance-moment
  4. [4] OpenAI Offers A New Policy Blueprint — Zvi's AI Roundups (2026-06-05)
  5. [5] Sam Altman's new blog about OpenAI's future path says by March-2028 a significant fraction of its own research will be d… — Rohan Paul Twitter (2026-06-08)
  6. [6] Three Labs With a Plan and A Memorandum — Zvi's AI Roundups (2026-06-09)
  7. [7] Anthropic calls for global pause in AI development before humans ... — reactive:rsi-governance-moment
  8. [8] OpenAI's latest official blog says the world may need a way to coordinate "slowing frontier development when needed." ht… — Rohan Paul Twitter (2026-06-09)
  9. [9] The company that just confidentially filed its S-1 for a trillion-dollar IPO published a blog post four days later askin... — reactive:rsi-governance-moment (2026-06-05)
  10. [10] Anthropic calls for global AI slowdown after $965B valuation. Critics claim it's just to hobble competition. — reactive:rsi-governance-moment
  11. [11] Anthropic has discovered the perfect pre-IPO narrative: its product is so powerful that it justifies a trillion-dollar v... — reactive:rsi-governance-moment (2026-06-07)
  12. [12] Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing — Import AI (2026-06-08)
  13. [13] RSI is the new AGI — and it's just as hard to pin down | TechCrunch — reactive:rsi-governance-moment
  14. [14] Anthropic just disclosed that Claude now writes more than 80% of the production code it merges. — Rohan Paul Twitter (2026-06-05)
  15. [15] @claudeai 8. A global pause or slowdown cannot be verified by anyone — reactive:rsi-governance-moment (2026-06-04)
  16. [16] Could a global slowdown in frontier AI model development really happen? — reactive:rsi-governance-moment (2026-06-04)
  17. [17] Anthropic just called for a global way to slow frontier AI because its own models may be approaching recursive self-impr… — Rohan Paul Twitter (2026-06-05)
  18. [18] Anthropic calls for pause of global AI development — reactive:rsi-governance-moment
  19. [19] [PDF] Democratic Governance of Frontier AI - OpenAI — reactive:rsi-governance-moment
  20. [20] https://t.co/ux9SUEPVWG — reactive:rsi-governance-moment (2026-06-05)
  21. [21] @unusual_whales Anthropic floated the idea of a global pause on frontier AI development right after securing $65 billion... — reactive:rsi-governance-moment (2026-06-05)
  22. [22] My research agenda and work — Alignment Forum (2026-06-05)
  23. [23] Anthropic Urges Global Pause in AI Development, Flags 'Self ... - WSJ — reactive:rsi-governance-moment
  24. [24] Anthropic Says Leading AI Labs May Need to Hit the Brakes — reactive:rsi-governance-moment
  25. [25] 'Human role narrowing': Anthropic calls for global AI slowdown as ... — reactive:rsi-governance-moment
  26. [26] Anthropic warns AI may soon begin recursive self-improvement — reactive:rsi-governance-moment
  27. [27] Anthropic Calls for Global Slowdown in AI Development - Reddit — reactive:rsi-governance-moment