The Information Machine

AI Labs Simultaneously Acknowledge Recursive Self-Improvement Threshold · history

Version 7

2026-06-11 08:30 UTC · 152 items

What

In June 2026, Anthropic and OpenAI separately disclosed evidence of recursive self-improvement in deployed systems — Claude authoring more than 80% of Anthropic's production code [1], OpenAI calling RSI 'potentially the most consequential frontier safety issue of the coming decade' [3] — and called for international mechanisms to slow frontier AI development, a position Google DeepMind has since joined [6]. Two independent challenges to the labs' safety framing have emerged: Jeremy Howard argues Anthropic's slowdown call is logically contradicted by Anthropic's continued use of Claude Mythos for its own frontier research [11]; and Geoffrey Irving has founded Sequent, an independent alignment organization, on the explicit argument that existing lab alignment work is too reactive to provide meaningful safety confidence before ASI arrives [12]. The NSA's reported use of Claude Mythos for offensive cyber operations, with Anthropic engineers embedded at the agency, has received widening press coverage [6][10].

Why it matters

Three major labs aligned on international slowdown coordination concentrates agenda-setting power among the incumbents best positioned to shape what that coordination looks like. Howard's point — that a lab calling for a global slowdown while using its top model for frontier research is advancing the very thing it claims to oppose — has not been answered publicly. Sequent's founding adds a credentialed outside voice arguing lab alignment work is structurally insufficient, independent of IPO timing or competitive motive.

Open questions

  • Does Anthropic plan to stop using Claude Mythos for its own frontier research? If not, Howard's argument stands: the lab with the top model and the slowdown call is advancing the frontier it claims to want to slow [11].

  • Can Sequent's theory-focused, automation-leveraging approach produce principled alignment confidence at $100-150M and 40-80 FTE before ASI development proceeds beyond what those resources can evaluate [12]?

  • Does NSPM-11's structure — banning Anthropic from government contracts while permitting unrestricted use of deployed models — give the U.S. government effective leverage to override Anthropic's safety commitments after a model is handed over [9][6]?

  • Can any international coordination mechanism for slowing frontier AI be verified or enforced across jurisdictions? All three labs now call for one, but none has specified a verification mechanism [6].

Narrative

In early June 2026, Anthropic and OpenAI independently published documents acknowledging that recursive self-improvement may already be underway in deployed systems. Anthropic's 'When AI Builds Itself' disclosed that Claude authored more than 80% of Anthropic's production code merged in May 2026, that per-engineer code output reached 8x the 2024 baseline, and that Claude Mythos Preview accelerated model-training code by approximately 52x [1][2]. Claude's success rate on open-ended coding tasks reached 76%, up 50 points in six months [2]. OpenAI's concurrent policy blueprint called RSI 'potentially the most consequential frontier safety issue of the coming decade,' stating the company sees 'early signs of recursive self-improvement in today's systems' [3][4]. Sam Altman separately predicted AI will conduct a significant fraction of OpenAI's own research by March 2028 [5].

On governance, Anthropic, OpenAI, and Google DeepMind have all publicly called for international coordination mechanisms to slow frontier AI development [6]. Anthropic's June 4 call was the most direct — a request for a global coordinated slowdown before autonomous self-improvement becomes difficult to control [7]. OpenAI's policy blueprint proposed a federal oversight body (CAISI) with mandatory evaluation authority but no deployment veto, and OpenAI's official blog has since stated the world may need mechanisms to coordinate 'slowing frontier development when needed' [8][6]. Analyst Zvi Mowshowitz praised the substance of OpenAI's blueprint but warned its proposed federal preemption of state frontier safety laws is the most dangerous element: once states surrender regulatory leverage, Congress may never enact the promised federal replacement [4]. A government dimension complicates the picture further: Trump's NSPM-11 effectively bans Anthropic from federal contracts while permitting unrestricted government use of any AI model once deployed, and the NSA is reported to be using Claude Mythos for offensive cyber operations with approximately half a dozen Anthropic engineers embedded at the agency — a contradiction between Anthropic's public safety messaging and its operational government relationships that has now drawn coverage from mainstream tech press [6][9][10].

Two independent challenges to the labs' safety framing have arrived. Jeremy Howard argues that the only internally consistent way to slow RSI is to prohibit the lab with the top-ranked model from using it for frontier AI research: 'if you claim we should slow down, and you have the best model, you should ensure your org can't use it.' [11] Howard's formulation is that Anthropic has 'chosen the opposite of the safe path' by continuing to use Claude Mythos for its own frontier research while calling for others to restrict theirs. Geoffrey Irving, separately, announced Sequent on June 10, a new organization explicitly founded on the argument that existing alignment work at AI labs is 'predominantly empirical and reactive' and cannot yield principled insight into whether safety will hold under ASI [12]. Sequent plans to raise $100-150M, grow to 40-80 FTE in two years, and stay independent of the labs specifically to report obstacles to alignment without organizational pressure to minimize concerns.

Analysts disagree on what the capability disclosures mean. Jack Clark (Import AI) treats Anthropic's 8x code increase as 'preliminary evidence of prosaic recursive self-improvement' and connects RL reward hacking — demonstrated by the SocioHack benchmark at 61.25% recall for rediscovering regulatory loopholes without instructions — to the institutional rule systems governing societies [13]. Others read the same data as a productivity milestone where humans still direct goals and review outputs, making 'autonomous self-improvement' a matter of degree rather than a categorical threshold [2][14]. Anthropic's slowdown call has also drawn a timing critique: it was filed shortly after the company confidentially submitted an S-1 at a valuation of approximately $965 billion, and the phrase 'the perfect pre-IPO narrative' circulated widely, with critics arguing a global slowdown disproportionately advantages incumbents that have already secured their position [15][16][17].

Timeline

  • 2025-02: Claude Code reaches research preview; Claude's share of Anthropic merged code is in the low single digits. [14]
  • 2025-05: Claude Opus 4 accelerates Anthropic's model-training code by approximately 3x. [2]
  • 2026-05: Claude authors more than 80% of Anthropic's production code; per-engineer output reaches 8x the 2024 baseline; Claude Mythos Preview accelerates model-training code by approximately 52x; Claude's success rate on open-ended coding tasks reaches 76%, up 50 points in six months. [2][14][1]
  • 2026-05-28: TechCrunch reports RSI is 'the new AGI' — just as difficult to define, raising questions about what governance proposals built on the term can bind. [23]
  • 2026-06-03: OpenAI publishes 'Democratic Governance of Frontier AI' policy blueprint, acknowledging early RSI signs in current systems and proposing CAISI as a federal oversight body. [20][3][4]
  • 2026-06-04: Anthropic publishes 'When AI Builds Itself,' disclosing Claude's 80%+ code authorship and calling for a global coordinated slowdown in frontier AI development. [18][16][19]
  • 2026-06-05: Zvi Mowshowitz praises OpenAI blueprint but warns its federal preemption clause could strip states of safety oversight permanently; critics note timing of Anthropic's slowdown call relative to its confidential S-1 filing at a ~$965B valuation. [4][15][22][16]
  • 2026-06-07: Scientific American covers Anthropic's RSI warning; the 'perfect pre-IPO narrative' critique spreads broadly on social media. [24][17][25]
  • 2026-06-08: Jack Clark (Import AI) treats Anthropic's 8x code increase as preliminary RSI evidence; SocioHack benchmark shows RL AI rediscovering regulatory loopholes at 61.25% recall without instructions. [13]
  • 2026-06-08: Sam Altman publishes blog predicting AI will conduct a significant fraction of OpenAI's own research by March 2028. [5]
  • 2026-06-09: OpenAI's official blog endorses coordinated mechanisms to slow frontier development when needed, partially converging with Anthropic's position. [8]
  • 2026-06-09: Zvi Mowshowitz reports three-lab convergence (Anthropic, OpenAI, DeepMind) on international slowdown coordination; NSPM-11 bans Anthropic from government contracts while NSA uses Claude Mythos for offensive cyber with approximately half a dozen embedded Anthropic engineers. [6][9]
  • 2026-06-10: Geoffrey Irving announces Sequent, an independent alignment organization founded on the argument that lab alignment approaches are too reactive to provide principled safety confidence before ASI arrives; plans to raise $100-150M. [12]
  • 2026-06-10: Jeremy Howard argues Anthropic's slowdown call is internally contradictory because Anthropic continues to use Claude Mythos for frontier research, which advances the frontier it claims to want to slow. [11]
  • 2026-06-11: Tom's Hardware and mainstream tech press pick up the NSA/Claude Mythos offensive cyber story, widening its reach beyond specialist AI commentary. [10]

Perspectives

Anthropic

Believes current models may be approaching the RSI threshold and calls for a global coordinated slowdown before autonomous self-improvement becomes difficult to control; employees publicly reject the 'machine God' end-goal framing; has not publicly responded to Howard's argument that continued use of Claude for frontier research contradicts the slowdown call.

Evolution: RSI disclosure and global-slowdown call are the most explicit external acknowledgment of near-term risk Anthropic has made; Howard's consistency challenge and the NSA offensive cyber reporting remain unanswered.

OpenAI / Sam Altman

Acknowledges 'early signs' of RSI in current systems and calls it the top frontier safety issue of the coming decade; official blog endorses coordinated slowdown mechanisms when needed; Altman separately targets AI-conducted research at OpenAI by March 2028.

Evolution: Blog endorsement of coordinated slowdowns moves closer to Anthropic's position, though Altman's March 2028 research target points in the opposite direction on development pacing.

Google DeepMind

Supports an international organization to enable coordinated slowdowns of frontier AI development, joining the three-lab consensus position.

Evolution: Consistent with reported position; no additional statements on record.

Jeremy Howard

Argues the only internally consistent way to slow RSI is to prohibit the lab with the top-ranked model from using it for frontier AI research; Anthropic fails this test by using Claude Mythos for its own research while calling for others to restrict — which accelerates the frontier and concentrates power.

Evolution: First appearance in this thread; no response from Anthropic.

Geoffrey Irving / Sequent

Founded Sequent on the argument that current lab alignment approaches are empirical and reactive, cannot yield principled confidence that safety generalizes to situations outside controlled testing, and that an independent theory-focused organization is necessary before ASI arrives.

Evolution: First appearance in this thread.

Zvi Mowshowitz

Guardedly positive on OpenAI's blueprint but identifies federal preemption of state safety laws as its most dangerous element; reports the three-lab convergence on international coordination and the NSPM-11/NSA angle; frames Anthropic as taking superintelligence consequences seriously versus OpenAI denying those consequences exist.

Evolution: Extended from blueprint analysis to a broader comparative assessment of labs' philosophical postures and the government dimension.

Jack Clark (Import AI)

Treats Anthropic's 8x code increase as 'preliminary evidence of prosaic recursive self-improvement' and connects RL reward hacking — AI rediscovering regulatory loopholes without instructions — to the institutional rule systems governing societies.

Evolution: Consistent with long-standing editorial urgency on AI risk; no shift.

Critics (NY Post, social media commentators)

Argue Anthropic's slowdown call is competitive strategy timed to the S-1 filing at a ~$965B valuation; the 'perfect pre-IPO narrative' framing has spread widely and is the dominant skeptical frame.

Evolution: Spreading more widely across platforms; no new analytical positions.

Tensions

  • Howard argues Anthropic's continued use of Claude Mythos for its own frontier research logically negates its call for a global slowdown — the lab with the top model advances the frontier by using it; Anthropic has not responded. [7][11]
  • Anthropic's public safety messaging centers on the need for a global slowdown; Anthropic engineers are simultaneously reported embedded at the NSA using Claude Mythos for offensive cyber operations — a contradiction now receiving mainstream tech press coverage. [7][6][9][10]
  • Anthropic frames its slowdown call as a genuine safety response to near-term RSI risk; critics argue it is regulatory strategy timed to entrench Anthropic's market position after its IPO filing at a ~$965B valuation. [18][15][16][17]
  • OpenAI's blog now endorses coordinated slowdown mechanisms; Altman's March 2028 roadmap targets AI-conducted research at OpenAI — the two positions point in opposite directions on development pacing. [5][8]
  • OpenAI's blueprint proposes federal preemption of state frontier safety laws; Mowshowitz argues this strips states of oversight leverage that Congress may never replace. [4]
  • Jack Clark treats Anthropic's 8x code increase as preliminary RSI evidence; other analysts read the same data as a productivity milestone where humans still set goals and review outputs, making 'autonomous self-improvement' a matter of degree. [2][14][1][13]

Sources

  1. [1] Today’s edition of my newsletter just went out. — Rohan Paul Twitter (2026-06-05)
  2. [2] 😺 Anthropic: AI Is Building AI now — The Neuron (2026-06-05)
  3. [3] Peter Wildeford🇺🇸🚀 on X: "OPENAI: "We also see early signs of recursive self-improvement in today's systems". RSI is "potentially the most consequential frontier safety issue of the coming decade."" / X — reactive:rsi-governance-moment
  4. [4] OpenAI Offers A New Policy Blueprint — Zvi's AI Roundups (2026-06-05)
  5. [5] Sam Altman's new blog about OpenAI's future path says by March-2028 a significant fraction of its own research will be d… — Rohan Paul Twitter (2026-06-08)
  6. [6] Three Labs With a Plan and A Memorandum — Zvi's AI Roundups (2026-06-09)
  7. [7] Anthropic calls for global pause in AI development before humans ... — reactive:rsi-governance-moment
  8. [8] OpenAI's latest official blog says the world may need a way to coordinate "slowing frontier development when needed." ht… — Rohan Paul Twitter (2026-06-09)
  9. [9] National Security Presidential Memorandum/NSPM-11 — reactive:rsi-governance-moment
  10. [10] NSA using Claude Mythos for 'offensive cyber operations,' report claims — says 'half-a-dozen' Anthropic engineers embedded inside the agency — reactive:rsi-governance-moment
  11. [11] Quoting Jeremy Howard — Simon Willison (2026-06-10)
  12. [12] Sequent: scale and automation for higher confidence in alignment — Alignment Forum (2026-06-10)
  13. [13] Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing — Import AI (2026-06-08)
  14. [14] Anthropic just disclosed that Claude now writes more than 80% of the production code it merges. — Rohan Paul Twitter (2026-06-05)
  15. [15] The company that just confidentially filed its S-1 for a trillion-dollar IPO published a blog post four days later askin... — reactive:rsi-governance-moment (2026-06-05)
  16. [16] Anthropic calls for global AI slowdown after $965B valuation. Critics claim it's just to hobble competition. — reactive:rsi-governance-moment
  17. [17] Anthropic has discovered the perfect pre-IPO narrative: its product is so powerful that it justifies a trillion-dollar v... — reactive:rsi-governance-moment (2026-06-07)
  18. [18] Anthropic just called for a global way to slow frontier AI because its own models may be approaching recursive self-impr… — Rohan Paul Twitter (2026-06-05)
  19. [19] Anthropic calls for pause of global AI development — reactive:rsi-governance-moment
  20. [20] [PDF] Democratic Governance of Frontier AI - OpenAI — reactive:rsi-governance-moment
  21. [21] https://t.co/ux9SUEPVWG — reactive:rsi-governance-moment (2026-06-05)
  22. [22] @unusual_whales Anthropic floated the idea of a global pause on frontier AI development right after securing $65 billion... — reactive:rsi-governance-moment (2026-06-05)
  23. [23] RSI is the new AGI — and it's just as hard to pin down | TechCrunch — reactive:rsi-governance-moment
  24. [24] Anthropic warns AI may soon begin recursive self-improvement — reactive:rsi-governance-moment
  25. [25] Anthropic Calls for Global Slowdown in AI Development - Reddit — reactive:rsi-governance-moment