AI Safety Advocacy Splits on US-China Cooperation vs. Domestic Controls · history
Version 3
2026-07-15 18:16 UTC · 33 items
What
The AI safety governance debate has extended into questions of whether stated ethical commitments hold under financial and political pressure. TurnTrout, a researcher who left Google DeepMind, documents the organization signing a classified Pentagon deal permitting 'any lawful government purpose' without restrictions on autonomous weapons or mass surveillance — while CEO Demis Hassabis publicly claimed 'nothing's changed about our principles' [8]. This contrasts directly with Anthropic ending Pentagon negotiations in February 2026 over the same issues [7]. OpenAI separately published advocacy for 'reverse federalism,' arguing California, New York, and Illinois AI safety laws should converge into a de facto national standard ahead of formal federal legislation, with the Trump administration targeting early August 2026 for a federal model-testing framework [9]. These developments sit alongside an ongoing split between Plan A's cooperative US-China governance approach and the domestic competitive controls — export restrictions and a possible open-weight model ban — favored by the current administration [2][6].
Why it matters
TurnTrout's account introduces empirical evidence that AI safety commitments can be abandoned under institutional pressure, which bears directly on whether voluntary industry governance or formal binding mechanisms are viable. If OpenAI's reverse federalism strategy takes hold, three states may set the practical terms of frontier AI safety regulation before any federal framework exists.
Open questions
Will the White House issue an executive order banning frontier-capability open-weight models within six months, and would it apply to non-US releases? [6]
Does Google DeepMind's classified Pentagon deal represent an industry-wide pattern of safety pledges weakening under financial pressure, or an outlier compared to Anthropic's maintained redlines? [8][7]
Can Plan A secure meaningful Chinese participation and verifiable compliance given current US-China strategic competition? [2]
Will OpenAI's 'reverse federalism' strategy produce coherent national AI safety standards before the Trump administration's August 2026 federal framework target is met? [9]
Narrative
The central debate in AI safety governance is between advocates of international cooperative controls and those pursuing domestic competitive strategies. The AI Futures Project's Plan A proposes that the US and China jointly control chip supply, audit data centers, and share research to slow superintelligent AI development [1]. Zvi Mowshowitz argues Plan A deserves serious engagement rather than dismissal, with the key dividing question being whether superintelligence arrives soon enough to justify Plan A's costs [2]. Vitalik Buterin defends Plan A against critics who call it naive, arguing they apply coordination skepticism to a cooperative pause but not to the assumption that an AI transition will simply go well by default — where humanity's hard power drops to zero if AIs outperform humans at every task [2]. The Trump administration has not engaged with the cooperative frame, governing AI through competitiveness: a January 2025 policy framework, export controls on model weights effective January 2026, preemption of state AI regulations, and a reported White House discussion of an executive order to ban open-weight frontier models outright [3][4][5][6].
The question of whether AI safety commitments hold under institutional pressure has taken concrete form. Anthropic ended Pentagon contract negotiations in February 2026 after the Defense Department demanded blanket 'anything lawful' usage rights, with Dario Amodei citing irreconcilable conflicts with Anthropic's redlines on mass surveillance and autonomous weapons targeting [7]. Google DeepMind took the opposite path: a former researcher writing as TurnTrout published a detailed account of leaving the organization after Google signed a classified Pentagon deal permitting 'any lawful government purpose' without binding restrictions on autonomous weapons or mass surveillance [8]. TurnTrout documents CEO Demis Hassabis removing specific prohibitions from Google's 2018 AI principles while publicly claiming 'nothing's changed about our principles,' and reports that the IASEAI publicly promised a member poll supporting Anthropic's stance but silently cancelled it without explanation [8]. The piece frames the divergence plainly: 'When profit and pressure met ethical commitment at Google DeepMind, pressure won and pledges lost. When profit and pressure met ethical commitment at Anthropic, ethics won' [8].
A separate domestic debate concerns open-weight frontier models. Nathan Lambert reports White House discussions of an executive order that could ban open-weight models above current frontier capability levels, and accuses Anthropic's campaign against Chinese model distillation of constituting regulatory capture — noting Anthropic would gain substantial commercial benefit if the targeted Chinese model makers were banned [6]. Lambert disputes the security rationale by pointing to Anthropic's Mythos model being accessed through unauthorized Discord channels during private beta, suggesting APIs are not meaningfully more secure than open weights [6]. OpenAI, for its part, published advocacy for 'reverse federalism': state-level AI safety laws in California, New York, and Illinois — requiring documented safety frameworks, public risk assessments, incident reporting, and independent audits — positioned as converging toward a de facto national standard ahead of federal legislation, with the Trump administration separately targeting early August 2026 for a cyber-evaluation framework [9]. A parallel argument on the Alignment Forum holds that political will, not research, is the main AI safety bottleneck, with a 3.6:1 researcher-to-advocate ratio in the US AI safety field and AI companies securing seven times as many European Commission meetings as civil society in 2023 [10].
Timeline
- 2025-01-01: Trump administration releases national AI policy framework centering competitiveness rather than safety regulation as the governing principle. [3][11]
- 2025-12-01: Trump signs executive order preempting state AI regulations to create a unified federal policy framework. [5]
- 2026-01-09: US model weight export controls take effect, prompting commentary that the US had become the world's most aggressive AI regulator on this issue. [4]
- 2026-02-26: Anthropic ends Pentagon contract negotiations after the Defense Department insists on blanket 'anything lawful' usage rights, citing irreconcilable conflicts with redlines on mass surveillance and autonomous weapons. [7]
- 2026-07-10: AI Futures Project publishes Plan A — a US-China cooperative pause on frontier AI — drawing attention as a safety-compatible growth alternative to domestic controls. [1]
- 2026-07-10: Zvi Mowshowitz publishes 'Plan B' analysis concluding the Trump administration will govern AI through ad hoc executive authority rather than formal licensing. [7]
- 2026-07-11: Zvi Mowshowitz publishes analysis of Plan A reactions including Vitalik Buterin's defense and Ryan Greenblatt's cautious endorsement. [2]
- 2026-07-11: Alignment Forum post argues political will — not research — is the main AI safety bottleneck, citing a 3.6:1 researcher-to-advocate ratio and AI industry's seven-to-one advantage in EU Commission meetings over civil society. [10]
- 2026-07-12: Nathan Lambert reports White House discussions of an executive order to ban frontier-capability open-weight models and accuses Anthropic of regulatory capture in its campaign against Chinese model distillation. [6]
- 2026-07-15: TurnTrout publishes account of leaving Google DeepMind over a classified Pentagon deal permitting unrestricted military AI use, documenting that Google's CEO removed safety prohibitions from its stated principles while publicly denying any change. [8]
- 2026-07-15: OpenAI publishes advocacy for 'reverse federalism,' supporting state AI safety laws in California, New York, and Illinois as a path to national standards ahead of formal federal legislation. [9]
Perspectives
AI Futures Project (Daniel Kokotajlo)
Advocates a US-China cooperative pause on frontier AI via joint chip supply controls, data center audits, and research sharing; frames safety and growth as compatible rather than competing goals.
Evolution: Plan A has moved from initial publication to generating substantive public debate, with named endorsements from Vitalik Buterin and Ryan Greenblatt.
Vitalik Buterin
Defends Plan A against critics who call it naive, arguing they apply coordination skepticism to a cooperative pause but not to the alternative — an unmanaged AI transition that concentrates power and eliminates human agency.
Evolution: Publicly aligned with Plan A's core premise; consistent in this thread.
Zvi Mowshowitz
Not endorsing Plan A but argues it deserves serious engagement; frames the key crux as whether superintelligence arrives soon enough to justify Plan A's costs; concludes the Trump administration will govern AI through ad hoc executive authority rather than formal licensing.
Evolution: Shifted from purely skeptical analyst to engaged interlocutor stress-testing Plan A's premises rather than dismissing them outright.
Anthropic (Dario Amodei)
Holds hard redlines against mass surveillance and autonomous weapons targeting; ended Pentagon negotiations rather than waive them; has publicly called for a coordinated, verifiable pause on frontier AI development.
Evolution: Now publicly advocating a coordinated pause while also campaigning against Chinese model distillation — a combination Nathan Lambert characterizes as regulatory capture. TurnTrout's account positions Anthropic as the institutional counterexample to Google's capitulation.
TurnTrout (former Google DeepMind researcher)
Argues AI safety pledges without binding enforcement are structurally inadequate; documents Google dropping its ethics principles under financial and political pressure while claiming otherwise; characterizes IASEAI as having failed to act when it counted.
Evolution: New voice in this thread; introduces empirical whistleblower evidence into an otherwise theoretical debate about whether voluntary commitments are sufficient.
Nathan Lambert / Open-Weight Advocates
Strongly opposed to any ban on open-weight frontier models; accuses Anthropic's anti-distillation campaign of regulatory capture; argues a unilateral US ban would be ineffective and that open models improve safety through broad access.
Evolution: Open-weight advocacy has sharpened from a diffuse position into a specific alarm about imminent executive action with a named target — Anthropic's lobbying.
OpenAI
Advocates 'reverse federalism' — state AI safety laws in California, New York, and Illinois converging into a de facto national standard — while supporting CAISI as durable federal evaluation capacity; warns against regulatory fragmentation and mission creep.
Evolution: New voice in this thread; takes a constructive-partner position distinct from both the cooperative international frame and the domestic-controls frame.
Trump Administration
Frames AI governance around US competitiveness; preempted state regulations; implemented model weight export controls; declined to build a formal licensing regime; reportedly in discussions about banning frontier-capability open-weight models; targeting early August 2026 for a federal model-testing framework focused on cyber evaluations.
Evolution: Reportedly moving toward more aggressive domestic regulatory action on open weights, while also building a federal testing framework that runs counter to OpenAI's preferred state-convergence timeline.
Tensions
- TurnTrout argues Google DeepMind dropped its AI ethics principles under financial and political pressure while Anthropic held its redlines; Google DeepMind CEO Demis Hassabis has publicly claimed 'nothing's changed about our principles.' [8][7]
- Nathan Lambert argues Anthropic's campaign against Chinese model distillation is regulatory capture serving commercial interests; Anthropic frames the same campaign as a legitimate safety concern about frontier-capability proliferation. [6][7]
- Plan A proponents and Vitalik Buterin argue a US-China cooperative pause is the necessary safety mechanism; the Trump administration treats China as a strategic competitor to contain through export controls, not a partner in cooperative governance. [2][1][3]
- Open-weight advocates argue frontier model weights should be publicly released and open access improves safety; the US government treats frontier open weights as a credible national security risk and is reportedly considering an executive order to ban them. [6][4]
- OpenAI argues state AI safety laws should converge into a de facto national standard ahead of federal legislation; the Trump administration preempted state AI regulations and is building its own federal framework targeting August 2026. [9][5]
- TurnTrout and Charbel-Raphaël argue voluntary pledges and research investments are insufficient without binding enforcement and political will; the mainstream AI safety field has historically prioritized technical research over advocacy and institutional accountability. [8][10]
Sources
- [1] 🟡 AI doom and bloom — Semafor Technology (2026-07-10)
- [2] Introduction for and Reactions to Plan A — Zvi's AI Roundups (2026-07-11)
- [3] Trump Administration Releases National AI Policy ... — reactive:ai-safety-governance-proposals
- [4] Ben Brooks on X: "Effective today, model weights are export controlled by Uncle Sam. This is a big deal. For all the smack talk about the EU, the US is now the world's most aggressive regulator of Expensive Maths. Here's my two cents on the model rule based on the released text (link below)." / X — reactive:ai-safety-governance-proposals
- [5] President Trump signs order attempting to block A.I. regulations at the state level — reactive:ai-safety-governance-proposals
- [6] 6 months to live for open models — Interconnects (2026-07-12)
- [7] AI #176 Part 2: Plan B — Zvi's AI Roundups (2026-07-10)
- [8] Why I Left Google DeepMind — Alignment Forum (2026-07-15)
- [9] The US is advancing AI safety through state and federal action — OpenAI Blog (2026-07-15)
- [10] The current bottleneck is political will, not research — Alignment Forum (2026-07-11)
- [11] Artificial Intelligence for the American People — reactive:ai-safety-governance-proposals