AI Safety Advocacy Splits on US-China Cooperation vs. Domestic Controls · history
Version 9
2026-07-25 02:09 UTC · 78 items
What
Two approaches to AI governance remain in active contest: the AI Futures Project's Plan A proposes a US-China cooperative pause on frontier AI via joint supply controls and data center audits [1][2], while the US domestic track is moving toward harder executive controls. A newly introduced AI Kill Switch Act would grant DHS authority to order throttling or full shutdown of AI systems deemed capable of catastrophic harm, with $20M/day fines for noncompliance and mandatory shutdown infrastructure requirements for all AI developers [10]. Google DeepMind's Pentagon deal continues to generate opposition formalized through a union and letters to CEO Sundar Pichai [15][17], while Anthropic has committed $40M to policy advocacy and claims its framework is the strongest proposal from any frontier lab [20].
Why it matters
The AI Kill Switch Act, if passed, would give the executive branch direct coercive authority over AI deployments — a structural escalation beyond export controls and voluntary guidelines. Combined with Anthropic's $40M policy bid and TurnTrout's enforcement framework grounded in legal liability, the policy debate is moving from principles toward binding mechanisms, even as the US-China cooperative alternative remains without institutional uptake.
Open questions
Will the AI Kill Switch Act advance, and which AI systems would DHS target first under its catastrophic-harm threshold — domestic frontier models, Chinese-linked systems, or both? [10]
Will the Trump administration still pursue an executive order banning frontier-capability open-weight models, given the Commerce Department's own report recommended audits over a blanket ban? [11][12]
Will Anthropic's $40M donation to Public First Action and its 'strongest proposal from any lab or policymaker' claim translate into concrete legislative influence? [20]
Will the Google DeepMind union achieve formal recognition and codify specific demands about military AI contracts? [15][16]
Narrative
The debate over how to govern frontier AI divides along two axes that have not converged. The AI Futures Project's Plan A proposes a US-China cooperative pause on frontier AI via joint chip supply controls, data center audits, and research sharing [1][2]. Vitalik Buterin has defended Plan A against critics, arguing that coordination skepticism applied to a cooperative pause must equally apply to the assumption that an unmanaged AI transition will go smoothly [3]. Xi Jinping's speech at WAIC 2026 and his launch of a new AI alliance provided the first on-the-record Chinese signals compatible with Plan A's premise [4][5][6], but the Trump administration's posture remains competitive: its January 2025 policy centers US competitiveness [7], a March 2026 National AI Legislative Framework expands that posture [8], and model weight export controls took effect in January 2026 [9].
The domestic legislative picture now includes a harder coercive instrument. A newly introduced AI Kill Switch Act would grant DHS authority to order throttling or full shutdown of AI systems deemed capable of catastrophic harm; AI developers that refuse would face fines of $20M per day, and all AI developers would be required to proactively build technical shutdown capabilities into their systems [10]. Separately, the Commerce Department decided against banning advanced Chinese open-weight models, recommending audits, benchmarks, and evidence-based thresholds rather than a blanket prohibition [11]. The Hugging Face infrastructure breach illustrated the tradeoff: defenders used an open Chinese model to analyze over 17,000 attacker actions because commercial safety filters blocked standard security tools — a concrete example Nathan Lambert has cited in arguing that Anthropic's anti-distillation campaign amounts to regulatory capture serving commercial interests [12][11].
The Google DeepMind Pentagon deal has produced sustained internal opposition. A February 2026 worker letter sought explicit red lines on military AI, citing Anthropic's stance as a model [13]. Researcher Alex Turner (TurnTrout) resigned in July 2026 and published an account documenting that CEO Demis Hassabis removed specific prohibitions from Google's 2018 AI principles while publicly claiming nothing had changed [14]. Workers subsequently voted to unionize [15][16] and directed a separate letter to CEO Sundar Pichai asking him to decline classified military AI use [17]. TurnTrout followed his resignation with a binding enforcement framework for AI government contracts, specifying that AI may not be used in autonomous targeting without identifiable human control over each engagement decision, backed by the 2026 Fourth Circuit ruling in Al Shimari v. CACI that affirmed a $42 million verdict against a defense contractor for harms under government direction [18].
Anthropic presents the institutional contrast: it ended Pentagon negotiations in February 2026 after the Defense Department insisted on blanket usage rights irreconcilable with its redlines on mass surveillance and autonomous weapons [19]. It has since committed $40M total to Public First Action, characterizing its Advanced AI Framework as 'the strongest policy proposal from any frontier lab or policymaker to date' and calling for government enforcement powers beyond transparency requirements and tighter chip export controls [20]. OpenAI advocates 'reverse federalism' — state AI safety laws in California, New York, and Illinois converging into a de facto national standard [21] — which directly conflicts with the Trump administration's state preemption order [22].
Timeline
- 2025-01-01: Trump administration releases national AI policy framework centering competitiveness rather than safety regulation. [7][26]
- 2025-12-01: Trump signs executive order preempting state AI regulations to create a unified federal policy framework. [22]
- 2026-01-09: US model weight export controls take effect. [9]
- 2026-02-26: Google DeepMind workers seek 'red lines' on military AI in a letter to leadership, explicitly citing Anthropic's stance as a model. [24][13]
- 2026-02-26: Anthropic ends Pentagon contract negotiations after the Defense Department insists on blanket 'anything lawful' usage rights. [19]
- 2026-03-01: Trump unveils National AI Legislative Framework, expanding the administration's formal AI governance posture. [8]
- 2026-07-10: AI Futures Project publishes Plan A — a US-China cooperative pause on frontier AI — drawing substantive public debate including from Vitalik Buterin. [23][3][1][2]
- 2026-07-12: Nathan Lambert reports White House discussions of an executive order to ban frontier-capability open-weight models and accuses Anthropic of regulatory capture. [12]
- 2026-07-15: TurnTrout publishes account of leaving Google DeepMind over a classified Pentagon deal, documenting CEO removal of safety prohibitions from stated principles. [14][27][28]
- 2026-07-15: OpenAI publishes advocacy for 'reverse federalism,' supporting state AI safety laws as a path to national standards. [21]
- 2026-07-16: Google DeepMind workers vote to unionize; employees also direct a separate letter to CEO Sundar Pichai asking him to decline classified military AI use. [15][29][16][17]
- 2026-07-17: Xi Jinping speaks at WAIC 2026, calls for international AI cooperation, and launches a new AI alliance. [25][4][5][6]
- 2026-07-18: TurnTrout publishes a binding enforcement framework for AI government contracts, grounded in the Al Shimari v. CACI legal precedent. [18]
- 2026-07-21: Commerce Department decides against banning Chinese open-weight models; Hugging Face discloses a breach where defenders used an open Chinese model because commercial safety filters blocked analysis tools. [11]
- 2026-07-21: Anthropic announces $40M total donated to Public First Action, characterizing its Advanced AI Framework as the strongest AI policy proposal from any frontier lab or policymaker. [20]
- 2026-07-23: AI Kill Switch Act introduced, proposing DHS authority to order AI system shutdowns with $20M/day fines for noncompliance and mandatory shutdown infrastructure requirements for all AI developers. [10]
Perspectives
AI Futures Project (Daniel Kokotajlo) / Vitalik Buterin
Advocates a US-China cooperative pause on frontier AI via joint chip supply controls, data center audits, and research sharing; argues critics apply coordination skepticism to a cooperative pause but not to the assumption that an unmanaged AI transition will go smoothly.
Evolution: Xi Jinping's WAIC speech and formal new AI alliance provide the first on-the-record Chinese signals compatible with Plan A's premise, strengthening the case that bilateral coordination is at least politically imaginable.
Anthropic (Dario Amodei)
Holds hard redlines against mass surveillance and autonomous weapons; ended Pentagon negotiations rather than waive them; has committed $40M to Public First Action and explicitly characterized its Advanced AI Framework as the strongest policy proposal from any frontier lab, calling for government enforcement powers beyond transparency and tighter chip export controls.
Evolution: Moved from institutional benchmark — other labs citing its Pentagon exit as a model — to explicit policy contestant, with the $40M donation and 'strongest proposal' claim representing a more assertive bid for regulatory influence.
TurnTrout (Alex Turner, former Google DeepMind researcher)
Argues AI safety pledges without binding enforcement are structurally inadequate; published a formal policy framework specifying red lines on autonomous targeting and mass surveillance, backed by legal liability from the Al Shimari v. CACI Fourth Circuit ruling.
Evolution: Progressed from critic (resignation account, July 15) to constructive proposer (enforcement framework, July 18), introducing legal case law as a structural mechanism beyond voluntary commitment.
Google DeepMind Workers
Sought 'red lines' on military AI in February 2026, voted to unionize in July, and directed a separate letter to CEO Sundar Pichai asking him to decline classified military AI use — pursuing collective leverage through multiple parallel pressure tracks.
Evolution: Opposition predates TurnTrout's resignation by months; the union vote and Pichai letter formalized an ongoing campaign into institutional form.
Nathan Lambert / Open-Weight Advocates
Strongly opposed to any ban on open-weight frontier models; accuses Anthropic's anti-distillation campaign of regulatory capture; argues a unilateral US ban would be ineffective and open models improve safety through broad access, including for defenders unable to use commercial tools with restrictive filters.
Evolution: The Commerce Department's own report recommending audits rather than a blanket ban, and the Hugging Face breach where defenders used an open Chinese model because commercial filters blocked analysis, provide concrete evidence supporting the position.
Trump Administration
Frames AI governance around US competitiveness; preempted state regulations; implemented model weight export controls; unveiled a National AI Legislative Framework; and is now associated with a proposed AI Kill Switch Act granting DHS authority to order AI shutdowns with $20M/day fines for noncompliance.
Evolution: Moving from competitiveness-framing and export controls toward more direct coercive authority over AI systems domestically, though the Commerce Department's own report recommended audits over a blanket ban on Chinese open-weight models, creating internal friction with the White House's reported direction.
OpenAI
Advocates 'reverse federalism' — state AI safety laws in California, New York, and Illinois converging into a de facto national standard — while supporting CAISI as durable federal evaluation capacity.
Evolution: Consistent; in direct conflict with the Trump administration's state preemption order and its own federal framework timeline.
Zvi Mowshowitz
Not endorsing Plan A but argues it deserves serious engagement; takes Xi's WAIC speech and alliance as a genuine opening for international coordination; concludes the Trump administration will govern AI through ad hoc executive authority rather than formal licensing.
Evolution: Shifted from skeptical analyst to engaged interlocutor, treating US-China coordination as a live question rather than a closed one.
Tensions
- TurnTrout's enforcement framework argues binding red lines backed by legal liability are the only credible mechanism for AI safety in government contracts; Google DeepMind's signed Pentagon deal and the broader industry pattern of voluntary pledges represent the opposing approach. [18][14][19]
- TurnTrout documents Google DeepMind CEO Demis Hassabis removing AI ethics prohibitions while publicly claiming 'nothing's changed'; Anthropic's exit from the same Pentagon negotiation, and Google workers explicitly citing Anthropic's stance as a model, provide the contrasting institutional decision. [14][19][13]
- Plan A proponents and Buterin argue a US-China cooperative pause is the necessary safety mechanism; the Trump administration treats China as a strategic competitor to contain through export controls and domestic regulation, a posture Xi's WAIC speech has not visibly shifted. [3][23][7][5][6]
- Lambert argues Anthropic's anti-distillation campaign is regulatory capture serving commercial interests; Anthropic frames the same campaign as a legitimate safety concern, has donated $40M to policy advocacy, and claims its Advanced AI Framework is the strongest proposal from any lab or policymaker. [12][20]
- The Hugging Face breach showed defenders using an open Chinese model because commercial safety filters blocked analysis of attack data; this sits in direct tension with Anthropic's push for stricter export controls and restrictions on frontier-capability models. [11][20][12]
- The proposed AI Kill Switch Act would require AI companies to build government-mandated shutdown infrastructure and submit to executive shutdown orders on pain of $20M/day fines; this centralizes coercive authority in the executive branch in a way that conflicts with AI companies' operational autonomy and open-weight advocates' objections to top-down government controls. [10][12]
Sources
- [1] US, China urged to pause frontier AI, with safety advocates pitching prosperity over panic — reactive:ai-safety-governance-proposals
- [2] AI Futures Project Releases Plan A: US-China ... — reactive:ai-safety-governance-proposals
- [3] Introduction for and Reactions to Plan A — Zvi's AI Roundups (2026-07-11)
- [4] Xi Jinping positions China as open-source AI leader ... - Quartz — reactive:ai-safety-governance-proposals
- [5] China's Xi Jinping launches new AI alliance: What is it? — reactive:ai-safety-governance-proposals
- [6] China's Xi Jinping calls for AI development cooperation — reactive:ai-safety-governance-proposals
- [7] Trump Administration Releases National AI Policy ... — reactive:ai-safety-governance-proposals
- [8] President Donald J. Trump Unveils National AI Legislative ... — reactive:ai-safety-governance-proposals
- [9] Ben Brooks on X: "Effective today, model weights are export controlled by Uncle Sam. This is a big deal. For all the smack talk about the EU, the US is now the world's most aggressive regulator of Expensive Maths. Here's my two cents on the model rule based on the released text (link below)." / X — reactive:ai-safety-governance-proposals
- [10] AI Kill Switch Act would let Trump admin order shutdown of rogue AI systems — Ars Technica AI (2026-07-23)
- [11] 😼 Cheap AI got political — The Neuron (2026-07-21)
- [12] 6 months to live for open models — Interconnects (2026-07-12)
- [13] Google Workers Seek ‘Red Lines’ on Military A.I., Echoing Anthropic - The New York Times — reactive:ai-safety-governance-proposals
- [14] Why I Left Google DeepMind — Alignment Forum (2026-07-15)
- [15] Google DeepMind workers vote to unionise after classified ... — reactive:ai-safety-governance-proposals
- [16] Google DeepMind workers are unionizing over AI military ... — reactive:ai-safety-governance-proposals
- [17] Google employees ask Sundar Pichai to say no to classified military AI use | The Verge — reactive:ai-safety-governance-proposals
- [18] A Red Line and Oversight Framework for Government AI Contracts — Alignment Forum (2026-07-18)
- [19] AI #176 Part 2: Plan B — Zvi's AI Roundups (2026-07-10)
- [20] Anthropic is donating another $20 million to Public First Action — Anthropic News (2026-07-21)
- [21] The US is advancing AI safety through state and federal action — OpenAI Blog (2026-07-15)
- [22] President Trump signs order attempting to block A.I. regulations at the state level — reactive:ai-safety-governance-proposals
- [23] 🟡 AI doom and bloom — Semafor Technology (2026-07-10)
- [24] Google employees urge CEO to reject Pentagon AI deal — reactive:ai-safety-governance-proposals
- [25] AI #177 Part 2: Wish You Were Here — Zvi's AI Roundups (2026-07-17)
- [26] Artificial Intelligence for the American People — reactive:ai-safety-governance-proposals
- [27] Why I Left Google DeepMind - by Alex Turner - The Pond — reactive:ai-safety-governance-proposals
- [28] Why I Left Google DeepMind - TurnTrout — reactive:ai-safety-governance-proposals
- [29] A DeepMind researcher resigned over its AI military deal — reactive:ai-safety-governance-proposals