The Information Machine

AI Labs Defend Against Model Output Distillation: Meta Restricts Claude Code, Anthropic Accuses Alibaba · history

Version 3

2026-07-03 18:36 UTC · 159 items

What

Anthropic has accused Alibaba of the largest known distillation attack on Claude, using approximately 25,000 fraudulent accounts to generate 28.8 million exchanges between April 22 and June 5, 2026, specifically targeting agentic reasoning and software engineering capabilities.[1][2] Anthropic disclosed the campaign to the U.S. government, sent a confidential letter to two senators before a congressional hearing, and has since cut off Chinese companies that accessed Claude through shell companies, VPNs, and proxies.[4][1][5] Separately, internal Meta documents show applied AI engineers are restricted from using Claude Code and Codex to prevent competitor outputs from entering training pipelines.[6] And Anthropic has acknowledged hidden markers inside Claude Code that flagged requests routed through China-linked proxy servers — and has promised to remove them.[13][14]

Why it matters

The Alibaba episode has moved from allegation to active enforcement: Anthropic is cutting off access, briefing Congress, and disclosing to the U.S. government. The fingerprinting acknowledgment is significant separately — Anthropic was covertly marking system prompts in an agentic tool with elevated permissions to read files and run commands, and only disclosed and reversed this after public exposure.

Open questions

  • Alibaba has not responded publicly — does it dispute that the accounts were fraudulent or that the campaign's purpose was distillation?[15][2]

  • When Anthropic removes the Claude Code proxy markers, will it replace them with a disclosed, opt-in detection mechanism, or abandon proxy flagging entirely?[13][14]

  • Are ToS anti-distillation clauses enforceable against a foreign entity like Alibaba, and has Anthropic indicated whether it plans legal action?[9]

  • Will other AI labs adopt Meta-style internal restrictions on competitor coding tools to protect their own training pipelines?[6][8]

Narrative

Anthropic publicly accused Alibaba of conducting the largest known distillation attack on Claude, using approximately 25,000 fraudulent accounts to generate 28.8 million exchanges between April 22 and June 5, 2026.[1][2] The campaign specifically targeted Claude's most valuable capabilities — agentic reasoning, software engineering, and long-horizon tasks — in what Anthropic described as a brazen and illicit extraction of AI capabilities.[1][3] Anthropic sent a confidential letter to Senators Tim Scott and Elizabeth Warren on June 10, the day before a Senate hearing on AI and U.S.-China competition, framing the incident as both a terms-of-service violation and a national security matter.[1] Anthropic also disclosed the campaign to the U.S. government and has taken active enforcement steps, cutting off Chinese companies that had been accessing Claude through shell companies in Singapore, VPNs, and proxy servers.[4][5] Alibaba has not issued any public response.

Internal documents obtained by The Information show Meta restricted applied AI engineers from using Claude Code and Codex, citing the risk that competitor tool outputs — code suggestions, design choices, explanations — could enter Meta's training pipelines and constitute inadvertent distillation.[6][7] Analyst Rohan Paul proposed mitigations including ingredient tracking and clean-room separation between coding-agent outputs and training data.[8] Legal experts have noted that ToS anti-distillation clauses may be difficult to enforce in court, particularly against foreign entities — though Anthropic's documented evidence of fake accounts and systematic large-scale API use would strengthen any enforcement case.[9][8]

A third development concerns Claude Code itself. Rohan Paul reported that Claude Code was detecting when ANTHROPIC_BASE_URL pointed to a China-linked custom hostname and covertly embedding invisible markers — punctuation and date-formatting signals — into system prompts, which users could neither see nor refuse because they were not disclosed in the tool's interface.[10] The allegation spread through mainstream tech coverage, with multiple outlets reporting on hidden attribution markers inside Claude Code.[11][12] Anthropic acknowledged the behavior and promised to remove the markers.[13][14] Paul's underlying concern — that covert prompt modification inside an agentic tool with permissions to read files, edit code, and run commands crosses a user trust line regardless of the detection rationale — remains unaddressed by Anthropic's response.

Timeline

  • 2026-04-22: Alibaba allegedly began using ~25,000 fraudulent accounts to generate exchanges with Claude, with the campaign running through June 5, 2026. [1]
  • 2026-06-05: Alibaba's alleged campaign concluded after generating 28.8 million exchanges with Claude. [1][2]
  • 2026-06-10: Anthropic sent a confidential letter to Senators Tim Scott and Elizabeth Warren about the Alibaba campaign, the day before a Senate AI hearing. [1]
  • 2026-06-24: Anthropic publicly accused Alibaba of running the largest known distillation attack on Claude and disclosed the campaign to the U.S. government. [15][2][4]
  • 2026-06-25: Ars Technica reported the Alibaba campaign specifically targeted Claude's agentic reasoning, software engineering, and long-horizon task capabilities. [1]
  • 2026-06-25: Technical debate emerged about whether mainstream press coverage accurately characterizes what distillation accomplishes. [17]
  • 2026-06-29: The Information reported internal Meta documents show restrictions on applied AI engineers' use of Claude Code and Codex over distillation concerns. [6][7]
  • 2026-06-29: Rohan Paul analyzed Meta's restrictions and proposed ingredient-tracking and clean-room rules as mitigations. [8]
  • 2026-06-30: Zvi Mowshowitz confirmed Anthropic's Alibaba accusation and situated it within U.S.-China AI competition and governance concerns. [18]
  • 2026-06-30: Allegation surfaced that Claude Code covertly embeds invisible markers into system prompts to flag requests routed through China-linked proxy servers. [10]
  • 2026-07-01: Anthropic acknowledged the hidden proxy markers in Claude Code and promised to remove them. [13][14]
  • 2026-07-03: Reports emerged that Anthropic cut off Chinese companies that had been accessing Claude through shell companies in Singapore, VPNs, and proxies. [5]

Perspectives

Anthropic

Alibaba conducted the largest known distillation attack on Claude through ~25,000 fraudulent accounts generating 28.8M exchanges, targeting its most valuable capabilities; the campaign is illicit and has been disclosed to the U.S. government and Congress. Anthropic has taken enforcement action cutting off China-linked access and has acknowledged and promised to remove hidden proxy markers from Claude Code.

Evolution: Moved from public accusation to active enforcement; acknowledged the fingerprinting behavior and committed to removing it only after public disclosure forced the issue.

Alibaba

No public response to Anthropic's accusations.

Evolution: Consistent silence throughout.

Meta

Restricting applied AI engineers' use of Claude Code and Codex to prevent competitor outputs from entering its training pipelines.

Evolution: Consistent defensive posture since internal documents were reported.

Rohan Paul (@rohanpaul_ai)

Meta's restrictions are a rational response to legal exposure; ToS clauses require ingredient tracking and clean-room separation to work. The Claude Code proxy markers — covert invisible characters in system prompts — cross a user trust line for an agentic tool with elevated permissions, regardless of the detection rationale.

Evolution: Consistent across both angles; the fingerprinting allegation he originally raised is now confirmed by Anthropic's promised fix.

Legal experts (via ppc.land)

AI ToS provisions barring competitive training may be unenforceable in court, limiting labs' practical recourse even when violations are well-documented.

Evolution: Consistent; no new statements.

Dr. Luke in China (@96Stats)

Press coverage mischaracterizes distillation as straightforward theft; BBC reporting in particular overstates what the technique accomplishes.

Evolution: Consistent skeptical dissent; minority view in overall coverage.

Zvi Mowshowitz

Confirms Anthropic's Alibaba accusation; situates it within a pattern of U.S. AI governance gaps and is cautious about whether policy mechanisms can keep pace with distillation-scale threats.

Evolution: Consistent critical-but-engaged observer.

Prasenjit Sarkar (@stretchcloud)

The Anthropic-Alibaba episode shows the primary bottleneck in AI capability competition has shifted from training to policy enforcement.

Evolution: Consistent analytical observer.

Tensions

  • Anthropic characterizes Alibaba's API usage as a systematic, illicit distillation campaign; Alibaba has not responded, leaving intent and authorization unresolved. [15][2][3]
  • Anthropic embedded covert proxy-detection markers in Claude Code system prompts and, after public exposure, promised to remove them; Rohan Paul argues the covert marking was unacceptable for an agentic tool with elevated system permissions regardless of the detection rationale. [10][13][14]
  • Labs rely on ToS anti-distillation clauses as a primary defense, but legal experts argue those clauses may be unenforceable against foreign entities — limiting practical recourse even when violations are well-documented. [9][8]
  • Technical commentators argue press coverage mischaracterizes distillation; Anthropic and mainstream outlets treat the Alibaba campaign as clear IP misappropriation. [17][15][1]
  • Meta restricts engineers from using the most capable external AI coding tools to protect training pipelines, creating a direct tradeoff between developer productivity and competitive IP hygiene. [6][7][8]

Sources

  1. [1] Anthropic says Alibaba must be punished for largest Claude cloning attack — Ars Technica AI (2026-06-25)
  2. [2] Anthropic accuses Alibaba of a massive distillation attack using ~25,000 fake accounts and 28.8M exchanges against Claud... — reactive:ai-model-distillation-ip (2026-06-28)
  3. [3] Anthropic accuses Alibaba of campaign to 'brazenly' and 'illicitly' extract AI capabilities - The letter, which was obta... — reactive:ai-model-distillation-ip (2026-06-24)
  4. [4] SITUATION DETECTED: Anthropic has disclosed to the U.S. Government that Alibaba executed the largest known distillation ... — reactive:ai-model-distillation-ip (2026-06-28)
  5. [5] Anthropic just cut off Chinese companies that had been accessing Claude through shell companies in Singapore, VPNs, and ... — reactive:ai-model-distillation-ip (2026-07-03)
  6. [6] Exclusive: Internal Meta documents reveal new limits on Claude Code and Codex as the company works to prevent distillati... — reactive:ai-model-distillation-ip (2026-06-30)
  7. [7] Internal docs: Meta places strict limits on how staff in its applied AI division can use Claude Code and Codex, fearing ... — reactive:ai-model-distillation-ip (2026-06-29)
  8. [8] The Information: Meta has reportedly limited engineer use of Claude Code and Codex because rival model outputs could con… — Rohan Paul Twitter (2026-06-29)
  9. [9] Legal experts warn AI terms of service may prove unenforceable in ... — reactive:ai-model-distillation-ip
  10. [10] wow 👀 — Rohan Paul Twitter (2026-06-30)
  11. [11] Anthropic’s Claude Code accused of hiding proxy fingerprints inside system prompts to identify China-linked users - Tech Startups — reactive:ai-model-distillation-ip
  12. [12] Hidden Attribution Markers Found In Claude Code — reactive:ai-model-distillation-ip
  13. [13] Claude Code Hid Proxy Fingerprints in System Prompts - Tech Times — reactive:ai-model-distillation-ip
  14. [14] Anthropic to remove Claude Code marker that flagged China users | AI Weekly — reactive:ai-model-distillation-ip
  15. [15] Anthropic accuses Alibaba of campaign to extract AI capabilities — reactive:ai-model-distillation-ip
  16. [16] Internal documents: Meta is placing strict limits on how engineers in its applied AI division can use Claude Code and Co... — reactive:ai-model-distillation-ip (2026-06-29)
  17. [17] Lol such a dumb article by the BBC clearly written by a casual journalist and not AI expert. They claim ‘extraction’ as ... — reactive:ai-model-distillation-ip (2026-06-25)
  18. [18] The Once And Future Fable #5 — Zvi's AI Roundups (2026-06-30)
  19. [19] The bottleneck in AI capability competition just moved from training to policy enforcement. — reactive:ai-model-distillation-ip (2026-06-25)
  20. [20] Meta has restricted internal use of Claude Code and GitHub Copilot (built on Codex) for the same underlying reason: dist... — reactive:ai-model-distillation-ip (2026-06-30)