AI Labs Defend Against Model Output Distillation: Meta Restricts Claude Code, Anthropic Accuses Alibaba
What's new in v5
Item 39878 (Ars Technica, July 6) adds concrete technical detail to the tracking story: the Claude Code markers monitored Chinese users' timezones, proxy configurations, and potential Chinese AI lab affiliations; the feature is characterized as "prompt steganography"; it was added in March 2026; and security researcher Thereallo is named as the person who first publicly exposed it. Ars Technica's Ashley Belanger frames the disclosure as directly contradicting Anthropic's anti-surveillance public stance — a sharper framing than previous coverage. Thereallo has been added as a new perspective voice. The other new items (39960–39969) are empty social media amplifiers with no substantive content.
What
Anthropic publicly accused Alibaba of the largest known distillation attack on Claude, using ~25,000 fraudulent accounts to generate 28.8 million exchanges between April 22 and June 5, 2026.[1][2] A separate controversy: Anthropic embedded hidden tracking code in Claude Code — characterized as "prompt steganography" — that monitored Chinese users' timezones, proxy configurations, and potential Chinese AI lab affiliations; security researcher Thereallo publicly exposed it, and Anthropic removed it after the disclosure.[6] Alibaba blocked Claude Code in response to the tracking disclosure, not the distillation accusations it has still not addressed publicly.[8] Meta separately restricted engineers from using Claude Code and Codex to protect its own training pipelines.[11][12]
Why it matters
Anthropic's covert monitoring of a specific user population — using an agentic tool with permissions to read files, edit code, and run commands — directly contradicts its public anti-surveillance positioning, and was removed only after public exposure.[6] The episode shows that technical enforcement against distillation can create trust and transparency problems that outlast the enforcement rationale: Alibaba's block of Claude Code is a consequence of the tracking disclosure, not the underlying accusation.
Open questions
Alibaba has not responded publicly to the distillation allegations — does it dispute that the accounts were fraudulent or that the campaign's purpose was distillation?[14][2]
Will Anthropic replace the removed prompt markers with a disclosed, opt-in detection mechanism, or abandon proxy-based detection entirely?[9][10]
Are ToS anti-distillation clauses enforceable against a foreign entity like Alibaba, and has Anthropic indicated whether it plans legal action?[13]
Will Alibaba's Claude Code block or the fingerprinting disclosure change how other Chinese organizations access or use Claude?[8][6]
Narrative
Anthropic publicly accused Alibaba of conducting the largest known distillation attack on Claude, using approximately 25,000 fraudulent accounts to generate 28.8 million exchanges between April 22 and June 5, 2026.[1][2] The campaign specifically targeted Claude's most capable functions — agentic reasoning, software engineering, and long-horizon tasks — which Anthropic described as a brazen and illicit extraction of its AI capabilities.[1][3] Anthropic sent a confidential letter to Senators Tim Scott and Elizabeth Warren on June 10, the day before a Senate AI hearing, framing the incident as both a terms-of-service violation and a national security matter.[1] Anthropic has also cut off Chinese companies that had been accessing Claude through shell companies in Singapore, VPNs, and proxy servers.[4][5] Alibaba has issued no public response to the distillation allegations.
A separate controversy emerged around Claude Code itself. Security researcher Thereallo publicly exposed that Claude Code had been using "prompt steganography" — shorthand markers hidden in prompts — to covertly monitor Chinese users' timezones, proxy configurations, and potential affiliations with Chinese AI labs.[6] An Anthropic engineer confirmed the tracker was added in March 2026 as an experiment to prevent account abuse by unauthorized resellers — who had been selling Claude pro access for as little as $12 against the $100 retail price — and to protect against distillation attacks.[6] Rohan Paul's reporting noted that Claude Code has permissions to read files, edit code, and run commands, making hidden prompt modification more consequential than ordinary app-level telemetry, and argued that invisible characters in agentic tools set a precedent making AI agents systematically hard to audit.[7][8] Anthropic acknowledged the behavior and promised to remove it, characterizing the feature as an experiment to stop API resellers and distillation, not targeted surveillance.[9][10] Ars Technica's Ashley Belanger framed the disclosure as a direct contradiction of Anthropic's public anti-surveillance stance.[6]
The tracking disclosure prompted Alibaba to block Claude Code internally, angering Chinese developers and security staff.[8] That block responds to the fingerprinting controversy rather than the distillation accusations — Alibaba still has not addressed the latter publicly. Internal Meta documents show the company separately restricted applied AI engineers from using Claude Code and Codex, citing the risk that competitor tool outputs could enter Meta's training pipelines and constitute inadvertent distillation.[11][12] Legal experts have noted that ToS anti-distillation clauses may be difficult to enforce against foreign entities, though Anthropic's documented evidence of fake accounts and systematic API use would strengthen any enforcement case.[13]
The episode exposes a structural problem in how AI labs try to detect and stop systematic API abuse: effective detection requires covert signals, because disclosed detection methods can be circumvented, but covert signals in agentic tools with elevated permissions undermine the user trust that makes those tools valuable. Anthropic's acknowledgment and removal of the markers resolves the immediate controversy but not the underlying tension between enforcement effectiveness and transparency.
Timeline
- 2026-03-01: Anthropic added hidden tracking code to Claude Code to monitor Chinese users' timezones, proxy configurations, and potential Chinese AI lab affiliations. [6]
- 2026-04-22: Alibaba allegedly began using ~25,000 fraudulent accounts to generate exchanges with Claude; the campaign ran through June 5, 2026. [1]
- 2026-06-05: Alibaba's alleged distillation campaign concluded after generating 28.8 million exchanges with Claude. [1][2]
- 2026-06-10: Anthropic sent a confidential letter to Senators Tim Scott and Elizabeth Warren about the Alibaba campaign, the day before a Senate AI hearing. [1]
- 2026-06-24: Anthropic publicly accused Alibaba of running the largest known distillation attack on Claude and disclosed the campaign to the U.S. government. [14][2][4]
- 2026-06-25: Ars Technica reported the Alibaba campaign specifically targeted Claude's agentic reasoning, software engineering, and long-horizon task capabilities. [1]
- 2026-06-25: Technical debate emerged about whether mainstream press coverage accurately characterizes what distillation accomplishes. [16]
- 2026-06-29: The Information reported internal Meta documents show restrictions on applied AI engineers' use of Claude Code and Codex over distillation concerns. [11][12]
- 2026-06-30: Rohan Paul reported that Claude Code covertly embeds invisible markers into system prompts to flag requests routed through China-linked proxy servers. [7]
- 2026-07-01: Anthropic acknowledged the hidden proxy markers in Claude Code and promised to remove them, describing the feature as an experiment to stop API resellers and distillation. [9][10]
- 2026-07-03: Reports emerged that Anthropic cut off Chinese companies that had been accessing Claude through shell companies in Singapore, VPNs, and proxies. [5]
- 2026-07-04: Alibaba blocked Claude Code internally after Chinese developers and security staff discovered Anthropic's proxy-detection tracking experiment. [8]
- 2026-07-06: Ars Technica published an investigative report framing the Claude Code tracker as contradicting Anthropic's anti-surveillance public stance, naming security researcher Thereallo as the person who exposed it. [6]
Perspectives
Anthropic
Alibaba conducted the largest known distillation attack on Claude through ~25,000 fraudulent accounts generating 28.8M exchanges; the campaign is illicit and has been disclosed to the U.S. government. The prompt markers in Claude Code were an experiment to stop API resellers and distillation, not targeted surveillance; Anthropic promised to remove them after public exposure.
Evolution: Moved from public accusation to active enforcement; acknowledged the fingerprinting behavior and reframed it as abuse-detection only after public disclosure forced the issue.
Alibaba
Has not publicly responded to the distillation allegations; blocked Claude Code internally after the fingerprinting disclosure angered developers and security staff.
Evolution: Moved from silence to a concrete defensive action — blocking Claude Code — though the action responds to the tracking controversy, not the distillation accusations.
Thereallo (security researcher)
Publicly exposed and condemned the hidden tracking code in Claude Code as a serious breach of user trust.
Evolution: First actor to publicly expose the tracker; Anthropic removed it following this disclosure.
Meta
Restricting applied AI engineers' use of Claude Code and Codex to prevent competitor outputs from entering its training pipelines.
Evolution: Consistent defensive posture since internal documents were reported.
Rohan Paul (@rohanpaul_ai)
The Claude Code proxy markers cross a user trust line: an agentic tool with permissions to read files, edit code, and run commands is not like ordinary telemetry, and invisible characters in prompts set a precedent making AI agents systematically hard to audit.
Evolution: Consistent; his fingerprinting allegation is confirmed by Anthropic's promised fix and Alibaba's block of Claude Code.
Legal experts (via ppc.land)
AI ToS provisions barring competitive training may be unenforceable in court, limiting labs' practical recourse even when violations are well-documented.
Evolution: Consistent; no new statements.
Dr. Luke in China (@96Stats)
Press coverage mischaracterizes distillation as straightforward theft; mainstream reporting overstates what the technique accomplishes.
Evolution: Consistent skeptical dissent; minority view in overall coverage.
Tensions
- Anthropic characterizes the prompt markers as a legitimate abuse-detection experiment; Thereallo and Rohan Paul argue covert monitoring of specific user populations in an agentic tool with elevated permissions is a serious breach of user trust regardless of the detection rationale, and contradicts Anthropic's public anti-surveillance positioning. [9][10][7][8][6]
- Anthropic characterizes Alibaba's API usage as a systematic, illicit distillation campaign; Alibaba has not responded to that accusation, leaving intent and authorization unresolved. [14][2][3]
- Labs rely on ToS anti-distillation clauses as a primary defense, but legal experts argue those clauses may be unenforceable against foreign entities — limiting practical recourse even when violations are well-documented. [13][15]
- Technical commentators argue press coverage mischaracterizes distillation; Anthropic and mainstream outlets treat the Alibaba campaign as clear IP misappropriation. [16][14][1]
- Meta restricts engineers from using the most capable external AI coding tools to protect training pipelines, creating a direct tradeoff between developer productivity and competitive IP hygiene. [11][12][15]
Status: active and growing
Sources
- [1] Anthropic says Alibaba must be punished for largest Claude cloning attack — Ars Technica AI (2026-06-25)
- [2] Anthropic accuses Alibaba of a massive distillation attack using ~25,000 fake accounts and 28.8M exchanges against Claud... — reactive:ai-model-distillation-ip (2026-06-28)
- [3] Anthropic accuses Alibaba of campaign to 'brazenly' and 'illicitly' extract AI capabilities - The letter, which was obta... — reactive:ai-model-distillation-ip (2026-06-24)
- [4] SITUATION DETECTED: Anthropic has disclosed to the U.S. Government that Alibaba executed the largest known distillation ... — reactive:ai-model-distillation-ip (2026-06-28)
- [5] Anthropic just cut off Chinese companies that had been accessing Claude through shell companies in Singapore, VPNs, and ... — reactive:ai-model-distillation-ip (2026-07-03)
- [6] Secret Claude tracker shocks users after Anthropic’s anti-surveillance stance — Ars Technica AI (2026-07-06)
- [7] wow 👀 — Rohan Paul Twitter (2026-06-30)
- [8] Alibaba blocked Claude Code after Anthropic’s tracking experiment angered Chinese developers and security staff. — Rohan Paul Twitter (2026-07-04)
- [9] Claude Code Hid Proxy Fingerprints in System Prompts - Tech Times — reactive:ai-model-distillation-ip
- [10] Anthropic to remove Claude Code marker that flagged China users | AI Weekly — reactive:ai-model-distillation-ip
- [11] Exclusive: Internal Meta documents reveal new limits on Claude Code and Codex as the company works to prevent distillati... — reactive:ai-model-distillation-ip (2026-06-30)
- [12] Internal docs: Meta places strict limits on how staff in its applied AI division can use Claude Code and Codex, fearing ... — reactive:ai-model-distillation-ip (2026-06-29)
- [13] Legal experts warn AI terms of service may prove unenforceable in ... — reactive:ai-model-distillation-ip
- [14] Anthropic accuses Alibaba of campaign to extract AI capabilities — reactive:ai-model-distillation-ip
- [15] The Information: Meta has reportedly limited engineer use of Claude Code and Codex because rival model outputs could con… — Rohan Paul Twitter (2026-06-29)
- [16] Lol such a dumb article by the BBC clearly written by a casual journalist and not AI expert. They claim ‘extraction’ as ... — reactive:ai-model-distillation-ip (2026-06-25)