The Information Machine

AI as Attack Tool and Attack Target: May 2026 Cybersecurity Moment · history

Version 25

2026-07-04 02:23 UTC · 656 items

What

AI operates simultaneously as an offensive tool and active attack surface. Sysdig's July 3 identification of JADEPUFFER — the first documented ransomware operation driven entirely by an LLM agent — puts autonomous AI attack capability in the present tense: the system generated 600+ purposeful payloads and chained attack steps adaptively without human planning [20]. Two frontier defense programs, Anthropic's Glasswing (~200 partners, 15+ countries) and OpenAI's Daybreak (GPT-5.5-Cyber at 85.6% on CyberGym, seven government agreements), remain deployed without common external audit standards [23][25], while prompt injection vulnerabilities have expanded from research demonstrations to documented attacks on named commercial products [14].

Why it matters

JADEPUFFER's appearance 10 days after the Five Eyes alliance warned autonomous AI attack capability could arrive 'within months' [21][22] narrows that gap to a question of degree rather than kind. The combination of demonstrated autonomous attack capability in the wild, growing frontier defense programs without common governance, and structurally contested prompt injection defenses leaves the question of what triggers meaningful oversight entirely open.

Open questions

  • JADEPUFFER exploited a missing-authentication vulnerability in Langflow [20] — does this represent an isolated targeting of an AI-specific attack surface, or does the autonomous adaptive chaining approach generalize to other open-source AI agent frameworks?

  • Sysdig attributes most JADEPUFFER damage to legacy security failures rather than novel AI capabilities [20], while the Five Eyes warned severe AI-enabled attacks could arrive 'within months' [21][22] — does JADEPUFFER constitute that threshold being crossed or a precursor to it?

  • Willison's experiment showed 0 successes in 6,000 injection attempts via anti-injection training [18], while role confusion research identifies style-based privilege parsing as the structural cause of defense failure [19] — does training-based resistance address the structural issue or only raise the difficulty bar?

  • Neither Glasswing nor Daybreak is subject to CAISI review [25][23][26] — do their deployment terms include circuit-breakers if capability thresholds are crossed, given JADEPUFFER suggests some thresholds may already be operational?

Narrative

The supply chain and infrastructure attacks that began in May 2026 established the scale of AI-enabled threat activity. TeamPCP's Mini Shai-Hulud campaign, launched May 11, compromised more than 1,000 SaaS environments [1], stole approximately 3,800–4,000 GitHub internal repositories via a poisoned VS Code extension [2][3], and breached 30 EU institutions via the Trivy container scanner [4]. Concurrent campaigns hit adjacent developer trust surfaces: an AntV ecosystem attack faked Sigstore provenance badges across 600+ packages [5], TrapDoor compromised 34+ packages across npm, PyPI, and crates.io [6], and 73 Microsoft-signed packages contained credential-stealing code that activated when developers opened them in AI coding agents [7]. AI-connected products were also targeted directly: a maximum-critical vulnerability in Microsoft 365 Copilot allowed 2FA code and email theft via prompt injection in third-party content [8], and Google's GTIG confirmed the first criminal AI-generated zero-day [9].

Prompt injection against AI systems expanded in scope and specificity through mid-2026. SafeBreach documented three separate bypasses of Google Gemini's defenses across voice injection, Calendar invites, and WhatsApp [10][11][12]. A June 30 research demonstration showed AI browsers can be tricked into accepting a false reality where safety guardrails no longer apply, enabling credential and code theft [13]; Brave subsequently documented indirect prompt injection in Perplexity Comet specifically [14], giving the AI browser attack class a named commercial product. Industry security teams at NVIDIA, Cisco, and HiddenLayer each characterized prompt injection as a structural problem that guardrails alone cannot solve [15][16][17]. Against this, Simon Willison's June 26 experiment showed 6,000 injection attempts failed against a Claude Opus 4.6 instance [18], but role confusion research published the same week identifies style-based privilege parsing — rather than prompt position — as the underlying structural cause of defense failure, with destyling reducing injection success from 61% to 10% [19].

On July 3, Sysdig reported JADEPUFFER as the first documented ransomware operation driven entirely by an LLM agent rather than scripted automation [20]. The attack exploited a missing-authentication vulnerability in Langflow, an open-source AI agent building tool. The LLM agent generated over 600 purposeful payloads and adaptively chained attack steps without human planning or retries. Distinctively, JADEPUFFER destroyed data without preserving a decryption key, making recovery impossible after payment. Sysdig attributed most damage to legacy failures — default credentials and weak service exposure — rather than novel AI capabilities, but identified the autonomous adaptive chaining as qualitatively different from prior scripted ransomware. JADEPUFFER arrived 10 days after the Five Eyes alliance issued a joint advisory warning that AI models capable of severe attacks on governments and businesses could arrive within months [21][22].

Two frontier programs are deployed for defense with contested governance. Anthropic's Project Glasswing operates with approximately 200 partners in 15+ countries across critical infrastructure sectors [23]; OpenAI's Daybreak, launched June 22, includes GPT-5.5-Cyber at 85.6% on CyberGym — the highest single-model score measured — and Codex Security, which has scanned 30 million commits across 30,000+ codebases [24][25]. OpenAI signed Trusted Access for Cyber agreements with seven governments including ENISA and restricts GPT-5.5-Cyber to verified defenders [25]. Neither program is subject to CAISI review, and governance critics argue two programs competing on capability while self-certifying access controls creates the gap CAISI was designed to prevent [26].

Timeline

  • 2026-05-05: NIST's CAISI formalized as US pre-deployment AI compliance gate with agreements covering Google, Microsoft, and xAI. [41]
  • 2026-05-11: TeamPCP launches Mini Shai-Hulud; 160+ npm and PyPI packages compromised; two OpenAI employee devices breached with code-signing certificates exfiltrated. [42][43][44]
  • 2026-05-11: Google GTIG intercepts the first confirmed criminal AI-generated zero-day targeting a hardcoded 2FA trust assumption. [9][45][46]
  • 2026-05-13: AISI evaluates Claude Mythos Preview as first AI to autonomously complete both UK offensive cyber ranges; autonomous AI cyber capability measured as doubling every 4.7 months. [42][47][40]
  • 2026-05-18: TeamPCP advertises Mistral AI source code — 450 repositories — for sale at $25,000; Mistral confirms impact. [48][49][50]
  • 2026-05-20: GitHub confirms theft of approximately 3,800–4,000 internal repositories via a poisoned Nx Console VS Code extension. [2][3][51]
  • 2026-05-24: CERT-EU confirms European Commission breach across 30 EU institutions via Trivy; Mandiant quantifies 1,000+ SaaS compromises; TrapDoor hits 34+ packages; AntV ecosystem fakes Sigstore badges across 600+ packages. [4][1][6][5]
  • 2026-06-02: Anthropic expands Project Glasswing to ~200 partners in 15+ countries; reports 10,000+ critical flaws found; releases Claude Security on Opus 4.8. [23]
  • 2026-06-04: SafeBreach documents third Gemini bypass via WhatsApp 'Fake Context Alignment'; Google DeepMind documents agents detecting malicious websites across six attack types. [12][52][10][11]
  • 2026-06-08: 73 Microsoft-signed packages with AI-agent-triggered credential stealers blocked; GitHub describes removal as a terms-of-service violation. [7]
  • 2026-06-12: Google sues Chinese Outsider Enterprise for using Gemini to run phishing-as-a-service: $88/week, 290+ templates, $1.9B in estimated losses, 3.87M stolen credit card numbers per FBI. [53][54][55]
  • 2026-06-16: Ars Technica reports maximum-critical M365 Copilot vulnerability allowed 2FA code theft via prompt injection; researchers characterize prompt injection as structurally unfixable in current LLMs. [8]
  • 2026-06-22: OpenAI launches Daybreak: GPT-5.5-Cyber achieves 85.6% on CyberGym (highest single-model score); Codex Security has scanned 30M+ commits; Trusted Access for Cyber signed with 7 governments including ENISA. [24][25][34][35][36]
  • 2026-06-22: Role confusion research: LLMs parse privilege from text style rather than prompt position; destyling drops injection success from 61% to 10%. [19]
  • 2026-06-23: Five Eyes alliance issues joint advisory — confirmed via official NSA statement — warning that AI models capable of severe attacks on governments and businesses could arrive within months. [21][22][56][57][58]
  • 2026-06-26: Willison reports 6,000 prompt injection attempts across ~2,000 participants failed against a Claude Opus 4.6 instance; credits anti-injection training while cautioning against treating the result as a production security guarantee. [18]
  • 2026-06-30: Researchers demonstrate 'false reality' attack on AI browsers enabling credential and code theft by disabling guardrails; Ars Technica argues the architectural flaw cannot be resolved through guardrail additions. [13]
  • 2026-07-01: Brave documents indirect prompt injection in Perplexity Comet, moving the AI browser attack class from research demonstration to a named commercial product. [14]
  • 2026-07-03: Sysdig identifies JADEPUFFER as the first documented ransomware operation driven entirely by an LLM agent: 600+ adaptive payloads, autonomous attack chaining without human planning, data destruction without a decryption key. [20]

Perspectives

Five Eyes Alliance

Issued a joint advisory, confirmed on NSA.gov, that AI models capable of severe attacks on governments and businesses could arrive within months, characterizing AI as making devastating cyberattacks far easier for malicious actors.

Evolution: The June 23 advisory's 'within months' framing is now directly adjacent to JADEPUFFER's July 3 appearance; no updated Five Eyes statement has addressed whether JADEPUFFER meets that threshold.

Sysdig

Identified JADEPUFFER as the first LLM-agent-driven ransomware: autonomous adaptive attack chaining without human planning, 600+ purposeful payloads, and data destruction without a recovery key; attributes most damage to legacy security failures rather than novel AI capabilities.

Evolution: New voice; first appearance in the thread.

Anthropic

Expanding Project Glasswing to ~200 partners under a proactive-defense rationale; FRT empirical data documents AI democratizing sophisticated post-compromise techniques.

Evolution: Consistent.

OpenAI

Launched Daybreak with GPT-5.5-Cyber (85.6% CyberGym, highest single-model score) and Codex Security; argues AI has shifted the bottleneck from discovering to patching vulnerabilities; restricts GPT-5.5-Cyber to verified defenders under Trusted Access for Cyber agreements with seven governments.

Evolution: Consistent since Daybreak's launch.

Simon Willison

His June 26 experiment — 6,000 failed injection attempts across ~2,000 participants — provides empirical evidence that anti-injection training works in controlled settings, while he advises against deploying systems where injection could cause irreversible damage.

Evolution: His experiment partially qualifies his prior 'perpetual whack-a-mole' framing: training-based resistance is measurably effective in controlled settings, but he stops short of endorsing current architectures as structurally fixed.

SafeBreach Labs

Has documented three successful bypasses of Google Gemini's prompt injection defenses: voice/audio injection, Google Calendar invites, and WhatsApp 'Fake Context Alignment' across six messaging platforms.

Evolution: Consistent; the pattern of repeated circumvention of individually patched defenses independently supports the view that per-vector guardrails are insufficient.

GitHub / Microsoft

Characterized theft of ~3,800 internal repositories as limited impact, described removal of 73 malicious packages as a terms-of-service violation, and patched a maximum-critical Copilot vulnerability enabling 2FA theft.

Evolution: Consistent; per-incident responses without addressing structural critiques.

AuthMind + Turing Institute CETAS

The CAISI voluntary framework evaluates only submitted models, not deployment programs; both Glasswing and Daybreak have expanded to major government partnerships without CAISI review, exactly the unaudited frontier deployment their governance critique describes.

Evolution: OpenAI's Daybreak launch with seven government partnerships, also outside CAISI review, extends the same governance gap to a second major actor, strengthening their critique.

Tensions

  • Sysdig identifies JADEPUFFER as qualitatively new — autonomous adaptive attack chaining without human planning — while attributing most damage to legacy security failures [20]; the Five Eyes advisory warned AI capable of severe attacks could arrive 'within months' [21][22], leaving unresolved whether JADEPUFFER constitutes that threshold being crossed or a precursor to it. [20][21][22]
  • Ars Technica and role confusion researchers argue current LLMs have no structural fix for prompt injection [8][19]; Brave's documentation of indirect prompt injection in Perplexity Comet [14] and the June 30 'false reality' attack demonstration [13] add a named commercial product and a new attack class to that critique; Willison's experiment showed 0 successes in 6,000 injection attempts via anti-injection training, while he maintains this is not a production safety guarantee [18]. [8][19][13][14][18]
  • OpenAI's Daybreak and Anthropic's Glasswing both claim structured defender-only access for frontier cyber AI, but neither is subject to CAISI review or a common external standard; governance critics argue two programs competing on capability while self-certifying access controls creates the gap CAISI was designed to prevent [25][23][26]. [25][23][26]
  • The Five Eyes alliance warns severe AI cyberattack capability could arrive 'within months' [21][22] and AISI measures autonomous capability doubling every 4.7 months [40] — but neither Glasswing nor Daybreak deployment terms publicly address what happens if that capability threshold is crossed during their current operating windows. [21][22][40]
  • Anthropic argues deploying Glasswing under controlled conditions is preferable to waiting while competitors deploy without safeguards; AuthMind and CETAS argue this expansion without CAISI review is precisely the unaudited frontier deployment their governance critique describes [23][26]. [23][26]

Sources

  1. [1] TeamPCP Supply Chain Campaign: Update 006 - CERT-EU Confirms European Commission Cloud Breach, Sportradar Details Emerge, and Mandiant Quantifies Campaign at 1,000+ SaaS Environments — reactive:ai-security-nexus
  2. [2] Nx Console 18.95.0 Incident: How TeamPCP Breached GitHub — reactive:ai-security-nexus
  3. [3] GitHub just confirmed that attackers stole about 3,800 internal repositories after a poisoned VS Code extension compromi… — Rohan Paul Twitter (2026-05-20)
  4. [4] European Commission cloud breach: a supply-chain compromise — reactive:ai-security-nexus
  5. [5] Mini Shai-Hulud Returns: 600+Malicious npm Packages Fake Sigstore Badges in AntV Ecosystem Attack — reactive:ai-security-nexus
  6. [6] TrapDoor Crypto Stealer Supply Chain Attack Hits 34 Packages... — reactive:ai-offensive-cyber
  7. [7] For the 2nd time in weeks, Microsoft packages laced with credential stealer — Ars Technica AI (2026-06-08)
  8. [8] Critical Copilot vulnerability allowed hackers to steal 2FA code from users — Ars Technica AI (2026-06-16)
  9. [9] Google Researchers Detect First AI-Built Zero-Day Exploit in Cyberattack - Bloomberg — reactive:ai-offensive-cyber
  10. [10] Exploiting Gemini via Prompt Injection | SafeBreach Original Research — reactive:ai-security-nexus
  11. [11] Invitation Is All You Need: Hacking Gemini | SafeBreach — reactive:ai-security-nexus
  12. [12] 😺 Google Gemini got hijacked via WhatsApp — The Neuron (2026-06-04)
  13. [13] New attack provides one more reason why AI browsers are a bad idea — Ars Technica AI (2026-06-30)
  14. [14] Agentic Browser Security: Indirect Prompt Injection in Perplexity Comet | Brave — reactive:ai-security-nexus
  15. [15] Securing Agentic AI: How Semantic Prompt Injections Bypass AI Guardrails | NVIDIA Technical Blog — reactive:ai-security-nexus
  16. [16] Prompt injection is the new SQL injection, and guardrails ... — reactive:ai-security-nexus
  17. [17] OpenAI Guardrails Bypass: The "Self-Policing" LLM ... — reactive:ai-security-nexus
  18. [18] What happened after 2,000 people tried to hack my AI assistant — Simon Willison (2026-06-26)
  19. [19] Prompt Injection as Role Confusion — Simon Willison (2026-06-22)
  20. [20] Ransomware has crossed from scripted automation to autonomous AI decision-making. — Rohan Paul Twitter (2026-07-03)
  21. [21] AI models capable of severe attacks on governments and businesses could arrive within months. — Rohan Paul Twitter (2026-06-23)
  22. [22] Five Eyes Cyber Security Agencies Statement — reactive:ai-security-nexus
  23. [23] Expanding Project Glasswing — Anthropic News (2026-06-02)
  24. [24] Patch the Planet: a Daybreak initiative to support open source maintainers — OpenAI Blog (2026-06-22)
  25. [25] Daybreak: Tools for securing every organization in the world — OpenAI Blog (2026-06-22)
  26. [26] When a Lab Withholds Its Best Model: What the Claude Mythos System Card Signals for Cybersecurity — reactive:ai-security-nexus
  27. [27] Democracy Now! - The intelligence alliance known as "Five... — reactive:ai-security-nexus
  28. [28] AI on pace to bypass cybersecurity systems in months, not years ... — reactive:ai-security-nexus
  29. [29] AI could breach government and business defenses in months, US ... — reactive:ai-security-nexus
  30. [30] Five Eyes Security Agencies Issue Urgent Warning On AI | 10 News — reactive:ai-security-nexus
  31. [31] What we learned mapping a year’s worth of AI-enabled cyber threats — Anthropic News (2026-06-03)
  32. [32] Gap: Anthropic mapped 832 banned accounts onto MITRE ATT&CK. AI in the back half of attacks jumped 8.9%; phishing dr... — reactive:ai-security-nexus (2026-06-14)
  33. [33] AI models capable of devastating attacks on governments ... — reactive:ai-security-nexus
  34. [34] OpenAI’s new GPT-5.5-Cyber just beat Mythos 5 on CyberGym. — Rohan Paul Twitter (2026-06-22)
  35. [35] OpenAI expands Daybreak program, updates GPT-5.5-Cyber, lands ... — reactive:ai-security-nexus
  36. [36] OpenAI Expands Daybreak With GPT-5.5-Cyber to Help ... — reactive:ai-security-nexus
  37. [37] OpenAI Help: Lockdown Mode — Simon Willison (2026-06-05)
  38. [38] Claude Mythos: What Does Anthropic's New Model Mean for the ... — reactive:ai-security-nexus
  39. [39] AISI: autonomous AI cyber capability now doubling every 4.7 months — reactive:ai-offensive-cyber
  40. [40] Our evaluation of Claude Mythos Preview's cyber capabilities — reactive:frontier-ai-cyber-capabilities
  41. [41] US government expands vetting of frontier AI models for security risks — reactive:ai-security-nexus
  42. [42] Our response to the TanStack npm supply chain attack — OpenAI Blog (2026-05-13)
  43. [43] Mini Shai-Hulud: TeamPCP compromette 160+ pacchetti npm e PyPI in un supply chain attack che ha colpito TanStack, Mistra... — reactive:ai-security-nexus (2026-05-19)
  44. [44] A Self-Spreading Supply Chain Attack Compromises TanStack npm ... — reactive:ai-security-nexus
  45. [45] Google Detects First AI-Generated Zero-Day Exploit - SecurityWeek — reactive:ai-offensive-cyber
  46. [46] Google spotted an AI-developed zero-day before attackers could use it | CyberScoop — reactive:ai-offensive-cyber
  47. [47] How fast is autonomous AI cyber capability advancing? — reactive:ai-offensive-cyber (2026-05-13)
  48. [48] TeamPCP vende repo Mistral AI dopo attacco TanStack su OpenAI — reactive:ai-security-nexus (2026-05-18)
  49. [49] Hackers threaten to leak Mistral files online — AI giant confirms breach, but not what data is involved | TechRadar — reactive:ai-offensive-cyber
  50. [50] TeamPCP Claims Sale of Mistral AI Repositories Amid Mini Shai ... — reactive:ai-security-nexus
  51. [51] GitHub Says 3,800 Repositories Breached—TeamPCP Hackers ... — reactive:ai-security-nexus
  52. [52] This Google DeepMind’s paper is a serious warning for anyone using autonomous agents today. — Rohan Paul Twitter (2026-06-04)
  53. [53] Google sues Chinese cybercrime network that used Gemini to automate scams — Ars Technica AI (2026-06-12)
  54. [54] Google Sues to Stop Chinese Cybercrime Group from Using Its A.I. — reactive:ai-security-nexus
  55. [55] 😺 Google sued the people spamming your phone — The Neuron (2026-06-16)
  56. [56] Five Eyes cybersecurity agencies warn of new AI models ... — reactive:ai-security-nexus
  57. [57] #Gravitas | Five Eyes intelligence alliance issued a joint ... — reactive:ai-security-nexus
  58. [58] 'Five Eyes' intelligence alliance warns that new AI models ... — reactive:ai-security-nexus