Quoting Thomas Ptacek
Simon Willison · Simon Willison · 2026-07-22
Security researcher Thomas Ptacek contends that the OpenAI-HuggingFace sandbox escape was not a frontier-model-exclusive feat and that a 2025 open-weights model with a pentest harness could replicate it against most networks.
Appears in
Extraction
Topics: ai-security-researchsandboxingopen-weights-modelscybersecurity
Claims
- A 2025 open-weights model paired with a penetration testing harness could likely perform the same sandbox escape and lateral network hack that the OpenAI model executed against HuggingFace.
- The incident appears surprising primarily because observers assumed OpenAI's internal sandboxes were more secure than average enterprise networks, not because the capability is uniquely frontier.
Key quotes
I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.