Claude Opus 5: The System Card
Zvi's AI Roundups · Zvi Mowshowitz · 2026-07-25
Zvi Mowshowitz analyzes the Claude Opus 5 system card, finding strong benchmark performance close to Mythos 5 at half the price and dramatically improved prompt-injection resistance, while criticizing Anthropic's claim that Opus 5 is their 'most aligned model to date' as dangerous conflation of proxy metrics with alignment itself.
Appears in
Extraction
Topics: claude-opus-5ai-safety-evaluationsagentic-safetyalignment-metricscyber-capabilities
Claims
- Claude Opus 5 is comparable to or ahead of Claude Fable 5 on many benchmarks while being faster and approximately half the price.
- Opus 5's cyber capabilities are stronger than Opus 4.8 but weaker than Mythos 5, particularly for multi-stage operations and binary exploitation.
- Opus 5 shows dramatically improved prompt-injection resistance, reducing computer-use attack success rates from roughly 7% to under 1%.
- Opus 5 now permits source-code vulnerability discovery at all access levels while continuing to block vulnerability discovery in compiled binaries.
- Anthropic's characterization of Opus 5 as their 'most aligned model to date' conflates benchmark scores with actual alignment, which critics argue is irresponsible and self-deceptive.
Key quotes
Our testing indicates that the cyber capabilities of Claude Opus 5 are generally stronger than those of Opus 4.8, but not as strong as those of Mythos 5.
Optimizing against metrics is practically unavoidable, but ontological confusion between benchmark scores and the North Star of alignment is not, and conflating them is one of the most irresponsible and destructive possible mistakes Anthropic can make in their position.
That is a big practical deal. Not getting hijacked via prompt injection is the key to unlocking the confidence to do a host of activities you otherwise can't do.