Claude Fable 5: Model Update, Safety Profile, Benchmarks, and Subscriber Trial Rollout · history
Version 2
2026-07-03 08:20 UTC · 162 items
What
Claude Fable 5 launched June 9, was suspended by US export controls June 12 after a reported jailbreak, and was globally restored June 30 after Anthropic disclosed the triggering vulnerability was replicable by Claude Opus 4.8, GPT-5.5, and Kimi K2.7—offering no unique uplift over already-deployed models.[6] The model returned with an updated cybersecurity classifier and a formal 4-tier content framework, but community benchmark data shows post-redeployment coding scores collapsed on BridgeBench (Debugging: 86.2 → 25.9).[9][8] Subscription access runs through July 7, then moves to per-usage billing due to capacity constraints.[11] Competitive pressure from OpenAI's GPT-5.6 Sol at half the price and open-source GLM 5.2 continues, while a June 9 critique by Nathan Lambert alleges Anthropic also uses hidden, undisclosed filters for frontier AI research tasks alongside its announced classifiers.[2]
Why it matters
The BridgeBench collapse is the first documented case of a US export control action measurably changing a frontier model's performance for ordinary users: safety classifiers added under regulatory pressure demonstrably degraded specific coding capabilities. If classifier-altered access becomes the norm for regulated frontier models, what subscribers receive may diverge substantially from what benchmark scores at launch promised.
Open questions
Anthropic says the updated classifier blocks the reported jailbreak in over 99% of cases, but BridgeBench shows Debugging scores fell from 86.2 to 25.9 — will Anthropic publish a technical account of what tradeoffs were made?[9][8]
The export control reversal happened in stages (partial for trusted US companies first, then full global) with no public explanation at either step — does the US government have any defined process for reviewing and lifting AI export controls, or were both decisions entirely discretionary?[6][4][5]
Anthropic says it intends to restore Fable 5 to subscriptions once capacity improves but gave no timeline — how long will it remain pay-per-use after July 7?[11]
Nathan Lambert argues Anthropic uses undisclosed filters for frontier AI research tasks (pretraining pipelines, ML accelerator design) that silently reduce capability — will Anthropic confirm, deny, or document these classifiers?[2]
Narrative
Anthropic launched Claude Fable 5 and the restricted-access Claude Mythos 5 on June 9, 2026, at $10/$50 per million input/output tokens, with classifiers that route queries touching cybersecurity, biology, chemistry, and model distillation to Claude Opus 4.8, and with Mythos 5 limited to vetted partners under 30-day data retention.[1] On the same day, Nathan Lambert published a detailed critique arguing that Anthropic also deploys hidden, undisclosed filters—applied through prompt modification, steering vectors, or PEFT—that silently reduce Fable 5's effectiveness for frontier AI research tasks such as pretraining pipeline development and ML accelerator design; Lambert characterized this covert capability reduction as competitive protection dressed as safety and as a form of misalignment distinct from the announced classifiers.[2]
Three days after launch, the US government issued an export control directive suspending Fable 5 and Mythos 5 for all foreign nationals, citing a jailbreak identified by Amazon researchers. Anthropic complied under legal obligation while publicly disputing both the technical grounds and the process.[3] Around June 27-28, the government partially reversed course, restoring Mythos access for "trusted" US companies before a full global reversal.[4][5] On June 30, the US fully lifted all restrictions without explanation, and Anthropic published its redeployment announcement.[6][7] In that announcement, Anthropic disclosed that its own testing found the triggering jailbreak—asking the model to read a codebase and identify vulnerabilities—was replicable by Claude Opus 4.8, GPT-5.5, and Kimi K2.7, providing no unique capability uplift over existing commercial deployments.[6] Anthropic also announced pre-release government access commitments for future frontier models, rapid jailbreak information sharing, and a collaboration with Amazon, Microsoft, Google, and Glasswing partners on an industry jailbreak severity framework.
The model that returned from suspension is measurably different. Anthropic published a formal 4-tier cybersecurity classification—prohibited, high-risk dual-use, low-risk dual-use, and benign—with classifiers targeting the first two categories, and acknowledged an intentional "safety margin" that also blocks some benign uses.[8] Community benchmark data showed the impact is substantial: BridgeBench scores for the redeployed Fable 5 dropped across all tested dimensions, with Debugging falling from 86.2 to 25.9, Refactoring from 73.6 to 38.4, and Hallucination resistance from 75.9 to 61.7.[9] In at least one documented case, classifiers routed 75% of a $321 coding session to Opus 4.8—$242 of the total cost—despite the work being routine coding.[10] Anthropic confirmed that the subscriber trial ends July 7, after which Fable 5 requires separate usage credits; Anthropic cited capacity constraints and stated its intention to restore Fable 5 to subscriptions without giving a timeline.[11]
On the capability side, Zvi Mowshowitz reported that Fable 5 completes 16.1% of remote projects at professional standard—roughly double the next-best model and up from Claude Opus 4.6's 4.2%—with AI automation of remote work having increased approximately 4x in five months.[12] Competitively, OpenAI's GPT-5.6 Sol launched at $5/$30 per million tokens, roughly half Fable 5's price, claiming 91.9% on TerminalBench 2.1.[13][14] China's open-source GLM 5.2 outperformed Fable 5 on Semgrep's security benchmark (39% vs 32% F1) at a fraction of the inference cost.[15] Dario Amodei told Congress that open-source AI represents the real national security threat, framing closed monitored deployment as the responsible policy approach[16]—a position observers noted sits in tension with the export control episode, which showed that closed-weight frontier models face government controls that open-weight models effectively circumvent.[17]
Timeline
- 2026-06-09: Anthropic launches Claude Fable 5 and Claude Mythos 5 at $10/$50 per million tokens with tiered safety classifiers and restricted biomedical/cyber research access under 30-day data retention. [1]
- 2026-06-09: Nathan Lambert publishes critique alleging Anthropic uses hidden, undisclosed filters for frontier AI research tasks alongside its announced classifiers, calling covert capability reduction a form of misalignment. [2]
- 2026-06-12: US government issues export control directive suspending Fable 5 and Mythos 5 for all foreign nationals; Anthropic complies but publicly disputes the technical basis and process. [3]
- 2026-06-25: Zhipu AI's open-source GLM 5.2 publicly compared against Fable 5, claiming competitive coding and design arena performance under an MIT license. [25][26]
- 2026-06-27: US government partially reverses the export directive, restoring Mythos 5 access for 'trusted' US companies before a full global reversal. [4][27][5]
- 2026-06-28: GLM 5.2 outscores Fable 5 on Semgrep security benchmark (39% vs 32% F1); OpenAI releases GPT-5.6 Sol at $5/$30 per million tokens claiming 91.9% on TerminalBench 2.1. [15][14][13]
- 2026-06-29: Dario Amodei tells Congress open-source AI is the real national security threat, framing closed frontier model deployment as the responsible policy approach. [16][17]
- 2026-06-30: US fully lifts export restrictions; Anthropic redeployment announcement discloses the triggering jailbreak was replicable by Opus 4.8, GPT-5.5, and Kimi K2.7, and commits to pre-release government access for future frontier models. [7][18][6]
- 2026-07-01: Subscriber trial begins: Pro/Max/Team/Enterprise users get Fable 5 for one week at 50% of remaining weekly limits; API users excluded at standard pricing. [19]
- 2026-07-01: HTML5 physics benchmark shows Fable 5 winning on quality but costing $3.12 vs GLM 5.2 at $0.08 and Opus 4.8 at roughly one-sixth the price. [20]
- 2026-07-02: Anthropic publishes 4-tier cybersecurity classification framework and proposed CJS jailbreak severity scoring system; launches HackerOne bug bounty for cyber jailbreaks. [8]
- 2026-07-02: BridgeBench scores for redeployed Fable 5 collapse: Debugging 86.2→25.9, Refactoring 73.6→38.4, Hallucination resistance 75.9→61.7. [9]
- 2026-07-02: Reports of classifiers routing 75% of a $321 coding session to Opus 4.8 ($242 of total cost) despite routine coding work. [10]
- 2026-07-02: Anthropic confirms subscription access to Fable 5 ends July 7, moving to per-usage credits due to capacity constraints, with no timeline for restoring subscription inclusion. [11][24]
- 2026-07-02: Zvi Mowshowitz reports Fable 5 completes 16.1% of remote projects at professional standard, double next-best and up from Opus 4.6's 4.2%, with remote work automation up roughly 4x in five months. [12]
- 2026-07-02: Developer community and Fast Company coverage reports widespread frustration with classifier over-triggering on legitimate coding and security tasks. [21][22]
Perspectives
Anthropic
Launched Fable 5/Mythos 5 as a deliberate capability-safety split; complied with the government suspension while publicly disputing it; redeployed with an updated classifier, a 4-tier cybersecurity framework, a proposed CJS jailbreak severity standard, and commitments to pre-release government access for future frontier models.
Evolution: More transparency-forward and collaborative post-redeployment, proposing formal industry standards and a HackerOne bounty program, while candidly acknowledging the new classifier intentionally blocks some benign uses as a safety margin.
US Government
Issued export control suspension citing a jailbreak with no public technical rationale; partially reversed for trusted US companies, then fully lifted restrictions around June 30 with no public explanation of what changed at either step.
Evolution: Moved from blanket suspension to staged partial reversal to full lift with no transparency at any point and no public legal or technical standard articulated.
Dario Amodei (Anthropic CEO)
Told Congress open-source AI is the real national security threat, framing closed and monitored frontier model deployment as the responsible policy approach.
Evolution: Consistent with Anthropic's safety-first positioning; the congressional testimony came while the government's export control on Anthropic's own model was still recent.
Nathan Lambert (Interconnects)
Argues Anthropic deploys hidden, undisclosed filters for frontier AI research tasks alongside its announced classifiers, framing the covert restriction as competitive protection dressed as safety and a form of misalignment; advocates open-source AI as the structural remedy.
Evolution: New voice in this thread; sharply distinguishes Anthropic's legitimate transparent policies from what he characterizes as covert manipulation of model intelligence.
Zvi Mowshowitz
Documents Fable 5's substantial capability lead (16.1% professional remote project completion, double next-best) while arguing the export control episode set a dangerous ad hoc governance precedent and that alignment remains unsolved given AI models routinely lie about completed tasks.
Evolution: New voice in this thread; combines capability data with governance and alignment critique.
Rohan Paul (@rohanpaul_ai)
Provides data-driven benchmark and cost analysis; documented BridgeBench score collapse post-redeployment and a case where classifier false positives routed 75% of a $321 coding session to Opus 4.8; frames the redeployed model as a permanent departure from the original promise of frontier access.
Evolution: Moved from neutral benchmark reporter to critical framing around 'permissioned intelligence' and the erosion of the social contract between AI labs and users.
Developer and power-user community
Impressed by Fable 5's benchmark performance and agentic capabilities but frustrated by high cost, classifier false positives on legitimate coding and security work, and the revelation that post-redeployment BridgeBench scores collapsed.
Evolution: Initial enthusiasm has given way to sustained frustration; post-redeployment concern centers on whether the model is permanently less capable than the version that launched June 9.
OpenAI
Released GPT-5.6 Sol at $5/$30 per million tokens claiming superiority on TerminalBench 2.1 (91.9%), positioning it as both higher-performing and cheaper than Fable 5.
Evolution: Consistent competitive posture; the Sol launch during Fable 5's suspension gave it additional visibility with users seeking alternatives.
Tensions
- Anthropic argues the export control-triggering jailbreak was low-severity and replicable by Claude Opus 4.8, GPT-5.5, and Kimi K2.7 with no unique Mythos-specific uplift; the US government provided no technical rebuttal and lifted restrictions without explanation. [6][3][7]
- Anthropic acknowledges an intentional 'safety margin' that blocks some benign uses; BridgeBench data shows Debugging scores dropped from 86.2 to 25.9 post-redeployment, and at least one user documented classifiers routing 75% of a routine $321 coding session to Opus 4.8. [8][9][10]
- Nathan Lambert argues Anthropic deploys undisclosed filters for frontier AI research tasks that silently reduce model capability—which he calls misalignment—while Anthropic's public documentation describes only its announced cybersecurity and biology classifiers. [2][8]
- At $10/$50 per million tokens, Fable 5 costs 2x GPT-5.6 Sol ($5/$30) and up to 39x GLM 5.2 in tested tasks, while both competitors claim benchmark parity or superiority on specific metrics. [13][20][15]
- Dario Amodei told Congress closed frontier deployment is responsible policy because open-source AI is the real national security threat; critics note the Fable 5 export control episode showed closed-weight models face government controls that open-weight models effectively circumvent. [16][17][3]
- Benchmark authority is contested: Fable 5 leads CursorBench and HTML5 physics quality tests, GPT-5.6 Sol leads TerminalBench 2.1, BridgeBench shows Fable 5's post-redeployment coding scores collapsed, and GLM 5.2 leads on Semgrep security. [24][14][9][15]
Sources
- [1] Claude Fable 5 and Claude Mythos 5 — Anthropic News (2026-06-09)
- [2] Claude Fable 5 and new AI safety fables — Interconnects (2026-06-09)
- [3] Statement on the US government directive to suspend access to Fable 5 and Mythos 5 — Anthropic News (2026-06-12)
- [4] Anthropic to Restore Mythos for 'Trusted' US Companies — Access Allowed After 2-Week Suspension — reactive:claude-fable-5-launch (2026-06-28)
- [5] The US government has partially reversed its June 12 export control order, allowing Anthropic to restore access to its C... — reactive:claude-fable-5-launch (2026-06-27)
- [6] Redeploying Fable 5 — Anthropic News (2026-06-30)
- [7] Summary of Anthropic’s “Redeploying Fable 5” announcement (June 30, 2026): — reactive:claude-fable-5-launch (2026-07-01)
- [8] More details on Fable 5’s cyber safeguards and our jailbreak framework — Anthropic News (2026-07-02)
- [9] Feels like an end of era, ordinary people will probably never again get upgraded frontier models. — Rohan Paul Twitter (2026-07-02)
- [10] This may be an extreme case but it still shows how quickly Fable 5 classifiers can reroute routine coding to Opus. — Rohan Paul Twitter (2026-07-02)
- [11] Current Fable-5's subscription access ends after July-07. — Rohan Paul Twitter (2026-07-02)
- [12] AI #175: The Fable Continues — Zvi's AI Roundups (2026-07-02)
- [13] GPT-5.6 Sol is priced at $5 input and $30 output per million tokens. Claude Fable 5 is priced at $10 input and $50 outpu... — reactive:claude-fable-5-launch (2026-06-29)
- [14] GPT-5.6 Sol Ultra scored 91.9 percent on Terminal-Bench 2.1. That is the highest score ever recorded on a command-line w... — reactive:claude-fable-5-launch (2026-06-28)
- [15] China's GLM-5.2 just outscored the banned Claude Code on Semgrep's security benchmark. 39% F1 vs 32%. Open-weight. MIT l... — reactive:claude-fable-5-launch (2026-06-28)
- [16] Dario Amodei told Congress this week that open source AI is the real threat. Once models are released, the Anthropic CEO... — reactive:claude-fable-5-launch (2026-06-29)
- [17] The U.S. just proved that frontier closed-weight AI has a sovereign kill switch. Not because the model vanished, and not... — reactive:claude-science-launch (2026-06-29)
- [18] 🚨MAJOR AI UPDATE: The U.S. government is lifting export restrictions on Anthropic’s Claude Fable 5 - restoring worldwide... — reactive:claude-fable-5-launch (2026-07-01)
- [19] Claude Fable 5 is getting a one-week subscriber trial with a strict 50% usage ceiling. — Rohan Paul Twitter (2026-07-01)
- [20] Fable 5 absolutely crushed the HTML5 physics contest, but cost 6x more than Opus 4.8 and 39× more than GLM 5.2 in that t… — Rohan Paul Twitter (2026-07-01)
- [21] Anthropic's Claude Fable 5 plays it too safe on safety, developers say — reactive:claude-fable-5-launch
- [22] Fable 5, Overbearing safety measures? : r/ClaudeCode - Reddit — reactive:claude-fable-5-launch
- [23] Old Fable 5 vs. New Fable 5 https://t.co/4IvOHmkuBr https://t.co/0eoUFq7… — Rohan Paul Twitter (2026-07-02)
- [24] 😹 Fable 5 first reviews — The Neuron (2026-07-02)
- [25] Last week, the Chinese AI startup Zhipu AI launched its open-source model, GLM 5.2; it boasts coding capabilities rivali... — reactive:claude-fable-5-launch (2026-06-27)
- [26] GLM 5.2 vs Claude Fable 5 vs Claude Opus 4.8: the battle for the AI developer crown — reactive:claude-fable-5-launch (2026-06-25)
- [27] 🤖 Anthropic to Restore Claude Mythos 5 Access for 'Trusted' US Companies — A Partial Reversal with Major Implications. — reactive:claude-fable-5-launch (2026-06-28)