GPT-5.6 Sol and Claude Fable 5 Establish a Two-Model Capability Frontier · history
Version 3
2026-07-12 08:06 UTC · 58 items
What
GPT-5.6 Sol (launched July 9) and Claude Fable 5 (launched June 9) now occupy a shared capability tier above all other frontier models, with Sol priced at $5/$30 per million tokens versus Fable 5's $10/$50 [1][2]. OpenAI claims Sol leads Fable 5 by 13.1 points on Agents' Last Exam at roughly one-quarter the cost, but Fable 5 outscored Sol on SWE-Bench Pro 80% to 64.6% — a benchmark OpenAI retracted on launch day [4][5]. Independent testers and hands-on comparison pieces continue to appear [9][10][12], with METR publishing an estimate of Sol's autonomous task time horizon [6]. The benchmark dispute remains unresolved: each lab's preferred metrics favor its own model.
Why it matters
The two models now compete on both capability and cost, with no neutral scorecard for choosing between them. Sol is substantially cheaper; Fable 5 outperforms on some benchmarks OpenAI chose to discredit. Meta's Muse Spark 1.1 at $0.80/M input tokens adds pricing pressure that may erode the two-model framing if it proves competitive.
Open questions
What specifically did METR find about Sol's autonomous task time horizon and any scheming or benchmark-manipulation behavior? Item 40484 references METR's estimate but extracted claims are not yet available [6][7].
Does Sol's 13.1-point Agents' Last Exam lead translate to real-world task advantage? Willison's personal testing found Sol not clearly superior to Fable for complex coding, and Mollick describes Fable as 'often smarter' [5][13].
How durable is the Fable-as-advisor + Sonnet-as-executor cost strategy, which reportedly achieves ~92% of Fable's benchmark performance at ~63% of the cost [8]?
Will the government-coordinated phased rollout remain OpenAI's release model, or does the full GPT-5.6 commercial launch signal a return to standard practices [14]?
Narrative
GPT-5.6 launched officially on July 9, 2026 as a three-tier family: Sol ($5/$30 per million tokens), Terra ($2.50/$15), and Luna ($1/$6) [1]. Claude Fable 5 is priced at $10/$50 per million tokens [2]. OpenAI's launch benchmarks place Sol 13.1 points ahead of Fable 5 on Agents' Last Exam (53.6 vs. 40.5) and 2.8 points ahead on the Artificial Analysis Coding Agent Index at less than half the output tokens [1]. Sol also scores 62.6% on OSWorld 2.0 and 73.5% on ExploitBench2, the latter a large jump from GPT-5.5's 47.9% [1]. A new 'ultra' mode runs four parallel agents by default for demanding tasks, and GPT-5.6 was simultaneously announced as the new default model in Microsoft 365 Copilot [1][3].
On the same day as the launch, OpenAI published an audit concluding that roughly 30% of SWE-Bench Pro tasks are broken across four failure categories, and retracted its prior recommendation of the benchmark [4]. The timing drew immediate notice: Fable 5 outscored Sol on SWE-Bench Pro 80% to 64.6%, and Simon Willison observed that the audit's publication 'may help explain why OpenAI chose to publish this article specifically calling out SWE-Bench Pro' on launch day [5]. OpenAI frames the audit as responsible transparency, citing frontier model pass rates on the public split rising from 23.3% to 80.3% in eight months as evidence of saturation [4]. Willison adds from personal testing that Sol is 'definitely very competent, though so far it hasn't struck me as better than Fable at the kind of complex coding tasks' he uses [5].
Anthropic's Claude Fable 5 launched June 9 alongside Claude Mythos 5 — the same underlying model with cybersecurity and biomedical safeguards removed for vetted research partners in Project Glasswing [2]. Mythos 5 conducted novel genomics research over more than a week in extended autonomous testing, training a model that outperformed a recent Science paper while being 100 times smaller [2]. This tiered access model — Mythos restricted to vetted partners under mandatory 30-day data retention — differs from OpenAI's general availability of Sol with layered but uniform safeguards [1]. METR has also published an estimate of Sol's autonomous task time horizon as part of its predeployment evaluation, providing an external benchmark for agentic capability that neither lab controls [6][7].
The competitive landscape includes Meta's Muse Spark 1.1 at $0.80 per million input tokens, applying direct pricing pressure on both labs [8]. A documented cost-optimization strategy using Fable 5 as a planning advisor with cheaper Sonnet executing tasks achieves roughly 92% of Fable's benchmark performance at about 63% of the cost, partially offsetting the price gap [8]. Hands-on comparison content from independent reviewers and aggregators continues to appear in volume [9][10][11][12], reflecting sustained practitioner interest in the two-model comparison, though these have not yet yielded extracted claims that shift the overall picture.
Timeline
- 2026-05-13: AISI published research on the pace of autonomous AI cyber capability advancement, providing context for interpreting both models' security evaluations. [16]
- 2026-06-09: Anthropic launched Claude Fable 5 and Claude Mythos 5, with Mythos offering the same model with cybersecurity and biomedical safeguards removed for vetted research partners. [2]
- 2026-06-26: OpenAI previewed GPT-5.6 Sol under a government-coordinated phased rollout, claiming SOTA on Terminal-Bench 2.1 for long-horizon coding. [14]
- 2026-06-26: METR published its predeployment evaluation summary of GPT-5.6 Sol, including an estimate of Sol's autonomous task time horizon. [7][6]
- 2026-07-08: OpenAI published a SWE-Bench Pro audit finding ~30% of tasks broken across four failure categories and retracting its prior recommendation of the benchmark. [4]
- 2026-07-09: GPT-5.6 officially launched as Sol/Terra/Luna, with Sol claiming a 13.1-point Agents' Last Exam lead over Fable 5 at roughly one-quarter the cost. [1]
- 2026-07-09: GPT-5.6 announced as the new default model powering Microsoft 365 Copilot across Word, Excel, PowerPoint, and related tools. [3]
- 2026-07-09: Zvi Mowshowitz synthesized early tester impressions, concluding both Sol and Fable 5 have opened a large capability gap over all other frontier models. [13]
- 2026-07-09: Simon Willison published independent analysis flagging the SWE-Bench audit timing and reporting Sol is competent but not clearly superior to Fable in his own testing. [5]
- 2026-07-09: The Neuron hosted a live real-world task comparison of GPT-5.6 Sol against Claude Fable 5. [15]
- 2026-07-10: Meta's Muse Spark 1.1 launched at $0.80/M input tokens; The Neuron reported widespread confusion over OpenAI's ChatGPT for Work rebranding. [8]
- 2026-07-11: Multiple independent comparison pieces and hands-on reviews of Sol vs. Fable 5 appeared across Medium, YouTube, and Substack, with no extracted claims shifting the existing picture. [9][10][12][17]
Perspectives
OpenAI
Sol leads Fable 5 on the benchmarks OpenAI considers valid, costs substantially less per token, and the SWE-Bench Pro audit is responsible transparency about a saturated benchmark, not strategic timing.
Evolution: Consistent with launch framing; no new statements this pass.
Anthropic
Fable 5 sets state-of-the-art across software engineering, vision, and scientific research; Mythos 5 enables the same capabilities for vetted researchers under strict access controls and mandatory 30-day data retention.
Evolution: Consistent; tiered safety access via Mythos 5 remains a distinct posture from OpenAI's general availability.
Simon Willison
The SWE-Bench Pro audit's launch-day timing reads as strategically convenient; Sol is competent but has not clearly beaten Fable in his personal complex-coding tests; the new API features are the most technically interesting part of the release.
Evolution: Consistent; the most detailed independent skeptical voice on benchmark framing.
Ethan Mollick (Wharton, via Zvi and The Neuron)
Both Sol and Fable constitute a genuine capability jump; Sol is reliable and diligent, Fable is often smarter but more self-directed — each suited to different work profiles.
Evolution: Consistent; also confirmed confusion about OpenAI's ChatGPT for Work rebranding.
Grant Harvey / The Neuron
OpenAI's product launch is capable but UX-confused from the Codex rebranding; with Meta offering frontier-class pricing, this is 'Anthropic's game to lose' on cost.
Evolution: Consistent; explicit concern about Anthropic's pricing premium relative to Sol and Muse Spark.
Zvi Mowshowitz
Bullish on both models' capabilities; concerned about agentic overreach, AI writing quality degradation, and educational integrity erosion; skeptical of AI job-creation claims.
Evolution: Consistent.
METR
Published an estimate of Sol's autonomous task time horizon as part of predeployment evaluation; tracking autonomous AI cyber capability advancement as a priority concern.
Evolution: The autonomous task time horizon estimate is now attributed to METR via item 40484, adding specificity to what was previously a gap in this thread.
Tensions
- OpenAI published the SWE-Bench Pro audit on launch day, framing it as responsible transparency [4]; Willison argues the timing may explain why OpenAI chose to discredit the benchmark on which Fable 5 outscored Sol 80% to 64.6% [5]. [4][5]
- OpenAI claims Sol leads Fable 5 by 13.1 points on Agents' Last Exam [1]; Willison's complex-coding tests found Sol not clearly superior to Fable, and Mollick characterizes Fable as 'often smarter' [5][13]. [1][5][13]
- Anthropic's tiered access model separates Fable 5 (general availability) from Mythos 5 (safeguards removed for vetted partners with mandatory data retention) [2]; OpenAI makes Sol generally available with layered but uniform safeguards [1] — which approach better balances capability access and misuse risk is unresolved. [2][1]
- Sol is priced at $5/$30 per million tokens versus Fable 5's $10/$50 with OpenAI claiming performance leads on its preferred benchmarks [1]; a Fable-as-advisor plus Sonnet-as-executor strategy achieves ~92% of Fable's benchmark performance at ~63% of the cost, partially offsetting the gap [8]. [1][8]
- OpenAI argued the government-coordinated phased release withholds tools from users and defenders and should not become the long-term default [14]; the arrangement's existence implies government partners held a different view. [14]
Sources
- [1] GPT-5.6: Frontier intelligence that scales with your ambition — OpenAI Blog (2026-07-09)
- [2] Claude Fable 5 and Claude Mythos 5 — Anthropic News (2026-06-09)
- [3] GPT-5.6 is now the preferred model in Microsoft 365 Copilot — OpenAI Blog (2026-07-09)
- [4] Separating signal from noise in coding evaluations — OpenAI Blog (2026-07-08)
- [5] The new GPT-5.6 family: Luna, Terra, Sol — Simon Willison (2026-07-09)
- [6] METR estimates GPT-5.6 Sol's autonomous task time horizon at ... — reactive:gpt-56-frontier-race
- [7] Summary of METR's predeployment evaluation of GPT-5.6 Sol — reactive:gpt-56-frontier-race
- [8] 😼 OpenAI's Super Thursday — The Neuron (2026-07-10)
- [9] GPT-5.6 Sol vs GPT-5.5 and Claude Fable 5: First Look | Medium — reactive:gpt-56-frontier-race
- [10] Fable 5 vs GPT 5.6 Sol: The Early Results - YouTube — reactive:gpt-56-frontier-race
- [11] GPT-5.6 Sol vs Claude Fable 5 - AI Model Comparison | OpenRouter — reactive:gpt-56-frontier-race
- [12] Sol vs. Fable: What the Two New AI Models Mean for People Who Don't Code — reactive:gpt-56-frontier-race
- [13] AI #176 Part 1: Doing It Live — Zvi's AI Roundups (2026-07-09)
- [14] Previewing GPT-5.6 Sol: a next-generation model — OpenAI Blog (2026-06-26)
- [15] 😺 LIVE now: GPT-5.6 Sol goes hands-on — The Neuron (2026-07-09)
- [16] How fast is autonomous AI cyber capability advancing? — reactive:ai-offensive-cyber (2026-05-13)
- [17] Claude Fable 5 Is INSANE – Hands-On With the BEST Model Yet! — reactive:claude-fable-5-mythos-launch