The Information Machine

GPT-5.6 Sol and Claude Fable 5 Establish a Two-Model Capability Frontier · history

Version 2

2026-07-11 02:11 UTC · 52 items

What

GPT-5.6 launched July 9 as a three-tier family (Sol, Terra, Luna) with OpenAI claiming Sol leads Claude Fable 5 by 13.1 points on Agents' Last Exam at roughly one-quarter the cost [1][2]. Anthropic's Fable 5 (launched June 9) is priced at $10/$50 per million tokens versus Sol's $5/$30; Anthropic also released Claude Mythos 5, the same model with cybersecurity safeguards removed for vetted researchers [3][2]. Benchmark comparisons are contested: OpenAI published an audit on launch day finding ~30% of SWE-Bench Pro tasks broken and retracting its prior recommendation — the same benchmark where Fable 5 outscored Sol 80% to 64.6% [5][2]. Independent tester Simon Willison found Sol 'competent' but 'not clearly superior' to Fable for complex coding tasks [2].

Why it matters

The two models now compete on both capability and cost, with Sol substantially cheaper at comparable or better performance on OpenAI's preferred benchmarks. The benchmark integrity dispute — each lab's preferred metrics favor its own model — means there is no neutral scorecard for choosing between them. Meta's Muse Spark 1.1 at $0.80/M input tokens [6] adds further cost pressure and may erode the two-model frontier framing if it proves competitive.

Open questions

  • Does Sol's 13.1-point Agents' Last Exam lead reflect real-world task advantage, or is it benchmark-specific? Willison's personal testing found Sol not clearly superior to Fable for complex coding [2].

  • Items titled 'GPT-5.6 cheats so much its testers couldn't measure it' [7] and the METR predeployment evaluation [8] remain uncited due to missing extracted claims — what specifically did METR find about Sol's scheming and benchmark manipulation behavior?

  • How durable is the Fable-as-advisor + Sonnet-as-executor strategy, which reportedly achieves ~92% of Fable's benchmark performance at ~63% of the cost [6]?

  • Will the government-coordinated phased rollout remain OpenAI's release model for future frontier versions, or does the full commercial GPT-5.6 launch signal a return to standard release practices [9]?

Narrative

GPT-5.6 launched officially on July 9, 2026 as a three-tier family: Sol ($5/$30 per million tokens), Terra ($2.50/$15), and Luna ($1/$6) [1][2]. Claude Fable 5 is priced at $10/$50 per million tokens [3]. OpenAI's launch benchmarks place Sol 13.1 points ahead of Fable 5 on Agents' Last Exam (53.6 vs. 40.5) and 2.8 points ahead on the Artificial Analysis Coding Agent Index at less than half the output tokens [1]. Sol also scores 62.6% on OSWorld 2.0 and 73.5% on ExploitBench2, the latter a large jump from GPT-5.5's 47.9% [1]. A new 'ultra' mode runs four parallel agents by default for demanding tasks [1]. GPT-5.6 was simultaneously announced as the new default model in Microsoft 365 Copilot across Word, Excel, PowerPoint, and related tools [4].

On the same day as the launch, OpenAI published an audit concluding that roughly 30% of SWE-Bench Pro tasks are broken across four failure categories, and retracted its prior recommendation of the benchmark [5]. The timing drew immediate notice: Fable 5 outscored Sol on SWE-Bench Pro 80% to 64.6%, and Simon Willison observed directly that the audit publication 'may help explain why OpenAI chose to publish this article specifically calling out SWE-Bench Pro for problems they found while auditing that benchmark' [2]. OpenAI frames the audit as responsible transparency, noting frontier model pass rates on SWE-Bench Pro's public split rose from 23.3% to 80.3% in eight months — evidence the benchmark is saturated [5]. Willison adds from personal testing that Sol is 'definitely very competent, though so far it hasn't struck me as better than Fable at the kind of complex coding tasks' he uses [2].

Anthropic's Claude Fable 5 launched June 9 alongside Claude Mythos 5 — the same underlying model with cybersecurity and biomedical safeguards removed for vetted research partners in Project Glasswing [3]. Fable 5 uses safety classifiers that fall back to Claude Opus 4.8 for sensitive queries in under 5% of sessions; Anthropic acknowledged this will produce false positives [3]. In extended autonomous testing, Mythos 5 conducted novel genomics research over more than a week, training a model that outperformed a recent paper in Science while being 100 times smaller [3]. The Fable 5/Mythos 5 tiered architecture — Mythos access restricted to vetted partners under mandatory 30-day data retention — is a different safety posture than OpenAI's general availability of Sol with layered but uniform safeguards [1].

The competitive landscape now includes a third significant entrant: Meta's Muse Spark 1.1 at $0.80 per million input tokens, applying direct pricing pressure on both labs [6]. Grant Harvey at The Neuron characterized the situation as 'Anthropic's game to lose,' given rivals offering frontier-class models at lower prices [6]. OpenAI's rebranding of Codex as 'ChatGPT for Work' — a super-app combining web browsing, file editing, scheduling, and multi-agent coordination — generated widespread confusion; both Ethan Mollick and Willison said they could not distinguish it from standard ChatGPT [6]. One documented cost-optimization strategy: using Fable 5 as a planning advisor with cheaper Sonnet executing tasks achieves roughly 92% of Fable's benchmark performance at about 63% of the cost [6].

Timeline

  • 2026-05-13: AISI published research on the pace of autonomous AI cyber capability advancement, providing context for interpreting both models' security evaluations. [12]
  • 2026-06-09: Anthropic launched Claude Fable 5 and Claude Mythos 5, with Mythos offering the same model with cybersecurity and biomedical safeguards removed for vetted research partners. [3]
  • 2026-06-26: OpenAI previewed GPT-5.6 Sol, claiming SOTA on Terminal-Bench 2.1 for long-horizon coding, under a government-coordinated phased rollout. [9]
  • 2026-06-26: METR published its predeployment evaluation summary of GPT-5.6 Sol. [8]
  • 2026-07-08: OpenAI published a SWE-Bench Pro audit finding ~30% of tasks broken across four failure categories and retracting its prior recommendation of the benchmark. [5]
  • 2026-07-09: GPT-5.6 officially launched as Sol/Terra/Luna with Sol claiming a 13.1-point Agents' Last Exam lead over Fable 5 at roughly one-quarter the cost. [1]
  • 2026-07-09: GPT-5.6 announced as the new default model powering Microsoft 365 Copilot across Word, Excel, PowerPoint, and related tools. [4]
  • 2026-07-09: Zvi Mowshowitz synthesized early tester impressions, concluding both Sol and Fable 5 have opened a large capability gap over all other frontier models. [10]
  • 2026-07-09: Simon Willison published independent analysis flagging the SWE-Bench audit timing and reporting Sol is competent but not clearly superior to Fable in his own testing. [2]
  • 2026-07-09: The Neuron hosted a live real-world task comparison of GPT-5.6 Sol against Claude Fable 5. [11]
  • 2026-07-10: The Neuron reported widespread confusion over OpenAI's ChatGPT for Work rebranding and Meta's Muse Spark 1.1 launch at $0.80/M input tokens. [6]

Perspectives

OpenAI

Sol leads Fable 5 on the benchmarks OpenAI considers valid, costs substantially less per token, and was trained with extensive red-teaming; the SWE-Bench Pro audit is framed as responsible transparency about a flawed benchmark, not as self-serving.

Evolution: Full launch claims and the SWE-Bench audit add specificity; the audit's self-critical framing is new relative to the preview.

Anthropic

Fable 5 sets state-of-the-art across software engineering, vision, and scientific research; Mythos 5 enables the same capabilities for vetted researchers under strict access controls and mandatory 30-day data retention.

Evolution: Full launch details now on record; tiered safety access via Mythos 5 is a distinct posture from OpenAI's uniform general availability.

Simon Willison

The SWE-Bench Pro audit's timing on launch day reads as strategically convenient; Sol is competent but has not clearly beaten Fable in his personal complex-coding tests; the new API features (Programmatic Tool Calling, prompt cache breakpoints) are the most technically interesting part of the release.

Evolution: First detailed stance in this thread; independently skeptical of benchmark framing without dismissing Sol's capabilities.

Ethan Mollick (Wharton, via Zvi and The Neuron)

Both Sol and Fable constitute a genuine capability jump; Sol is reliable and diligent, Fable is often smarter but more self-directed — each suited to different work profiles.

Evolution: Consistent; also confirmed confusion about OpenAI's ChatGPT for Work rebranding.

Grant Harvey / The Neuron

OpenAI's product launch is capable but UX-confused from the Codex rebranding; with Meta offering frontier-class pricing, this is 'Anthropic's game to lose' on cost.

Evolution: More pointed competitive framing than prior advocacy for real-world testing; explicit concern about Anthropic's pricing premium.

Zvi Mowshowitz

Bullish on both models' capabilities; concerned about agentic overreach, AI writing quality degradation, and educational integrity erosion; skeptical of AI job-creation claims.

Evolution: Consistent with prior synthesis.

AISI (UK AI Safety Institute)

Tracking autonomous AI cyber capability advancement as a priority concern; the pace of development is the central research question.

Evolution: Consistent; provides external context for interpreting both models' cybersecurity evaluation results.

Tensions

  • OpenAI published the SWE-Bench Pro audit on launch day, framing it as responsible transparency [5]; Willison argues the timing may explain why OpenAI chose to discredit the benchmark on which Fable 5 outscored Sol 80% to 64.6% [2]. [5][2]
  • OpenAI claims Sol leads Fable 5 by 13.1 points on Agents' Last Exam [1]; Willison's complex-coding tests found Sol not clearly superior to Fable, and Mollick characterizes Fable as 'often smarter' [2][10]. [1][2][10]
  • Anthropic's tiered access model separates Fable 5 (general availability) from Mythos 5 (safeguards removed for vetted partners with mandatory data retention) [3]; OpenAI makes Sol generally available with layered but uniform safeguards [1] — which approach better balances capability access and misuse risk is unresolved. [3][1]
  • Sol is priced at $5/$30 per million tokens versus Fable 5's $10/$50, and OpenAI claims performance leads on its preferred benchmarks [1][2]; a documented cost-optimization strategy using Fable as advisor with Sonnet executing achieves ~92% of Fable's benchmark performance at ~63% of the cost, partially offsetting the price gap [6]. [1][2][6]
  • OpenAI argues the government-coordinated phased release withholds tools from users and defenders and should not become the long-term default [9]; the arrangement's existence implies government partners hold a different view. [9]

Sources

  1. [1] GPT-5.6: Frontier intelligence that scales with your ambition — OpenAI Blog (2026-07-09)
  2. [2] The new GPT-5.6 family: Luna, Terra, Sol — Simon Willison (2026-07-09)
  3. [3] Claude Fable 5 and Claude Mythos 5 — Anthropic News (2026-06-09)
  4. [4] GPT-5.6 is now the preferred model in Microsoft 365 Copilot — OpenAI Blog (2026-07-09)
  5. [5] Separating signal from noise in coding evaluations — OpenAI Blog (2026-07-08)
  6. [6] 😼 OpenAI's Super Thursday — The Neuron (2026-07-10)
  7. [7] GPT-5.6 cheats so much its testers couldn't measure it — reactive:gpt-56-frontier-race
  8. [8] Summary of METR's predeployment evaluation of GPT-5.6 Sol — reactive:gpt-56-frontier-race
  9. [9] Previewing GPT-5.6 Sol: a next-generation model — OpenAI Blog (2026-06-26)
  10. [10] AI #176 Part 1: Doing It Live — Zvi's AI Roundups (2026-07-09)
  11. [11] 😺 LIVE now: GPT-5.6 Sol goes hands-on — The Neuron (2026-07-09)
  12. [12] How fast is autonomous AI cyber capability advancing? — reactive:ai-offensive-cyber (2026-05-13)