The Information Machine

Sakana AI Fugu Ultra: Multi-Model Orchestration Layer Launch and Early Benchmarks · history

Version 7

2026-06-29 18:55 UTC · 159 items

What

Sakana AI's Fugu Ultra, launched June 22, 2026, is an orchestration layer routing tasks across GPT-5, Claude Sonnet 4, and Gemini 2.5 Pro via a 7B RL-trained coordinator behind a single API endpoint.[2][1] Sakana's self-reported benchmarks claim parity with Anthropic's Fable 5 and Mythos; no independent peer-reviewed evaluation has published.[2][25] Coverage divides between enthusiast amplification treating Fugu as a frontier model release, a critical reading that it is an orchestration layer over other labs' models, and an affirmative counter-thesis — stated explicitly by Prasenjit Sarkar on June 28 — that 'The model is no longer the product. The orchestration layer is.'[11] A wave of secondary analysis from 36kr, Towards AI, Medium, and ayautomate.com shows the story spreading internationally without introducing new arguments.[21][22][23][24]

Why it matters

If orchestration over existing APIs genuinely matches frontier model performance, it suggests AI capabilities can be assembled without new training runs, affecting both development economics and the practical reach of weight-based export restrictions. Whether Fugu's trained routing meaningfully differs from existing model-routing services — architecturally and in practice — remains the central unresolved question.

Open questions

  • Will informal tests — Reddit personal benchmarks [18], YouTube comparisons [19][20], LinkedIn battle tests [26], and the Sakana Fugu vs Fable 5 benchmarks on ayautomate.com [24] — converge on consistent conclusions, or remain too scattered to be meaningful?

  • Does Rohan Paul's description of Fugu-Ultra constructing a distinct multi-model workflow for each query [14] describe a substantively novel capability, or is it what existing routing services already do?

  • If Anthropic's Mythos model is subject to export restrictions [13], does that validate or complicate Sakana's claim that API routing avoids export control risk?

  • Does the 17x cost premium over alternatives like GLM 5.2 reflect genuine capability differences, or primarily the overhead of routing through multiple frontier APIs simultaneously? [16][17]

Narrative

Sakana AI, the Tokyo-based lab co-founded by former Google Brain researchers, launched Fugu and Fugu Ultra on June 22, 2026. The system is not a new large language model but an orchestration layer built around a 7B parameter coordinator model trained with reinforcement learning.[1] That coordinator decomposes incoming tasks into subtasks and routes each to one of a pool of models — currently GPT-5, Claude Sonnet 4, and Gemini 2.5 Pro — then synthesizes outputs behind a single OpenAI-compatible endpoint.[2][3] Sakana's self-reported technical report, now on arXiv, claims Fugu Ultra matches Fable 5 and Mythos on most standard benchmarks, supported by a 500-user beta showing progress on fully automated data science and cybersecurity tasks.[2][4][5] The underlying research — 'Learning to Orchestrate Agents in Natural Language with the Conductor,' accepted at ICLR 2026 — provides the academic grounding for the architecture.[6]

Critical reception divides along consistent lines. Danny Livshits argued on launch day that Fugu is being misrepresented as a frontier model when it is an orchestration layer over other labs' models;[7] paddo.dev published a dedicated post titled 'A Multi-Agent System Sold as a Model: Sakana's Fugu,'[8] and @kelterix amplified the same observation independently.[9] Against this, @jumperz offered an affirmative counter-reading — 'fugu might be one of the first AI products where the model is not really the product... but the orchestration is'[10] — and Prasenjit Sarkar stated the thesis directly on June 28: 'The model is no longer the product. The orchestration layer is.'[11] Sakana marketed Fugu Ultra as delivering frontier capability 'without the risk of export controls' because it routes API calls rather than distributing model weights;[12] a Windows Forum discussion titled 'Mythos Export Ban Sparks Sakana and 360: AI Security Becomes Geopolitical Dependency' added unresolved context to how that framing holds in practice.[13]

Rohan Paul's June 28 analysis of the technical report adds architectural specificity. The coordinator is trained from data to learn which model performs best for each task type — including fine-grained distinctions like routing debugging-heavy code problems to debugging-specialized models — rather than applying handcrafted rules.[14] Fugu-Ultra constructs a distinct multi-model workflow for each individual query at inference time rather than applying a fixed pipeline, which Rohan Paul contrasts with most existing multi-model systems that rely on majority voting or static per-domain routing.[14] Skeptics continue to ask how trained routing differs substantively from model-routing services already on the market.[15]

Cost remains a parallel concern. Live coding tests found Fugu Ultra produces the richest UI output at roughly 17x the cost of alternatives, with GLM 5.2 performing comparably on overall metrics at a fraction of the price;[16] Hemant argues the cost story is more significant than the benchmark story.[17] Informal independent testing has accumulated — a Reddit user found Fugu Ultra better than Fable in a personal benchmark[18] alongside YouTube comparisons[19][20] — but no peer-reviewed external evaluation has published. Coverage has now spread to Chinese-language international tech press, with 36kr asking 'Has an AI Model Surpassing Claude Mythos Been Born?'[21] and to Towards AI, Medium, and ayautomate.com,[22][23][24] indicating widening international awareness without new substantive arguments.

Timeline

  • 2026-06-22: Sakana AI launches Fugu and Fugu Ultra: a 7B RL-trained coordinator routing tasks across GPT-5, Claude Sonnet 4, and Gemini 2.5 Pro behind a single OpenAI-compatible endpoint. [2][41][1]
  • 2026-06-22: Sakana's self-reported technical report claims Fugu Ultra matches Fable 5 and Mythos on most benchmark evaluations. [2][42][29]
  • 2026-06-22: Live coding test finds Fugu Ultra produces the richest UI output at ~17x the cost of alternatives; GLM 5.2 performs close on overall metrics. [16]
  • 2026-06-22: Danny Livshits argues media coverage misrepresents Fugu as a frontier model when it is an orchestration layer over existing models. [7]
  • 2026-06-22: Peter Wildeford publicly expresses skepticism about Fugu Ultra's benchmark claims. [25]
  • 2026-06-22: VentureBeat publishes technical detail on how Sakana trained the 7B RL conductor to orchestrate multiple frontier APIs. [1]
  • 2026-06-23: Sakana cites 500-user beta results, claiming meaningful progress on fully automated data science and cybersecurity tasks. [4]
  • 2026-06-23: Sakana markets Fugu Ultra as delivering frontier capability 'without the risk of export controls'; Chris Albon quotes the line publicly with implied skepticism. [12][43]
  • 2026-06-23: Hemant argues Fugu Ultra's billing is more notable than its benchmark story, framing cost as the primary practical concern. [17]
  • 2026-06-23: Multiple accounts circulate threads titled 'THE GEOPOLITICAL ANGLE NOBODY IS TALKING ABOUT,' amplifying the export controls framing. [44][33][34][45]
  • 2026-06-25: Sakana technical report published on arXiv; Requesty.ai publishes a reverse engineering of the multi-agent orchestration architecture. [5][46]
  • 2026-06-25: YouTube comparisons against Fable 5 and GPT-5.5, Reddit personal benchmarks, and a LinkedIn battle test appear as informal independent testing accumulates. [19][20][26]
  • 2026-06-25: @kelterix posts that 'the AI that just beat Anthropic's Fable 5 on benchmarks isn't actually a model,' giving the 'not a model' framing independent social amplification. [9]
  • 2026-06-26: Paddo.dev publishes 'A Multi-Agent System Sold as a Model: Sakana's Fugu,' giving the critical framing a dedicated long-form post. [8]
  • 2026-06-26: @jumperz posts that Fugu 'might be one of the first AI products where the model is not really the product... but the orchestration is,' reframing the 'not a model' critique as a product thesis. [10]
  • 2026-06-26: A Reddit user in r/LLMDevs adds Fugu Ultra to a personal benchmark and reports it better than Fable, one of the first informal external benchmark results. [18]
  • 2026-06-27: Jay.TL describes Fugu as 'the cleverest packaging trick in AI, a single API over swappable frontier models,' framing it as clever abstraction rather than a frontier advance. [40]
  • 2026-06-28: Rohan Paul analyzes the Fugu technical report, noting the coordinator learns routing from data rather than handcrafted rules and constructs distinct per-query multi-model workflows at inference time. [14]
  • 2026-06-28: Prasenjit Sarkar states explicitly: 'The model is no longer the product. The orchestration layer is,' independently reinforcing @jumperz's earlier product-thesis reframing. [11]
  • 2026-06-29: Coverage spreads to Chinese-language tech press (36kr), Towards AI, Medium, and ayautomate.com, indicating widening international awareness without new substantive arguments. [21][22][23][24]

Perspectives

Sakana AI

A small RL-trained coordinator reaching frontier benchmark parity by orchestrating existing models is a viable alternative to training ever-larger base models; API routing also sidesteps export control risk by distributing no weights.

Evolution: Export controls framing was added June 23; otherwise consistent with the lab's prior research direction toward learned coordination over scale.

Danny Livshits / @kelterix / paddo.dev

Fugu is an orchestration layer over other labs' models, not a frontier model, and marketing it as one misrepresents what Sakana actually built.

Evolution: Livshits stated this on launch day; by June 25-26 the 'not actually a model' framing gained independent amplification through separate accounts and a dedicated blog post.

@jumperz / Prasenjit Sarkar (orchestration-as-product thesis)

The orchestration itself is the product — Fugu may be one of the first AI products where the model is incidental and the routing logic is the actual value. Sarkar stated this directly on June 28: 'The model is no longer the product. The orchestration layer is.'

Evolution: @jumperz introduced this affirmative reframing on June 26; Sarkar made an identical, independent statement June 28, giving the thesis a second named voice.

Rohan Paul / Prasenjit Sarkar (technical analysis)

The trained routing approach is a meaningful departure from static multi-model systems; per-query dynamic workflow construction and the single API abstraction are the central technical achievements.

Evolution: Prasenjit Sarkar has been consistent since launch; Rohan Paul added architectural specificity on June 28 by summarizing the technical report approvingly.

Peter Wildeford

Skeptical that Fugu Ultra's benchmark claims reflect genuine frontier-level performance.

Evolution: Consistent; no new statement since launch day.

Hemant (@heman10x)

Cost is the more durable concern: Fugu Ultra produces the richest output at ~17x the cost of alternatives, making billing the more interesting story than benchmarks.

Evolution: Consistent since June 23.

Tech media (VentureBeat, GovInfoSecurity, 36kr, Towards AI, DataCamp)

Cover the launch as substantively novel; international outlets including 36kr and Towards AI frame it as a potential frontier model alternative, while GovInfoSecurity frames it as a strategic bet on orchestration over frontier model training.

Evolution: Coverage has expanded from enterprise security outlets to mainstream tutorial sites and international tech press, including Chinese-language outlets, with framing consistently qualified on stronger performance claims.

Enthusiast amplifiers (Julian Goldie SEO, Digital Ledger, @p0lybender, Jay.TL)

Frame Fugu variously as Japan matching frontier AI without US export controls, as evidence orchestration is the next development lever, or as 'the cleverest packaging trick in AI.'

Evolution: Amplification continued through June 27; Jay.TL's 'packaging trick' framing is somewhat more skeptical than earlier enthusiast posts, though still celebratory.

Tensions

  • Sakana claims Fugu Ultra matches Fable 5 and Mythos on most benchmarks [2]; Peter Wildeford disputes those claims [25]; all formal benchmark data remains self-reported with no independent peer-reviewed verification. [2][25]
  • Danny Livshits, @kelterix, and paddo.dev argue Fugu is an orchestration layer being sold as a frontier model [7][9][8]; enthusiast accounts and most social media coverage treat it as equivalent to a frontier model release [36]. [7][9][8][36]
  • @jumperz and Prasenjit Sarkar argue the orchestration is the product and the framing debate misses the point [10][11]; critics argue the 'not a model' framing is precisely the misrepresentation that matters [8]. [10][11][8]
  • Sakana markets Fugu Ultra as avoiding export control risk because it routes API calls rather than distributing weights [12]; the Windows Forum thread on a 'Mythos Export Ban' adds unresolved context to whether that framing holds practically [13]. [12][13]
  • Fugu Ultra produces the richest output in practical coding tests [16] but at ~17x the cost of alternatives; Hemant argues the cost story is more significant than the benchmark story [17]. [16][17]
  • Rohan Paul argues the coordinator's trained routing — learning from data rather than applying handcrafted rules — is a meaningful departure from existing multi-model systems [14]; critics ask how it differs substantively from routing services already on the market [15]. [14][15]

Sources

  1. [1] How Sakana trained a 7B model to orchestrate GPT, Claude and ... — reactive:sakana-fugu-ultra
  2. [2] Sakana AI has unveiled Fugu Ultra, an orchestration layer that assembles and routes subtasks across a pool of models th… — Rohan Paul Twitter (2026-06-22)
  3. [3] The detail people are skipping in Sakana's Fugu launch: it isn't a framework you wire up, it's a single OpenAI-compatibl... — reactive:sakana-fugu-ultra (2026-06-23)
  4. [4] Sakana AI on X: "Benchmarks tell only part of the story. Fugu’s real value shows up in long, messy, real-world workflows. During our beta with 500 users, we saw Fugu Ultra drive meaningful progress in fully automated tasks from data science to complete cybersecurity assessments. Our early users https://t.co/lbTOOJYqIJ" / X — reactive:sakana-fugu-ultra
  5. [5] Sakana Fugu Technical Report - arXiv — reactive:sakana-fugu-ultra
  6. [6] Sakana AI on X: "Introducing our new work: “Learning to Orchestrate Agents in Natural Language with the Conductor” accepted at #ICLR2026 https://t.co/31QhVGCSzq What if we trained an AI not to solve problems directly, but to act as a manager that delegates tasks to a diverse team of other AIs? https://t.co/1BcqayXSGl" / X — reactive:sakana-fugu-ultra
  7. [7] Everyone sharing Sakana's Fugu launch presenting it as if the lab shipped a frontier model. They shipped an orchestratio... — reactive:sakana-fugu-ultra (2026-06-22)
  8. [8] A Multi-Agent System Sold as a Model: Sakana's Fugu — reactive:sakana-fugu-ultra
  9. [9] THE AI THAT JUST BEAT ANTHROPIC'S FABLE 5 ON BENCHMARKS ISN'T ACTUALLY A MODEL. — reactive:sakana-fugu-ultra (2026-06-25)
  10. [10] fugu might be one of the first AI products where the model is not really the product... but the orchestration is. — reactive:sakana-fugu-ultra (2026-06-26)
  11. [11] The model is no longer the product. The orchestration layer is. — reactive:sakana-fugu-ultra (2026-06-28)
  12. [12] Chris Albon on X: ""Our ‘Fugu Ultra’ model matches the performance of Fable and Mythos, delivering frontier capability without the risk of export controls." me: https://t.co/o5DVG278a5" / X — reactive:sakana-fugu-ultra
  13. [13] Mythos Export Ban Sparks Sakana and 360: AI Security Becomes ... — reactive:sakana-fugu-ultra
  14. [14] Sakana Fugu Technical Report — Rohan Paul Twitter (2026-06-28)
  15. [15] How is this different from what a model-routing company like Perplexity is already doing from the past 3 years? — reactive:sakana-fugu-ultra (2026-06-22)
  16. [16] Sakana Fugu Ultra just beat the other models on visual polish in a live trading-desk coding test, got close to GLM 5.2, … — Rohan Paul Twitter (2026-06-22)
  17. [17] Fugu Ultra’s benchmark story is less interesting than its bill. — reactive:sakana-fugu-ultra (2026-06-23)
  18. [18] Added the Sakana Fugu Ultra model to my personal benchmark (better than Fable) : r/LLMDevs — reactive:sakana-fugu-ultra
  19. [19] Sakana Fugu Ultra BEATS Fable 5 & GPT-5.5? (Fully Tested) — reactive:sakana-fugu-ultra
  20. [20] Added the Sakana Fugu Ultra model to my personal benchmark ... — reactive:sakana-fugu-ultra
  21. [21] Has an AI Model Surpassing Claude Mythos Been Born? — reactive:sakana-fugu-ultra
  22. [22] What's Better than Mythos 5?. Sakana's Fugu is a multi-agent… — reactive:sakana-fugu-ultra
  23. [23] Sakana Fugu Might Be the Most Important AI Release Nobody Noticed — reactive:sakana-fugu-ultra
  24. [24] Sakana Fugu vs Fable 5: Benchmarks & Verdict (2026) — reactive:sakana-fugu-ultra
  25. [25] I really do not believe that 'Fugu Ultra' "matches the performance of ... — reactive:sakana-fugu-ultra
  26. [26] I Battle Tested Sakana Fugu's Fable Killer Fugu just went viral ... — reactive:sakana-fugu-ultra
  27. [27] The detail worth sitting with from Sakana AI's Fugu launch isn't the benchmark line, it's the AutoResearch run. Fugu Ult... — reactive:sakana-fugu-ultra (2026-06-22)
  28. [28] A developer testing Sakana's new Fugu system left one line worth more than the benchmark grid: on code review, "where ot... — reactive:sakana-fugu-ultra (2026-06-22)
  29. [29] Sakana AI's Fugu orchestrates multiple LLMs to match Anthropic's ... — reactive:sakana-fugu-ultra
  30. [30] Sakana AI Bets on Agent Orchestration Over Frontier Models — reactive:sakana-fugu-ultra
  31. [31] Sakana AI Launches Sakana Fugu: An Orchestration Model That ... — reactive:sakana-fugu-ultra
  32. [32] Sakana Fugu: Features, Benchmarks, and How It Works - DataCamp — reactive:sakana-fugu-ultra
  33. [33] 6. THE GEOPOLITICAL ANGLE NOBODY IS TALKING ABOUT — reactive:sakana-fugu-ultra (2026-06-23)
  34. [34] Sakana AI just launched a Fable-killer that bypasses geopolitical restrictions — reactive:sakana-fugu-ultra (2026-06-23)
  35. [35] 🐡 Meet Fugu: Japan’s Clever Hack Around AI Export Controls — reactive:sakana-fugu-ultra (2026-06-25)
  36. [36] THE FABLE KILLER IS HERE, AND IT’S ORCHESTRATING A GLOBAL REVOLUTION — reactive:sakana-fugu-ultra (2026-06-25)
  37. [37] https://t.co/3fqWt3dV41 — reactive:sakana-fugu-ultra (2026-06-27)
  38. [38] https://t.co/EMx9DoV93S — reactive:sakana-fugu-ultra (2026-06-27)
  39. [39] https://t.co/Exy5GSWgOU — reactive:sakana-fugu-ultra (2026-06-27)
  40. [40] Sakana Fugu: the cleverest packaging trick in AI. A single API over swappable frontier models. — reactive:sakana-fugu-ultra (2026-06-27)
  41. [41] Sakana Fugu: One Model to Command Them All — reactive:sakana-fugu-ultra
  42. [42] Sakana Fugu Ultra Beats Fable on Benchmarks — reactive:sakana-fugu-ultra
  43. [43] 🚨 NEW ALPHA: Japan just matched Claude Fable without US export controls. — reactive:sakana-fugu-ultra (2026-06-23)
  44. [44] 6. THE GEOPOLITICAL ANGLE NOBODY IS TALKING ABOUT — reactive:sakana-fugu-ultra (2026-06-24)
  45. [45] 3. Frontier Benchmarks & Geopolitical Arbitrage — reactive:sakana-fugu-ultra (2026-06-23)
  46. [46] Inside Sakana Fugu Ultra: We Reverse Engineered Its Multi Agent ... — reactive:sakana-fugu-ultra