On Monday, Sakana AI published benchmark scores that put its new Fugu Ultra shoulder-to-shoulder with Anthropic's Fable 5. By Wednesday, the people who actually ran it were reporting 30-minute waits and a verdict that landed harder than any leaderboard: "fine, but not Fable."
Two days. That's how long the number survived contact with real use.
The bet, and the receipt
Fugu is a real bet, and an interesting one. Sakana's pitch is to stop training frontier models and start conducting them. A 7B "conductor" routes each task across a swappable pool of frontier LLMs behind one OpenAI-compatible endpoint — and on paper it works. Fugu Ultra posted frontier-level marks, GPQA-Diamond 95.5 and LiveCodeBench 93.2, level with Fable 5 and Mythos Preview. Frontier performance without training a frontier model is a genuinely good story.
Then Ethan Mollick ran it. His shader and interactive-scene coding tests double as an informal community benchmark, and on his "Harbor Town" 3D simulation, Fugu Ultra was "incredibly slow" — runs taking 30 minutes — and the output, while "fine," "does not match Fable in real use." Hacker News compressed the whole launch into one line: "a premium model router with a very good marketing story." Sakana's published scores have not been independently reproduced.
A benchmark is confidence. Real use is the audit.
Here's the part worth slowing down on. The problem isn't that Fugu used many models. It's that it hid them.
Fugu's design choice — its actual selling point — is that model selection, delegation, verification, and synthesis happen internally and invisibly. One endpoint, one number, no seams. Which is exactly why a 30-minute run is so hard to diagnose: you can't see which models it called, why it stalled, or where the synthesis lost the thread. You get a confident score, and when real use contradicts it, there's nothing to look at.
That's the trade. A benchmark is confidence. Real use is the audit. And the audit is precisely the thing the design abstracts away.
The research has been circling this all year
This is the cleaner version of a point the literature keeps making. A June paper, The Confident Liar (arXiv:2606.10296), found that a model's stated confidence barely tracks whether its reasoning is actually sound. A benchmark is that same confidence one level up — a number that says "trust me" without showing the work. Another paper this month (arXiv:2606.04223) argued that collapsing a multi-model process into a single answer throws away the signal you'd need to trust it. Fugu does both at once: it manufactures system-level confidence and collapses the deliberation into one endpoint.
Sakana isn't alone in the move. Salesforce shipped Agentforce Multi-Agent Orchestration nine days earlier, and the sharpest criticism was the "seam problem" — Atlas routes by reading each subagent's plain-language description, and "no single agent holds a global view… there is no single place where end-to-end logic is reviewed." Two production answers to "use more than one model," nine days apart, and both make the same choice: route the work, hide the reasoning, return one result. Salesforce removes the reviewer. Sakana abstracts the reviewer into an endpoint. Fugu's stumble is the first widely-cited receipt for what that abstraction costs.
The honest version isn't a dunk
Mollick's result doesn't prove multi-model loses. It proves that hidden multi-model can't be audited when it underperforms. "More models" was never automatically better, and Mollick just showed it can be slower and no smarter. That concession matters, because the council bet was always narrower and stranger than "more models win." It's that on the decisions that matter, you want to watch the models disagree — not trust a hidden router's score.
Watch for the demand that follows. "Show me the routing, show me the deliberation" is the practitioner reflex this kind of result trains, and it points at the one surface orchestration products are built to hide.
That surface is the whole design difference. When a council runs on Shingikai, the deliberation is the product, not the thing the product buries. You watch it stream live. Red Team vs. Blue Team forces the disagreement into the open instead of smoothing it away. Chairperson Synthesis decides what to do with a split — it doesn't pretend the split never happened. Same raw ingredient as Fugu, several models on one question, opposite philosophy about whether you get to see the part that went wrong.
And to be fair to the other side: not every question needs that. Sometimes one good answer is enough and the council is overhead you don't want, which is why Quick Take exists. The argument was never "always convene a council." It's "when wrong is expensive, you should be able to see why the answer is what it is."
Sakana built something real — learned orchestration, frontier marks without a frontier model. But the moment the number was wrong, there was nothing to inspect. That's the tax on hiding the deliberation behind one score.
A council's whole pitch is that you don't have to pay it. You can watch it work.
Try it free — no signup. shingik.ai