Video: "Sakana Fugu Ultra BEATS Fable 5 & GPT-5.5? (Fully Tested)" by Julian Goldie on YouTube.
What Sakana Fugu actually does
Fugu is not a new model. It is a routing layer. You send a prompt through Sakana's API, the system distributes it across a panel of models — closed and open-source — runs them in parallel, then synthesises a single reply. The idea is similar to OpenRouter's Fusion: rather than betting on one model getting the right answer, you run several and merge the best of each.
The difference Sakana is pitching is mainly cost and independence. Where Fusion draws its panel from OpenRouter's existing model catalogue, Sakana built its own. Fugu Ultra is the premium tier — a larger panel, more synthesis passes. Fugu without the Ultra suffix is the standard version, priced lower and aimed at lighter tasks.
The benchmark numbers: what to make of them
SWE-Bench Pro is a coding benchmark that presents AI systems with real-world software engineering tasks taken from open-source GitHub repositories. These are not toy problems — they require reading existing code, understanding context, and producing a working fix. It is one of the more credible tests for practical coding ability.
Fugu Ultra scored 73.7 on SWE-Bench Pro. Opus 4.8 scored 69.2. GPT-5.5 scored 58.6. Gemini 3.1 Pro scored 54.2. That is a meaningful gap at the top, not a rounding error. Worth knowing: the benchmark tests coding specifically. If your AI agent workload is mainly writing, summarising or question-answering rather than code, these numbers are background reading rather than a buying signal.
What the cost difference means in practice
Fusion on OpenRouter charges per token routed across its model panel. Sakana is claiming Fugu Ultra runs at roughly 25% of that figure for equivalent prompts. There is also a flat-rate subscription option for businesses running high volumes — which is relevant if you have an AI agent making hundreds of API calls a day.
To be fair, "equivalent prompts" is doing some work in that comparison. The two systems use different model panels and different synthesis methods, so the output quality is not identical even when the task description matches. What this does suggest is that if you are currently running Fusion and paying full OpenRouter token rates, Fugu Ultra is worth a direct comparison on your actual workloads rather than just the benchmark table.
What's genuinely useful versus what to treat carefully
The benchmark lead is real on coding tasks, and the cost argument is plausible. Those two things together make Fugu Ultra worth a proper look for teams building or running AI agents that do substantial code work. The flat-rate subscription is specifically useful if your usage is predictable and high-volume — the sort of pattern you see in automated pipelines or agent loops running overnight tasks.
What to treat carefully: Sakana AI is newer and considerably less established than Anthropic or OpenAI. API reliability, uptime guarantees, and support quality are things you find out over months of use, not from benchmark tables. The panel composition — which models are actually included — is also worth checking before committing, since the benchmark score reflects a specific set of models that could change with updates. In practice, this is a tool for AI-native builders first, not a drop-in swap for existing setups.
Where this connects to NordSys
If tracking this kind of AI-agent news makes you wonder whether one could actually run inside your business, that's exactly what our AI Agents do — named, briefed and managed for you, no setup fee, from £6 a day.
See our AI Agents →