Back to Blog
Tools & Resources6 min readSeptember 13, 2026

A Cheaper Multi-Model Swarm Beat Frontier Models on Benchmark

Sakana AI reports Fugu Ultra v2 scored 48.3 on Chartography versus 27.3 for Opus 5, using a swarm of smaller models instead of one frontier model.

Emma Watson

Emma Watson

Growth at NeedBase

Sakana AI released two new orchestration systems on 11 September 2026 โ€” Fugu Max and Fugu Ultra v2 โ€” and reported that Fugu Ultra v2 scored 48.3 on a benchmark called Chartography, against 27.3 for Anthropic's Opus 5 and 29.5 for Claude Fable 5 on the same test.

That's a large gap, on one benchmark, from a system that coordinates a pool of smaller models rather than routing everything through a single large one. It's also, according to Sakana, achieved without Claude Fable 5, Claude Fable 5.1 or GPT-6 Astra anywhere in the underlying model pool โ€” so the result isn't a smaller wrapper quietly calling a bigger frontier model and taking credit for its output.

What Fugu Max and Fugu Ultra v2 actually are

Both are orchestration systems: instead of sending a task to one large frontier model and waiting for a single response, they coordinate a swarm of smaller models against the task, combining or arbitrating between their outputs. Fugu Ultra v2 specifically ships with a 1-million-token context window and is priced at $5 per million input tokens and $30 per million output tokens, according to Sakana's release. That pricing sits well below what a comparable context window from a single frontier model typically costs, which is the point: the swarm approach is being pitched on cost as much as on capability.

The Chartography number, and why it needs a hedge

The 48.3 versus 27.3 and 29.5 scores come from Sakana's own release and from coverage by MarkTechPost and AlphaSignal reporting on that release โ€” they are not yet independently reproduced by a third-party evaluator. That doesn't make the number meaningless, but it does mean it should be read as Sakana's reported score on its own benchmark run, not as an audited, third-party-verified result. Chartography is also a single benchmark, not a general capability measure, and Sakana's own framing is that the swarm approach is better suited to tasks that decompose into pieces smaller models can handle well, not a claim that Fugu Ultra v2 outperforms Opus 5 or Claude Fable 5 across the board.

Worth noting too: the model pool underlying Fugu no longer includes Claude Fable 5, Claude Fable 5.1 or GPT-6 Astra. That's a meaningful detail because it heads off the obvious objection โ€” that a smaller-model orchestrator is really just a thin routing layer in front of a frontier model, benefiting from that model's capability while claiming a cost advantage that isn't real. Sakana's release states the pool doesn't include those models, which is why the comparison is being read as swarm-versus-single-frontier-model rather than frontier-model-wrapped-cheaply.

What this means if you're running a small team

The practical takeaway isn't "switch to Fugu" โ€” it's that orchestrating several cheaper models against a well-scoped task can, in at least this reported case, beat a single expensive frontier-model API call on cost and on at least one benchmark. That's a genuine data point in favour of a category worth watching: small-model swarms as a way to cut inference spend on tasks that break down into independently solvable pieces, rather than defaulting to the largest available model for every call.

The caveat is doing your own comparison before you act on it. A benchmark score, even a large gap like 48.3 versus 27.3, doesn't tell you how Fugu Ultra v2 performs on your specific task โ€” your prompts, your data shape, your failure modes. If your workload involves a well-defined task that decomposes naturally (parsing structured documents, extracting fields across many independent records, running the same narrow classification repeatedly), that's exactly the shape of problem swarm orchestration is reportedly suited to, and it's worth a direct side-by-side test: run a representative sample of your actual task through both your current frontier-model setup and Fugu Ultra v2, and compare cost per successful output, not just raw benchmark score.

What to actually do this week

Pick one task you currently send to a single large frontier model on a recurring basis, and check whether it decomposes into smaller independent sub-tasks โ€” if it does, it's a candidate for this category. Get access to Fugu Ultra v2 or Fugu Max and run a real cost-per-task comparison against your current setup using your own data, not a public benchmark. If the cost difference holds up on your workload, that's a stronger signal than the Chartography number by itself, since it accounts for the parts of your task Sakana's benchmark doesn't cover. Either way, treat this as one benchmark result from one vendor's own report, worth testing, not yet worth treating as settled.

The bottom line

Sakana AI's Fugu Ultra v2, released 11 September 2026, reportedly scored 48.3 on the Chartography benchmark against 27.3 for Opus 5 and 29.5 for Claude Fable 5, using a swarm of smaller models rather than a single frontier model, at $5/$30 per million input/output tokens. Treat that as Sakana's own reported figure on one benchmark, and test it against your own decomposable task before assuming the cost and capability advantage generalises.

Found this useful?

Share it with a founder who needs it.

Ready to launch your product?

Join thousands of makers who launched on NeedBase.

Submit Your Product โ†’