Model Reviews

Sakana Fugu Measured: $57.06 a Thousand Tasks (and $2.19 From the Same Vendor)

Sakana Fugu is new to our set, and it arrived holding a record we did not expect anyone to take. sakana/fugu-ultra-v2 scored 9 out of 9 on our executed Python benchmark at a derived $57.06 per 1,000 tasks — the most expensive result we have ever measured, priced at list on 2026-09-15. It did not get there by thinking: it spends 7 reasoning tokens per call. It got there on 60,645 input tokens across nine short prompts. The same vendor's sakana/fugu-max scored the identical 9 out of 9 for $2.19, same day, same nine tasks.

DataLLM Lab article cover: Sakana Fugu Measured: $57.06 a Thousand Tasks (and $2.19 From the Same Vendor)

We have measured 82 entries. The intuitive story about an expensive endpoint is that it costs a lot because it thinks a lot. Not at the top of our table: Sakana's record-setting endpoint spends 7 reasoning tokens per call, and the second-priciest 9/9 result in our set, GPT-6 Astra Pro at $35.44, spends 111. Both got expensive on token volume instead, which changes what the record means.

The two results

Metricsakana/fugu-ultra-v2sakana/fugu-max
Score9/99/9
Derived cost / 1,000 tasks$57.06$2.19
Mean latency23.7s12s
Reasoning tokens per call7424
Tokens across the suite60,645 in / 7,011 out4,466 in / 1,791 out
List price in / out per 1M$5 / $30$2 / $6
Context window1,000,0001,000,000
Rank among the 56 at 9/956th cheapest · 52nd fastest26th cheapest · 37th fastest
Run date / price date2026-09-16 / 2026-09-152026-09-16 / 2026-09-15

Both prices are derived from measured token counts at list price captured 2026-09-15. Divide them and you get the headline: $57.06 ÷ $2.19 = 26.05. One vendor, one benchmark run, one day, and a 26-fold spread with no difference in score.

Where $57.06 actually comes from

Our nine prompts are short. A function signature, a spec, no example tests. Models that clear the suite land in the high hundreds of input tokens for the whole run: GPT-6 Astra used 604, Ling 3.0 Flash VL used 741, Claude Fable 5.1 used 914. fugu-ultra-v2 used 60,645, on the same nine prompts: 60,645 ÷ 604 = 100.4.

Input tokens consumed across the same nine promptsEvery model below scored 9/9 on the identical nine tasks. Only the input volume differs.Sakana fugu-ultra-v260,645GPT-6 Astra Pro17,011Sakana fugu-max4,466Claude Fable 5.1914Ling 3.0 Flash VL741GPT-6 Astra604One scale throughout: 0.008 px per input token — 1,000 input tokens = 8 px. Counts are suite totals from runs measured 2026-09-16.
Six models, one identical set of nine prompts, and a hundredfold spread in what the endpoint counted as input.

Now weight those token counts by their list prices, captured 2026-09-15. The arithmetic is one line each and it is exact:

The larger half of this bill is the half that costs six times less per token — $30 ÷ $5 = 6, exactly. That is the whole finding. An endpoint that charges $5 per million input tokens is not expensive unless something makes it eat sixty thousand of them, and on prompts this short, something did. We cannot see inside the endpoint, so we cannot tell you whether that is a system preamble, an internal scaffold, or retried context being re-billed. We can only tell you the meter read 60,645.

This is not unique to Sakana. The second-priciest 9/9 in our set, GPT-6 Astra Pro at $35.44 per 1,000 tasks — also priced 2026-09-15 — shows the same shape: 17,011 input tokens against 2,977 out, where the plain GPT-6 Astra on identical prompts used 604 in. Two vendors, two top-of-line endpoints, the same inflation on the cheap side of the price card. If you are modelling spend from a price sheet, that is the failure mode to plan for: the rate card was honest and the token count was the surprise.

For contrast, the most reasoning-heavy model that clears our suite spends 1,327 reasoning tokens per call. fugu-ultra-v2 spends 7. Whatever the $57.06 bought, it was not deliberation.

fugu-max is the one that makes sense

The cheaper Sakana endpoint has a normal cost shape. Same list-price capture date of 2026-09-15, same exact arithmetic:

Price-weighted sidefugu-ultra-v2fugu-max
Input60,645 × $5 = 303,2254,466 × $2 = 8,932
Output7,011 × $30 = 210,3301,791 × $6 = 10,746
Which side dominatesInputOutput
Derived cost / 1,000 tasks$57.06$2.19

fugu-max pays most of its bill for tokens it generated, which is what a normal API call looks like. Its input total is still high for prompts this short — 4,466 against Ling's 741 — but the gap to its sibling is large: 60,645 ÷ 4,466 = 13.58. At 12 seconds mean it is 37th fastest of the 56 models at 9/9, and 26th cheapest. Mid-table on both axes, and it solved every task.

So the practical recommendation for anyone evaluating Sakana Fugu is short: on work shaped like ours, fugu-max returns the identical score for a 26th of the cost. We wrote about the vendor's orchestration story in our earlier Fugu review; this is the first time we have had our own token counts for their endpoints, and the token counts are the part that matters.

What nine Python functions cannot tell you

Here is the honest ceiling on all of this: nine self-contained Python functions cannot separate a frontier model from a competent small one. Fifty-six models in our set score 9 out of 9. The suite saturates. Ling 3.0 Flash VL clears it for $0.07 per 1,000 tasks, priced 2026-09-15 — the same perfect score as the $57.06 endpoint.

That does not make the $57.06 meaningless. It means the number answers one question and not another. It cannot tell you whether fugu-ultra-v2 is better at long-horizon agent work, at multi-file refactors, or at anything its vendor built it for. It can tell you, precisely, what an endpoint charges to do work this size — and that a nine-prompt errand on fugu-ultra-v2 costs more than the same errand on any other endpoint we have metered. If your workload is mostly small, well-specified calls, that is the number that governs your bill, and it moves.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived from measured token counts at list price on the date stated beside each figure — it is not a billing statement, and it does not include any discount, cache rate or commitment you may have. Runs go through OpenRouter.

An API-layer failure is recorded separately from a wrong answer. A run that does not complete all nine is marked excluded and carries no score at all, comparable to nothing: ibm-granite/granite-4.2-8b in this same sweep completed only 7 of 9 after retries, so it is excluded and appears nowhere in any ranking on this page. We have been burned before by treating a plumbing failure as a model failure. Full method on the methodology page.

What we did not measure

FAQ

How much does Sakana Fugu cost? At list on 2026-09-15, fugu-ultra-v2 is $5 per million input tokens and $30 output; fugu-max is $2 in and $6 out. On our suite that worked out to a derived $57.06 and $2.19 per 1,000 tasks respectively.

Is fugu-ultra-v2 better than fugu-max? Not on anything we measured. Both scored 9 out of 9. fugu-max was also faster at the mean, 12s against 23.7s.

Why is fugu-ultra-v2 so expensive if its rate card is not extreme? Because it consumed 60,645 input tokens across nine short prompts. Volume, not rate.

Where did these numbers come from? Our own runs through OpenRouter on 2026-09-16, scored by executing the generated code. Costs are derived from measured tokens at list price, not from an invoice.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.