The Cheapest LLM That Works (and the 1,902x Spread Above It)
The cheapest LLM that works — meaning it solves all nine tasks in our executed Python suite — now costs $0.03 per 1,000 tasks. That is Solar Mini 4, priced 2026-10-02, and it is also the fastest model to score 9 out of 9, at a 2s mean. The most expensive model that earned the identical 9 out of 9, Fugu Ultra v2, costs $57.06, priced 2026-09-15. Divide them: $57.06 ÷ $0.03 = 1,902. 1,902 times the cost, for the same nine passing functions. That is not an argument for the cheap model. It is an argument that nine short Python functions are a floor test, and that anyone quoting a 9/9 as a capability ranking — including us, if we are careless — is reading it wrong.
Our set now holds 109 entries, of which 93 are usable and 75 scored 9 out of 9. This piece is about those 75 and nothing else — every model ranked below solved every task. The only variable left is what it cost.
Dated note: when this page was first published, the floor was $0.07, held on 2026-09-16 by Ling 3.0 Flash VL, and the spread was $57.06 ÷ $0.07 = 815.14. Both figures are superseded by the 2026-10-02 sweep.
The spread
The top bar deserves a second look. Fugu Ultra v2 is not the most expensive 9/9 because it thinks hardest — it emitted 7 reasoning tokens per call, close to none. It is the most expensive because it consumed 60,645 input tokens across the suite against 7,011 out, at $5 in / $30 out per 1M priced on 2026-09-15. Solar Mini 4 read 1,009 input tokens on the same nine prompts; 60,645 ÷ 1,009 = 60.10. Its per-token price, $5 in, is half the $10 in that GPT-6 Astra charges, and its derived cost is still $57.06 ÷ $8.19 = 6.97 times Astra's. The family is covered in the Fugu review; the short version is that a model can be cheap on the price card and expensive in practice because of what it decides to read.
Where the floor is now
The floor for a paid endpoint that clears the suite is $0.03, held by Solar Mini 4, priced 2026-10-02 at $0.05 in / $0.2 out per 1M, with a 2s mean and 0 reasoning tokens per call. It ranks 1st of 75 on cost and 1st of 75 on speed. Second is MiMo-V2.6-Flash at $0.04, priced 2026-10-02, at a 5s mean. Third is the old floor, Ling 3.0 Flash VL at $0.07, priced 2026-09-15 — reviewed in the Ling 3.0 Flash VL write-up.
Some entries look as if they belong in this conversation and do not count, and the reasons are the point of the phrase “that works”:
- space-bunny-alpha, $0 — a free stealth endpoint. Free endpoints are barred from cost rankings, because a $0 list price is a promotional posture, not a measurement. It is also excluded with no valid score: across two runs on 2026-10-02, roman_to_int returned provider_unavailable (502) both times, so neither run completed.
- Codestral 2508, $0.11, and Qwen3 Coder Flash, $0.12, both priced 2026-10-02, scored a valid 8/9. Both missed parse_csv_line.
- Command A Plus, $1.71, priced 2026-10-02, is excluded: parse_csv_line came back as an empty body with finish_reason=length, so only 8 of 9 tasks were scored. That is an incomplete run, not an 8/9 and not a near-miss of a 9/9; its cost means nothing until it is re-run.
Price does not buy the pass, either. GLM-5.3-Prime cost $11.52, priced 2026-10-02, and scored 7/9, missing token_bucket and parse_csv_line — while plain Qwen3 Coder cleared all nine at $0.13 and its Plus sibling scored 8/9 at $0.41, all on the same day.
Third-party context on the new floor, read 2026-10-02: Artificial Analysis describes Solar Mini 4 as a proprietary reasoning model from Upstage with weights not released, and reports it as a heavy token user on its own index — costlier per task there than GPT-6 Luna. On our suite it emitted no reasoning tokens at all. We cannot tell from our data whether that gap is the workload, the default reasoning setting on the endpoint we called, or both.
The rebuilt leaderboard
Ranks are positions within the 75 models at 9/9; gaps in the rank column are models not listed here.
| Cost rank | Model | Derived cost / 1k tasks | Priced on | Mean latency | Speed rank |
|---|---|---|---|---|---|
| 1 | Solar Mini 4 | $0.03 | 2026-10-02 | 2s | 1 |
| 2 | MiMo-V2.6-Flash | $0.04 | 2026-10-02 | 5s | 25 |
| 3 | Ling 3.0 Flash VL | $0.07 | 2026-09-15 | 4.4s | 16 |
| 4 | DeepSeek V3.2 | $0.08 | 2026-07-30 | 7.1s | 38 |
| 5 | Qwen3 Coder Next | $0.1 | 2026-07-17 | 7s | 37 |
| 8 | Qwen3 Coder | $0.13 | 2026-10-02 | 2.7s | 6 |
| 9 | GPT-6 Luna | $0.16 | 2026-10-02 | 4.9s | 23 |
| 11 | Devstral 2512 | $0.23 | 2026-10-02 | 2.1s | 2 |
| 13 | GLM-5.3-Flash | $0.34 | 2026-08-31 | 24.9s | 71 |
| 18 | GPT-5.4 Mini | $0.53 | 2026-07-30 | 2.3s | 3 |
| 34 | Claude Sonnet 5.5 | $1.95 | 2026-10-02 | 3.4s | 9 |
| 48 | Claude Opus 5.5 | $4.03 | 2026-10-02 | 4.8s | 20 |
| 55 | MiMo-V2.6-Pro-UltraSpeed | $4.8 | 2026-10-02 | 2.5s | 5 |
| 66 | Claude Fable 5.1 | $8.09 | 2026-09-15 | 6.8s | 36 |
| 68 | GPT-6 Astra | $8.19 | 2026-09-15 | 5.6s | 28 |
| 72 | Qwen3.8 Max Prime | $12.41 | 2026-10-02 | 15.5s | 62 |
| 74 | GPT-6 Astra Pro | $35.44 | 2026-09-15 | 8.4s | 43 |
| 75 | Fugu Ultra v2 | $57.06 | 2026-09-15 | 23.7s | 70 |
First, the cheapest and the fastest are now the same model — but below the top row, cost rank and speed rank still barely track each other. MiMo-V2.6-Flash is 2nd cheapest and 25th fastest. MiMo-V2.6-Pro-UltraSpeed is 5th fastest and 55th cheapest. GLM-5.3-Flash is 13th cheapest and 71st fastest; it pays for its $0.34 with a 24.9s mean per call. GPT-5.4 Mini, which held the speed crown in earlier versions of this page, is now 3rd fastest.
Second, an exact division worth doing out loud: $8.19 ÷ $0.03 = 273. GPT-6 Astra costs 273 times what the floor costs and returns the same nine passing functions.
Same list price, different bill
The more useful pattern is that models sharing an identical price card land far apart on derived cost, because cost is set by how many tokens a model consumes, not by the per-million rate.
| Model | List price in / out per 1M | Tokens across the suite | Derived cost / 1k tasks | Priced on |
|---|---|---|---|---|
| GPT-6.1 Sol | $2 / $10 | 604 in / 1,329 out | $1.61 | 2026-10-02 |
| Claude Sonnet 5.5 | $2 / $10 | 914 in / 1,569 out | $1.95 | 2026-10-02 |
| GPT-6 Sol | $2 / $10 | 604 in / 1,704 out | $2.03 | 2026-10-02 |
| GPT-6 Sol Pro | $2 / $10 | 16,898 in / 3,322 out | $7.45 | 2026-10-02 |
| Claude Fable 5.1 | $10 / $50 | 914 in / 1,273 out | $8.09 | 2026-09-15 |
| GPT-6 Astra | $10 / $50 | 604 in / 1,353 out | $8.19 | 2026-09-15 |
| GPT-6 Astra Pro | $10 / $50 | 17,011 in / 2,977 out | $35.44 | 2026-09-15 |
Four models on the $2 / $10 card cost $1.61, $1.95, $2.03 and $7.45. Sol Pro is not a different price card; it read 16,898 input tokens on the nine prompts where Sol read 604, and $7.45 ÷ $2.03 = 3.67. It is the same shape we found in the Astra review, where Astra Pro read 17,011 against Astra's 604. Luna follows suit: 604 in for GPT-6 Luna, 17,806 for Luna Pro. A Pro tier on this family starts every request with input you did not write. Quoting a model's per-million price tells you very little about what a workload will cost. Put your own token counts into the cost calculator before you trust a price card.
One more caution on dates. Every figure is priced on the date stored with its run, and those dates span 2026-07-17 to 2026-10-02. Some list prices have moved since: GLM-5.3-Flash was priced at $0.075 in / $0.25 out per 1M on 2026-08-31, and the list today reads $0.15 in / $0.5 out. We do not restate old costs at new prices, which is why its row still says $0.34. Prices move more than the leaderboard does.
Why 9/9 is a floor, not a ranking
Here is the part any cost leaderboard has to say plainly, or it is selling something. Nine self-contained Python functions cannot separate a frontier model from a competent small one. They are short, unambiguous, and solutions to problems like them are all over every model's training data. A 9/9 means a model clears a floor. It does not rank anything above that floor.
What the suite does detect is failure to clear the floor, and that is common enough to publish: of 109 entries, 93 are usable and 75 reach 9/9. The token-budget failure mode behind many drop-outs is collected in the piece on models that cannot finish inside a budget.
So read this leaderboard as a shortlist filter, not a verdict. If your workload is short, well-specified functions of the kind our suite contains, the evidence says the floor is usable and you can spend $0.03 to $0.53 per 1,000 tasks instead of $57.06. For long-context refactoring, agent loops or ambiguous specs, this data says nothing. The wider cheap end is in our cheap coding roundup and the cheapest LLM API guide.
How these numbers were produced
Nine Python tasks, each given as a function signature plus a spec with no example tests. The generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why an incomplete run is excluded rather than scored low. Cost is derived from measured token counts at list price on the date shown for each row — a calculation, not a billing statement. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.
What we did not measure
- Anything above the floor. No long context, no multi-file work, no tool use, no agent loops, no ambiguous requirements. The suite cannot distinguish the 75 models that passed it.
- Repeat runs. One scored attempt per task. The one-cent gap between ranks 1 and 2 is inside the noise of that design, and so is a one-position swap anywhere in the table.
- Reasoning settings. We call each endpoint with its defaults. Solar Mini 4 returned 0 reasoning tokens per call; a third party measuring it on a harder index saw heavy reasoning. Our $0.03 is a price for this workload at these defaults, not for the model in every mode.
- The entries that are not usable. 93 of 109 are usable. The excluded runs — Command A Plus, space-bunny-alpha, Seed 2.0 Code, and KAT-Coder-Pro V2.5, which returned HTTP 400 on every call and was never measured — have not been re-run, so the set skews toward models that behave well under a 4000-token cap.
- Prices on a single day. The leaderboard is a composite of several pricing days. Qwen3 Coder Next's derived cost predates the stored run-date price field entirely.
- Routers. A typesafe/jev-router probe on 2026-10-02 passed all nine at $1.01 per 1,000 tasks, but that figure is cost reported by the API per call, a different method from ours, so it is not in the table.
- What anyone is actually billed. We multiply measured tokens by published rates, ignoring caching, batch tiers, provider routing and negotiated pricing. The floor has moved from DeepSeek V3.2 to Ling 3.0 Flash VL to Solar Mini 4 since July; it will move again.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab