Benchmarks

The Cheapest LLM That Works (and the 1,902x Spread Above It)

The cheapest LLM that works — meaning it solves all nine tasks in our executed Python suite — now costs $0.03 per 1,000 tasks. That is Solar Mini 4, priced 2026-10-02, and it is also the fastest model to score 9 out of 9, at a 2s mean. The most expensive model that earned the identical 9 out of 9, Fugu Ultra v2, costs $57.06, priced 2026-09-15. Divide them: $57.06 ÷ $0.03 = 1,902. 1,902 times the cost, for the same nine passing functions. That is not an argument for the cheap model. It is an argument that nine short Python functions are a floor test, and that anyone quoting a 9/9 as a capability ranking — including us, if we are careless — is reading it wrong.

DataLLM Lab article cover: The Cheapest LLM That Works (and the 1,902x Spread Above It)

Our set now holds 109 entries, of which 93 are usable and 75 scored 9 out of 9. This piece is about those 75 and nothing else — every model ranked below solved every task. The only variable left is what it cost.

Dated note: when this page was first published, the floor was $0.07, held on 2026-09-16 by Ling 3.0 Flash VL, and the spread was $57.06 ÷ $0.07 = 815.14. Both figures are superseded by the 2026-10-02 sweep.

The spread

Same score, 1,902x apart in costDerived cost per 1,000 tasks. All eleven models below scored 9/9 on the identical nine executed Python tasks.Solar Mini 4$0.03MiMo-V2.6-Flash$0.04Ling 3.0 Flash VL$0.07DeepSeek V3.2$0.08GPT-6 Luna$0.16GPT-5.4 Mini$0.53Claude Sonnet 5.5$1.95Claude Opus 5.5$4.03GPT-6 Astra$8.19GPT-6 Astra Pro$35.44Fugu Ultra v2$57.06$0.1$1$10Logarithmic axis, one scale throughout: 150 px per 10× in cost, origin at $0.01, so bar width = 150 × log10(cost ÷ $0.01).Each cost is priced on that model's own stored pricing date, from 2026-07-30 to 2026-10-02.
A linear axis would render the bottom nine models as a smear against the top two bars. The log axis is the only honest way to draw a 1,902x range.

The top bar deserves a second look. Fugu Ultra v2 is not the most expensive 9/9 because it thinks hardest — it emitted 7 reasoning tokens per call, close to none. It is the most expensive because it consumed 60,645 input tokens across the suite against 7,011 out, at $5 in / $30 out per 1M priced on 2026-09-15. Solar Mini 4 read 1,009 input tokens on the same nine prompts; 60,645 ÷ 1,009 = 60.10. Its per-token price, $5 in, is half the $10 in that GPT-6 Astra charges, and its derived cost is still $57.06 ÷ $8.19 = 6.97 times Astra's. The family is covered in the Fugu review; the short version is that a model can be cheap on the price card and expensive in practice because of what it decides to read.

Where the floor is now

The floor for a paid endpoint that clears the suite is $0.03, held by Solar Mini 4, priced 2026-10-02 at $0.05 in / $0.2 out per 1M, with a 2s mean and 0 reasoning tokens per call. It ranks 1st of 75 on cost and 1st of 75 on speed. Second is MiMo-V2.6-Flash at $0.04, priced 2026-10-02, at a 5s mean. Third is the old floor, Ling 3.0 Flash VL at $0.07, priced 2026-09-15 — reviewed in the Ling 3.0 Flash VL write-up.

Some entries look as if they belong in this conversation and do not count, and the reasons are the point of the phrase “that works”:

Price does not buy the pass, either. GLM-5.3-Prime cost $11.52, priced 2026-10-02, and scored 7/9, missing token_bucket and parse_csv_line — while plain Qwen3 Coder cleared all nine at $0.13 and its Plus sibling scored 8/9 at $0.41, all on the same day.

Third-party context on the new floor, read 2026-10-02: Artificial Analysis describes Solar Mini 4 as a proprietary reasoning model from Upstage with weights not released, and reports it as a heavy token user on its own index — costlier per task there than GPT-6 Luna. On our suite it emitted no reasoning tokens at all. We cannot tell from our data whether that gap is the workload, the default reasoning setting on the endpoint we called, or both.

The rebuilt leaderboard

Ranks are positions within the 75 models at 9/9; gaps in the rank column are models not listed here.

Cost rankModelDerived cost / 1k tasksPriced onMean latencySpeed rank
1Solar Mini 4$0.032026-10-022s1
2MiMo-V2.6-Flash$0.042026-10-025s25
3Ling 3.0 Flash VL$0.072026-09-154.4s16
4DeepSeek V3.2$0.082026-07-307.1s38
5Qwen3 Coder Next$0.12026-07-177s37
8Qwen3 Coder$0.132026-10-022.7s6
9GPT-6 Luna$0.162026-10-024.9s23
11Devstral 2512$0.232026-10-022.1s2
13GLM-5.3-Flash$0.342026-08-3124.9s71
18GPT-5.4 Mini$0.532026-07-302.3s3
34Claude Sonnet 5.5$1.952026-10-023.4s9
48Claude Opus 5.5$4.032026-10-024.8s20
55MiMo-V2.6-Pro-UltraSpeed$4.82026-10-022.5s5
66Claude Fable 5.1$8.092026-09-156.8s36
68GPT-6 Astra$8.192026-09-155.6s28
72Qwen3.8 Max Prime$12.412026-10-0215.5s62
74GPT-6 Astra Pro$35.442026-09-158.4s43
75Fugu Ultra v2$57.062026-09-1523.7s70

First, the cheapest and the fastest are now the same model — but below the top row, cost rank and speed rank still barely track each other. MiMo-V2.6-Flash is 2nd cheapest and 25th fastest. MiMo-V2.6-Pro-UltraSpeed is 5th fastest and 55th cheapest. GLM-5.3-Flash is 13th cheapest and 71st fastest; it pays for its $0.34 with a 24.9s mean per call. GPT-5.4 Mini, which held the speed crown in earlier versions of this page, is now 3rd fastest.

Second, an exact division worth doing out loud: $8.19 ÷ $0.03 = 273. GPT-6 Astra costs 273 times what the floor costs and returns the same nine passing functions.

Same list price, different bill

The more useful pattern is that models sharing an identical price card land far apart on derived cost, because cost is set by how many tokens a model consumes, not by the per-million rate.

ModelList price in / out per 1MTokens across the suiteDerived cost / 1k tasksPriced on
GPT-6.1 Sol$2 / $10604 in / 1,329 out$1.612026-10-02
Claude Sonnet 5.5$2 / $10914 in / 1,569 out$1.952026-10-02
GPT-6 Sol$2 / $10604 in / 1,704 out$2.032026-10-02
GPT-6 Sol Pro$2 / $1016,898 in / 3,322 out$7.452026-10-02
Claude Fable 5.1$10 / $50914 in / 1,273 out$8.092026-09-15
GPT-6 Astra$10 / $50604 in / 1,353 out$8.192026-09-15
GPT-6 Astra Pro$10 / $5017,011 in / 2,977 out$35.442026-09-15

Four models on the $2 / $10 card cost $1.61, $1.95, $2.03 and $7.45. Sol Pro is not a different price card; it read 16,898 input tokens on the nine prompts where Sol read 604, and $7.45 ÷ $2.03 = 3.67. It is the same shape we found in the Astra review, where Astra Pro read 17,011 against Astra's 604. Luna follows suit: 604 in for GPT-6 Luna, 17,806 for Luna Pro. A Pro tier on this family starts every request with input you did not write. Quoting a model's per-million price tells you very little about what a workload will cost. Put your own token counts into the cost calculator before you trust a price card.

One more caution on dates. Every figure is priced on the date stored with its run, and those dates span 2026-07-17 to 2026-10-02. Some list prices have moved since: GLM-5.3-Flash was priced at $0.075 in / $0.25 out per 1M on 2026-08-31, and the list today reads $0.15 in / $0.5 out. We do not restate old costs at new prices, which is why its row still says $0.34. Prices move more than the leaderboard does.

Why 9/9 is a floor, not a ranking

Here is the part any cost leaderboard has to say plainly, or it is selling something. Nine self-contained Python functions cannot separate a frontier model from a competent small one. They are short, unambiguous, and solutions to problems like them are all over every model's training data. A 9/9 means a model clears a floor. It does not rank anything above that floor.

What the suite does detect is failure to clear the floor, and that is common enough to publish: of 109 entries, 93 are usable and 75 reach 9/9. The token-budget failure mode behind many drop-outs is collected in the piece on models that cannot finish inside a budget.

So read this leaderboard as a shortlist filter, not a verdict. If your workload is short, well-specified functions of the kind our suite contains, the evidence says the floor is usable and you can spend $0.03 to $0.53 per 1,000 tasks instead of $57.06. For long-context refactoring, agent loops or ambiguous specs, this data says nothing. The wider cheap end is in our cheap coding roundup and the cheapest LLM API guide.

How these numbers were produced

Nine Python tasks, each given as a function signature plus a spec with no example tests. The generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why an incomplete run is excluded rather than scored low. Cost is derived from measured token counts at list price on the date shown for each row — a calculation, not a billing statement. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

What we did not measure

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.