Benchmark

Gemini Cost Per Task: Five Models Measured, $0.36 to $28.58 per 1,000 Tasks (And Why Reasoning Tokens Are the Whole Story)

Gemini cost per task spans 79x inside one vendor's own family on our executed nine-task Python harness. Five Gemini models have now been through it. Four of them spend 933 to 2,671 reasoning tokens per call and land in the expensive half of our 45-model set — ranks 1, 2, 5 and 7 by measured cost per 1,000 tasks — despite list prices between $1.25 and $2 per 1M input, which is the middle of the market, not the top. The fifth spends zero reasoning tokens and cost $0.36 per 1,000 tasks, which puts it 39th of 45. Same vendor, same nine tasks, same harness. The mechanism is arithmetic, not opinion: reasoning tokens bill at the output rate, and for the four heavy models that spend accounts for 87% to 93% of what each one cost us. The sharpest single line: Gemini 2.5 Pro cost $28.58 per 1,000 tasks and scored 6/9. DeepSeek V3.2 scored 9/9 at $0.08. That is a 357x cost difference with the cheaper model scoring better.

Bar chart of measured cost per 1,000 tasks - five Gemini models against DeepSeek V3.2, Mistral Medium 3.5 and Claude Haiku 4.5

One expensive model is a data point. Four from one family, on one harness, with the same cause behind each, is a mechanism. A fifth from the same family that avoids the cause and lands 22x cheaper than the next Gemini is what turns the mechanism into a finding.

Four of the five Gemini models we have run spend 933 to 2,671 reasoning tokens per call and sit in the expensive half of our 45-model set. The fifth spends zero and cost $0.36 per 1,000 tasks. The thinkers cost 22x to 79x more per 1,000 completed tasks than the one that does not think — $8.02 and $28.58 against $0.36 — on identical prompts at temperature 0.

List prices do not predict any of this. Gemini 2.5 Pro lists at $1.25 / $10 per 1M — the identical rate to GPT-5 — and cost us 2.6x what GPT-5 cost. Gemini 3.6 Flash lists at $1.50 / $7.50 — the identical rate to Mistral Medium 3.5 — and cost us 9.2x what Mistral Medium 3.5 cost. The rate card does not explain the cheap end either: Gemini 3 Flash Preview's $3 output rate is 2.5x below Gemini 3.6 Flash's $7.50, but its bill was 22x below. Same catalog, different bill, because the token counts are different and the token counts are not printed anywhere.

This page is about one thing only: measured cost per completed short coding task. It is not a verdict on Gemini as a family. The scope limits are in the fairness section and they are real ones, not a disclaimer.

The five Gemini models we ran

These five went through the harness between 2026-07-28 and 2026-08-06, each priced at rates captured from the live catalog on the date shown:

Those five, and only those five. We have not run Gemini 3 Pro, and we have not run any of the Lite tiers. Nothing below generalises past this list.

One immediate oddity worth noting before the wider comparison: within the two mid Flash models, the newer one is both cheaper and more accurate. Gemini 3.5 Flash cost 26% more than Gemini 3.6 Flash and scored 8/9 against 9/9. Generation numbers moved in the buyer's favour there, which is not always how it goes — see the Gemini 3.6 Flash review and the Gemini 3.5 Flash review for each run on its own.

Where they land: four at ranks 1, 2, 5 and 7 of 45, one at 39th

Sort all 45 usable models by measured cost per 1,000 tasks, most expensive first. Gemini 2.5 Pro is 1st. Gemini 3.1 Pro Preview is 2nd. Gemini 3.5 Flash is 5th. Gemini 3.6 Flash is 7th. Those four are inside the seven most expensive results in the set. The three non-Gemini models interleaved with them are GPT-5 at $11.00, Claude Opus 5 Fast at $10.20 and GPT-5.5 at $8.83. Gemini 3 Flash Preview is 39th of 45 — the 7th cheapest model we have run as of 2026-08-06. One family holds both ends of the same list.

ModelScoreMeasured cost / 1k tasksMean latencyReasoning tokens / taskList price in / out per 1MPriced at
Gemini 2.5 Pro6/9$28.5825.3 s2,671$1.25 / $102026-07-30
Gemini 3.1 Pro Preview9/9$14.7010.8 s1,094$2 / $122026-07-29
Gemini 3.5 Flash8/9$10.087.3 s977$1.50 / $92026-07-30
Gemini 3.6 Flash9/9$8.026.5 s933$1.50 / $7.502026-07-28
Gemini 3 Flash Preview8/9$0.362.3 s0$0.50 / $32026-08-06
GPT-5 · reference9/9$11.0022.9 s924$1.25 / $102026-07-30
Mistral Medium 3.5 · reference9/9$0.872.9 s0$1.50 / $7.502026-07-17
Claude Haiku 4.5 · reference9/9$0.943.7 s0$1 / $52026-07-29
DeepSeek V3.2 · reference9/9$0.087.1 s0$0.27 / $0.402026-07-30
Measured cost per 1,000 tasks: five Gemini models against three reference modelsNine executed Python tasks, temperature 0, one scored attempt each. Score shown after each cost.DeepSeek V3.2$0.08 · 9/9Gemini 3 Flash Preview$0.36 · 8/9Mistral Medium 3.5$0.87 · 9/9Claude Haiku 4.5$0.94 · 9/9Gemini 3.6 Flash$8.02 · 9/9Gemini 3.5 Flash$10.08 · 8/9Gemini 3.1 Pro Preview$14.70 · 9/9Gemini 2.5 Pro$28.58 · 6/9One scale throughout: 18 px per dollar. Cost is measured token counts multiplied by list price on the date in the table above, not a vendor invoice.
Chart: DataLLM Lab. Scores, latencies and token counts are measured on our executed nine-task benchmark; cost is those token counts multiplied by each model's list price on the date shown in the table above. Method: our methodology. Full run: the coding cost benchmark.

Read the list-price column next to the cost column and the mismatch is the whole story. Not one of these five Gemini models is expensive on paper. $1.25 / $10, $1.50 / $7.50, $1.50 / $9, $2 / $12 and $0.50 / $3 are ordinary mid-market rates in this set — cheaper on output than Claude Sonnet 4.6 at $3 / $15, cheaper than GPT-5.4 at $2.50 / $15, and far cheaper than GPT-5.5 at $5 / $30 (all three rates from our 2026-07-29 capture). Two of those three cost us less per completed task than any of the four reasoning-heavy Geminis, by a wide margin: Claude Sonnet 4.6 at $2.22 and GPT-5.4 at $1.69, both 9/9. The exception is GPT-5.5 at $8.83 (priced 2026-07-17), which sits just above Gemini 3.6 Flash. It lists at four times Gemini 3.6 Flash's output rate — $30 against $7.50 — and still landed within a dollar of it, because it spent 176 reasoning tokens per task against Gemini 3.6 Flash's 933. That is the same mechanism running the other way.

Reasoning tokens bill at the output rate

Rank all 45 models by reasoning tokens per task instead of by cost and the picture resolves. Gemini 2.5 Pro is the heaviest reasoning-token spender in the set at 2,671 per task — nearly double the next model, MiMo v2.5 at 1,402. Three more Gemini models sit 5th, 6th and 7th, at 1,094, 977 and 933 — behind MiMo v2.5, GLM-5.1 at 1,327 and Qwen3.7-Max at 1,236, which take 2nd, 3rd and 4th. The fifth Gemini sits at the far end of the same ranking, tied at zero.

So the honest version of the claim is not that Gemini holds the top four reasoning slots. It holds 1st, 5th, 6th and 7th of 45, with three non-Gemini models in between, and it also holds one of the zeroes. Fourteen of the 45 models emitted no reasoning tokens at all, and Gemini 3 Flash Preview is one of them. That is the point of this page: the split runs through the family, not around it.

Reasoning tokens are billed at the output rate. That single fact converts the ranking above into the cost ranking in the previous section. Here is the arithmetic, model by model — reasoning tokens per task multiplied by 1,000 tasks multiplied by the published output rate:

ModelReasoning tokens / taskOutput rate per 1MReasoning billing / 1k tasksTotal measured cost / 1k tasksShare
Gemini 3 Flash Preview0$3$0.00$0.360%
Gemini 3.6 Flash933$7.50$7.00$8.0287%
Gemini 3.5 Flash977$9$8.79$10.0887%
Gemini 3.1 Pro Preview1,094$12$13.13$14.7089%
Gemini 2.5 Pro2,671$10$26.71$28.5893%

Between 87% and 93% of what each of the four reasoning models cost us was reasoning. The visible answer — the Python function that actually got graded — is the small remainder. You are not mostly paying for code there. You are mostly paying for deliberation about nine functions that eleven zero-reasoning models solved completely without any, and that three more — one of them a Gemini — got eight of nine on.

That table is derived arithmetic, so we checked it against raw counts where we have them. For three of the five models our run record stores the total token split. Gemini 2.5 Pro returned 25,648 output tokens across the nine tasks, of which 24,039 were reasoning — 94% of output. Gemini 3.5 Flash returned 9,983 output tokens, of which 8,793 were reasoning — 88%. Gemini 3 Flash Preview returned 971 output tokens, of which none were reasoning. All three match the derived shares to within a point, which is why we are comfortable printing the derived figures for Gemini 3.6 Flash and 3.1 Pro Preview, where the split is not stored.

The comparison that lands hardest. DeepSeek V3.2 returned 1,265 output tokens across all nine tasks and scored 9/9. Gemini 2.5 Pro returned 25,648 and scored 6/9. That is 20.3x the output for three fewer correct answers, on identical prompts, at temperature 0. Gemini 3 Flash Preview, from the same vendor as the 25,648, returned 971 and scored 8/9. All three figures are measured, not modelled.

The fifth Gemini: zero reasoning tokens, $0.36

Gemini 3 Flash Preview ran on 2026-08-06: 8/9, $0.36 per 1,000 tasks, 2.3 s mean latency, zero reasoning tokens per task, priced at $0.50 / $3 captured the same day. It missed parse_csv_line.

Three things about that result matter more than the headline number.

It is 22x cheaper than the next Gemini and 79x cheaper than the most expensive one. $8.02 and $28.58 against $0.36. Nothing about the vendor changed between those runs. The token behaviour did.

The rate card explains only a small slice of that gap. Gemini 3 Flash Preview returned 588 input and 971 output tokens across the nine tasks. Hold its token counts fixed and re-price them at Gemini 3.6 Flash's $1.50 / $7.50 and you get $0.91 per 1,000 tasks — still 8.8x below Gemini 3.6 Flash's measured $8.02. That $0.91 is arithmetic on our measured token counts at a published rate, not a run we performed. The cheaper sticker price accounts for roughly 2.5x of the 22x. Verbosity accounts for the rest.

It is fast, and it is not perfect. 2.3 s is the joint-second-fastest mean latency of the 45 models we have run as of 2026-08-06; only Llama 4 Scout at 1.5 s was quicker, and it also scored 8/9. The miss is parse_csv_line, which is the most-missed task in our whole set — 9 of the 45 models failed it, against 3 for flatten and 1 each for lcs_len, roman_to_int and valid_parentheses. Missing the hardest task is not the same as being weak everywhere, but 8/9 is 8/9 and we do not round it up.

One caveat that belongs right here rather than buried below: google/gemini-3-flash-preview is a preview id. Preview endpoints get re-tuned, re-priced and retired without notice. This measurement describes what that id did on 2026-08-06 at that day's rates and nothing more — the general problem we lay out in LLM price volatility.

Two price-controlled pairs that isolate verbosity

The obvious objection to everything above is that maybe Gemini is just priced higher and reasoning tokens are a red herring. Our set contains two pairs that rule that out, because in each pair the list price is identical to the cent and only the token behaviour differs.

Pair one — $1.50 / $7.50. Gemini 3.6 Flash and Mistral Medium 3.5 carry exactly the same list rate, both verified in our price capture. Both scored 9/9. Gemini 3.6 Flash measured $8.02 per 1,000 tasks with 933 reasoning tokens per task. Mistral Medium 3.5 measured $0.87 with zero. Same rate card, 9.2x the bill, same score. Mistral was also 3.6 s faster. One disclosure on the dates: the $8.02 is priced at 2026-07-28 rates and the $0.87 at 2026-07-17, so the two costs are not struck on the same day — but both models carry the identical $1.50 / $7.50 in our 2026-07-29 capture, which is what makes the pair a control at all.

Pair two — $1.25 / $10. Gemini 2.5 Pro and GPT-5 carry exactly the same list rate, and this pair is cleaner still: both were priced on 2026-07-30. GPT-5 measured $11.00 at 924 reasoning tokens per task and scored 9/9. Gemini 2.5 Pro measured $28.58 at 2,671 and scored 6/9. Same rate card, 2.6x the bill. The cost ratio is 2.60x and the reasoning-token ratio is 2.89x — close enough that reasoning volume alone explains almost the entire gap, with the rest sitting in visible output.

Two pairs, price held constant, cost moving by 9.2x and 2.6x — and one within-family case where holding the price constant still leaves an 8.8x gap. Verbosity is a price input, and it is not on the rate card. Anyone budgeting from a per-token price without a measured token count is estimating the wrong number — which is the same trap we walk through for multi-call workloads in what AI agents actually cost. For your own token mix, the cost calculator does the arithmetic.

$28.58 against $0.08, with the cheaper model scoring better

Set the two extremes of our set side by side.

Gemini 2.5 Pro: 6/9, $28.58 per 1,000 tasks, 25.3 s per call.
DeepSeek V3.2: 9/9, $0.08 per 1,000 tasks, 7.1 s per call.

That is 357x the cost, for three fewer correct answers and 3.6x the wait. Run a thousand tasks of this size a day for a year and the projection is roughly $10,430 against $29 — a projection from our measured per-task figures at the stated rates, not a bill anyone sent us.

A second contrast, because DeepSeek is an easy model to wave away as a special case. Claude Haiku 4.5 scored 9/9 at $0.94 per 1,000 tasks in 3.7 s. Against Gemini 2.5 Pro's 6/9 at $28.58 in 25.3 s, that is 30x cheaper and 6.8x faster with a better score, from a mainstream frontier-lab model with no exotic pricing. The full head-to-head across both vendors is in Gemini vs Claude, and the Haiku run on its own is in the Haiku 4.5 review.

The best-scoring Gemini we ran does not escape this either. Gemini 3.6 Flash matched DeepSeek V3.2's 9/9 and cost 100x more — $8.02 against $0.08. And Gemini 3.1 Pro Preview, at $14.70, is the most expensive perfect score in the 45-model set as of 2026-08-06; the cheapest perfect score is DeepSeek V3.2's $0.08, a spread of 184x across models that produced identical output quality on these nine tasks. The DeepSeek versus Gemini comparison takes that pair further. Gemini 3 Flash Preview at $0.36 is the only Gemini that lands anywhere near the cheap models in that comparison, and it does so one answer short of them.

Where this article would be easy to write unfairly

Five things could make the numbers above mean less than they appear to. We would rather state them ourselves.

1. Gemini 2.5 Pro is an older model, and its list price is not what a buyer would pay today. It is the headline villain of this page because it is the extreme in our data, but nobody starting a project in August 2026 is choosing 2.5 Pro on purpose. Judge the current tiers on the current tiers: Gemini 3.6 Flash at $8.02, Gemini 3.1 Pro Preview at $14.70 and Gemini 3 Flash Preview at $0.36 are the numbers that describe today's catalog, and all three are better results than 2.5 Pro's. The generational trend inside our own data is clearly downward on cost and upward on score.

2. Five models is not the Gemini family. We ran Gemini 3 Flash Preview, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.1 Pro Preview and Gemini 2.5 Pro. We have not run Gemini 3 Pro. We have not run any Lite tier. The Lite tiers list far cheaper — Gemini 3.1 Flash Lite at $0.25 / $1.50 in our 2026-07-29 capture — and after the $0.36 result the Lite question is more open, not less: a cheaper model that also reasons less could land somewhere completely different again. We do not know, because we did not run it, and a family-wide claim from five members would be exactly the kind of overreach this site exists to avoid. Current rates for the tiers we did not test are in the Gemini API pricing guide.

3. Nine short Python functions is a narrow test, and it is arguably the worst possible showcase for the reasoning-heavy tiers. Gemini is a 1M-context multimodal family. Our harness sends a few hundred tokens, asks for one function, and never touches long context, vision, multi-turn tool use or anything agentic. Heavy reasoning on a bounded function spec is overkill; the same behaviour on a 400,000-token document might be exactly what you want and worth every token. This finding is specifically about cost per completed short coding task. It is not a capability ranking, and reading it as one would be a mistake.

4. Gemini 2.5 Pro scored 7/9 on an earlier run of ours and 6/9 on the re-run. Both at temperature 0, both on the same nine tasks. We use 6/9 throughout this page because that is the figure in the current data file, but the honest reading is that this model is unstable on the boundary tasks rather than reliably a 6. Temperature 0 is not determinism, and a one-point swing on a nine-item test is a reminder that nine tasks is nine data points. We did not store token counts for the earlier run, so we cannot say whether the cost moved with the score — the $28.58 describes the 6/9 re-run and nothing else. The re-check that produced both figures is written up in the content-filter incident.

5. Gemini 3 Flash Preview is a preview id, and one run is one run. Preview endpoints change behaviour without a version bump. The zero-reasoning-token result is what google/gemini-3-flash-preview did on 2026-08-06 with plain requests and no thinking-budget parameter. If Google turns default reasoning on for that id tomorrow, the $0.36 goes with it, and we would have to re-run to find out.

What we are not claiming. Not that Gemini is a bad family — one of its models is among the cheapest we have run. Not that reasoning tokens are waste in general. Not that these five results predict Gemini 3 Pro, the Lite tiers, or any workload outside single-turn Python. Only this: on our nine executed tasks, four Gemini models spent 933 to 2,671 reasoning tokens per call and cost $8.02 to $28.58 per 1,000 completed tasks, a fifth spent none and cost $0.36, and reasoning-token billing explains the gap between them.

What to actually do with this

If your workload is bounded, well-specified code generation, the four reasoning-heavy tiers are not the rational pick on our data. DeepSeek V3.2 at $0.08, Qwen3 Coder Next at $0.10, DeepSeek Chat at $0.10 and GPT-5.4 mini at $0.53 all scored 9/9. So did Claude Haiku 4.5 at $0.94 in 3.7 s. You are choosing between $0.08 and $8.02 for the same nine results. More options at that tier are in the cheap coding model roundup, and the wider field sorted by axis is in the AI coding ranking.

If you want Gemini specifically, the choice is now a real one. Gemini 3 Flash Preview is the cheapest and fastest of the five at $0.36 and 2.3 s, and it scores 8/9. Gemini 3.6 Flash is the cheapest 9/9 Gemini at $8.02 and 6.5 s. You are paying 22x for one extra correct answer out of nine — worth it if that ninth task is the one you cannot get wrong, indefensible if it is not. Gemini 3.5 Flash costs 26% more than 3.6 Flash for one fewer correct answer, and on our data there is no reason to prefer it. The Gemini 3.1 Pro Preview review covers the Pro tier on its own terms, and the older Pro line is dissected in Gemini 2.5 Pro vs Claude.

If you are committed to a reasoning tier and the volume is batchable, look at the batch rates. Our 2026-07-29 capture shows Gemini 3.6 Flash at $0.75 / $3.75 in batch mode — exactly half the synchronous rate. Applied to our measured token counts that would put it near $4.01 per 1,000 tasks. We did not run batch mode, so treat that as arithmetic on a published rate, not a measurement of ours. It still leaves it above every zero-reasoning model in the table, including its own family's.

Whatever you pick, measure token counts, not just prices. That is the transferable lesson here. Two models on the same rate card differed by 9.2x on identical work, and two models from the same vendor differed by 22x. A per-token price tells you the multiplier, never the multiplicand, and the multiplicand is where reasoning-heavy models hide their cost.

Compare Gemini against DeepSeek on one key

One OpenAI-compatible endpoint, 300+ models, one key. Swap the model id between Gemini 3 Flash Preview, Gemini 3.6 Flash, Gemini 3.1 Pro Preview and DeepSeek V3.2, send your own prompts, and read the reasoning-token counts yourself.

How these numbers were produced

Nine executed Python tasks: two_sum, valid_parentheses, merge_intervals, roman_to_int, lcs_len, flatten, top_k_words, token_bucket, parse_csv_line.

The model receives a function signature and a prose spec. It never sees the assertions. The returned code runs against hidden assertions in an isolated python3 -I subprocess with a 12-second timeout. Every assertion passes or the task fails — no partial credit, no human grader, no LLM judge.

Temperature 0, max_tokens 4000, no thinking-budget or effort parameter set. One scored attempt per task. The harness retries only on an API error, never on a wrong answer, which is why a miss stayed a miss.

Cost is derived, not invoiced. It is the token counts the API reported multiplied by that model's list price on a stated date. The five Gemini models are priced at 2026-07-28, 2026-07-29, 2026-07-30 and 2026-08-06 rates as marked in the table. List prices move, so an undated cost figure is not a fact — a habit we argue for in LLM price volatility. Recompute before you act on any of it.

Calls go through OpenRouter's OpenAI-compatible endpoint, deliberately not through the DataLLM Lab gateway. Nothing on this page depends on our infrastructure and you do not have to be our customer to reproduce it.

On the counting. Our core sweep was 13 models run in one sitting. The remaining models, including all five Gemini models, ran later on the same harness under the same settings, bringing the set to 45 usable results. Two further entries are excluded from every figure on this page. Claude Fable 5 is excluded because four of its nine tasks came back empty under a content filter rather than wrong, leaving only five scorable tasks. Qwen3.5-397B-A17B is excluded because token_bucket came back empty with finish_reason=length — it exhausted our 4,000-token ceiling on reasoning without emitting an answer, which is our constraint, not a defect in the model. Neither has a quotable score. Where this page says 45 models, that is the combined set. Of those 45, 33 scored 9/9.

What we did not measure

Not measured at all: long-context behaviour, multi-file refactoring, agentic or multi-turn tool use, any non-Python language, vision, and anything run locally. The harness is single-turn and API-hosted. Every capability the Gemini family is actually built around sits in that list, starting with the 1M window. A nine-function coding score predicts nothing about 900,000-token retrieval — the failure mode we cover in context rot.

We did not tune thinking budgets. Google exposes controls over reasoning effort on some tiers. We sent plain requests with no such parameter, so every reasoning-token count here is the default behaviour, not a floor and not a ceiling. That cuts both ways: the 2,671 on Gemini 2.5 Pro is a default, and so is the zero on Gemini 3 Flash Preview. A configured budget would very likely move these costs, and we have no data on how far. That is a genuine caveat on the whole finding, not a footnote.

Batch mode is untested. The $4.01 figure in the previous section is arithmetic on a published batch rate captured 2026-07-29, applied to our synchronous token counts. We did not run it. The $0.91 re-pricing in the zero-reasoning section is arithmetic of the same kind.

A known scoring artefact: the 4,000-token ceiling can truncate a very verbose model mid-answer, and a truncated answer scores as a miss. That matters more here than on most pages, because four of these five models are the verbose end of our set. We cannot rule out that some of Gemini 2.5 Pro's three misses were truncation rather than incorrect logic. The ceiling penalises verbosity, which is a real production cost, but it is not the same thing as being wrong.

Not tested, and never claimed as ours: Gemini 3 Pro, any Gemini Lite tier, any Gemini image or Nano Banana variant, and any Gemini model other than the five named at the top of this page. If a first-party number for any of those appears anywhere on this site, it is an error.

FAQ

Why are some Gemini models expensive if their list prices are mid-range?

Because reasoning tokens bill at the output rate, and four of the five Gemini models we ran spend heavily on them. Reasoning billing accounts for 87% to 93% of those four models' measured cost per 1,000 tasks on our nine executed Python tasks. Gemini 2.5 Pro emitted 2,671 reasoning tokens per task, Gemini 3.1 Pro Preview 1,094, Gemini 3.5 Flash 977 and Gemini 3.6 Flash 933. The list rates — $1.25 to $2 per 1M input — are ordinary; the token volume is not. The fifth model, Gemini 3 Flash Preview, emitted none and cost $0.36.

Which Gemini model was cheapest per completed task?

Gemini 3 Flash Preview, at a measured $0.36 per 1,000 tasks with an 8/9 score and 2.3 s mean latency, priced at 2026-08-06 rates. It emitted zero reasoning tokens and returned 971 output tokens across all nine tasks. If you need 9/9, the cheapest Gemini that delivered it is Gemini 3.6 Flash at $8.02 — 22x more money for one more correct answer out of nine.

Does the $0.36 result contradict the rest of this page?

No — it is the same mechanism producing the opposite result. The claim here has never been that Gemini is expensive. It is that reasoning tokens drive cost per completed task. Four Gemini models spend 933 to 2,671 reasoning tokens per call and cost $8.02 to $28.58 per 1,000 tasks. One spends zero and costs $0.36. Within a single vendor's own family, the models that think a lot cost 22x to 79x more per completed task than the one that does not.

Is a 357x cost difference really fair to Gemini 2.5 Pro?

The arithmetic is exact — $28.58 against DeepSeek V3.2's $0.08, both priced 2026-07-30, with DeepSeek scoring 9/9 against 2.5 Pro's 6/9. But 2.5 Pro is an older model and nobody should be starting a new project on it. The fairer current-generation comparison is Gemini 3.6 Flash at $8.02, which is 100x DeepSeek V3.2 rather than 357x, or Gemini 3 Flash Preview at $0.36, which is 4.5x it at one score lower. All three ratios are real; the last two are the ones that describe today's catalog.

Did Gemini 2.5 Pro really only score 6 out of 9?

On the run in our current data file, yes — it missed roman_to_int, flatten and parse_csv_line. An earlier run of ours scored it 7/9, also at temperature 0 on the same nine tasks. We publish the later figure because it is the one in the data file, but the honest reading is that the model is unstable on those boundary tasks rather than reliably a 6. We also cannot rule out that a miss was a truncation at our 4,000-token ceiling, given how verbose this model is.

Would setting a thinking budget lower these costs?

Probably, but we have no data on it. Every run here used plain requests with no thinking-budget or effort parameter, so the reasoning-token counts are default behaviour — including the zero on Gemini 3 Flash Preview, which is a default and not a configuration we applied. If your provider exposes a budget control, measuring your own token counts with it set is the obvious next step, and it is a measurement we have not made, so do not take a number from this page as a prediction of what you would see with one configured.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.