Gemini 3.1 Pro Review: 9/9 Correct, and the Most Expensive Result We Have Measured
We ran Gemini 3.1 Pro — catalog id google/gemini-3.1-pro-preview — on our executed coding harness on 2026-07-29. It scored a perfect 9/9 and cost $14.70 per 1,000 tasks at list prices captured that same day. That is the most expensive result we have ever measured: 66% above GPT-5.5's $8.83 and 147x above Qwen3 Coder Next's $0.10 — that one priced 2026-07-17 — for the identical 9/9 score. The cause is a single measured number — 1,094 reasoning tokens per task, the highest of any model we have run — and reasoning tokens bill at the output rate. At $2 / $12 per 1M tokens it lists below GPT-5.4, Claude Sonnet 4.6 and GPT-5.6 Sol on output, and all three finished cheaper. The rate is not what put it at the top.
This review reports one run of one model on one narrow harness. The headline number is unflattering, so the boundaries of the test belong before the number, not after it.
First, what this benchmark cannot see
Gemini 3.1 Pro did not miss a task. Nine out of nine, executed against hidden assertions, no partial credit. Nothing below is a claim that it is weak at code.
Its context window is 1,048,576 tokens. Our nine tasks are short, self-contained Python functions with a signature and a prose spec. The longest prompt in the set would not fill a thousandth of that window. A model built for long-context and multimodal work is being measured here on none of it.
We did not test long-context reasoning, multi-file refactoring, agentic or multi-turn tool use, any language other than Python, or vision — the capabilities a Pro-tier Gemini is most plausibly sold on. If Gemini 3.1 Pro earns its price, it earns it somewhere this harness does not look.
So the conclusion here is narrow and stated as such: on short, well-specified Python tasks, its thinking budget makes it very expensive per job, and if that is your workload the money buys nothing our harness can see. That is a real finding about a real workload. It is not a verdict on the model.
One more label: google/gemini-3.1-pro-preview is a preview id in the catalog. Preview ids get replaced, renamed and repriced. Everything below is true of that id at those prices on 2026-07-29.
What we measured
Run date 2026-07-29. Prices captured 2026-07-29. Same harness, same nine tasks, same settings as every other model on this site.
- Score: 9/9. No missed tasks.
- Measured cost: $14.70 per 1,000 tasks, at list prices on 2026-07-29. That is $0.0147 per task.
- Mean latency: 10.8 s per task. Of the 21 models we have run, 14 were faster.
- Reasoning tokens: 1,094 per task. The highest figure in our entire set.
- List price: $2 input / $12 output per 1M tokens, verified 2026-07-29.
Cost here is computed, not invoiced: the token counts the API reported multiplied by that model's list price on the stated date. We call it measured cost and never a bill.
Eight models, the same 9/9, eight different bills
These eight models all ran on the same harness after the original 13-model sweep. All eight scored 9/9 — not one of them missed a task. Seven were run and priced on 2026-07-29; Gemini 3.6 Flash was run and priced on 2026-07-28, marked in its row. Sorted by measured cost, cheapest first.
| Model | Score | Measured cost / 1k tasks | Latency | Reasoning tokens / task | List price in / out per 1M |
|---|---|---|---|---|---|
| Claude Haiku 4.5 | 9/9 | $0.94 | 3.7 s | 0 | $1 / $5 |
| GPT-5 mini | 9/9 | $1.53 | 15.2 s | 555 | $0.25 / $2 |
| GPT-5.4 | 9/9 | $1.69 | 3.6 s | 0 | $2.50 / $15 |
| Claude Sonnet 4.6 | 9/9 | $2.22 | 4.9 s | 0 | $3 / $15 |
| Grok 4.5 | 9/9 | $2.93 | 6.6 s | 289 | $2 / $6 |
| GPT-5.6 Sol | 9/9 | $4.98 | 6.6 s | 58 | $5 / $30 |
| Gemini 3.6 Flash | 9/9 | $8.02 · priced 07-28 | 6.5 s | 933 | $1.50 / $7.50 |
| Gemini 3.1 Pro | 9/9 | $14.70 | 10.8 s | 1,094 | $2 / $12 |
Read the last two columns together and the whole review is there. The list-price column does not predict the cost order at all — Gemini 3.1 Pro's $12 output rate is below GPT-5.4's $15, below Claude Sonnet 4.6's $15 and below GPT-5.6 Sol's $30, and it still produced the largest bill by a wide margin. The cost column spans 15.6x. The reasoning-token column is what separates them.
Where the $14.70 goes
Reasoning tokens are billed at the output rate. That one billing fact plus one measured count explains almost the entire figure.
1,094 reasoning tokens at $12 per 1M output tokens is $0.01313 per task, or $13.13 per 1,000 tasks. The full measured cost is $14.70 per 1,000 tasks at 2026-07-29 list prices. So roughly 89% of the bill is thinking, and the visible answer plus the prompt account for the remaining $1.57.
Run that arithmetic across all eight models and the pattern is not subtle.
| Model | Reasoning tokens / task | Output list price / 1M | Reasoning-token cost / 1k tasks | Share of measured cost |
|---|---|---|---|---|
| Claude Haiku 4.5 | 0 | $5 | $0.00 | 0% |
| GPT-5.4 | 0 | $15 | $0.00 | 0% |
| Claude Sonnet 4.6 | 0 | $15 | $0.00 | 0% |
| GPT-5.6 Sol | 58 | $30 | $1.74 | 35% |
| Grok 4.5 | 289 | $6 | $1.73 | 59% |
| GPT-5 mini | 555 | $2 | $1.11 | 73% |
| Gemini 3.6 Flash | 933 | $7.50 | $7.00 | 87% |
| Gemini 3.1 Pro | 1,094 | $12 | $13.13 | 89% |
Every figure in the third column is measured reasoning tokens multiplied by the verified output list price on 2026-07-29 — 2026-07-28 for the Gemini 3.6 Flash row. The share column is that product divided by the model's measured cost per 1,000 tasks on the same date.
GPT-5.6 Sol is the control case. Its output rate is $30 per 1M, two and a half times Gemini 3.1 Pro's $12 — the most expensive output rate in this group. It emitted 58 reasoning tokens and finished the same nine tasks for $4.98 per 1,000 tasks against Gemini 3.1 Pro's $14.70, both priced 2026-07-29. The cheap-per-token model paid three times as much, because it thought nineteen times as long.
That is the general rule this site keeps running into: what you pay is tokens burned multiplied by rate, and the rate is the number vendors publish while the token count is the number that actually moves. If you want to run it on your own token mix, the cost calculator does the arithmetic.
A two-model Gemini pattern, and its limits
We have now run two Gemini models on this harness. Both landed in the top three most expensive per job out of 21 models, and neither did it from a top-tier sticker price.
- Gemini 3.1 Pro — $2 / $12 list, 1,094 reasoning tokens, $14.70 per 1,000 tasks. Most expensive of the 21.
- Gemini 3.6 Flash — $1.50 / $7.50 list, 933 reasoning tokens, $8.02 per 1,000 tasks. Third most expensive, behind GPT-5.5 at $8.83.
They are also first and second on reasoning tokens across everything we have run. No other model in our set exceeded 732. Two models, neither with a top-tier output rate, two of the three worst bills, both driven by the same mechanism. On this evidence the family thinks a great deal, and thinking bills at the output rate.
Now the limit, stated as plainly as the pattern. Two models is not a family. Google ships many Gemini ids and we have run exactly these two. We have not run Gemini 3.5 Flash — it lists at $1.50 / $9 as of 2026-07-29 and we have no measured number for it — nor any 2.5-series model, nor the flash-lite tier, nor any batch variant. A two-point pattern is a reason to check your own token counts before you commit, not a law about Google models. If a first-party benchmark number for any Gemini other than 3.1 Pro and 3.6 Flash appears on this site, it is an error.
The Gemini 3.6 Flash review has the per-task reasoning-token breakdown for the sibling model, including a single task that burned 2,615 reasoning tokens on its own. Broader framing is in Gemini against Claude and GPT-5 against Gemini 3.
What $0.94 and $1.69 bought instead
The useful comparison costs nothing to state, because both models were run and priced on the same day as Gemini 3.1 Pro. No date mismatch, no price drift, same nine tasks.
Claude Haiku 4.5 scored the same 9/9 for $0.94 per 1,000 tasks with zero reasoning tokens. That is 15.6x cheaper, priced 2026-07-29. Its output list rate is $5 against Gemini 3.1 Pro's $12 — a 2.4x difference. The rate explains 2.4x of the gap. The other 6.5x is tokens.
GPT-5.4 scored the same 9/9 for $1.69 per 1,000 tasks in 3.6 s. Gemini 3.1 Pro took 10.8 s — exactly 3x longer — for 8.7x the money, both priced 2026-07-29. And GPT-5.4's list price is higher on both sides: $2.50 / $15 against $2 / $12. A model that costs 25% more per token finished the same work for less than an eighth of the price, because it emitted no reasoning tokens at all.
Widen it past the same-date group and the ratio grows. Qwen3 Coder Next scored 9/9 for $0.10 per 1,000 tasks — but that figure is priced 2026-07-17, in the original 13-model sweep, and Qwen3 Coder Next's list input price moved from $0.11 to $0.12 per 1M in the twelve days since. The 147x headline therefore straddles two pricing dates and you should read it as an order-of-magnitude fact, not a precise ratio. The 66% gap to GPT-5.5's $8.83 is cleaner: GPT-5.5's list price was $5 / $30 on both dates.
Projected forward — a projection at list prices, not a bill anyone received — a steady thousand tasks a day of roughly this size works out to about $5,366 a year on Gemini 3.1 Pro against about $343 a year on Claude Haiku 4.5, at 2026-07-29 prices, for output that scored identically on this set. Put either inside an agent loop that makes twenty calls per run and the same arithmetic multiplies again; AI agent traps walks through that. The cheap end of the field is covered in the cheap coding model roundup, and the full field is ranked four different ways in the 2026 AI coding ranking.
When Gemini 3.1 Pro is still the right call
Three honest cases, none of which our harness can confirm or deny.
You need the context window. 1,048,576 tokens is far beyond what any model in our set is being asked to handle here. If your job is a whole repository or a very long document in one shot, cost per short function is the wrong metric entirely and this page does not answer your question.
You need multimodal input. We tested text-in, code-out. Vision is not in this harness at any point.
Your tasks are hard enough that thinking pays. Our nine tasks are not. Ten of the original 13 models scored 9/9 and 18 of 21 overall did, which tells you the ceiling was reached by nearly everyone. On a workload where models actually diverge on correctness, 1,094 reasoning tokens might buy a right answer instead of a wrong one. Here it bought a right answer that all seven cheaper models on this page also produced.
The failure mode to avoid is the one this data does show: routing short, well-specified, high-volume calls to a heavy reasoning model because it is the flagship. On that traffic the thinking is pure cost.
How these numbers were produced
Nine executed Python tasks: two_sum, valid_parentheses, merge_intervals, roman_to_int, lcs_len, flatten, top_k_words, token_bucket, parse_csv_line.
The model receives a function signature and a prose spec. It never sees the assertions. The returned code runs against hidden assertions in an isolated python3 -I subprocess with a 12-second timeout. Every assertion passes or the task fails — no partial credit, no human grader, no LLM judge.
Temperature 0, max_tokens 4000, one scored attempt per task. The harness retries only on an API error, never on a wrong answer, so a miss stays a miss.
Calls go through OpenRouter's OpenAI-compatible endpoint, deliberately not through the DataLLM Lab gateway. Nothing here depends on our infrastructure and you do not have to be our customer to reproduce it.
Cost is derived. It is the reported token counts multiplied by that model's list price on a stated date. It is a measured cost, never a billed invoice. This matters more than it sounds: 49 of roughly 396 models in the catalog changed price in the twelve days to 2026-07-29. Every dollar figure on this page is only true as of its pricing date, which is why each one carries a date.
Gemini 3.1 Pro was run on 2026-07-29 and priced 2026-07-29. The original sweep was 13 models run in one sitting and priced 2026-07-17; the other eight models, Gemini 3.1 Pro included, were run later on the same harness. The full sweep is in the coding cost benchmark.
What we did not measure
Not measured at all: long-context reasoning, multi-file refactoring, agentic or multi-turn tool use, any non-Python language, and vision. The harness is single-turn, it does not run agents, and it does not use tools. For a model whose headline feature is a million-token window, that is a large blind spot and we are not going to pretend otherwise.
A known scoring artefact: the 4,000-token ceiling can truncate a very verbose model mid-answer, and a truncated answer scores as a miss. Gemini 3.1 Pro was not truncated — it scored 9/9 — but a model averaging 1,094 reasoning tokens per task is operating closer to that ceiling than a model emitting zero.
One run, one date. These are single measurements at temperature 0, not distributions. We did not run it twice and we do not report error bars we did not compute.
Preview id caveat: google/gemini-3.1-pro-preview is a preview listing. Its behaviour and its price can change without the id changing.
Measure it on your own tasks
One OpenAI-compatible endpoint, 300+ models, one key. Swap the model id, run your real prompts, compare the token counts — the only cost figure that predicts your bill is the one from your own workload.
FAQ
Is Gemini 3.1 Pro bad at coding?
No. It scored 9/9 on our executed nine-task Python benchmark — a perfect score, code run against assertions it never saw. The finding is about cost, not capability: it cost $14.70 per 1,000 tasks at list prices on 2026-07-29, the most expensive result we have measured, because it averaged 1,094 reasoning tokens per task. On short, well-specified Python functions that thinking budget buys nothing our harness can detect. On long-context, multimodal or agentic work — none of which we tested — it may buy a great deal.
Why is it more expensive than GPT-5.5 when its list price is lower?
Because reasoning tokens bill at the output rate. Gemini 3.1 Pro lists at $2 / $12 per 1M tokens and GPT-5.5 at $5 / $30 — two and a half times the output rate. But Gemini 3.1 Pro emitted 1,094 reasoning tokens per task against GPT-5.5's 176. At $12 per 1M, those 1,094 tokens alone cost $13.13 per 1,000 tasks, about 89% of the $14.70 total. GPT-5.5 came in at $8.83. Sticker price ranks the rate; the bill is the rate multiplied by the tokens.
What is the cheapest model that also scored 9/9?
Qwen3 Coder Next, at $0.10 per 1,000 tasks — but that figure is priced 2026-07-17, from the original 13-model sweep, and its list input price has since moved from $0.11 to $0.12 per 1M. Among models run and priced on the same day as Gemini 3.1 Pro, 2026-07-29, the cheapest 9/9 was Claude Haiku 4.5 at $0.94 per 1,000 tasks with zero reasoning tokens — 15.6x cheaper for the identical score.
Does this mean every Gemini model is expensive per job?
We cannot say that, and we will not. We have run exactly two Gemini models: 3.1 Pro at $14.70 and 3.6 Flash at $8.02. Both landed in the top three most expensive per job out of 21 models, and both are the top two on reasoning tokens, so the pattern is consistent across everything we have measured. But two models is not a family. We have no first-party number for Gemini 3.5 Flash, any 2.5-series model, the flash-lite tier or any batch variant. Check your own token counts.
Is $14.70 what I would actually be billed?
No. It is a measured cost: the token counts the API reported multiplied by the model's list price captured on 2026-07-29. It is not a vendor invoice. Your bill will differ with prompt length, caching, discounts, provider routing and any price change since that date — 49 of roughly 396 catalog models changed price in the twelve days to 2026-07-29. The ratios between models on the same date are the durable part; the absolute dollars are dated.
Why is it slower than models that cost a tenth as much?
It averaged 10.8 s per task, and 14 of the 21 models we have run were faster. GPT-5.4 finished the same nine tasks in 3.6 s — exactly 3x quicker — at $1.69 per 1,000 tasks, priced the same day. Generating 1,094 reasoning tokens before the answer takes wall-clock time as well as money. The three fastest models we have measured are Llama 4 Scout at 1.5 s, GPT-5.4 mini at 2.3 s and Gemini 3 Flash Preview at 2.3 s, and all three emitted zero reasoning tokens.
DataLLM Lab