Gemini 3.7 Flash Review: Same Score as 3.6 Flash, at One Sixth the Cost
Gemini 3.7 Flash scored 9 out of 9 on our executed Python benchmark at $1.40 per 1,000 tasks and a 5.7-second mean. Three weeks earlier we measured Gemini 3.6 Flash on the identical suite: the same 9 out of 9, at $8.02. Google halved the list price between the two generations — $0.75 and $3.75 per million tokens down to $0.38 and $1.88. But the measured cost fell 83%, not 50%. A price cut explains about half of the saving; the model getting less talkative explains the rest. That gap between the published discount and the real one is the whole point of measuring instead of reading price tables.
Correction, October 3, 2026: single-model benchmark latency is a mean, not a median. Labels were corrected; measured values and the original run dates are unchanged. Measurement details.
Google ships Flash models quickly enough that the interesting question is no longer whether the new one is better, but by how much and along which axis. We have measured four Gemini models on the same nine tasks, so this one comes with its own baseline.
The result
| Metric | Gemini 3.7 Flash |
|---|---|
| Score | 9/9 |
| Measured cost / 1,000 tasks | $1.40 |
| Mean latency | 5.7s |
| Reasoning tokens per call | 609 |
| Output tokens on the suite | 6,605 |
| List price in / out | $0.38 / $1.88 per 1M |
| Context window | 1,048,576 tokens |
| Measured | 2026-08-22 |
A clean pass. Nine tasks, nine solved, no retries needed.
3.6 Flash to 3.7 Flash
| Gemini 3.6 Flash | Gemini 3.7 Flash | Change | |
|---|---|---|---|
| List price in / out | $0.75 / $3.75 | $0.38 / $1.88 | −49% |
| Score | 9/9 | 9/9 | same |
| Measured cost / 1k tasks | $8.02 | $1.40 | −83% |
| Mean latency | 6.5s | 5.7s | −12% |
| Reasoning tokens | 933 | 609 | −35% |
| Measured on | 2026-07-28 | 2026-08-22 | — |
Read the price row against the cost row. Google cut the sticker roughly in half. Halving the price, on its own, halves the bill. The bill actually fell to one sixth — so roughly two thirds of the saving came from somewhere other than the discount. It came from the model emitting less: reasoning spend alone dropped 35%, and reasoning bills at the output rate.
We keep finding this. Qwen3.8-Max raised its list price 36% and still cut the bill 27%. GPT-5.1-Codex-Max lists below GPT-5.2-Codex and costs 3.09x more. List price and invoice are only loosely related, and the loose part is token volume — the variable that also dominates agent bills in our agent model comparison.
The Gemini ladder, measured
Four Gemini models, identical tasks, identical conditions:
The Gemini 2.5 Pro row is the one worth staring at: $28.58 per 1,000 tasks for 6 out of 9. Twenty times the price of 3.7 Flash for a worse result on the same work. Our Gemini cost-per-task page covers that spread in full.
Is it the Gemini to pick?
For coding work of this shape, yes — it is the cheapest Gemini in our set that clears all nine tasks, and it is faster than the generation it replaces.
Outside the Gemini family, the picture changes. GPT-5.4 mini also scored 9/9, at $0.53 and 2.3 seconds. Among the 43 models in our set that scored 9/9, Gemini 3.7 Flash ranks 15th cheapest. It is a good Gemini, not a cheap model in absolute terms — and the million-token context is the reason you would pay the difference. Our 3.6 Flash review has the previous generation in full, and Gemini vs Claude covers the cross-vendor call.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived — measured input and output token counts multiplied by the list price on the date shown, not a billing statement. The two generations were priced 25 days apart and Google moved the price between them, which is precisely why every figure needs a date attached. Runs go through OpenRouter. Full method on the methodology page.
What we did not measure
- The million-token context. Our prompts are short. Nothing here tests long-context recall, which is a large part of what a 1M window is for.
- Multimodal input. Text only. No images, audio or video.
- Anything but Python, and only nine self-contained functions at that.
- Repeat runs. One scored attempt per task. Single-run figures, not averages, and we do not average across runs.
- Gemini 3.8 Flash, which is reported to be in internal testing and is not callable, so it appears nowhere in our data.
FAQ
How much does Gemini 3.7 Flash cost? It lists at $0.38 per million input tokens and $1.88 output as of 2026-08-22. On our nine executed tasks that worked out to $1.40 per 1,000 tasks.
Is Gemini 3.7 Flash better than 3.6 Flash? Same score, 9 out of 9 for both. It is 83% cheaper per finished task, 12% faster, and emits 35% fewer reasoning tokens.
How fast is Gemini 3.7 Flash? 5.7 seconds mean on our tasks.
Is it the cheapest model that scores 9/9? No. It ranks 15th cheapest among the 43 models in our set that scored 9 out of 9.
Does it beat Gemini 2.5 Pro? On this suite, comprehensively: 9/9 at $1.40 against 6/9 at $28.58.
DataLLM Lab