Gemini 3.8 Flash: Separating Token Growth from OpenRouter Rate Changes
Gemini 3.8 Flash scored 9 out of 9 on our executed Python benchmark at $3.87 per 1,000 tasks and a 10.2-second mean. When we measured Gemini 3.7 Flash on 2026-08-22 it cost $1.40. That is 2.8x in three weeks, and it is two separate things that most coverage will merge into one. Our dated OpenRouter catalogue snapshots show doubled Flash list rates — 3.7 Flash was $0.375 and $1.875 per million in August and lists at $0.75 and $3.75 today, the same as 3.8. And the new model emits 39% more output tokens. Hold the price constant and 3.8 is 38% more expensive than 3.7. Let the price move too and you get 2.8x. Both numbers are true; only one of them is about the model.
A model that costs 2.8x its predecessor three weeks later is an eye-catching headline and a misleading one. The number is real. Most of it is not about the new model.
The result
| Metric | Gemini 3.8 Flash |
|---|---|
| Score | 9/9 |
| Measured cost / 1,000 tasks | $3.87 |
| Mean latency | 10.2s |
| Reasoning tokens per call | 901 |
| Tokens across the suite | 588 in / 9,168 out |
| List price in / out | $0.75 / $3.75 per 1M |
| Context window | 1,048,576 |
| Measured | 2026-09-16 |
Among the 54 models in our set at 9 out of 9 it ranks 33rd cheapest and 31st fastest — squarely mid-field, which is a step down for a line whose whole identity is being the cheap fast one.
Splitting the price rise from the token rise
We measured 3.7 Flash on 2026-08-22, when it listed at $0.375 and $1.875. Today both 3.7 and 3.8 list at $0.75 and $3.75. So to compare the models rather than the calendar, price both token counts at today's identical rate:
| Gemini 3.7 Flash | Gemini 3.8 Flash | Difference | |
|---|---|---|---|
| Score | 9/9 | 9/9 | none |
| Output tokens across the suite | 6,605 | 9,168 | +39% |
| Reasoning tokens per call | 609 | 901 | +48% |
| Mean latency | 5.7s | 10.2s | +79% |
| Cost at today's shared price | $2.80 | $3.87 | +38% |
| Cost we actually recorded | $1.40 (at Aug prices) | $3.87 | 2.8x |
Neither half is hidden, but neither is on the model card either. This is the clearest example we have measured of why every price needs a date attached: a figure we published three weeks ago is now wrong by 2x, through no change in the thing it described.
Four Flash generations, measured
| Model | Score | Measured cost / 1k | Latency | Reasoning tokens | Measured on |
|---|---|---|---|---|---|
| Gemini 3 Flash Preview | 8/9 | $0.36 | 2.3s | 0 | 2026-08-06 |
| Gemini 3.7 Flash | 9/9 | $1.40 | 5.7s | 609 | 2026-08-22 |
| Gemini 3.8 Flash | 9/9 | $3.87 | 10.2s | 901 | 2026-09-16 |
| Gemini 3.6 Flash | 9/9 | $8.02 | 6.5s | 933 | 2026-07-28 |
Read down the reasoning column. Flash got cheap between 3.6 and 3.7 by thinking less — 933 tokens down to 609 — and 3.8 gives most of that back at 901. Latency followed the same path: 6.5s, then 5.7s, now 10.2s. On this suite 3.8 Flash is the slowest Flash we have measured and the extra thinking bought nothing, because 3.7 already scored 9 out of 9.
Is it worth running
- Over 3.7 Flash, on this evidence: no. Same score, 38% more at the same rate, and 79% slower. If your provider still serves 3.7, keep it.
- As a cheap tier: it no longer is one. At $3.87 it sits above GLM-5.3-Flash at $0.34 and DeepSeek V3.2 at $0.08, both 9 out of 9.
- Where it still wins: the 1,048,576-token context and Google-native integration, neither of which this suite tests.
- If you budgeted from our August figure, rebuild it. The rate doubled.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived from measured token counts at the list price on each run's date. The $2.80 figure in the comparison table is the recorded 3.7 Flash token count repriced at today's rate — arithmetic on a real measurement, not a second run, and labelled as such. Runs go through OpenRouter. Full method on the methodology page.
What we did not measure
- A fresh 3.7 Flash run at today's price. We repriced the August token count instead; if 3.7's behaviour has drifted, that figure would move.
- The 1,048,576-token context. Our prompts are short.
- Multimodal input. Text only.
- Batch pricing. A
gemini-3.8-flash:batchendpoint exists at a lower rate and is not what we measured. - Repeat runs. One scored attempt per task.
FAQ
How much does Gemini 3.8 Flash cost? $0.75 per million input tokens and $3.75 output as of 2026-09-15. On our nine tasks that came to $3.87 per 1,000 tasks.
Is Gemini 3.8 Flash better than 3.7 Flash? Not on our suite. Both scored 9 out of 9; 3.8 cost 38% more at the same price and took 79% longer.
Did Google raise the price of Gemini Flash? Yes. 3.7 Flash listed at $0.375 and $1.875 when we measured it on 2026-08-22 and lists at $0.75 and $3.75 today.
Why did your cost figure change 2.8x? About half is the list price doubling and about half is 3.8 emitting 39% more output tokens.
Is Gemini 3.8 Flash still a cheap model? Relative to the field, no. GLM-5.3-Flash scored the same 9 out of 9 at $0.34.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab