Grok 4.6 Review: Same Price as 4.5, Same Score, 70% Bigger Bill
Grok 4.6 scored 9 out of 9 on our executed Python benchmark at $4.99 per 1,000 tasks and a 12.7-second median. Grok 4.5 scored the same 9 out of 9 at $2.93 in 6.6 seconds — and both list at exactly $2.00 in, $6.00 out. Identical sticker, identical score, and the newer model costs 70% more and takes 92% longer. The whole difference is reasoning tokens: 628 against 289, up 117%. On this suite the extra thinking bought nothing, because 4.5 already solved everything. Grok 4.6 landed on Google Cloud Vertex AI this week to very little coverage; this page is what it does on our tasks.
A point release that keeps the price and changes the behaviour is the hardest kind to evaluate from a spec sheet, because the spec sheet is identical. You have to run it.
The result
| Metric | Grok 4.6 |
|---|---|
| Score | 9/9 |
| Measured cost / 1,000 tasks | $4.99 |
| Median latency | 12.7s |
| Reasoning tokens per call | 628 |
| Output tokens on the suite | 6,680 |
| List price in / out | $2.00 / $6.00 per 1M |
| Context window | 500,000 tokens |
| Measured | 2026-08-22 |
A clean sweep of the suite. Nothing failed, nothing needed a retry.
4.6 against 4.5, at the same price
This is the comparison that matters, because the two are priced identically and a buyer reading a price page has no way to tell them apart.
| Grok 4.5 | Grok 4.6 | Change | |
|---|---|---|---|
| List price in / out | $2.00 / $6.00 | $2.00 / $6.00 | none |
| Context window | 500,000 | 500,000 | none |
| Score | 9/9 | 9/9 | none |
| Measured cost / 1k tasks | $2.93 | $4.99 | +70% |
| Median latency | 6.6s | 12.7s | +92% |
| Reasoning tokens | 289 | 628 | +117% |
| Measured on | 2026-07-29 | 2026-08-22 | — |
Reasoning tokens bill at the output rate and you cannot see them in advance. Grok 4.6 more than doubled its reasoning spend and converted none of it into a better score, because there was no headroom — 4.5 already went nine for nine. On a harder suite the extra thinking might pay for itself. On this one it is pure overhead, and it is the same pattern we keep measuring: GPT-5.1-Codex-Max costs 3.09x its sibling at a lower list price, and Qwen3.8-Max raised its price 36% and still cut the bill by getting more concise. Reasoning volume, not price per token, is what moves the invoice — the same finding that dominates agent costs across many turns.
The Grok lineup we have measured
| Model | Score | Measured cost / 1k | Latency | Reasoning tokens | Context |
|---|---|---|---|---|---|
| Grok 4.20 | 8/9 | $0.55 | 2s | 0 | 2,000,000 |
| Grok 4.3 | 8/9 | $1.75 | 8.4s | 482 | 1,000,000 |
| Grok 4.5 | 9/9 | $2.93 | 6.6s | 289 | 500,000 |
| Grok 4.6 | 9/9 | $4.99 | 12.7s | 628 | 500,000 |
Grok 4.20 is the outlier worth noticing: $0.55 per 1,000 tasks, a 2-second median, zero reasoning tokens and a 2,000,000-token window — nine times cheaper and six times faster than 4.6 — but it missed parse_csv_line, so 8/9. That is a genuine wrong answer, not a harness failure. Grok 4.3 also went 8/9, dropping flatten. Of the four, only 4.5 and 4.6 cleared the suite, and 4.5 did it in half the time for 41% less.
Note also what is not here: Grok 5 does not exist yet, and the specifications circulating for it are mostly unsourced — we went through what xAI has actually said in the Grok 5 release date page.
Which Grok to run
- Grok 4.5 for coding work of this shape. Same score as 4.6, same price page, 41% cheaper in practice and nearly twice as fast.
- Grok 4.20 if cost and speed dominate and 8/9 is acceptable — it is the cheapest and fastest Grok we have measured, with the largest context at 2,000,000 tokens.
- Grok 4.3 if you want a 1,000,000-token window, double what 4.5 and 4.6 offer, and can live with 8/9.
- Grok 4.6 if you have measured it on your own harder workload and found the extra reasoning earns its keep. We could not surface that on nine self-contained Python functions, and we would not pay 70% on the assumption that it exists.
Outside xAI, the gap is wide: GPT-5.4 mini scored the same 9/9 at $0.53 in 2.3 seconds. Grok is not competing on cost per finished task. Our Grok API pricing page tracks the family's list prices and the 4.5 review has the previous generation in full.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived — measured input and output token counts multiplied by the list price on the date shown, not a billing statement. The two Grok generations were measured 24 days apart at the same list price, which is unusual enough to be worth stating; prices normally move. Runs go through OpenRouter. Full method on the methodology page.
What we did not measure
- The 500,000-token context. Our prompts are short; nothing here tests long-context behaviour.
- Anything harder than nine self-contained Python functions. If 4.6's extra reasoning pays off, it pays off on work this suite does not contain.
- Grok 4.6 on Vertex AI. We called it through OpenRouter. Latency and price on Google Cloud will differ.
- Repeat runs. One scored attempt per task. Single-run figures, not averages, and we do not average across runs.
- Grok Bot, Grok Voice and the consumer tiers — different products. The message limit page covers the consumer quotas.
FAQ
Is Grok 4.6 better than Grok 4.5? Not on our nine tasks. Both scored 9 out of 9; 4.6 cost 70% more and took 92% longer at an identical list price.
How much does Grok 4.6 cost? It lists at $2.00 per million input tokens and $6.00 output as of 2026-08-22. On our suite that came to $4.99 per 1,000 tasks.
Why is Grok 4.6 more expensive if the price is the same? It emits 117% more reasoning tokens, and reasoning bills at the output rate.
How fast is Grok 4.6? 12.7 seconds median on our tasks, against 6.6 for Grok 4.5.
What is Grok 4.6's context window? 500,000 tokens — the same as Grok 4.5, half of Grok 4.3's and a quarter of Grok 4.20's.
DataLLM Lab