Model Reviews

Grok 4.6 Review: Same Price as 4.5, Same Score, 70% Bigger Bill

Grok 4.6 scored 9 out of 9 on our executed Python benchmark at $4.99 per 1,000 tasks and a 12.7-second median. Grok 4.5 scored the same 9 out of 9 at $2.93 in 6.6 seconds — and both list at exactly $2.00 in, $6.00 out. Identical sticker, identical score, and the newer model costs 70% more and takes 92% longer. The whole difference is reasoning tokens: 628 against 289, up 117%. On this suite the extra thinking bought nothing, because 4.5 already solved everything. Grok 4.6 landed on Google Cloud Vertex AI this week to very little coverage; this page is what it does on our tasks.

Chart comparing measured cost and latency for Grok 4.6 against Grok 4.5 at an identical list price

A point release that keeps the price and changes the behaviour is the hardest kind to evaluate from a spec sheet, because the spec sheet is identical. You have to run it.

The result

MetricGrok 4.6
Score9/9
Measured cost / 1,000 tasks$4.99
Median latency12.7s
Reasoning tokens per call628
Output tokens on the suite6,680
List price in / out$2.00 / $6.00 per 1M
Context window500,000 tokens
Measured2026-08-22

A clean sweep of the suite. Nothing failed, nothing needed a retry.

4.6 against 4.5, at the same price

This is the comparison that matters, because the two are priced identically and a buyer reading a price page has no way to tell them apart.

Grok 4.5Grok 4.6Change
List price in / out$2.00 / $6.00$2.00 / $6.00none
Context window500,000500,000none
Score9/99/9none
Measured cost / 1k tasks$2.93$4.99+70%
Median latency6.6s12.7s+92%
Reasoning tokens289628+117%
Measured on2026-07-292026-08-22
Same sticker, same score, everything else worseGrey = Grok 4.5. Blue = Grok 4.6. Both scored 9/9 on the identical nine tasks.List output price per 1M$6.00$6.00 — identicalMeasured cost per 1,000 tasks$2.93$4.99 — up 70%Median latency6.6s12.7s — up 92%Each row has its own scale, sized so the larger value fills 240 px. Compare within a row, not between rows.
Nothing on the price page distinguishes these two. Everything in the measurement does.

Reasoning tokens bill at the output rate and you cannot see them in advance. Grok 4.6 more than doubled its reasoning spend and converted none of it into a better score, because there was no headroom — 4.5 already went nine for nine. On a harder suite the extra thinking might pay for itself. On this one it is pure overhead, and it is the same pattern we keep measuring: GPT-5.1-Codex-Max costs 3.09x its sibling at a lower list price, and Qwen3.8-Max raised its price 36% and still cut the bill by getting more concise. Reasoning volume, not price per token, is what moves the invoice — the same finding that dominates agent costs across many turns.

The Grok lineup we have measured

ModelScoreMeasured cost / 1kLatencyReasoning tokensContext
Grok 4.208/9$0.552s02,000,000
Grok 4.38/9$1.758.4s4821,000,000
Grok 4.59/9$2.936.6s289500,000
Grok 4.69/9$4.9912.7s628500,000

Grok 4.20 is the outlier worth noticing: $0.55 per 1,000 tasks, a 2-second median, zero reasoning tokens and a 2,000,000-token window — nine times cheaper and six times faster than 4.6 — but it missed parse_csv_line, so 8/9. That is a genuine wrong answer, not a harness failure. Grok 4.3 also went 8/9, dropping flatten. Of the four, only 4.5 and 4.6 cleared the suite, and 4.5 did it in half the time for 41% less.

Note also what is not here: Grok 5 does not exist yet, and the specifications circulating for it are mostly unsourced — we went through what xAI has actually said in the Grok 5 release date page.

Which Grok to run

Outside xAI, the gap is wide: GPT-5.4 mini scored the same 9/9 at $0.53 in 2.3 seconds. Grok is not competing on cost per finished task. Our Grok API pricing page tracks the family's list prices and the 4.5 review has the previous generation in full.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived — measured input and output token counts multiplied by the list price on the date shown, not a billing statement. The two Grok generations were measured 24 days apart at the same list price, which is unusual enough to be worth stating; prices normally move. Runs go through OpenRouter. Full method on the methodology page.

What we did not measure

FAQ

Is Grok 4.6 better than Grok 4.5? Not on our nine tasks. Both scored 9 out of 9; 4.6 cost 70% more and took 92% longer at an identical list price.

How much does Grok 4.6 cost? It lists at $2.00 per million input tokens and $6.00 output as of 2026-08-22. On our suite that came to $4.99 per 1,000 tasks.

Why is Grok 4.6 more expensive if the price is the same? It emits 117% more reasoning tokens, and reasoning bills at the output rate.

How fast is Grok 4.6? 12.7 seconds median on our tasks, against 6.6 for Grok 4.5.

What is Grok 4.6's context window? 500,000 tokens — the same as Grok 4.5, half of Grok 4.3's and a quarter of Grok 4.20's.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.