Grok 4.7 Review: 9/9 at $4.70 (and 4.8x the Input of Grok 4.6)
Grok 4.7 scored 9 out of 9 on our executed Python benchmark at $4.70 per 1,000 tasks (priced 2026-10-02) and a 10-second mean. The number that matters is not on any price page: across the identical nine prompts, Grok 4.7 consumed 11,761 input tokens; Grok 4.6 consumed 2,437. That is 4.8x the input between adjacent versions, on the same $2 in / $6 out rate card, for prompts we did not change by a character. Grok 4.7 thinks less — 239 reasoning tokens per call against 628 — and writes 53% less, so the bill barely moved: $4.70 against Grok 4.6's $4.99 (priced 2026-08-22). Grok 4.5 still clears the same suite for $2.93 (derived 2026-07-29).
A point release on an unchanged rate card usually means an unchanged bill shape. Grok 4.7 kept the price and the score and rebuilt the shape underneath: far more tokens in, far fewer out.
The result
| Metric | Grok 4.7 | Grok 4.6 | Grok 4.5 |
|---|---|---|---|
| Score | 9/9 | 9/9 | 9/9 |
| Measured cost / 1,000 tasks | $4.70 | $4.99 | $2.93 |
| Priced on | 2026-10-02 | 2026-08-22 | 2026-07-29 |
| Mean latency | 10s | 12.7s | 6.6s |
| Reasoning tokens per call | 239 | 628 | 289 |
| Input tokens across the suite | 11,761 | 2,437 | not stored |
| Output tokens across the suite | 3,131 | 6,680 | not stored |
| List price in / out per 1M | $2 / $6 | $2 / $6 | $2 / $6 today |
| Context window | 500,000 | 500,000 | 500,000 |
| Rank among 75 at 9/9: cost / speed | 54th / 48th | 57th / 55th | 44th / 34th |
Grok 4.5's $2.93 was derived on 2026-07-29, before our results file stored a run-date price; its token counts were not kept either, which is why two cells read not stored. Today's list price for 4.5 is the same $2 / $6.
The input jump
Every model in our suite receives the same nine prompts, byte for byte. A model's input count across them should therefore move only if the provider adds something before our text, or counts the same text differently. Between Grok 4.6 on 2026-08-22 and Grok 4.7 on 2026-10-02 it moved from 2,437 to 11,761 — 9,324 extra input tokens, or 4.8x (11,761 ÷ 2,437).
Per prompt, that is roughly 271 tokens on 4.6 (2,437 ÷ 9) against roughly 1,307 on 4.7 (11,761 ÷ 9). Our prompts did not grow. Something the provider sends with them did.
We have seen this pattern before, inside one family rather than between versions. On the same nine prompts GPT-6 Sol read 604 input tokens and GPT-6 Sol Pro read 16,898; in our GPT-6 Astra review a single-request probe showed Astra Pro billing 1,721 prompt tokens for a 65-token message. What is new with Grok is that the growth arrived in a version bump, not a separately named premium tier. Nothing in the model ID tells you the prompt got bigger.
Two cautions. First, we have not run the single-request probe on Grok, so we know the size of the growth and not its cause — a larger provider-side system prompt is the likeliest explanation, a tokenizer change is a possible contributor, and we cannot separate them from totals. Second, Grok 4.6 was already reading more than the non-Grok models we cite here: GPT-6 Sol read 604 on the same prompts, Claude Opus 5.5 read 914. Tokenizers differ between vendors, so those cross-vendor counts are indicative only. The Grok-to-Grok comparison is the clean one.
What 4.7 traded away, and what it took on
On every other axis Grok 4.7 got leaner. Reasoning fell from 628 to 239 tokens per call, 62% fewer. Output across the suite fell from 6,680 to 3,131, 53% fewer. That is a striking result next to the launch positioning: OpenRouter's launch post (read 2026-10-02) says the model works longer on hard tasks. On nine easy, self-contained functions it worked noticeably less than its predecessor, which is a sensible thing for a model to do on easy work.
Here is why the bill hardly moved. Output costs 3x input on this rate card ($6 ÷ $2). Grok 4.7 shed 3,549 output tokens and took on 9,324 input tokens; at those two prices, the saving and the new cost nearly cancel. Net, measured cost fell from $4.99 to $4.70, about 6%.
The shape of the bill changed more than its size. On Grok 4.6, output tokens outnumbered input 2.7 to 1 (6,680 ÷ 2,437) at three times the price, so input was a minor line item. On Grok 4.7, input outnumbers output 3.76 to 1 (11,761 ÷ 3,131) — more than the 3x price gap — so input is now the larger half of the bill. If you budget Grok by output volume, as most people budget reasoning models, 4.7 will surprise you. If your own prompts are short, a fixed per-call overhead weighs proportionally more; on long prompts it fades into noise.
Grok 4.7 against 4.6 and 4.5
| Change | 4.6 → 4.7 | 4.5 → 4.7 |
|---|---|---|
| Measured cost / 1,000 tasks | $4.99 → $4.70, about 6% lower | $2.93 → $4.70, about 60% higher |
| Mean latency | 12.7s → 10s, about 21% faster | 6.6s → 10s, about 52% slower |
| Reasoning tokens per call | 628 → 239 | 289 → 239 |
| Score | 9/9 → 9/9 | 9/9 → 9/9 |
Against 4.6, Grok 4.7 is a modest improvement on cost, latency and output: cheaper, faster, quieter, same score — with the input jump as the one axis that went the other way. In our Grok 4.6 review the complaint was that 4.6 cost 70% more than 4.5 for identical results; 4.7 claws back a little of that and no more.
Against 4.5 the verdict from August stands. Grok 4.5 remains the cheapest and fastest of the three Grok versions on this suite, at $2.93 (derived 2026-07-29) and 6.6 seconds. Among the 75 models in our set that score 9 out of 9, Grok 4.7 ranks 54th on cost and 48th on speed; Grok 4.5 ranks 44th and 34th. For pricing across the whole Grok line, see the Grok API pricing guide.
From rumour to endpoint
Grok 4.7 is now a real, callable endpoint, and its third-party specs settle partly. What we could confirm on 2026-10-02: MarkTechPost and Decrypt both report a release date of September 21, 2026; xAI's own Grok 4.7 developer page lists a 500,000-token context window and $2 / $6 per million tokens, which matches what we priced. OpenRouter lists the model as x-ai/grok-4.7.
What we could not confirm: launch coverage, including Decrypt, states a 2.1 trillion parameter count and supplemental training data drawn from SpaceX. Decrypt does not attribute either claim, and xAI's developer page, read the same day, gives no parameter count and no training-data detail. Treat both as unconfirmed. We have been here before with Grok sizes: our Grok 5 ledger found the most repeated parameter figure had no named source at all. And whatever the parameter count is, our suite cannot see it — only its cost and its behaviour.
Is it worth $4.70?
On nine self-contained Python functions, plainly no. Solar Mini 4 scored the same 9 out of 9 at $0.03 per 1,000 tasks in 2 seconds (priced 2026-10-02), the cheapest and fastest model to clear the suite as of 2026-10-02 — about 157x cheaper than Grok 4.7 ($4.70 ÷ $0.03). Nine functions of this size cannot separate a frontier model from a competent small one, and we will not pretend they can. For the cheap end of the field, our cheap coding roundup has the list.
What the suite can tell you is what changed between versions at a fixed price. If you already run Grok 4.6, moving to 4.7 lowered our measured cost and latency slightly. If you run Grok 4.5 and your work looks like ours, 4.7 costs more and answers slower for the same score. If your work is the long, hard kind xAI says it built 4.7 for, our numbers are silent on it — measure that yourself, and watch your input-token line when you do.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; the harness reconnects only on an API error, never on a wrong answer. Cost is derived — measured input and output token counts multiplied by the list price on the run date (2026-10-02 for Grok 4.7) — not a billing statement, and prices move. Input-token totals are the counts the provider reported back for each call. Runs go through OpenRouter, deliberately not through the DataLLM Lab gateway, so the numbers do not depend on our infrastructure. Full method on the methodology page.
What we did not measure
- The cause of the input jump. We measured totals across nine prompts, not a single-request usage payload. System prompt, tokenizer change, or both — we cannot say yet.
- Caching. Our derived cost charges every input token at the full list rate. If the added prefix is cached on repeat calls, your bill on a busy workload would be lower than ours.
- Grok 4.6 today. Its 2,437 comes from 2026-08-22. We did not re-run it on 2026-10-02, so if xAI changed 4.6's prompt in the meantime, the gap could be smaller now.
- Grok 4.5's token counts, which were not stored. We cannot tell whether input grew across three versions or only the last two.
- Long-running agentic work, which is the stated reason Grok 4.7 exists, and the 500,000-token context. Our prompts are short and single-turn.
- Repeat runs. One scored attempt per task; single-run figures, not averages. A 6% cost difference to 4.6 is small enough to sit near run-to-run variation.
FAQ
How much does Grok 4.7 cost? $2 per million input tokens and $6 output as of 2026-10-02, the same as Grok 4.6. On our nine tasks that worked out to $4.70 per 1,000 tasks.
Is Grok 4.7 better than Grok 4.6? On our suite it is slightly cheaper ($4.70 against $4.99) and faster (10s against 12.7s) at the same 9 out of 9, but it reads 4.8x more input tokens on identical prompts.
Does Grok 4.7 have 2.1 trillion parameters? That figure appears in launch coverage without attribution, and xAI's developer page gives no parameter count. We treat it as unconfirmed.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab