Model Reviews

Grok 4.7 Review: 9/9 at $4.70 (and 4.8x the Input of Grok 4.6)

Grok 4.7 scored 9 out of 9 on our executed Python benchmark at $4.70 per 1,000 tasks (priced 2026-10-02) and a 10-second mean. The number that matters is not on any price page: across the identical nine prompts, Grok 4.7 consumed 11,761 input tokens; Grok 4.6 consumed 2,437. That is 4.8x the input between adjacent versions, on the same $2 in / $6 out rate card, for prompts we did not change by a character. Grok 4.7 thinks less — 239 reasoning tokens per call against 628 — and writes 53% less, so the bill barely moved: $4.70 against Grok 4.6's $4.99 (priced 2026-08-22). Grok 4.5 still clears the same suite for $2.93 (derived 2026-07-29).

DataLLM Lab article cover: Grok 4.7 Review: 9/9 at $4.70 (and 4.8x the Input of Grok 4.6)

A point release on an unchanged rate card usually means an unchanged bill shape. Grok 4.7 kept the price and the score and rebuilt the shape underneath: far more tokens in, far fewer out.

The result

MetricGrok 4.7Grok 4.6Grok 4.5
Score9/99/99/9
Measured cost / 1,000 tasks$4.70$4.99$2.93
Priced on2026-10-022026-08-222026-07-29
Mean latency10s12.7s6.6s
Reasoning tokens per call239628289
Input tokens across the suite11,7612,437not stored
Output tokens across the suite3,1316,680not stored
List price in / out per 1M$2 / $6$2 / $6$2 / $6 today
Context window500,000500,000500,000
Rank among 75 at 9/9: cost / speed54th / 48th57th / 55th44th / 34th

Grok 4.5's $2.93 was derived on 2026-07-29, before our results file stored a run-date price; its token counts were not kept either, which is why two cells read not stored. Today's list price for 4.5 is the same $2 / $6.

The input jump

Every model in our suite receives the same nine prompts, byte for byte. A model's input count across them should therefore move only if the provider adds something before our text, or counts the same text differently. Between Grok 4.6 on 2026-08-22 and Grok 4.7 on 2026-10-02 it moved from 2,437 to 11,761 — 9,324 extra input tokens, or 4.8x (11,761 ÷ 2,437).

Per prompt, that is roughly 271 tokens on 4.6 (2,437 ÷ 9) against roughly 1,307 on 4.7 (11,761 ÷ 9). Our prompts did not grow. Something the provider sends with them did.

Same nine prompts. Grok 4.7 reads 4.8x more and writes half as much.Blue = Grok 4.7. Grey = earlier Grok versions. Every model scored 9/9 on the identical suite.INPUT TOKENS ACROSS THE NINE PROMPTSGrok 4.62,437Grok 4.711,761OUTPUT TOKENS ACROSS THE SUITEGrok 4.66,680Grok 4.73,131 — down 53%MEASURED COST PER 1,000 TASKS · ALL 9/9Grok 4.5$2.93Grok 4.6$4.99Grok 4.7$4.70Scales: tokens 0.05 px per token (11,761 x 0.05 = 588.05 px); cost 100 px per dollar ($4.99 = 499 px).Cost derived at list price on each run date: 4.5 2026-07-29, 4.6 2026-08-22, 4.7 2026-10-02.
The input bar is the story. The cost bar barely notices it, for reasons below.

We have seen this pattern before, inside one family rather than between versions. On the same nine prompts GPT-6 Sol read 604 input tokens and GPT-6 Sol Pro read 16,898; in our GPT-6 Astra review a single-request probe showed Astra Pro billing 1,721 prompt tokens for a 65-token message. What is new with Grok is that the growth arrived in a version bump, not a separately named premium tier. Nothing in the model ID tells you the prompt got bigger.

Two cautions. First, we have not run the single-request probe on Grok, so we know the size of the growth and not its cause — a larger provider-side system prompt is the likeliest explanation, a tokenizer change is a possible contributor, and we cannot separate them from totals. Second, Grok 4.6 was already reading more than the non-Grok models we cite here: GPT-6 Sol read 604 on the same prompts, Claude Opus 5.5 read 914. Tokenizers differ between vendors, so those cross-vendor counts are indicative only. The Grok-to-Grok comparison is the clean one.

What 4.7 traded away, and what it took on

On every other axis Grok 4.7 got leaner. Reasoning fell from 628 to 239 tokens per call, 62% fewer. Output across the suite fell from 6,680 to 3,131, 53% fewer. That is a striking result next to the launch positioning: OpenRouter's launch post (read 2026-10-02) says the model works longer on hard tasks. On nine easy, self-contained functions it worked noticeably less than its predecessor, which is a sensible thing for a model to do on easy work.

Here is why the bill hardly moved. Output costs 3x input on this rate card ($6 ÷ $2). Grok 4.7 shed 3,549 output tokens and took on 9,324 input tokens; at those two prices, the saving and the new cost nearly cancel. Net, measured cost fell from $4.99 to $4.70, about 6%.

The shape of the bill changed more than its size. On Grok 4.6, output tokens outnumbered input 2.7 to 1 (6,680 ÷ 2,437) at three times the price, so input was a minor line item. On Grok 4.7, input outnumbers output 3.76 to 1 (11,761 ÷ 3,131) — more than the 3x price gap — so input is now the larger half of the bill. If you budget Grok by output volume, as most people budget reasoning models, 4.7 will surprise you. If your own prompts are short, a fixed per-call overhead weighs proportionally more; on long prompts it fades into noise.

Grok 4.7 against 4.6 and 4.5

Change4.6 → 4.74.5 → 4.7
Measured cost / 1,000 tasks$4.99 → $4.70, about 6% lower$2.93 → $4.70, about 60% higher
Mean latency12.7s → 10s, about 21% faster6.6s → 10s, about 52% slower
Reasoning tokens per call628 → 239289 → 239
Score9/9 → 9/99/9 → 9/9

Against 4.6, Grok 4.7 is a modest improvement on cost, latency and output: cheaper, faster, quieter, same score — with the input jump as the one axis that went the other way. In our Grok 4.6 review the complaint was that 4.6 cost 70% more than 4.5 for identical results; 4.7 claws back a little of that and no more.

Against 4.5 the verdict from August stands. Grok 4.5 remains the cheapest and fastest of the three Grok versions on this suite, at $2.93 (derived 2026-07-29) and 6.6 seconds. Among the 75 models in our set that score 9 out of 9, Grok 4.7 ranks 54th on cost and 48th on speed; Grok 4.5 ranks 44th and 34th. For pricing across the whole Grok line, see the Grok API pricing guide.

From rumour to endpoint

Grok 4.7 is now a real, callable endpoint, and its third-party specs settle partly. What we could confirm on 2026-10-02: MarkTechPost and Decrypt both report a release date of September 21, 2026; xAI's own Grok 4.7 developer page lists a 500,000-token context window and $2 / $6 per million tokens, which matches what we priced. OpenRouter lists the model as x-ai/grok-4.7.

What we could not confirm: launch coverage, including Decrypt, states a 2.1 trillion parameter count and supplemental training data drawn from SpaceX. Decrypt does not attribute either claim, and xAI's developer page, read the same day, gives no parameter count and no training-data detail. Treat both as unconfirmed. We have been here before with Grok sizes: our Grok 5 ledger found the most repeated parameter figure had no named source at all. And whatever the parameter count is, our suite cannot see it — only its cost and its behaviour.

Is it worth $4.70?

On nine self-contained Python functions, plainly no. Solar Mini 4 scored the same 9 out of 9 at $0.03 per 1,000 tasks in 2 seconds (priced 2026-10-02), the cheapest and fastest model to clear the suite as of 2026-10-02 — about 157x cheaper than Grok 4.7 ($4.70 ÷ $0.03). Nine functions of this size cannot separate a frontier model from a competent small one, and we will not pretend they can. For the cheap end of the field, our cheap coding roundup has the list.

What the suite can tell you is what changed between versions at a fixed price. If you already run Grok 4.6, moving to 4.7 lowered our measured cost and latency slightly. If you run Grok 4.5 and your work looks like ours, 4.7 costs more and answers slower for the same score. If your work is the long, hard kind xAI says it built 4.7 for, our numbers are silent on it — measure that yourself, and watch your input-token line when you do.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; the harness reconnects only on an API error, never on a wrong answer. Cost is derived — measured input and output token counts multiplied by the list price on the run date (2026-10-02 for Grok 4.7) — not a billing statement, and prices move. Input-token totals are the counts the provider reported back for each call. Runs go through OpenRouter, deliberately not through the DataLLM Lab gateway, so the numbers do not depend on our infrastructure. Full method on the methodology page.

What we did not measure

FAQ

How much does Grok 4.7 cost? $2 per million input tokens and $6 output as of 2026-10-02, the same as Grok 4.6. On our nine tasks that worked out to $4.70 per 1,000 tasks.

Is Grok 4.7 better than Grok 4.6? On our suite it is slightly cheaper ($4.70 against $4.99) and faster (10s against 12.7s) at the same 9 out of 9, but it reads 4.8x more input tokens on identical prompts.

Does Grok 4.7 have 2.1 trillion parameters? That figure appears in launch coverage without attribution, and xAI's developer page gives no parameter count. We treat it as unconfirmed.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.