Model Reviews

Granite 4.2 8B: An Incomplete Run (and Why We Publish No Score)

IBM Granite 4.2 8B emitted 1,483 reasoning tokens per call — more thinking per call than GLM 5.1 at 1,327, which is the heaviest reasoner among the models in our set that did clear the suite — and 12,033 output tokens across the suite, the largest output total among the six small models we chart below. We are publishing no score for it. Two of the nine tasks failed at the API layer after retries, seven were scored, and seven scored tasks is not our benchmark. So this is the article about what an incomplete run is actually worth: something, but not a number you can rank.

DataLLM Lab article cover: Granite 4.2 8B: An Incomplete Run (and Why We Publish No Score)

Almost every row in our set that misses full marks misses because a model wrote a function that failed a hidden assert. Granite 4.2 8B did not write a wrong function. Two of its nine calls simply never came back, after retries — and those two are why this page carries a table of observations instead of a score.

What the run recorded

FieldIBM Granite 4.2 8B
StatusExcluded — incomplete run
Tasks scored7 of 9
Failures2, at the API layer, after retries
Reasoning tokens per call1,483
Tokens across the suite512 in / 12,033 out
Mean latency12.8s
Derived cost / 1,000 tasks$0.43 — see the warning below
List price in / out$0.06 / $0.25 per 1M, captured 2026-09-15
Context window131,072
Run date2026-09-16

Hold onto one caveat before reading any of that: we do not know how many calls produced the 12,033 output tokens. Two tasks died at the API layer, and whatever they emitted first — possibly nothing at all — is not something our record separates out. If those two contributed nothing, that total came out of seven calls, where every other row on this page had nine. Under that reading the verbosity below gets worse, not better.

Seven of nine is not seven out of nine

Our harness records two different things when a task fails to produce a passing function. One is a wrong answer: the model returned code, the code ran in the sandbox, a hidden assert failed. The other is an API-layer failure: no usable body came back after retries. In a spreadsheet they look almost identical. They mean opposite things. A wrong answer is evidence about the model. A transport failure is evidence about the day.

We learned to keep those apart the expensive way. For a stretch, filtered empty responses were being scored as wrong answers — the harness saw no passing function and wrote down a miss, quietly defaming every model that tripped a content filter. The fix was to record the failure class, not just the miss. The rule that came with it is the one being applied here: any run carrying a recorded API-layer failure is marked excluded and publishes no score.

So Granite 4.2 8B does not get a 7/9. It also does not get the number that sits in the raw field, which is 7/7 — and that is the most misleading figure on this page. It says every task that came back passed, which is true and worth knowing, but it is 100% of a set that the model itself partly selected by not returning. Printing it beside a 9/9 would invite exactly the comparison it cannot support. We would rather leave a hole in the table than fill it with a number that fits the column.

Where the token budget goes

What the seven completed calls do tell us is how Granite 4.2 spends a budget, and the answer is: nearly all of it, on thinking.

Output tokens across the nine-task suiteGranite 4.2 8B is the dashed bar: seven of nine tasks scored, so its total is not a like-for-like total.IBM Granite 4.2 8B12,033 · no scoreGLM 5.3 Flash11,884Mercury 2.59,882DeepSeek V4.1 Flash5,217Ling 3.0 Flash VL3,299GPT-5.4 Mini969One scale throughout: 3.75 px per 100 output tokens. Each model measured on its own run date; dates in the table below.
The longest bar on the chart belongs to the one model here with nothing to show for it.
ModelReasoning tok/callOutput tok, suiteScoreDerived cost / 1k tasksPriced on
IBM Granite 4.2 8B1,48312,033none — 7 of 9 scored$0.432026-09-15
GLM 5.3 Flash1,21211,8849/9$0.342026-08-31
Mercury 2.59899,8829/9$0.172026-09-15
DeepSeek V4.1 Flash4785,2179/9$0.362026-09-15
Ling 3.0 Flash VL2343,2999/9$0.072026-09-15
GPT-5.4 Mini09699/9$0.532026-07-30

Every call in this suite runs with max_tokens at 4000. Three times 1,483 is 4,449, which is more than 4,000 — so on any endpoint where reasoning tokens are charged against that cap, Granite 4.2 is spending more than a third of its allowance before the first line of code appears. We cannot verify from outside how this particular endpoint accounts for reasoning, and we are not going to pretend otherwise. What we can say is that a model budgeting like this has very little runway left for a long answer.

The instructive comparison is GLM 5.3 Flash, one row down: 1,212 reasoning tokens per call, 11,884 output tokens across the suite, and it finished all nine. Verbosity is not itself a failure mode. But we have written up the models that run out of budget and return empty bodies, and models that live this close to the ceiling are the ones that end up in that article. To be explicit, because it matters: Granite's two failures are recorded at the API layer, not as length truncation. We are pointing at a neighbourhood, not at a cause.

Why $0.43 is not a price you can rank

Granite 4.2 8B was priced at $0.06 in / $0.25 out per 1M on 2026-09-15, and the derived cost from that run is $0.43 per 1,000 tasks. GLM 5.3 Flash was priced at $0.075 in / $0.25 out on 2026-08-31 and lands at $0.34. Identical output price, cheaper input price, higher derived cost — because what you actually pay for is tokens emitted, and Granite emitted 12,033 to GLM's 11,884.

That is the honest half of the comparison. Here is the dishonest half, which is why the figure carries a warning: those two costs are not derived over the same number of scored tasks, and they were priced two weeks apart. A derived cost from an incomplete run tells you what that run consumed. It does not tell you what nine tasks would have cost, because we never got nine. Ranking it against a complete run would be arithmetic with a missing denominator, and list prices move underneath all of this anyway.

What we would reach for meanwhile

Nothing here is a verdict on Granite 4.2. If you need a small model with a measured result today, the set has options: Ling 3.0 Flash VL at $0.07 per 1,000 tasks priced 2026-09-15, 9/9 and 4.4s mean; GPT-5.4 Mini at $0.53 priced 2026-07-30, 9/9, 0 reasoning tokens and the fastest 9/9 in the set at 2.3s; GLM 5.3 Flash at $0.34 priced 2026-08-31 if you want the reasoning-heavy shape done by something that finished. The cheap-coding roundup has the rest of that field.

And the caveat that applies to all of them, Granite included: nine self-contained Python functions cannot separate a frontier model from a competent small one. Fifty-six models in our set score 9 out of 9, spanning $0.07 to $57.06 per 1,000 tasks at the prices each was measured on. The suite is good at catching models that cannot do the basics and good at pricing the ones that can. It is not a ranking of intelligence, and a model that clears it has only shown you the floor.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is the entire reason this run is excluded rather than scored low. Cost is derived from measured token counts at the list price captured on the date shown, not a billing statement. Runs go through OpenRouter. Full method on the methodology page.

What we did not measure

FAQ

What did Granite 4.2 score on your benchmark? Nothing we will publish. Seven of nine tasks were scored, two failed at the API layer after retries, and an incomplete run carries no score in our data.

Is Granite 4.2 bad at Python, then? We have no evidence for that. Every task that returned a body passed its hidden asserts. We just cannot tell you what the other two would have done.

How much does Granite 4.2 8B cost? $0.06 per million input tokens and $0.25 output, captured 2026-09-15. The $0.43 per 1,000 tasks derived from our run is not comparable to a complete run.

Why is it so verbose? 1,483 reasoning tokens per call and 12,033 output tokens across the suite is the largest output total among the six small models we chart above. We observed it; we cannot explain it from outside.

Will you re-run it? Yes. Until then the row stays excluded and this page stays a set of observations.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.