Model Reviews

Inkling Small: Same Score as Inkling, Half the Bill, Twice the Wait

Inkling Small scored 9 out of 9 on our executed Python benchmark at $0.94 per 1,000 tasks and an 8.9-second mean. We ran Inkling, its 3.5x larger sibling, on the identical suite the same day: also 9 out of 9, at $2.04 in 4.2 seconds. So the small model is 54% cheaper and 112% slower, and it gets to the same answer by emitting 72% more reasoning tokens — 665 against 386. Artificial Analysis reported that Inkling Small lands within a point of Inkling on their intelligence index at less than a third of the parameters. This page is the other half of that sentence: what the parameter saving costs you operationally, in dollars and seconds, on work that actually executes.

Correction, October 3, 2026: single-model benchmark latency is a mean, not a median. Labels were corrected; measured values and the original run dates are unchanged. Measurement details.

Chart comparing measured cost and latency between Inkling and Inkling Small, both scoring 9 of 9

“Nearly as good at a quarter the size” is the standard claim for a small sibling model, and it is usually argued on benchmark points. Points are not the whole trade. A smaller model that reaches the same answer by thinking longer costs you time, and reasoning time is billed.

What Inkling Small is

Third-party facts, from the vendor and the launch coverage:

One thing sources disagree on: the context window. The published model card is quoted at 256K, while OpenRouter serves the paid endpoint at 1,048,576 tokens and a free variant at 262,144. We did not test long context either way, so we are reporting the disagreement rather than resolving it. Our Inkling review flagged a similar sourcing problem on that model's benchmark figures.

What it scored

MetricInkling Small
Score9/9
Measured cost / 1,000 tasks$0.94
Mean latency8.9s
Reasoning tokens per call665
Output tokens on the suite6,819
List price in / out$0.45 / $1.20 per 1M
Measured2026-08-22

Small against Inkling

InklingInkling SmallChange
Total / active parameters975B / 41B276B / 12B−72% total
List price in / out$0.95 / $4.05$0.45 / $1.20−53% / −70%
Score9/99/9none
Measured cost / 1k tasks$2.04$0.94−54%
Mean latency4.2s8.9s+112%
Reasoning tokens386665+72%
Output tokens4,3776,819+56%

The trade is latency, not quality

Half the bill, twice the wait, same nine out of nineGrey = Inkling (975B/41B). Blue = Inkling Small (276B/12B). Identical nine tasks, same day.Measured cost per 1,000 tasks$2.04$0.94 — 54% cheaperMean latency4.2s8.9s — 112% slowerReasoning tokens per call386665 — thinks 72% harderEach row has its own scale, sized so the larger value fills 240 px. Compare within a row, not between rows.
The parameter saving did not cost a point on our suite. It cost 4.7 seconds a task.

Read the three rows together. The small model is cheaper per token and spends more tokens, and the price cut wins — but only on money. On wall-clock it loses outright, because reasoning tokens are generated serially and there are 72% more of them.

That distinction decides where each one belongs. In a batch job, 8.9 seconds against 4.2 is invisible and the 54% saving is the whole story. In an interactive loop, or an agent firing dozens of calls per task, doubling per-call latency is the story and the saving is a rounding error — the dynamic we worked through in our agent model comparison.

It is also a useful counter-example to the pattern we have been finding all week. Grok 4.6 and GLM-5.3 both spend more reasoning than their predecessors and deliver nothing extra for it. Inkling Small spends more reasoning too — but it is doing so from a much smaller model at a much lower token price, which is a trade that actually buys you something.

Which one to run

Against the wider open field, neither is cheap: DeepSeek V3.2 scored 9/9 at $0.08 and Qwen3 Coder Next at $0.10, both roughly ten times cheaper than Inkling Small on the same tasks. What you are buying from Thinking Machines is multimodal input, the effort dial and Apache 2.0 weights, not the lowest cost per finished task. Our open-source guide has the wider field.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. Cost is derived — measured token counts multiplied by the list price captured 2026-08-22, not a billing statement. Both models were run the same day at the same settings, which is what makes the comparison worth anything. Runs go through OpenRouter. Full method on the methodology page, and prices move, so the date matters.

What we did not measure, starting with the effort dial

FAQ

How big is Inkling Small? 276B total parameters with 12B active, against Inkling's 975B total and 41B active.

Is Inkling Small as good as Inkling? On our nine executed Python tasks, yes — both scored 9 out of 9. It got there 112% slower.

How much does Inkling Small cost? It lists at $0.45 per million input tokens and $1.20 output as of 2026-08-22. On our suite that came to $0.94 per 1,000 tasks, against $2.04 for Inkling.

What licence is Inkling Small under? Apache 2.0, the same as Inkling.

Why is the smaller model slower? It emits 72% more reasoning tokens to reach the same answer, and those are generated serially.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.