Model Review

Inkling Review: Thinking Machines First Open-Weights Model (July 2026)

Thinking Machines Lab released Inkling on 15 July 2026: a 975B-total / 41B-active Mixture-of-Experts model under Apache 2.0, with up to 1M-token context and a continuous thinking-effort dial. Almost every article about it restates the same spec sheet. This one does something the coverage did not: it tags every headline number as vendor-run or externally sourced, flags a benchmark figure that two sources disagree on, and converts the effort dial into dollars.

Inkling by Thinking Machines Lab: vendor-reported versus independently measured benchmark numbers

What Thinking Machines actually shipped

Inkling landed on Wednesday 15 July 2026, with coverage running into the 16th. It is the first open-weights model from Thinking Machines Lab, the company founded by former OpenAI CTO Mira Murati, which TechCrunch reports has roughly 200 employees after a round of departures in early 2026 including two co-founders who left for OpenAI in January.

The architecture, as described on the model card and in the launch write-ups: a 66-layer decoder-only transformer with sparse Mixture-of-Experts feed-forward layers, 975B total parameters and 41B active per token. Each token is routed to 6 of 256 experts, plus 2 shared experts that fire on every token. Context runs up to 1M tokens for the downloadable weights, while the hosted Tinker offering exposes 64K and 256K options. Pretraining covered 45 trillion tokens of text, images, audio and video. It accepts text, image and audio input and emits UTF-8 text only — no image or audio generation.

The license is Apache 2.0, confirmed on both the model card and the Hugging Face repo. That is a real permissive grant: commercial use, modification, fine-tuning and redistribution, with none of the usage carve-outs that make some so-called open models awkward for commercial deployment. Apache 2.0 weights are still not open source in the OSI sense, because the training data and pipeline are not published — a distinction we work through model by model in the best open source LLMs of 2026, ranked and licensed.

The distinctive feature is the thinking effort control: a continuous parameter that scales how much the model reasons before answering. Thinking Machines says effort was trained in during reinforcement learning by varying the system message and the per-token cost, and the published sweep range runs 0.2 to 0.99. In the Hugging Face transformers integration it surfaces as named levels — none, minimal, low, medium, high, xhigh, max — with published numeric mappings of none at 0, low at 0.2, medium at 0.7, high at 0.9, and xhigh and max at 0.99. Note that none sits at 0, below the swept 0.2 floor, so the dial is not literally bounded at 0.2.

The most useful thing in the launch materials is a sentence from the vendor itself. Thinking Machines writes that Inkling is not the strongest overall model available today, open or closed, and positions it instead as a base for customization: multimodal, token-efficient, and fine-tunable on Tinker from day one with a limited-time 50% discount and a playground that is free for a limited time. TechCrunch frames Tinker as the company's revenue path. Hold that thought — it explains a lot about the release.

Vendor numbers vs independent numbers

Here is the problem with the Inkling coverage. Thinking Machines published a model-card benchmark table. Artificial Analysis published an independent evaluation. Press coverage blended the two, and several outlets — plus the research brief that started this article — attributed the 77.6% SWE-bench Verified score to Artificial Analysis. AA never reported a SWE-bench Verified figure for Inkling. That number is the vendor's own, measured at effort 0.99 and temperature 1.0.

It cuts the other way too, and almost nobody noticed: Thinking Machines states in its own launch post that it relies on externally reported evaluations — naming Artificial Analysis — for HLE, GPQA Diamond, GDPval, τ³ Banking, AA-Omniscience and MMMU Pro, for both its own model and its competitors. So parts of the "vendor table" are not vendor-run at all, and the clean two-column story most coverage told is wrong in both directions.

So the first original asset here is not another spec table. It is the same numbers with a provenance column attached.

MetricInkling resultWho produced itWhat it actually tells you
AA Intelligence Index41Independent — Artificial AnalysisAA's framing: the new leading U.S. open weights model. AA comparison points: Nemotron 3 Ultra 38, Gemma 4 31B 29, gpt-oss-120b 24.
Output tokens per AA Index task~25KIndependent — Artificial AnalysisAgainst GLM-5.2 (max) 43K, Kimi K2.6 38K, DeepSeek v4 Pro (max) 37K. The efficiency story survives independent measurement.
GDPval-AA v2 Elo1238Independent — Artificial AnalysisAhead of Kimi K2.6 at 1190 and DeepSeek v4 Flash (max) at 1189 on AA's agentic suite.
τ³ Banking24%Independent — Artificial AnalysisNarrow lead over DeepSeek v4 Flash (max) 23% and Kimi K2.6 21%. Close enough that ranking here is not a buying signal.
AA-Omniscience+2Independent — Artificial AnalysisThe counterweight. Ahead of Nemotron 3 Ultra at -1, behind the leading open-weights field. Accuracy 40%, hallucination rate 63% on that benchmark's construct.
SWE-bench Verified77.6%Vendor-run — TM model card, effort 0.99Not an AA figure. In Thinking Machines' own table: Nemotron 3 Ultra 70.7%, Kimi K2.5 76.8%, GLM 5.2 80.0%, Kimi K2.6 80.2%, DeepSeek V4 Pro 80.6%, Gemini 3.1 Pro 80.6%, GPT 5.6 Sol 82.2%, Claude Fable 5 95.0%.
AIME 202697.1%Vendor-run — effort 0.99Maximum-effort setting. TM's competitor figures in the same table come from externally reported evaluations, not from TM re-running them.
GPQA Diamond87.2%Externally sourced — TM says it uses Artificial Analysis figures for this rowNot a vendor-run number despite sitting in a vendor table. Kimi K2.6 91.1%, Gemini 3.1 Pro and GPT 5.6 Sol 94.1%.
Terminal Bench 2.163.8%Vendor-run — internal coding harnessTM states its Inkling numbers here were produced with an internal coding harness while external models' figures are self-reported, and that rollouts with web-search solution contamination were scored 0.
VoiceBench / MMMU Pro (Standard 10)91.4% / 73.5%Mixed — VoiceBench vendor-run; MMMU Pro externally sourced per TMThe multimodal claim. Gemini 3.1 Pro is listed at 94.3% on VoiceBench and 82.0% on MMMU Pro.
Serving price, 64K tier$1.87 in / $0.374 cached / $4.68 out per 1MThird-party listing — AA, July 2026256K tier doubles to $3.74 / $0.748 / $9.36. Explicitly launch-period pricing.

Table: DataLLM Lab. Figures as published July 2026. Provenance column is ours; the underlying numbers are from the Thinking Machines model card and launch post and the Artificial Analysis launch evaluation.

Read the table top to bottom and a different story emerges from the launch headlines. The independently verified wins are about efficiency and agentic behaviour, not raw capability. The headline capability numbers that are genuinely vendor-run — SWE-bench Verified, AIME, Terminal Bench — were produced at maximum effort, with TM's own harness in the Terminal Bench case, against competitor figures TM took from those competitors' reports. That is a structurally flattering comparison, and Thinking Machines is more candid about it than most: it publishes effort-sweep curves for Inkling on Terminal Bench 2.1, HLE and IFBench, where competitors are represented by single externally reported results. A curve will always look good against a dot.

The Nemotron figure nobody checked

The most repeated line in the coverage is some version of "Inkling's 77.6% on SWE-bench Verified beats Nemotron 3 Ultra's 71.9%." Go to the source table and that comparison number is not there.

SourceNemotron 3 Ultra, SWE-bench VerifiedStatus
Thinking Machines model card comparison table70.7%What TM actually published alongside Inkling
NVIDIA's own Nemotron 3 Ultra materials71.9%The number circulating in secondary press
Inkling (TM, effort 0.99)77.6%Vendor-run, same table

Table: DataLLM Lab. Cross-source reconciliation, July 2026.

A 1.2-point gap does not change who is ahead. It matters because of what it reveals: cross-vendor benchmark numbers are not apples-to-apples even on a benchmark as standardised as SWE-bench Verified. Harness differences, scaffold differences, checkpoint and quantisation differences all move the score by a point or two. Thinking Machines said as much about its own Terminal Bench runs. If a 1.2-point discrepancy can survive into near-universal repetition, so can a 5-point one, in either direction. Our take on Nemotron in isolation is in the Nemotron 3 Ultra review, and the broader field is ranked in the best open source LLMs of 2026.

The practical rule: never let a single cross-vendor benchmark decide a model choice. Run your own eval on your own harness, on your own tasks, and treat leaderboards as a shortlist generator only.

The effort dial is a cost dial

Every article about Inkling describes the effort parameter. None of them price it. That is the gap, because output tokens are what you actually pay for, and a knob that controls output token volume is a knob that controls your bill.

Output tokens per AA Intelligence Index task Token averages measured by Artificial Analysis; dollar figures modeled by DataLLM Lab Inkling 25K tokens · $117 modeled DeepSeek v4 Pro (max) 37K tokens · $173 modeled Kimi K2.6 38K tokens · $178 modeled GLM-5.2 (max) 43K tokens · $201 modeled 0 43K output tokens
Token averages: Artificial Analysis, July 2026. Dollar column is modeled, not measured — it applies Inkling's own AA-listed 64K-tier output price of $4.68 per 1M tokens to all four token counts, per 1,000 tasks, to isolate the effect of verbosity alone. Rivals are billed at their own rates, so this is a token-volume comparison, not a vendor price comparison. Chart: DataLLM Lab

Hold price constant and the arithmetic is blunt: 25K output tokens per task at $4.68 per 1M is about $0.117 a task, or $117 per thousand tasks. At 38K tokens the same work costs about $178, and at 43K about $201. Inkling's independently measured verbosity advantage is worth roughly 40% of the output bill against the most verbose comparison point in AA's set — before you touch the effort dial at all.

Now add the dial. Because effort directly scales generated tokens, moving from a high setting to a medium one on tasks that do not need deep reasoning cuts the output bill proportionally. The buying question stops being "is Inkling ranked above model X" and becomes "what is the lowest effort setting that still passes my eval?" That is a far more tractable question, and it is the one nobody in the launch coverage asked.

The effort dial: what it costs vs what it buys Illustrative shape only — no per-effort token-and-score table was published Tokens billed Task accuracy none 0 low 0.2 medium 0.7 high 0.9 xhigh / max 0.99 Vendor-run benchmark scores for Inkling were reported at the far right of this axis: effort 0.99.
Level-to-value mapping as documented in the Hugging Face transformers integration; the full level list also includes minimal, whose numeric value is not published. Curve shapes are illustrative: Thinking Machines published effort sweeps on Terminal Bench 2.1, HLE and IFBench but no per-effort token-and-score table, so the exact curvature for your workload is unknown until you measure it. Chart: DataLLM Lab

A workable procedure: build a 50-task eval from your real traffic, run it at low, medium and high, and record pass rate and mean output tokens at each. You will usually find one or two task classes that need high effort and a long tail that does not. Route accordingly. We walk through the same measurement method for coding workloads in our LLM coding cost benchmark.

Test the effort dial without a new vendor contract

DataLLM Lab is an OpenAI-compatible gateway with 300+ models behind one key, so you can run the same eval across open-weights and closed frontier models by changing a string. Point your existing SDK at https://www.datallmlab.com/v1 and start comparing.

Open weights you probably cannot self-host

Apache 2.0 on a 975B-parameter model sounds like sovereignty. Check the hardware bill before you plan around it. Thinking Machines lists a minimum of roughly 2 TB of aggregated VRAM for the full BF16 checkpoint — 8x NVIDIA B300 or 16x H200 — and roughly 600 GB for the NVFP4-quantised checkpoint, or 4x B300 / 8x H200. Computerworld and MarkTechPost both repeat those configurations. Community GGUF quantisations exist for llama.cpp, but the vendor's own floor is the number to plan against.

For the overwhelming majority of teams, "open weights" here means: you can legally fine-tune it, redistribute your derivative, and audit the architecture — but you will rent the compute. Distribution reflects that. Weights sit on Hugging Face; hosted APIs come from Together AI, Fireworks, Modal, Databricks and Baseten among others; runtime support exists in SGLang, vLLM, llama.cpp and Hugging Face Transformers; and NVIDIA Build lists it, with BF16 requiring Hopper or later and NVFP4 targeting Blackwell. Inkling was itself trained on NVIDIA GB300 NVL72 systems.

This is also the quiet business logic. Tinker, the fine-tuning platform, is where TechCrunch says the revenue is meant to come from — training, fine-tuning, and a cut of the hosting ecosystem around the model. Giving away weights that almost nobody can serve at scale, while selling the fine-tuning and hosting layer, is a coherent strategy — and it is worth naming, because VentureBeat notes the company has not explained how it will cover its training costs, and TechCrunch reports a roughly $50B round said to be in progress as of November 2025 that had stalled by January, with the company declining to discuss funding since. Those are reported figures about a private company, not confirmed facts.

One more framing worth attributing rather than endorsing: VentureBeat built its coverage around low cost and resistance to censorship, reflecting Thinking Machines' emphasis on epistemics — calibration, instruction following, and answering directly on topics that may be subject to censorship. TM says the model showed strong patterns of censorship non-compliance when evaluated by Cognition on its Propaganda and Censorship Eval; that is a third-party eval, but it reaches you through the vendor. On safety, the model card's conclusion that Inkling did not present risk of material uplift beyond what is already available in the open-weight ecosystem is a vendor summary of its own testing, and TM acknowledges residual risks including occasional compliance with role-play and indirectly framed prompts on harmful topics. The published FORTRESS figures — 78.0% adversarial, 95.9% benign — are likewise from the model card. If licensing and distribution policy is your angle, see export controls and open weights.

Who should actually use Inkling

Take the vendor at its word. Inkling is not the strongest model available, and it is not trying to be. It is a customization base with three properties that are genuinely hard to find together: permissive licensing, native multimodal input across text, images and audio, and independently confirmed token efficiency.

A decision rule, in order:

  1. You need the best raw capability and cost is secondary. Inkling is the wrong answer. Thinking Machines' own table puts several closed models well ahead on SWE-bench Verified, and three open-weights models — DeepSeek V4 Pro at 80.6%, Kimi K2.6 at 80.2% and GLM 5.2 at 80.0% — ahead of Inkling's 77.6%.
  2. You need a fine-tunable base with a clean license and multimodal input. Strong fit. Apache 2.0 plus audio and image understanding plus day-one Tinker support is the actual pitch, and it is a good one.
  3. Your workload is high-volume and verbosity-sensitive. Strong fit. The ~25K-token average is independently measured by Artificial Analysis, not a marketing claim, and the effort dial gives you a lever most models do not expose.
  4. Your workload is knowledge-heavy and hallucination-sensitive. Be careful. AA's Omniscience result places Inkling at +2, ahead of Nemotron 3 Ultra at -1 but behind the leading open-weights field, with 40% accuracy and a 63% hallucination rate on that benchmark's construct. That metric measures behaviour on deliberately hard knowledge probes — it is not a claim that the model is wrong 63% of the time in general — but it is the honest counterweight to the launch narrative.
  5. You want a drop-in for a Chinese open-weights model on leaderboard grounds alone. Re-read the U.S. qualifier. AA's framing was the new leading U.S. open weights model, and that scoping is doing real work.

Independent hands-on colour, for calibration: Simon Willison ran Inkling through the API to generate an SVG of a pelican on a bicycle, then fed the rendering back for description. The model described colours and composition accurately but called its own drawing a stylized white bird resembling a stork or seagull. He frames the release as a new viable contender in the U.S. open-weights ecosystem alongside NVIDIA Nemotron and Gemma 4 — which is, stripped of launch-day framing, exactly what it is.

As for Inkling-Small: 276B total / 12B active, announced as a preview, weights not released at launch while TM finishes testing. The claim that it matches or exceeds the larger model on many benchmarks is vendor-only and unverified. If it ships and holds up, it is the more interesting model for most readers, because 12B active is a serving proposition rather than a datacentre one. Until then it is an announcement. For where it would land against the current field, see our Kimi K3 review and the 2026 open-weights roundup.

FAQ

Is Inkling the best open-weights model available?

No. Artificial Analysis called it the new leading U.S. open weights model, measuring it at 41 on the AA Intelligence Index. That U.S. qualifier matters: in Thinking Machines' own comparison table, DeepSeek V4 Pro scores 80.6% on SWE-bench Verified, Kimi K2.6 80.2% and GLM 5.2 80.0%, against Inkling's 77.6%. Thinking Machines itself states that Inkling is not the strongest overall model available today, open or closed.

What license is Inkling released under?

Apache 2.0, confirmed on the Thinking Machines model card and the Hugging Face repository. That permits commercial use, modification, fine-tuning and redistribution, which is a genuinely permissive grant compared with the custom community licenses used by several other open-weights releases.

Did Artificial Analysis measure the 77.6% SWE-bench Verified score?

No, and this is the most common error in the launch coverage. The 77.6% figure comes from the Thinking Machines model card, run at effort 0.99. The Artificial Analysis launch article does not report a SWE-bench Verified figure for Inkling at all. What AA independently contributed was the Intelligence Index score of 41, token-efficiency measurements, GDPval-AA Elo, τ³ Banking, AA-Omniscience and a serving price table. Confusingly, Thinking Machines also states that it relies on externally reported Artificial Analysis evaluations for several rows of its own table, including GPQA Diamond, HLE, GDPval, τ³ Banking, AA-Omniscience and MMMU Pro.

What hardware do I need to self-host Inkling?

Thinking Machines lists a minimum of roughly 2 TB of aggregated VRAM for the full BF16 checkpoint, as 8x NVIDIA B300 or 16x H200. The NVFP4-quantised checkpoint drops the floor to roughly 600 GB, as 4x B300 or 8x H200. Those configurations are repeated in Computerworld and MarkTechPost coverage of the release. Either way, self-hosting is a datacentre decision, not a workstation one.

What does the thinking effort dial actually change?

It changes how many reasoning tokens the model generates before answering. Thinking Machines trained effort in during reinforcement learning by varying the system message and the per-token cost, and the Hugging Face transformers integration exposes it as named levels: none, minimal, low, medium, high, xhigh and max. The published numeric mappings cover none at 0, low at 0.2, medium at 0.7, high at 0.9, and xhigh and max at 0.99. Because output tokens are billed, the effort dial is functionally a cost dial.

Can I download Inkling-Small yet?

Not as of publication. Inkling-Small, at 276B total and 12B active parameters, was announced as a preview alongside the main release, with Thinking Machines saying it is finishing testing and will release the full weights once that work is complete. The claim that it matches or exceeds the larger model on many benchmarks is vendor-only and has not been independently verified. Check the Hugging Face organisation page before planning around it.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.