Model Reviews

Ling 3.1 Flash: Nine of Nine Where Ling 3.0 Flash Could Not Finish (and Why Its $0 Is Not a Rank)

Ling 3.1 Flash scored 9 out of 9 on our executed Python benchmark on 2026-10-03, with a 6.4-second mean and 464 reasoning tokens per call. That makes it the first text-only Ling Flash model we have run that completes the suite: Ling 3.0 Flash failed to finish in two runs and is excluded from our data. It does not get the 59% token diet that made Ling 3.0 Flash VL work, though — Ling 3.1 Flash still thinks almost twice as hard as the VL variant. And there is one number we deliberately do not publish: a cost rank. Its only OpenRouter listing is priced at $0 in and $0 out, which describes a launch trial, not a model.

DataLLM Lab article cover: Ling 3.1 Flash: Nine of Nine Where Ling 3.0 Flash Could Not Finish (and Why Its $0 Is Not a Rank)

The interesting thing about a point release is rarely its headline number. It is which old failure it fixes and which old habit it keeps. Ling 3.1 Flash fixes one and keeps the other.

The result

MetricLing 3.1 Flash
Score9/9
Mean latency6.4s
Reasoning tokens per call464
Tokens across the suite741 in / 5,351 out
Price used for cost$0 in / $0 out per 1M (captured 2026-10-03)
Derived cost / 1,000 tasks$0 — not ranked, see below
Context window262,144
Measured2026-10-03

All nine generated functions passed their hidden asserts on the first scored attempt, with no API-layer failures. That is the same score as the 75 priced models on our board that clear the suite; as a free listing it sits outside that ranked set. It is not near the top on speed either: Solar Mini 4 is still the cheapest and fastest model to score 9 out of 9 as of 2026-10-03, at $0.03 per 1,000 tasks and a 2s mean, and it answers with 0 reasoning tokens. Ling 3.1 Flash takes 6.4 seconds and spends 464 reasoning tokens per call to reach the same nine passes.

Three Ling Flash runs, side by side

We have now run three InclusionAI Flash models on the identical nine prompts. The text-only Ling 3.0 Flash is excluded: run 1 had one API-layer failure, and run 2 had one API-layer failure and one wrong answer. Its API-layer failures were empty bodies at our 4,000-token max_tokens cap — it spent the budget reasoning and never emitted the function. Re-tested at a 16k cap it does answer, using 9,902 tokens; that story is in our piece on models that cannot finish inside a token budget. The vision-language Ling 3.0 Flash VL went the other way and scored 9 out of 9, as we covered in the Ling 3.0 Flash VL review.

Ling 3.1 FlashLing 3.0 Flash VLLing 3.0 Flash
Status9/99/9excluded — no comparable score
Reasoning tokens per call464234565
Input tokens across the suite741741645
Output tokens across the suite5,3513,2995,599
Mean latency6.4s4.4s2.7s (incomplete run)
Measured2026-10-032026-09-162026-08-22
Ling 3.1 Flash finishes, but it still thinks like 3.0 FlashSame nine executed Python tasks, temperature 0, 4,000-token cap per call.REASONING TOKENS PER CALLLing 3.1 Flash464 · 9/9Ling 3.0 Flash VL234 · 9/9Ling 3.0 Flash565 · excludedOUTPUT TOKENS ACROSS THE NINE TASKSLing 3.1 Flash5,351Ling 3.0 Flash VL3,299Ling 3.0 Flash5,599 · excludedScales: top 0.8 px per reasoning token; bottom 0.08 px per output token (80 px per 1,000).Every width = value × scale. Dashed outline = excluded run, shown for tokens only, no score.
Ling 3.1 Flash sits much closer to the excluded 3.0 Flash on token spend than to the disciplined VL variant.

Read the chart carefully, because the obvious story is wrong. The obvious story is that 3.1 Flash completes the suite because it learned to think less. It did, but only a little: 101 fewer reasoning tokens per call than Ling 3.0 Flash (565 − 464), and 248 fewer output tokens across the suite (5,599 − 5,351). Against the VL variant it spends 230 more reasoning tokens per call (464 − 234), nearly twice as many, and 2,052 more output tokens across the suite (5,351 − 3,299), about 62% more.

So the fix is not in the average. It is in the tail. Ling 3.0 Flash failed because individual calls ran into the 4,000-token cap with nothing written; every one of Ling 3.1 Flash's nine calls came back with a function that passed. Our data sheet records suite totals and per-call reasoning, not the token count of each task, so we cannot say how close its hardest call came to the cap. What we can say is that at a 4,000-token budget, the text model in this family now finishes — and that it pays roughly as many tokens to do so as the model that did not.

Input is the other detail worth noting: 741 tokens across the suite, identical to the VL variant and higher than 3.0 Flash's 645, on identical prompts. That is consistent with 3.1 Flash sharing the VL variant's prompt formatting, but it is an observation, not something InclusionAI has told us.

Why a 9/9 at $0 is not ranked on cost

Our cost column is derived: measured tokens multiplied by the list price captured on the run date. For Ling 3.1 Flash that price was $0 in and $0 out per 1M on 2026-10-03, so the arithmetic returns $0 per 1,000 tasks. Taken at face value, that would put it at the top of every cost table on this site, above Solar Mini 4's $0.03.

We mark it free and leave it out of cost ranks instead, for three reasons:

One small trap for people chasing free models. Our probe of OpenRouter's openrouter/free router scored 8 out of 9, at a 7.722s mean, and that router sends requests only to :free model ids. Ling 3.1 Flash's id carries no :free suffix even though it is priced at zero, and the router did not serve it on any of the nine tasks. If you want it, call it by name. For the wider picture of what is free and on what terms, see our free LLM API guide.

What InclusionAI says, and who says it

Everything in this section is third-party, not measured by us, and was read on 2026-10-03.

ClaimTierSource
About 560B total parameters, about 25B active per token, up to 1M-token contextConfirmedAnt Ling's own X account. We read the post text as indexed by search; the post itself would not load for us.
Self-reported scores: 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, 65.35 on HealthBench ProfessionalConfirmed as vendor claimsSame Ant Ling post. Confirmed that Ant says it, not that the numbers hold — we cannot re-run them.
Plans to open-source the modelConfirmed as a planSame post, which says soon and names no date. Our own check of InclusionAI's Hugging Face organisation on 2026-10-03 found Ling 3.0 repositories and no Ling 3.1 repository.
Listed 2026-10-02; 262,144 context; 32,768 max output; $0 in / $0 outCatalogueOpenRouter's model list, read directly. The 262,144 matches our run record.
Launched 2026-09-30 as a two-week free trial capped at 256K context; larger window and open weights to follow the trialAttributedTechNode, citing IT Home. Not stated in the vendor post we could see.
Aimed at agent tasks, search, office software and specialist applicationsAttributedTechNode
The trial's end date and the price after itNot publishedNo vendor or named outlet we found gives either.

Notice the gap between the 1M context in the announcement and the 262,144 that is actually served. TechNode's report explains it as a trial limit. Until the larger window appears on an endpoint you can call, plan around 262,144.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why Ling 3.0 Flash is excluded rather than given a partial score. Cost is derived from measured token counts at the list price captured on the run date — $0 on 2026-10-03 for Ling 3.1 Flash, $0.06 in and $0.18 out per 1M on 2026-09-15 for Ling 3.0 Flash VL, which has since repriced. Router figures quoted above are a different method: the API-reported usage.cost, summed. Runs go through OpenRouter, not through the DataLLM Lab gateway, so anyone can reproduce them without being our customer. Full method on the methodology page.

What we did not measure

If you are choosing a cheap model to put into production today, our cheapest-9/9 comparison ranks only models with a price you can still expect to pay next month. Ling 3.1 Flash will join that ranking when it has one.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.