Ling 3.1 Flash: Nine of Nine Where Ling 3.0 Flash Could Not Finish (and Why Its $0 Is Not a Rank)
Ling 3.1 Flash scored 9 out of 9 on our executed Python benchmark on 2026-10-03, with a 6.4-second mean and 464 reasoning tokens per call. That makes it the first text-only Ling Flash model we have run that completes the suite: Ling 3.0 Flash failed to finish in two runs and is excluded from our data. It does not get the 59% token diet that made Ling 3.0 Flash VL work, though — Ling 3.1 Flash still thinks almost twice as hard as the VL variant. And there is one number we deliberately do not publish: a cost rank. Its only OpenRouter listing is priced at $0 in and $0 out, which describes a launch trial, not a model.
The interesting thing about a point release is rarely its headline number. It is which old failure it fixes and which old habit it keeps. Ling 3.1 Flash fixes one and keeps the other.
The result
| Metric | Ling 3.1 Flash |
|---|---|
| Score | 9/9 |
| Mean latency | 6.4s |
| Reasoning tokens per call | 464 |
| Tokens across the suite | 741 in / 5,351 out |
| Price used for cost | $0 in / $0 out per 1M (captured 2026-10-03) |
| Derived cost / 1,000 tasks | $0 — not ranked, see below |
| Context window | 262,144 |
| Measured | 2026-10-03 |
All nine generated functions passed their hidden asserts on the first scored attempt, with no API-layer failures. That is the same score as the 75 priced models on our board that clear the suite; as a free listing it sits outside that ranked set. It is not near the top on speed either: Solar Mini 4 is still the cheapest and fastest model to score 9 out of 9 as of 2026-10-03, at $0.03 per 1,000 tasks and a 2s mean, and it answers with 0 reasoning tokens. Ling 3.1 Flash takes 6.4 seconds and spends 464 reasoning tokens per call to reach the same nine passes.
Three Ling Flash runs, side by side
We have now run three InclusionAI Flash models on the identical nine prompts. The text-only Ling 3.0 Flash is excluded: run 1 had one API-layer failure, and run 2 had one API-layer failure and one wrong answer. Its API-layer failures were empty bodies at our 4,000-token max_tokens cap — it spent the budget reasoning and never emitted the function. Re-tested at a 16k cap it does answer, using 9,902 tokens; that story is in our piece on models that cannot finish inside a token budget. The vision-language Ling 3.0 Flash VL went the other way and scored 9 out of 9, as we covered in the Ling 3.0 Flash VL review.
| Ling 3.1 Flash | Ling 3.0 Flash VL | Ling 3.0 Flash | |
|---|---|---|---|
| Status | 9/9 | 9/9 | excluded — no comparable score |
| Reasoning tokens per call | 464 | 234 | 565 |
| Input tokens across the suite | 741 | 741 | 645 |
| Output tokens across the suite | 5,351 | 3,299 | 5,599 |
| Mean latency | 6.4s | 4.4s | 2.7s (incomplete run) |
| Measured | 2026-10-03 | 2026-09-16 | 2026-08-22 |
Read the chart carefully, because the obvious story is wrong. The obvious story is that 3.1 Flash completes the suite because it learned to think less. It did, but only a little: 101 fewer reasoning tokens per call than Ling 3.0 Flash (565 − 464), and 248 fewer output tokens across the suite (5,599 − 5,351). Against the VL variant it spends 230 more reasoning tokens per call (464 − 234), nearly twice as many, and 2,052 more output tokens across the suite (5,351 − 3,299), about 62% more.
So the fix is not in the average. It is in the tail. Ling 3.0 Flash failed because individual calls ran into the 4,000-token cap with nothing written; every one of Ling 3.1 Flash's nine calls came back with a function that passed. Our data sheet records suite totals and per-call reasoning, not the token count of each task, so we cannot say how close its hardest call came to the cap. What we can say is that at a 4,000-token budget, the text model in this family now finishes — and that it pays roughly as many tokens to do so as the model that did not.
Input is the other detail worth noting: 741 tokens across the suite, identical to the VL variant and higher than 3.0 Flash's 645, on identical prompts. That is consistent with 3.1 Flash sharing the VL variant's prompt formatting, but it is an observation, not something InclusionAI has told us.
Why a 9/9 at $0 is not ranked on cost
Our cost column is derived: measured tokens multiplied by the list price captured on the run date. For Ling 3.1 Flash that price was $0 in and $0 out per 1M on 2026-10-03, so the arithmetic returns $0 per 1,000 tasks. Taken at face value, that would put it at the top of every cost table on this site, above Solar Mini 4's $0.03.
We mark it free and leave it out of cost ranks instead, for three reasons:
- It is the only price there is. In OpenRouter's catalogue (read 2026-10-03),
inclusionai/ling-3.1-flashlists at $0 in and $0 out with no paid variant alongside it. A free tier next to a paid one tells you the model's market price; a free tier on its own tells you only the launch terms. - The launch terms are temporary. TechNode reported on 2026-09-30, citing IT Home (read 2026-10-03), that the release runs as a two-week free trial. Nobody we found has published what it will cost afterwards. A rank that holds for a fortnight and then silently becomes wrong is worse than no rank — see how often prices move even without a trial ending.
- Zero hides the token bill. At any shared per-token price, Ling 3.1 Flash would cost more than the VL variant on this suite, because it produces about 62% more output tokens. A $0 sticker erases exactly the behaviour that made its predecessor fail.
One small trap for people chasing free models. Our probe of OpenRouter's openrouter/free router scored 8 out of 9, at a 7.722s mean, and that router sends requests only to :free model ids. Ling 3.1 Flash's id carries no :free suffix even though it is priced at zero, and the router did not serve it on any of the nine tasks. If you want it, call it by name. For the wider picture of what is free and on what terms, see our free LLM API guide.
What InclusionAI says, and who says it
Everything in this section is third-party, not measured by us, and was read on 2026-10-03.
| Claim | Tier | Source |
|---|---|---|
| About 560B total parameters, about 25B active per token, up to 1M-token context | Confirmed | Ant Ling's own X account. We read the post text as indexed by search; the post itself would not load for us. |
| Self-reported scores: 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, 65.35 on HealthBench Professional | Confirmed as vendor claims | Same Ant Ling post. Confirmed that Ant says it, not that the numbers hold — we cannot re-run them. |
| Plans to open-source the model | Confirmed as a plan | Same post, which says soon and names no date. Our own check of InclusionAI's Hugging Face organisation on 2026-10-03 found Ling 3.0 repositories and no Ling 3.1 repository. |
| Listed 2026-10-02; 262,144 context; 32,768 max output; $0 in / $0 out | Catalogue | OpenRouter's model list, read directly. The 262,144 matches our run record. |
| Launched 2026-09-30 as a two-week free trial capped at 256K context; larger window and open weights to follow the trial | Attributed | TechNode, citing IT Home. Not stated in the vendor post we could see. |
| Aimed at agent tasks, search, office software and specialist applications | Attributed | TechNode |
| The trial's end date and the price after it | Not published | No vendor or named outlet we found gives either. |
Notice the gap between the 1M context in the announcement and the 262,144 that is actually served. TechNode's report explains it as a trial limit. Until the larger window appears on an endpoint you can call, plan around 262,144.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why Ling 3.0 Flash is excluded rather than given a partial score. Cost is derived from measured token counts at the list price captured on the run date — $0 on 2026-10-03 for Ling 3.1 Flash, $0.06 in and $0.18 out per 1M on 2026-09-15 for Ling 3.0 Flash VL, which has since repriced. Router figures quoted above are a different method: the API-reported usage.cost, summed. Runs go through OpenRouter, not through the DataLLM Lab gateway, so anyone can reproduce them without being our customer. Full method on the methodology page.
What we did not measure
- Whether 560B parameters buy anything here. Nine self-contained Python functions cannot separate a frontier model from a competent small one: 75 priced models on our board score 9 out of 9, and the cheapest and fastest of them, Solar Mini 4, does it with no reasoning tokens at all. A pass from Ling 3.1 Flash says it clears the bar, not how far above it it sits.
- The vendor's benchmarks. FrontierSWE, GDPVal-AA and HealthBench Professional are Ant's numbers. Our suite tests none of that.
- Long context or the promised 1M window. Our prompts are short, and the served window is 262,144.
- The post-trial model. If pricing, serving or the checkpoint changes when the trial ends, this run describes the trial endpoint only.
- Repeat runs. One scored attempt per task. Ling 3.0 Flash failed differently on its two runs, so a single clean 9/9 from its successor is encouraging, not conclusive.
If you are choosing a cheap model to put into production today, our cheapest-9/9 comparison ranks only models with a price you can still expect to pay next month. Ling 3.1 Flash will join that ranking when it has one.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab