Ling 3.0 Flash VL: Seven Cents, Nine of Nine, and a Twist
Ling 3.0 Flash VL scored 9 out of 9 on our executed Python benchmark at $0.07 per 1,000 tasks and a 4.4-second mean. When we measured it on 2026-09-16 that made it the cheapest of the 56 models in our set that cleared the suite — a position Solar Mini 4 took on 2026-10-02 at $0.03, edging past DeepSeek V3.2 at $0.08, which had held that position since July. The twist is which variant did it. We ran the text-only Ling 3.0 Flash twice and it failed to finish both times — it exhausts its token budget on reasoning and returns empty bodies. The vision-language variant of the same family completes the suite, uses 59% fewer reasoning tokens, and does it for seven cents.
A model with vision bolted on usually costs more than its text-only sibling and does the text part slightly worse. This one inverts both halves of that expectation.
The result
| Metric | Ling 3.0 Flash VL |
|---|---|
| Score | 9/9 |
| Measured cost / 1,000 tasks | $0.07 |
| Mean latency | 4.4s |
| Reasoning tokens per call | 234 |
| Tokens across the suite | 741 in / 3,299 out |
| List price in / out | $0.06 / $0.18 per 1M |
| Context window | 131,072 |
| Measured | 2026-09-16 |
Of the 56 models that had scored 9 out of 9 as of 2026-09-16 it ranked 1st on cost and 10th on latency. That combination — cheapest and top-ten fastest — is not one we have seen before; the cheap end of the field is usually cheap because it is slow, or fast because it barely thinks.
The text-only variant cannot finish
We ran Ling 3.0 Flash, the text-only sibling, twice in August. Both runs are marked excluded in our data, because neither completed:
| Ling 3.0 Flash (text) | Ling 3.0 Flash VL | |
|---|---|---|
| Run 1 | 1 API-layer failure — 8 of 9 scored | 9 of 9, first attempt |
| Run 2 | 1 API-layer failure and 1 wrong answer — 7 of 8 | |
| Reasoning tokens per call | 565 | 234 |
| Output tokens across the suite | 5,599 | 3,299 |
| Status in our data | excluded | 9/9 |
The text variant's failures were finish_reason=length with an empty body: it spends its entire 4,000-token allowance thinking and never emits an answer. We gave it a fairness re-test at a 16,000-token cap and it did produce a working function — using 9,902 tokens. That write-up is in the piece on models that cannot finish inside a token budget.
The VL variant does not do this. It uses 59% fewer reasoning tokens and 41% fewer output tokens than its text-only sibling, on identical prompts, and never came close to the cap. We cannot tell you why from outside — two variants of one family, and the one carrying extra modality is the disciplined one. What we can tell you is that if you tried Ling 3.0 Flash, hit empty responses, and wrote the family off, you tried the wrong variant.
The new floor
Be honest about the size of that lead: one cent per thousand tasks, on a single scored run per task. We do not average across runs, so a re-run could reorder the top two. The useful reading is not “Ling wins” but that the floor for a model that solves every task now sits at seven to ten cents, and that several unrelated vendors are clustered there.
Against the top of the market that gap is the whole story. GPT-6 Astra scored the same 9 out of 9 at $8.19 — 117 times more — and Astra Pro at $35.44. Our suite cannot tell you what a flagship buys on harder work, but it can tell you precisely what it does not buy on work this size.
What seven cents does not buy
- Context. 131,072 tokens, against a million or more for most of the frontier. The text-only Ling 3.0 Flash carries 262,144 — twice the VL variant. If you need long context, this is the wrong end of the market.
- Any evidence on vision. It is a vision-language model and our suite is text-only. We measured the thing it is not primarily built for.
- A margin you can rely on. One cent ahead of DeepSeek V3.2 is inside single-run noise.
- Long-horizon behaviour. Nine self-contained functions say nothing about multi-turn sessions — and its sibling's token blowouts are a reason to watch that carefully.
For the wider cheap end, our cheap coding roundup and the open-source guide have the field.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why the text-only variant is excluded rather than scored. Cost is derived from measured token counts at the list price captured 2026-09-15, not a billing statement, and prices move. Runs go through OpenRouter. Full method on the methodology page.
What we did not measure
- Vision, which is the model's reason for existing.
- The 131,072-token context. Our prompts are short.
- Why the VL variant is disciplined and the text variant is not. We observed the difference; we cannot explain it.
- The free variant. A
:freetier exists and is a different endpoint from the one we priced. - Repeat runs. One scored attempt per task, which is exactly why we call a one-cent lead a cluster rather than a win.
FAQ
Is Ling 3.0 Flash VL the cheapest LLM that passes your benchmark? It was as of 2026-09-16, at $0.07 per 1,000 tasks. Solar Mini 4 overtook it on 2026-10-02 at $0.03.
How much does Ling 3.0 Flash VL cost? $0.06 per million input tokens and $0.18 output as of 2026-09-15.
Is the text-only Ling 3.0 Flash cheaper? On paper yes, but it did not complete our suite in either of two runs, so we publish no score for it.
Is it fast? 4.4 seconds mean, 10th fastest of the 56 models that scored 9 out of 9.
What is the catch? A 131,072-token context, half what its text-only sibling offers and a fraction of the frontier.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab