New LLM Models, September 2026: What Fourteen Releases Actually Cost
September 2026 put fourteen new or newly-callable models in front of us, and we ran every one through the same nine executed Python tasks rather than reading the launch posts. Eleven scored 9 out of 9. The cheapest of those charged $0.07 per 1,000 tasks; the most expensive charged $57.06. That is an 815x spread for an identical score, which is the single most useful thing a benchmark this narrow can tell you: on work this size the money has almost nothing to do with whether the answer is right. The three results that were not 9 out of 9 turned out to be more interesting than most of the ones that were.
This is the index for our September sweep. Every model below was called through the same harness on 2026-09-16, and every figure carries the list price it was derived from. Each row links to the full write-up.
All fourteen, measured
| Model | Score | Measured cost / 1k tasks | Latency | Reasoning tokens |
|---|---|---|---|---|
| Ling 3.0 Flash VL | 9/9 | $0.07 | 4.4s | 234 |
| Mercury 2.5 | 9/9 | $0.17 | 2.4s | 989 |
| DeepSeek V4.1-Flash | 9/9 | $0.36 | 15.6s | 478 |
| Sakana Fugu Max | 9/9 | $2.19 | 12s | 424 |
| Gemini 3.8 Flash | 9/9 | $3.87 | 10.2s | 901 |
| Qwen3.8-Max-0902 | 9/9 | $4.21 | 25.2s | 546 |
| Claude Fable 5.1 | 9/9 | $8.09 | 6.8s | 0 |
| GPT-6 Astra | 9/9 | $8.19 | 5.6s | 47 |
| GPT-6 Astra Pro | 9/9 | $35.44 | 8.4s | 111 |
| Sakana Fugu Ultra V2 | 9/9 | $57.06 | 23.7s | 7 |
| NEX N2.5 Pro | 9/9 | free | 66.8s | 523 |
| and the three that did not clear the suite | ||||
| Schematron V2 Turbo | 0/9 | $0.03 | 5.3s | 0 |
| IBM Granite 4.2 8B | excluded | — | 12.8s | 1,483 |
| Meta Muse Spark 1.3 | excluded | — | — | — |
The flagships arrived priced identically
OpenAI and Anthropic shipped two weeks apart into the same bracket — $10 per million input tokens and $50 output, to the cent. GPT-6 Astra measured $8.19; Claude Fable 5.1 measured $8.09. Both 9 out of 9. That gap is 1.2%, which on a single scored attempt per task is not a gap at all, and we say so rather than picking a winner from noise — the head-to-head works through why.
Two footnotes on those two. Astra settles a naming question that was genuinely open three weeks ago: the model ships as gpt-6-astra, so the class-versus-number question resolved to GPT-6. And Fable 5.1 completes a suite its predecessor never finished — four of nine tasks came back as empty bodies from a content filter, which is the failure that forced us to separate an API-layer failure from a wrong answer in the first place.
The floor moved down to seven cents
Ling 3.0 Flash VL took the cheapest position at $0.07, one cent ahead of DeepSeek V3.2, which had held it since July. One cent on a single run is a cluster, not a victory — but the cluster itself is the point, and the full leaderboard has the field. The twist is that Ling's text-only sibling cannot finish the suite at all; the vision variant is the disciplined one.
The three that did not score 9 of 9
- Schematron V2 Turbo scored 0 of 9 — our first genuine zero. It answered every task and got every one wrong. It is a structured-extraction specialist, so this is a scope report rather than a defect report, and it is the clearest illustration we have of what domain specialisation actually costs you outside the domain.
- IBM Granite 4.2 8B is excluded — seven of nine tasks scored, two failed at the API layer. A score with a different denominator is not comparable to a 9/9, so we publish none.
- Meta Muse Spark 1.3 is excluded — not for any model reason. The endpoint requires an age confirmation on the calling account that ours does not carry, so it never ran.
Three things a price page will not tell you
The sweep surfaced three cost mechanisms that no rate card exposes, and each one got its own page because each one is worth more than a model review.
- You are billed for prompt you did not write. GPT-6 Astra Pro consumed 17,011 input tokens across nine short prompts against plain Astra's 604, on an identical rate card. A single-request probe billed 1,721 prompt tokens for a 65-token message.
- Batch endpoints are not reliably cheaper. Of 85 base models publishing a
:batchprice, 70 are exactly half. Nine cost more than the standard endpoint they queue behind. - Prices went up, quietly. Google doubled the Gemini Flash line, which makes a figure we published three weeks ago wrong by 2x through no change in the model. This is why every number on this site carries the date it was priced on — see price volatility.
A fourth, adjacent: we resolved all 19 of the catalogue's -latest aliases and found three carrying a listed price their target does not honour.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer. Cost is derived — measured token counts multiplied by the list price captured 2026-09-15, not a billing statement. A model served free is recorded but barred from cost rankings. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.
What this sweep cannot tell you
- Whether a flagship is worth its price. Nine self-contained Python functions cannot separate a frontier model from a competent small one. Eleven models scored identically here; that is a floor test, not a capability ranking, and reading it as a ranking is the main way this data gets misused.
- Anything about long-horizon agentic work, which is what the expensive tiers are sold on and what our single-turn suite is structurally blind to.
- Anything about context windows, several of which exceed a million tokens. Our prompts are short.
- Anything about vision, audio or tool use. Text-only, Python-only.
- Stability. One scored attempt per task, no averaging across runs. A one-cent lead is inside the noise and we treat it that way.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab