Benchmarks

New LLM Models, September 2026: What Fourteen Releases Actually Cost

September 2026 put fourteen new or newly-callable models in front of us, and we ran every one through the same nine executed Python tasks rather than reading the launch posts. Eleven scored 9 out of 9. The cheapest of those charged $0.07 per 1,000 tasks; the most expensive charged $57.06. That is an 815x spread for an identical score, which is the single most useful thing a benchmark this narrow can tell you: on work this size the money has almost nothing to do with whether the answer is right. The three results that were not 9 out of 9 turned out to be more interesting than most of the ones that were.

DataLLM Lab article cover: New LLM Models, September 2026: What Fourteen Releases Actually Cost

This is the index for our September sweep. Every model below was called through the same harness on 2026-09-16, and every figure carries the list price it was derived from. Each row links to the full write-up.

All fourteen, measured

ModelScoreMeasured cost / 1k tasksLatencyReasoning tokens
Ling 3.0 Flash VL9/9$0.074.4s234
Mercury 2.59/9$0.172.4s989
DeepSeek V4.1-Flash9/9$0.3615.6s478
Sakana Fugu Max9/9$2.1912s424
Gemini 3.8 Flash9/9$3.8710.2s901
Qwen3.8-Max-09029/9$4.2125.2s546
Claude Fable 5.19/9$8.096.8s0
GPT-6 Astra9/9$8.195.6s47
GPT-6 Astra Pro9/9$35.448.4s111
Sakana Fugu Ultra V29/9$57.0623.7s7
NEX N2.5 Pro9/9free66.8s523
and the three that did not clear the suite
Schematron V2 Turbo0/9$0.035.3s0
IBM Granite 4.2 8Bexcluded—12.8s1,483
Meta Muse Spark 1.3excluded———

The flagships arrived priced identically

OpenAI and Anthropic shipped two weeks apart into the same bracket — $10 per million input tokens and $50 output, to the cent. GPT-6 Astra measured $8.19; Claude Fable 5.1 measured $8.09. Both 9 out of 9. That gap is 1.2%, which on a single scored attempt per task is not a gap at all, and we say so rather than picking a winner from noise — the head-to-head works through why.

Two footnotes on those two. Astra settles a naming question that was genuinely open three weeks ago: the model ships as gpt-6-astra, so the class-versus-number question resolved to GPT-6. And Fable 5.1 completes a suite its predecessor never finished — four of nine tasks came back as empty bodies from a content filter, which is the failure that forced us to separate an API-layer failure from a wrong answer in the first place.

The floor moved down to seven cents

Eight models, one identical score, an 815x spreadMeasured cost per 1,000 tasks. Log scale — a linear axis would collapse the left half into the margin.Ling 3.0 Flash VL$0.07Mercury 2.5$0.17DeepSeek V4.1-Flash$0.36Sakana Fugu Max$2.19Gemini 3.8 Flash$3.87Claude Fable 5.1$8.09Sakana Fugu Ultra V2$57.06Log scale: 145 px per decade, origin $0.05. Compare orders of magnitude, not bar lengths.
Every bar on this chart solved all nine tasks.

Ling 3.0 Flash VL took the cheapest position at $0.07, one cent ahead of DeepSeek V3.2, which had held it since July. One cent on a single run is a cluster, not a victory — but the cluster itself is the point, and the full leaderboard has the field. The twist is that Ling's text-only sibling cannot finish the suite at all; the vision variant is the disciplined one.

The three that did not score 9 of 9

Three things a price page will not tell you

The sweep surfaced three cost mechanisms that no rate card exposes, and each one got its own page because each one is worth more than a model review.

A fourth, adjacent: we resolved all 19 of the catalogue's -latest aliases and found three carrying a listed price their target does not honour.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer. Cost is derived — measured token counts multiplied by the list price captured 2026-09-15, not a billing statement. A model served free is recorded but barred from cost rankings. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

What this sweep cannot tell you

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.