New LLM Models, October 2026: Twenty-Seven Releases on One Harness
Between 2026-09-16 and 2026-10-02 the catalogue gained enough new text models that we ran 27 of them through the same nine executed Python tasks, plus one router. Nineteen scored a clean 9 out of 9, at measured costs from $0.03 to $12.41 per 1,000 tasks — a 414x spread for an identical score. The headline result is a double record: Upstage Solar Mini 4 became the first model in our set to be both the cheapest and the fastest to clear the suite. The most useful finding is less flattering to the vendors: every OpenAI -pro tier we measured consumed roughly 28 times the input tokens of its standard sibling on the same prompts.
This is the index for our October sweep. Every model was called through the same harness on 2026-10-02 and every cost carries the list price captured that day. Each row links to its full write-up. The September sweep covers the fourteen models before these.
The nineteen that scored 9 of 9
| Model | Measured cost / 1k tasks | Latency | Reasoning tokens |
|---|---|---|---|
| Solar Mini 4 | $0.03 | 2s | 0 |
| MiMo v2.6 Flash | $0.04 | 5s | 9 |
| Qwen3 Coder | $0.13 | 2.7s | 0 |
| GPT-6 Luna | $0.16 | 4.9s | 201 |
| Devstral 2512 | $0.23 | 2.1s | 0 |
| MiMo v2.6 Pro | $0.29 | 9.5s | 169 |
| GPT-5.1-Codex-Mini | $0.45 | 4.4s | 151 |
| GPT-6 Luna Pro | $0.49 | 7.3s | 370 |
| GLM-5.3-FlashX | $1.36 | 12.9s | 963 |
| GPT-6.1 Sol | $1.61 | 5.9s | 48 |
| Claude Sonnet 5.5 | $1.95 | 3.4s | 31 |
| GPT-6 Sol | $2.03 | 5.2s | 85 |
| Claude Opus 5.5 | $4.03 | 4.8s | 33 |
| Fireworks Ember-1 | $4.63 | 4.1s | 166 |
| Grok 4.7 | $4.70 | 10s | 239 |
| MiMo v2.6 Pro UltraSpeed | $4.80 | 2.5s | 377 |
| GPT-6 Sol Pro | $7.45 | 5.9s | 158 |
| Aion 3.5 | $8.10 | 28.5s | 1,095 |
| Qwen3.8-Max-Prime | $12.41 | 15.5s | 875 |
A new floor, and a new speed record
Solar Mini 4 scored 9 out of 9 at $0.03 in 2.0 seconds with zero reasoning tokens. Before it, the cheapest and the fastest perfect scorers were different models, and both records changed hands this month. MiMo v2.6 Flash at $0.04 sits a cent behind. A one-run lead of a cent or a few tenths of a second is a cluster, not a coronation — the full cost leaderboard treats it that way.
Two older pages carried the previous records without a date on them. We corrected both and added dated update notes, rather than leaving a superlative that stopped being true.
Same price, same score: the pairings
Vendors keep converging on identical rate cards, which makes the measured differences the only differences. GPT-6 Sol and Claude Sonnet 5.5 both list at $2 and $10 and both scored 9/9, at $2.03 and $1.95 — a tie on cost, with Sonnet 1.8 seconds quicker. Inside OpenAI's own line, GPT-6.1 Sol did the same work for $1.61 at the identical list price: the newer point release is the cheaper one. And Qwen3 Coder Plus scored lower than base Qwen3 Coder while costing about three times as much.
The Pro tiers and the hidden prompt
Across the identical nine prompts, the standard GPT-6 models consumed 604 input tokens. Their Pro variants consumed 16,898 (Sol Pro) and 17,806 (Luna Pro), in line with Astra Pro's 17,011 last month. That is a fixed provider-side prompt on every request, and on short prompts it is most of the bill — written up as a pattern in the GPT-6 Pro piece. Grok 4.7 shows the same shape between versions: 2,437 input tokens on Grok 4.6, 11,761 on 4.7.
A router that spends where it thinks it should
We sent the same nine tasks to Jev Router. It passed all nine and used four different models to do it, for a summed reported $1.01 per 1,000 tasks. Its one decision to send a task to Claude Opus 5.5 accounted for 56% of that bill. The catalogue lists it at minus one million dollars per million tokens, which is a router placeholder, not a price.
The eight that did not get a clean 9 of 9
- GLM-5.3 Prime scored 7 of 9 at $11.52 — a complete run with two genuine wrong answers, from the most expensive GLM tier.
- Qwen3 Coder Plus and Flash and Codestral 2508 each scored 8 of 9, all wrong on the same CSV-parsing task.
- Seed 2.0 Code and Command A Plus are excluded: each ran out of its 4,000-token budget on one task and returned an empty body.
- Space Bunny Alpha is excluded after two incomplete runs on an unstable stealth endpoint. We publish no score and make no guess at who built it.
- KAT-Coder-Pro v2.5 never ran: its only provider returned HTTP 400 to every request, including a trivial one.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer. Cost is derived — measured token counts multiplied by the list price captured 2026-10-02 — not a billing statement, except for the router, where we summed the per-call cost the API reported. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page, and prices move.
What this sweep cannot tell you
- Whether a flagship is worth its price. Nineteen models scored identically. Nine self-contained Python functions are a floor test, not a capability ranking.
- Long-horizon and agentic work, which is what the expensive tiers and the router are sold on, and where cache loss on a model switch would show up. Our suite is single-turn.
- Context windows, vision, audio and tool use. Text-only, Python-only, short prompts.
- Stability. One scored attempt per task. A one-cent lead is inside the noise.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab