Benchmarks

New LLM Models, October 2026: Twenty-Seven Releases on One Harness

Between 2026-09-16 and 2026-10-02 the catalogue gained enough new text models that we ran 27 of them through the same nine executed Python tasks, plus one router. Nineteen scored a clean 9 out of 9, at measured costs from $0.03 to $12.41 per 1,000 tasks — a 414x spread for an identical score. The headline result is a double record: Upstage Solar Mini 4 became the first model in our set to be both the cheapest and the fastest to clear the suite. The most useful finding is less flattering to the vendors: every OpenAI -pro tier we measured consumed roughly 28 times the input tokens of its standard sibling on the same prompts.

DataLLM Lab article cover: New LLM Models, October 2026: Twenty-Seven Releases on One Harness

This is the index for our October sweep. Every model was called through the same harness on 2026-10-02 and every cost carries the list price captured that day. Each row links to its full write-up. The September sweep covers the fourteen models before these.

The nineteen that scored 9 of 9

ModelMeasured cost / 1k tasksLatencyReasoning tokens
Solar Mini 4$0.032s0
MiMo v2.6 Flash$0.045s9
Qwen3 Coder$0.132.7s0
GPT-6 Luna$0.164.9s201
Devstral 2512$0.232.1s0
MiMo v2.6 Pro$0.299.5s169
GPT-5.1-Codex-Mini$0.454.4s151
GPT-6 Luna Pro$0.497.3s370
GLM-5.3-FlashX$1.3612.9s963
GPT-6.1 Sol$1.615.9s48
Claude Sonnet 5.5$1.953.4s31
GPT-6 Sol$2.035.2s85
Claude Opus 5.5$4.034.8s33
Fireworks Ember-1$4.634.1s166
Grok 4.7$4.7010s239
MiMo v2.6 Pro UltraSpeed$4.802.5s377
GPT-6 Sol Pro$7.455.9s158
Aion 3.5$8.1028.5s1,095
Qwen3.8-Max-Prime$12.4115.5s875
Twelve October models, one identical scoreMeasured cost per 1,000 tasks, all 9/9. Log scale — a linear axis would hide the left half.Solar Mini 4$0.03MiMo v2.6 Flash$0.04Qwen3 Coder$0.13GPT-6 Luna$0.16Devstral 2512$0.23GPT-6.1 Sol$1.61Claude Sonnet 5.5$1.95GPT-6 Sol$2.03Claude Opus 5.5$4.03Grok 4.7$4.70GPT-6 Sol Pro$7.45Qwen3.8-Max-Prime$12.41Log scale: 150 px per decade, origin $0.02. Bar widths computed, not drawn by eye.
Every bar solved all nine tasks. The axis is about money, not ability.

A new floor, and a new speed record

Solar Mini 4 scored 9 out of 9 at $0.03 in 2.0 seconds with zero reasoning tokens. Before it, the cheapest and the fastest perfect scorers were different models, and both records changed hands this month. MiMo v2.6 Flash at $0.04 sits a cent behind. A one-run lead of a cent or a few tenths of a second is a cluster, not a coronation — the full cost leaderboard treats it that way.

Two older pages carried the previous records without a date on them. We corrected both and added dated update notes, rather than leaving a superlative that stopped being true.

Same price, same score: the pairings

Vendors keep converging on identical rate cards, which makes the measured differences the only differences. GPT-6 Sol and Claude Sonnet 5.5 both list at $2 and $10 and both scored 9/9, at $2.03 and $1.95 — a tie on cost, with Sonnet 1.8 seconds quicker. Inside OpenAI's own line, GPT-6.1 Sol did the same work for $1.61 at the identical list price: the newer point release is the cheaper one. And Qwen3 Coder Plus scored lower than base Qwen3 Coder while costing about three times as much.

The Pro tiers and the hidden prompt

Across the identical nine prompts, the standard GPT-6 models consumed 604 input tokens. Their Pro variants consumed 16,898 (Sol Pro) and 17,806 (Luna Pro), in line with Astra Pro's 17,011 last month. That is a fixed provider-side prompt on every request, and on short prompts it is most of the bill — written up as a pattern in the GPT-6 Pro piece. Grok 4.7 shows the same shape between versions: 2,437 input tokens on Grok 4.6, 11,761 on 4.7.

A router that spends where it thinks it should

We sent the same nine tasks to Jev Router. It passed all nine and used four different models to do it, for a summed reported $1.01 per 1,000 tasks. Its one decision to send a task to Claude Opus 5.5 accounted for 56% of that bill. The catalogue lists it at minus one million dollars per million tokens, which is a router placeholder, not a price.

The eight that did not get a clean 9 of 9

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer. Cost is derived — measured token counts multiplied by the list price captured 2026-10-02 — not a billing statement, except for the router, where we summed the per-call cost the API reported. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page, and prices move.

What this sweep cannot tell you

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.