Model Reviews

Solar Mini 4: Cheapest and Fastest at Once (With Zero Reasoning Tokens)

Solar Mini 4 scored 9 out of 9 on our executed Python benchmark at $0.03 per 1,000 tasks — priced at its 2026-10-02 list — with a 2s mean latency. As of 2026-10-02 that makes it 1st cheapest and 1st fastest of the 75 models in our set that clear the suite, the same model at the top of both lists. The part we did not expect: it emitted 0 reasoning tokens per call. Third-party coverage describes Solar Mini 4 as a reasoning model that spends heavily on thinking. On the endpoint we called, it did not think at all, and got everything right anyway.

DataLLM Lab article cover: Solar Mini 4: Cheapest and Fastest at Once (With Zero Reasoning Tokens)

Cheap models are usually cheap because they are slow, and fast models are usually fast because a big provider throws hardware at them. Solar Mini 4 sits at the top of both of our lists at once.

The result

MetricSolar Mini 4 (upstage/solar-mini4)
Score9/9
Measured cost / 1,000 tasks$0.03 (priced 2026-10-02)
Mean latency2s
Reasoning tokens per call0
Tokens across the suite1,009 in / 938 out
Priced at, in / out$0.05 / $0.2 per 1M (2026-10-02)
Context window (catalogue)524,288
Measured2026-10-02
Rank among 75 models at 9/91 cheapest · 1 fastest

Two things produce the cost lead, and both are visible in the table. The rate card is low — $0.05 in and $0.2 out per million as of 2026-10-02 — and the model is terse: 938 output tokens across all nine tasks, with nothing spent on hidden reasoning. Its input count of 1,009 is actually higher than most of the models in this article for the same nine prompts (OpenAI models read 604), but input is the cheap side of the bill.

The cheapest 9/9, ranked

Before this sweep the cost position belonged to Ling 3.0 Flash VL at $0.07, which held it as of 2026-09-16. In the same 2026-10-02 sweep that measured Solar Mini 4, Xiaomi's MiMo V2.6 Flash came in at $0.04 — it would have taken the spot on its own, and lost it by one cent. These are the top five by rank on our sheet, each with the date its cost was priced:

Cost rankModelCost / 1k tasksPriced onLatencyReasoning / call
1Solar Mini 4$0.032026-10-022s0
2MiMo V2.6 Flash$0.042026-10-025s9
3Ling 3.0 Flash VL$0.072026-09-154.4s234
4DeepSeek V3.2$0.082026-07-307.1s0
5Qwen3 Coder Next$0.12026-07-177s0

Two of those rows carry a warning. Ling 3.0 Flash VL's list price has moved since its run: it was priced at $0.06 in / $0.18 out on 2026-09-15 and lists at $0.021 / $0.06 today. Its $0.07 uses the run-date price, and we have not re-run it, so we do not know where it would land now. DeepSeek V3.2's list has also moved since its run. Rankings built on different price dates are a snapshot, not a standing order — prices move, and they have moved twice in this table alone.

The fastest 9/9, ranked

On latency, Solar Mini 4 displaces GPT-5.4 mini, whose 2.3s mean made it our fastest 9/9 until this sweep. Mistral's Devstral 2512, also measured on 2026-10-02, slots in between at 2.1s.

Speed rankModelMean latencyCost / 1k tasksPriced on
1Solar Mini 42s$0.032026-10-02
2Devstral 25122.1s$0.232026-10-02
3GPT-5.4 mini2.3s$0.532026-07-30
5MiMo V2.6 Pro UltraSpeed2.5s$4.82026-10-02
6Qwen3 Coder2.7s$0.132026-10-02

Rank 4 is held by a model outside the extract this article draws on, so we leave the row out rather than guess it. The more useful observation is what sits just outside this table: Qwen3 Coder Plus and Qwen3 Coder Flash also ran at a 2s mean on 2026-10-02, and Codestral 2508 at 2.5s. All three scored 8/9, each missing parse_csv_line. At this speed, fast and correct is not a given; Solar Mini 4 is notable for being both.

Top of both lists: the cheapest and the fastest models that clear the suiteAll ten bars are models that scored 9/9 on the identical nine executed Python tasks.MEASURED COST PER 1,000 TASKS · COST RANKS 1–5Solar Mini 4$0.03MiMo V2.6 Flash$0.04Ling 3.0 Flash VL$0.07DeepSeek V3.2$0.08Qwen3 Coder Next$0.1MEAN LATENCY · SPEED RANKS 1, 2, 3, 5, 6Solar Mini 42sDevstral 25122.1sGPT-5.4 mini2.3sMiMo V2.6 Pro UltraSpeed2.5sQwen3 Coder2.7sScales: cost 4,000 px per dollar; latency 150 px per second. Speed rank 4 is not in our extract.
The same model leads both charts. In both, the lead is one step of the rounding.

A cluster, not a coronation

Read the size of the lead before you read the rank. On cost, Solar Mini 4 is one cent per thousand tasks ahead of MiMo V2.6 Flash — $0.03 against $0.04, both priced 2026-10-02, and both figures rounded to the cent. On speed it is one tenth of a second ahead of Devstral 2512. Each task gets one scored attempt and we do not average across runs, so a re-run could reorder either pair. Latency in particular depends on which upstream provider served the call and how busy it was at that moment.

The honest reading is that, as of 2026-10-02, the floor for a model that solves every task in this suite now sits at a few cents per thousand tasks and about two seconds, and several unrelated vendors cluster there. Solar Mini 4 is the first name on both lists today. It is not proven to be a different class from the models one row below it.

Against the top of the market the price gap is wide. GPT-6 Astra scored the same 9 out of 9 at $8.19 per 1,000 tasks, priced 2026-09-15. Nine self-contained Python functions cannot separate a frontier model from a competent small one — that is a limit of this suite, not a verdict that the two are equal. Our cheapest-that-works analysis and cheap coding roundup cover the wider field.

A reasoning model that did not reason

Third-party facts first. Artificial Analysis, in its launch write-up (read 2026-10-02), describes Solar Mini 4 as a proprietary reasoning model from Upstage with weights not released, built as a mixture-of-experts with a small active parameter count, and reports that on its own index it spends so many reasoning tokens per task that it costs more per task than GPT-6 Luna despite similar per-token prices. Korean outlet Digital Today (read 2026-10-02) also reports the mixture-of-experts design. The two sources give different launch dates, so we leave the date out. Artificial Analysis also quotes a per-token price and a context window that differ from the catalogue entry we priced against; the figures in our tables are the catalogue's, captured 2026-10-02.

Our measurement points the other way. On the endpoint we called, Solar Mini 4 produced 0 reasoning tokens per call and 938 output tokens across the whole suite. GPT-6 Luna, on our same nine tasks, emitted 201 reasoning tokens per call and cost $0.16 per 1,000 tasks, priced 2026-10-02. We have seen this pattern before: behaviour is set by the endpoint and its defaults, not by the weights alone. Our most likely explanation is that reasoning is not switched on by default where we called it, but we did not test that, and we will not present a guess as a finding. What we can say is narrow: called the way we call every model, it did not need to think to pass these nine tasks.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer. Cost is derived from measured token counts multiplied by the list price captured on each model's run date — 2026-10-02 for Solar Mini 4 — and is not a billing statement. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

What we did not measure

If latency is your constraint more than cost, our latency guide covers what you control besides model choice.

FAQ

Is Solar Mini 4 the cheapest model that passes your benchmark? As of 2026-10-02, yes: $0.03 per 1,000 tasks, one cent below MiMo V2.6 Flash.

Is Solar Mini 4 the fastest? As of 2026-10-02, yes among the 75 models at 9/9: a 2s mean, one tenth of a second ahead of Devstral 2512.

How much does Solar Mini 4 cost? $0.05 per million input tokens and $0.2 output on the catalogue entry we priced on 2026-10-02.

Does it use reasoning tokens? Not on the endpoint we called: 0 per call. Third-party benchmarks describe it as a reasoning model, so your configuration matters.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.