Step 3.5 Flash: A Cheap Model That Cannot Finish Inside a 4,000-Token Budget
Step 3.5 Flash lists at $0.10 in and $0.30 out per million tokens — cheaper per output token than DeepSeek V3.2 at $0.40. On one task from our benchmark, the identical prompt at temperature 0, DeepSeek V3.2 answered in 166 output tokens with zero reasoning. Step 3.5 Flash used 11,802 output tokens, of which 10,675 were reasoning. That is 71x the tokens and 53x the cost, from the cheaper sticker. Worse: at our standard 4,000-token cap it does not answer at all — it burns the entire budget thinking and returns an empty body. This page is what that behaviour looks like measured, and why we excluded the model from our comparison table rather than publishing a partial score.
A cheap price per token is only cheap if the model stops talking. This is the clearest case of that we have measured.
What happened in the benchmark
We ran Step 3.5 Flash through our nine executed Python tasks twice. Neither run completed.
- First run: 7 of 9 tasks returned an answer. Two failed with
finish_reason=lengthand an empty body after retries. - Second run, same settings: worse. Only 5 of 9 returned an answer.
An empty body with finish_reason=length means the model spent its entire max_tokens allowance on reasoning and never emitted a response. Our cap is 4,000 tokens, which 58 other models in our set find comfortable.
We treat that as a failure to answer, not a wrong answer — a distinction that matters, and one we got wrong once before in a way that nearly cost a strong model five points. That story is in the content filter post.
Giving it room, to be fair
Before publishing any of this, there was an obvious objection to settle: is the model incapable, or is our budget just too small for it? Those deserve very different write-ups. So we re-ran parse_csv_line — one of the tasks it failed — at a 16,000-token cap, four times our standard.
It succeeded. Here is what that cost:
| Run | finish_reason | Output tokens | Reasoning tokens | Answer |
|---|---|---|---|---|
| Step 3.5 Flash, 4,000 cap | length | 4,000 | 3,981 | empty |
| Step 3.5 Flash, 16,000 cap | stop | 11,802 | 10,675 | 735 characters |
So it is not incapable. Given room, it writes the function. It simply needs eleven thousand eight hundred tokens to write a CSV parser, and 90% of those are reasoning tokens you never see and always pay for.
The three-model comparison
Same prompt, same task, temperature 0, all run 2026-08-22:
Now put the list prices next to it. Step 3.5 Flash bills output at $0.30 per million, DeepSeek V3.2 at $0.40. On this single task:
| Model | Output tokens | Output price / 1M | Cost for this one task | Versus DeepSeek |
|---|---|---|---|---|
| DeepSeek V3.2 | 166 | $0.40 | $0.0000664 | — |
| Ling 3.0 Flash | 9,902 | $0.06 | $0.0005941 | 8.9x |
| Step 3.5 Flash | 11,802 | $0.30 | $0.0035406 | 53.3x |
The cheaper price per token loses by a factor of 53, and Ling 3.0 Flash — whose $0.02 and $0.06 is the third-cheapest combined list price among the 420 priced models in the OpenRouter catalogue on 2026-08-22 — still loses by a factor of nearly nine. Volume beats price, and it is not close.
Ling 3.0 Flash does the same thing
Ling 3.0 Flash failed our suite twice too. First run: one task failed at the API layer. Second run: one API failure and one genuinely wrong answer. On the fairness re-test it did produce a working parse_csv_line — using 9,902 tokens, 8,343 of them reasoning.
Two different vendors, the same failure shape. Both are marketed as fast, cheap tiers. Both spend tokens like a frontier reasoning model and price like a budget one, which is a combination that reads well on a pricing page and badly on an invoice.
Why we excluded them instead of scoring them
Both models are marked excluded in our benchmark data rather than carrying a score. The reason is arithmetic, not judgement: Step 3.5 Flash completed 5 of 9 scored tasks on its second run and Ling 3.0 Flash 8 of 8 on its first. Printing “5/5” or “8/8” in a column where every other row says “9/9” implies a comparison that does not exist, because the denominators differ.
That is the same call we made on an earlier Qwen model with the identical failure. It is not a verdict on the models — the data above is the verdict, and it is more informative than a score would have been.
What this means if you were going to use them
- Do not size your
max_tokensfrom other models' behaviour. A 4,000-token cap is generous for most of our set and fatal for these two. If you cap, you will get empty responses; if you do not cap, you will get the bill. - Budget from measured output tokens, not from the price page. A price per million tokens tells you nothing until you know how many millions.
- In an agent loop this compounds. Per-call reasoning multiplied by every turn is what actually sets a harness bill — the case we worked through in our agent model comparison and again in the DeepSeek Harness write-up.
- If you want cheap and predictable, our measured floor is DeepSeek V3.2 at $0.08 per 1,000 tasks with 9/9 and zero reasoning tokens. The cheap coding roundup has the rest of that end of the market.
This is the fourth time this week our measurements have contradicted a price page. GPT-5.1-Codex-Max lists below its sibling and costs 3.09x more. Qwen3.8-Max raised its price 36% and cut the bill 27%. Grok 4.6 matches Grok 4.5's sticker exactly and costs 70% more. The list price is not the bill.
How these numbers were produced
The benchmark is nine Python tasks, each a function signature plus a spec and no example tests; generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout, at temperature 0 and max_tokens 4000, one scored attempt per task. The fairness re-test in this article used the identical prompt string from that harness, at temperature 0, with the cap raised to 16,000; token counts are read from the provider's usage payload. Costs here are for a single task and are derived from list prices captured 2026-08-22 — not a billing statement, and prices move. Runs go through OpenRouter. Full method on the methodology page.
What we did not measure
- A full scored run at 16,000 tokens. We raised the cap for one task to answer the fairness question, not to produce a comparable score. Our published scores all use the same 4,000-token cap, and changing it for one model would break that.
- Whether the reasoning is useful anywhere. On this task it was not — DeepSeek V3.2 got the same result in 166 tokens. On a harder problem the extra thinking might pay.
- Any effort or reasoning control. If either model exposes a way to turn thinking down, we did not use it, and it would change everything on this page.
- Anything but Python, and only nine self-contained functions.
FAQ
Is Step 3.5 Flash cheap? Per token, yes — $0.10 in and $0.30 out. Per finished task, no. On one CSV parser it cost 53 times what DeepSeek V3.2 cost for the same result.
Why does Step 3.5 Flash return empty responses? It exhausts max_tokens on reasoning before emitting an answer. At a 4,000-token cap it returned finish_reason=length with an empty body; at 16,000 it finished, using 11,802 tokens.
How many reasoning tokens does Step 3.5 Flash use? 10,675 of its 11,802 output tokens on this task — about 90%.
Did you score Step 3.5 Flash? No. It completed 5 of 9 scored tasks on its second run, so we marked it excluded rather than print a score with a different denominator from every other model.
What about Ling 3.0 Flash? Same failure shape. Its $0.02 and $0.06 is the third-cheapest combined list price in the 420-model OpenRouter catalogue on 2026-08-22, and it still cost 8.9 times DeepSeek V3.2 on this task, using 9,902 tokens.
DataLLM Lab