Benchmarks

Step 3.5 Flash: A Cheap Model That Cannot Finish Inside a 4,000-Token Budget

Step 3.5 Flash lists at $0.10 in and $0.30 out per million tokens — cheaper per output token than DeepSeek V3.2 at $0.40. On one task from our benchmark, the identical prompt at temperature 0, DeepSeek V3.2 answered in 166 output tokens with zero reasoning. Step 3.5 Flash used 11,802 output tokens, of which 10,675 were reasoning. That is 71x the tokens and 53x the cost, from the cheaper sticker. Worse: at our standard 4,000-token cap it does not answer at all — it burns the entire budget thinking and returns an empty body. This page is what that behaviour looks like measured, and why we excluded the model from our comparison table rather than publishing a partial score.

Bar chart of output tokens spent on one CSV parsing task by three models, from 166 to 11,802

A cheap price per token is only cheap if the model stops talking. This is the clearest case of that we have measured.

What happened in the benchmark

We ran Step 3.5 Flash through our nine executed Python tasks twice. Neither run completed.

An empty body with finish_reason=length means the model spent its entire max_tokens allowance on reasoning and never emitted a response. Our cap is 4,000 tokens, which 58 other models in our set find comfortable.

We treat that as a failure to answer, not a wrong answer — a distinction that matters, and one we got wrong once before in a way that nearly cost a strong model five points. That story is in the content filter post.

Giving it room, to be fair

Before publishing any of this, there was an obvious objection to settle: is the model incapable, or is our budget just too small for it? Those deserve very different write-ups. So we re-ran parse_csv_line — one of the tasks it failed — at a 16,000-token cap, four times our standard.

It succeeded. Here is what that cost:

Runfinish_reasonOutput tokensReasoning tokensAnswer
Step 3.5 Flash, 4,000 caplength4,0003,981empty
Step 3.5 Flash, 16,000 capstop11,80210,675735 characters

So it is not incapable. Given room, it writes the function. It simply needs eleven thousand eight hundred tokens to write a CSV parser, and 90% of those are reasoning tokens you never see and always pay for.

The three-model comparison

Same prompt, same task, temperature 0, all run 2026-08-22:

One CSV parser. Same prompt. 166 tokens versus 11,802.All three returned a working function. The difference is entirely in how long they thought about it first.DeepSeek V3.2166 tokens · 0 reasoningLing 3.0 Flash9,902 tokens · 8,343 reasoningStep 3.5 Flash11,802 tokens · 10,675 reasoningOne scale throughout: 0.0506 px per token. Step 3.5 Flash and Ling 3.0 Flash were given a 16,000-token cap; DeepSeek V3.2 ran at 4,000 and did not need it.Task: parse_csv_line. Prompt identical across all three, temperature 0, measured 2026-08-22.
Reasoning tokens bill at the output rate, so the grey and black portions are the invoice.

Now put the list prices next to it. Step 3.5 Flash bills output at $0.30 per million, DeepSeek V3.2 at $0.40. On this single task:

ModelOutput tokensOutput price / 1MCost for this one taskVersus DeepSeek
DeepSeek V3.2166$0.40$0.0000664—
Ling 3.0 Flash9,902$0.06$0.00059418.9x
Step 3.5 Flash11,802$0.30$0.003540653.3x

The cheaper price per token loses by a factor of 53, and Ling 3.0 Flash — whose $0.02 and $0.06 is the third-cheapest combined list price among the 420 priced models in the OpenRouter catalogue on 2026-08-22 — still loses by a factor of nearly nine. Volume beats price, and it is not close.

Ling 3.0 Flash does the same thing

Ling 3.0 Flash failed our suite twice too. First run: one task failed at the API layer. Second run: one API failure and one genuinely wrong answer. On the fairness re-test it did produce a working parse_csv_line — using 9,902 tokens, 8,343 of them reasoning.

Two different vendors, the same failure shape. Both are marketed as fast, cheap tiers. Both spend tokens like a frontier reasoning model and price like a budget one, which is a combination that reads well on a pricing page and badly on an invoice.

Why we excluded them instead of scoring them

Both models are marked excluded in our benchmark data rather than carrying a score. The reason is arithmetic, not judgement: Step 3.5 Flash completed 5 of 9 scored tasks on its second run and Ling 3.0 Flash 8 of 8 on its first. Printing “5/5” or “8/8” in a column where every other row says “9/9” implies a comparison that does not exist, because the denominators differ.

That is the same call we made on an earlier Qwen model with the identical failure. It is not a verdict on the models — the data above is the verdict, and it is more informative than a score would have been.

What this means if you were going to use them

This is the fourth time this week our measurements have contradicted a price page. GPT-5.1-Codex-Max lists below its sibling and costs 3.09x more. Qwen3.8-Max raised its price 36% and cut the bill 27%. Grok 4.6 matches Grok 4.5's sticker exactly and costs 70% more. The list price is not the bill.

How these numbers were produced

The benchmark is nine Python tasks, each a function signature plus a spec and no example tests; generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout, at temperature 0 and max_tokens 4000, one scored attempt per task. The fairness re-test in this article used the identical prompt string from that harness, at temperature 0, with the cap raised to 16,000; token counts are read from the provider's usage payload. Costs here are for a single task and are derived from list prices captured 2026-08-22 — not a billing statement, and prices move. Runs go through OpenRouter. Full method on the methodology page.

What we did not measure

FAQ

Is Step 3.5 Flash cheap? Per token, yes — $0.10 in and $0.30 out. Per finished task, no. On one CSV parser it cost 53 times what DeepSeek V3.2 cost for the same result.

Why does Step 3.5 Flash return empty responses? It exhausts max_tokens on reasoning before emitting an answer. At a 4,000-token cap it returned finish_reason=length with an empty body; at 16,000 it finished, using 11,802 tokens.

How many reasoning tokens does Step 3.5 Flash use? 10,675 of its 11,802 output tokens on this task — about 90%.

Did you score Step 3.5 Flash? No. It completed 5 of 9 scored tasks on its second run, so we marked it excluded rather than print a score with a different denominator from every other model.

What about Ling 3.0 Flash? Same failure shape. Its $0.02 and $0.06 is the third-cheapest combined list price in the 420-model OpenRouter catalogue on 2026-08-22, and it still cost 8.9 times DeepSeek V3.2 on this task, using 9,902 tokens.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.