Benchmarks

Reasoning Effort: Low vs Medium vs High on Seven Models (Measured)

Reasoning effort is how you tell a thinking model how hard to think. It is the setting that replaced temperature as the knob that changes anything on GPT-5 and later. We ran our nine executed coding tasks on seven models at four settings each: no setting, low, medium and high. That is 28 runs. Low scored 9 out of 9 on all seven models. High never passed a task that low had failed, and on the same tasks it cost between 0.96 and 10.28 times as much. And “no setting” is not the same as medium. On Gemini 3.8 Flash the default used 6,389 reasoning tokens and low used zero.

DataLLM Lab article cover: Reasoning Effort: Low vs Medium vs High on Seven Models (Measured)

Every provider now has a version of this setting. OpenAI calls it reasoning_effort, Google has a thinking budget, and Anthropic and xAI have their own controls. OpenRouter maps all of them to one field, reasoning: {"effort": "low" | "medium" | "high"}. That made it possible to change only this setting and hold everything else fixed. Here is what changed.

What the setting does

A reasoning model writes hidden reasoning tokens before its answer. You pay for them at the output price, even though you never see them. Effort tells the model roughly how many to spend. More reasoning can fix mistakes on hard problems. On easy ones it is just a bigger bill and a longer wait. You can see the count in usage.completion_tokens_details.reasoning_tokens. Our hidden token investigation covers the other token cost you cannot see, the system prompts some models prepend.

28 runs in one table

Cost is the usage.cost the API reported, summed over the nine tasks and scaled to 1,000 tasks. Reasoning tokens are summed over the nine.

ModelDefaultLowMediumHigh
GPT-6 Sol9/9 · 693 tok · $1.919/9 · 171 · $1.299/9 · 829 · $2.079/9 · 1,405 · $2.73
GPT-6.1 Sol9/9 · 538 · $1.799/9 · 194 · $1.388/9 · 539 · $1.559/9 · 1,213 · $2.48
GPT-6 Luna8/9 · 1,111 · $0.119/9 · 509 · $0.089/9 · 811 · $0.109/9 · 1,577 · $0.14
Claude Sonnet 5.59/9 · 225 · $1.919/9 · 0 · $1.639/9 · 0 · $1.649/9 · 264 · $1.93
Grok 4.79/9 · 3,863 · $4.319/9 · 1,398 · $2.489/9 · 3,437 · $3.749/9 · 4,530 · $4.58
Gemini 3.8 Flash9/9 · 6,389 · $3.179/9 · 0 · $0.508/9 · 8,302 · $3.928/9 · 11,191 · $5.12
DeepSeek V4.1 Flash9/9 · 4,314 · $0.309/9 · 4,337 · $0.329/9 · 2,359 · $0.189/9 · 5,071 · $0.31

Bold marks the setting that cost least for each model. Low cost least on six of the seven. All four runs that missed a task missed the same one, parse_csv_line, our quoted-CSV parser, which also trips many models on our main leaderboard.

High effort cost 0.96x to 10.28x as much as low, for no extra passesCost at effort high ÷ cost at effort low, same nine tasks, API-reported. Low scored 9/9 on all seven models.Gemini 3.8 Flash10.28× · $0.50 → $5.12 per 1kGPT-6 Sol2.12× · $1.29 → $2.73 per 1kGrok 4.71.84× · $2.48 → $4.58 per 1kGPT-6.1 Sol1.80× · $1.38 → $2.48 per 1kGPT-6 Luna1.72× · $0.08 → $0.14 per 1kClaude Sonnet 5.51.19× · $1.63 → $1.93 per 1kDeepSeek V4.1 Flash0.96× · $0.32 → $0.31 per 1k1× (no difference)Scale: width = ratio × 50 px. One run per setting, 2026-10-03, via OpenRouter.
The bar to the left of the dashed line is the one model where high effort came out slightly cheaper than low.

Low never lost a task

On this suite, more thinking did not mean more passes. Low scored 9 out of 9 on all seven models. The four misses came at default (GPT-6 Luna), medium (GPT-6.1 Sol and Gemini 3.8 Flash) and high (Gemini 3.8 Flash). High did fix GPT-6 Luna's miss at default, but low had already fixed it, at $0.08 instead of $0.14.

One run per setting cannot show that low is more accurate than high. A single miss on parse_csv_line is within the run-to-run noise we see on that task. What it does show is the absence of a benefit. If high effort made a real difference on these tasks, seven models and 28 runs should have shown it somewhere. They did not.

Gemini 3.8 Flash: 6.36x cheaper at low

The largest effect was on Gemini 3.8 Flash. At low it wrote zero reasoning tokens, scored 9 out of 9, and cost $0.50 per 1,000 tasks. Its default spent 6,389 reasoning tokens for the same score at $3.17, which is 6.36 times as much. Its high setting spent 11,191, cost $5.12, which is 10.28 times low, and missed a task. Low also cut median latency from 10.922s to 5.181s.

For comparison, our standard benchmark of Gemini 3.8 Flash measured $3.87 per 1,000 tasks, derived at its 2026-09-15 list price of $0.75 / $3.75. That run sends no effort setting, so it ran at the default. If you use Gemini 3.8 Flash for this kind of work, low effort is the biggest saving on this page. Our Gemini 3.8 Flash review has the rest of its numbers.

Default is not medium

The documentation often implies that leaving effort unset means medium. Our token counts say otherwise:

Do not assume what the default is. Set the effort you want and read reasoning_tokens to confirm it took effect.

Two models where effort barely mattered

Claude Sonnet 5.5 wrote no reasoning tokens at either low or medium, and only 264 at high. All four settings scored 9 out of 9 and fell between $1.63 and $1.93 per 1,000 tasks. High cost 1.19 times low. On short tasks Sonnet 5.5 hardly reasons at any setting, so the setting has little to work with.

DeepSeek V4.1 Flash did not follow the expected order. Low used 4,337 reasoning tokens, almost the same as the default's 4,314. Medium used 2,359, the fewest of the four. High used 5,071. So medium cost least, at $0.18 per 1,000 tasks, and high was 0.96 times the cost of low. From outside we cannot tell whether low is passed to DeepSeek's serving providers or dropped along the way. The token counts look the same as no setting at all. Our DeepSeek V4.1 Flash review has its standard run.

How to choose a setting

  1. Start at low for extraction, classification, short code and anything with a checker. Move up only for a case you have seen low get wrong.
  2. Read the reasoning token count on a sample of real requests at each setting. The count shows whether the setting did anything.
  3. Raise max_tokens when you raise effort. Reasoning counts against the same limit. Our Chat Completions test found eight reasoning models that returned an empty answer when the limit was too small.
  4. Price it on your own traffic. On our suite the range was 0.96x to 10.28x. Longer, harder prompts will be different.

How we tested

Our standard harness: nine Python tasks, each a function signature plus a spec, with no example tests. Generated code runs against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Requests sent temperature 0 and max_tokens 4000, with one scored attempt per task. Note that GPT-5.4 and GPT-6 drop the temperature value, as our temperature test shows. The only change between runs was reasoning: {"effort": ...}, which was absent for “default”. All 28 runs went through OpenRouter on 2026-10-03, not through the DataLLM Lab gateway, and cost $0.4649 in total. Costs here are API-reported usage.cost. They are not derived from list price, so they can differ from the derived figures on our model pages. API failures would have been recorded separately from wrong answers; none occurred. See the methodology page for the full method.

What this cannot tell you

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.