Model Reviews

Doubao Seed Code Review: Eight Answers, One Empty Body (and No Score)

Our Doubao Seed Code review has no score in it, and that is the finding. ByteDance's Seed 2.0 Code (bytedance-seed/seed-2.0-code) answered 8 of our 9 executed Python tasks on 2026-10-02. On the ninth, token_bucket, it returned an empty body with finish_reason=length: it spent the whole 4,000-token allowance thinking and never wrote the function. Across the run it emitted 2,046 reasoning tokens per call — more than GLM-5.1 at 1,327, the most reasoning-heavy model to score 9/9 in our set as of 2026-10-02 — and took a 31.1-second mean. A coding model that cannot finish inside the budget is not a wrong answer and not a right one, so we marked it excluded.

DataLLM Lab article cover: Doubao Seed Code Review: Eight Answers, One Empty Body (and No Score)

Most model reviews end on a number. This one ends on a denominator, because the denominator is what went wrong.

What the run actually produced

MetricSeed 2.0 Code
Status in our dataexcluded — incomplete run
Tasks scored8 of 9 (all 8 passed)
Task with no answertoken_bucket — empty body, finish_reason=length
Reasoning tokens per call2,046
Mean latency31.1s
Tokens across the suite705 in / 17,366 out
Priced at$0.5 in / $3 out per 1M, captured 2026-10-02
Derived cost of this run / 1,000 tasks$6.56 (priced 2026-10-02; not ranked)
Context window262,144
Run date2026-10-02

The tempting summary is “8/8, so it is perfect on what it finished.” We do not print that, and you should not read it that way. Every model in our comparison that carries a score answered all nine tasks. An 8/8 beside a 9/9 implies a comparison with a different denominator — and the missing task is not random. It is the one the model found hardest to stop thinking about. Excluding it from the denominator flatters the model on exactly the dimension that failed.

The $6.56 is shown for the same reason we show the token counts: it is what this run consumed. It does not go into our cost ranking, because the run that produced it did not finish.

The shape of a model that hits the cap

An empty body with finish_reason=length has one meaning: the model reached max_tokens while still in its reasoning phase, so nothing was ever emitted as an answer. Our cap is 4,000 tokens for every model. Here is how much of it Seed 2.0 Code spends on reasoning per call, next to the models around it:

Reasoning tokens per call, against the 4,000-token capIdentical nine Python tasks. Blue bars are excluded runs; grey bars scored 9/9.max_tokens cap: 4,000Seed 2.0 Code · excluded2,046GLM-5.1 · 9/91,327Command A+ · excluded1,095GLM-5.3 FlashX · 9/9963GPT-5.1 Codex Mini · 9/9151Devstral 2512 · 9/90Qwen3 Coder · 9/90One scale throughout: 0.15 px per token, bars start at x = 230. Width = tokens × 0.15:2,046 → 306.9 px · 1,327 → 199.05 px · 1,095 → 164.25 px · 963 → 144.45 px · 151 → 22.65 px · cap 4,000 → 600 px
The cap is the dashed line. A per-call figure at half of it leaves no room for the hardest task.

2,046 is a per-call figure, not a ceiling. If the typical call already burns about half the budget on reasoning, the hardest call in the suite will burn more — and on token_bucket it burned all of it. That is an inference from the shape, not something the sheet records call by call, but the outcome is consistent with it. Note too that 2,046 is about 1.5x GLM-5.1's 1,327 (2,046 ÷ 1,327), and GLM-5.1 is the heaviest thinker in our set that still finished all nine.

We have seen this exact failure before. Step 3.5 Flash and the text-only Ling 3.0 Flash both returned empty bodies at the same cap, and when we raised it for one task, Step 3.5 Flash finished — after spending most of its output on reasoning. On the same day as Seed 2.0 Code, Cohere's Command A+ failed the same way on parse_csv_line, at 1,095 reasoning tokens per call. Two vendors, one day, one failure shape.

Against the other coding models, same day

Seed 2.0 Code is sold as a coding specialist, so the fair comparison is the other coding-branded models we ran on 2026-10-02, all priced the same day:

ModelStatusMean latencyReasoning / callOutput tokens, suiteDerived cost / 1k (2026-10-02)
Seed 2.0 Codeexcluded — 8 of 9 scored31.1s2,04617,366$6.56, not ranked
Devstral 25129/92.1s0911$0.23
Qwen3 Coder9/92.7s01,010$0.13
GPT-5.1 Codex Mini9/94.4s1511,964$0.45
Codestral 25088/9 (missed parse_csv_line)2.5s0950$0.11
KAT-Coder-Pro V2.5excluded — not run, HTTP 400 on every call————

Codestral's 8/9 and Seed's 8 of 9 scored look alike and mean different things. Codestral answered all nine and got one wrong — a scored result. Seed answered eight and produced nothing on the ninth — an incomplete one.

The token row is the clearest contrast. Devstral 2512 cleared the full suite on 911 output tokens; Seed 2.0 Code used 17,366 on eight scored tasks, about 19x as many (17,366 ÷ 911). Its 31.1-second mean is slower than Aion 3.5's 28.5s, and Aion 3.5 ranks 75th fastest of the 75 models that scored 9/9. For reference, the cheapest and fastest model to clear the suite as of 2026-10-02 is Solar Mini 4, at $0.03 per 1,000 tasks and 2s, with 0 reasoning tokens.

What Doubao Seed Code is, per third parties

Everything in this section is third-party and unverified by us; all sources read 2026-10-02.

One caution on naming. We tested the OpenRouter endpoint. We cannot confirm that it serves the same weights or settings as the Doubao-Seed-Code model inside TRAE or on Volcano Engine's own API, so read this as a review of seed-2.0-code on OpenRouter, not of every product sold under the Doubao name.

What an incomplete run does tell you

Our suite is nine self-contained Python functions. It cannot separate a frontier coding model from a competent small one — Devstral, Qwen3 Coder and Solar Mini 4 all clear it. What it can show is whether a model finishes cheaply and predictably on small work, and here the answer is no.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure — including an empty body with finish_reason=length — is recorded separately from a wrong answer; we learned why that separation matters in the content filter post. Cost is derived from measured token counts at the list price captured 2026-10-02, not a billing statement. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.

What we did not measure

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.