Doubao Seed Code Review: Eight Answers, One Empty Body (and No Score)
Our Doubao Seed Code review has no score in it, and that is the finding. ByteDance's Seed 2.0 Code (bytedance-seed/seed-2.0-code) answered 8 of our 9 executed Python tasks on 2026-10-02. On the ninth, token_bucket, it returned an empty body with finish_reason=length: it spent the whole 4,000-token allowance thinking and never wrote the function. Across the run it emitted 2,046 reasoning tokens per call — more than GLM-5.1 at 1,327, the most reasoning-heavy model to score 9/9 in our set as of 2026-10-02 — and took a 31.1-second mean. A coding model that cannot finish inside the budget is not a wrong answer and not a right one, so we marked it excluded.
Most model reviews end on a number. This one ends on a denominator, because the denominator is what went wrong.
What the run actually produced
| Metric | Seed 2.0 Code |
|---|---|
| Status in our data | excluded — incomplete run |
| Tasks scored | 8 of 9 (all 8 passed) |
| Task with no answer | token_bucket — empty body, finish_reason=length |
| Reasoning tokens per call | 2,046 |
| Mean latency | 31.1s |
| Tokens across the suite | 705 in / 17,366 out |
| Priced at | $0.5 in / $3 out per 1M, captured 2026-10-02 |
| Derived cost of this run / 1,000 tasks | $6.56 (priced 2026-10-02; not ranked) |
| Context window | 262,144 |
| Run date | 2026-10-02 |
The tempting summary is “8/8, so it is perfect on what it finished.” We do not print that, and you should not read it that way. Every model in our comparison that carries a score answered all nine tasks. An 8/8 beside a 9/9 implies a comparison with a different denominator — and the missing task is not random. It is the one the model found hardest to stop thinking about. Excluding it from the denominator flatters the model on exactly the dimension that failed.
The $6.56 is shown for the same reason we show the token counts: it is what this run consumed. It does not go into our cost ranking, because the run that produced it did not finish.
The shape of a model that hits the cap
An empty body with finish_reason=length has one meaning: the model reached max_tokens while still in its reasoning phase, so nothing was ever emitted as an answer. Our cap is 4,000 tokens for every model. Here is how much of it Seed 2.0 Code spends on reasoning per call, next to the models around it:
2,046 is a per-call figure, not a ceiling. If the typical call already burns about half the budget on reasoning, the hardest call in the suite will burn more — and on token_bucket it burned all of it. That is an inference from the shape, not something the sheet records call by call, but the outcome is consistent with it. Note too that 2,046 is about 1.5x GLM-5.1's 1,327 (2,046 ÷ 1,327), and GLM-5.1 is the heaviest thinker in our set that still finished all nine.
We have seen this exact failure before. Step 3.5 Flash and the text-only Ling 3.0 Flash both returned empty bodies at the same cap, and when we raised it for one task, Step 3.5 Flash finished — after spending most of its output on reasoning. On the same day as Seed 2.0 Code, Cohere's Command A+ failed the same way on parse_csv_line, at 1,095 reasoning tokens per call. Two vendors, one day, one failure shape.
Against the other coding models, same day
Seed 2.0 Code is sold as a coding specialist, so the fair comparison is the other coding-branded models we ran on 2026-10-02, all priced the same day:
| Model | Status | Mean latency | Reasoning / call | Output tokens, suite | Derived cost / 1k (2026-10-02) |
|---|---|---|---|---|---|
| Seed 2.0 Code | excluded — 8 of 9 scored | 31.1s | 2,046 | 17,366 | $6.56, not ranked |
| Devstral 2512 | 9/9 | 2.1s | 0 | 911 | $0.23 |
| Qwen3 Coder | 9/9 | 2.7s | 0 | 1,010 | $0.13 |
| GPT-5.1 Codex Mini | 9/9 | 4.4s | 151 | 1,964 | $0.45 |
| Codestral 2508 | 8/9 (missed parse_csv_line) | 2.5s | 0 | 950 | $0.11 |
| KAT-Coder-Pro V2.5 | excluded — not run, HTTP 400 on every call | — | — | — | — |
Codestral's 8/9 and Seed's 8 of 9 scored look alike and mean different things. Codestral answered all nine and got one wrong — a scored result. Seed answered eight and produced nothing on the ninth — an incomplete one.
The token row is the clearest contrast. Devstral 2512 cleared the full suite on 911 output tokens; Seed 2.0 Code used 17,366 on eight scored tasks, about 19x as many (17,366 ÷ 911). Its 31.1-second mean is slower than Aion 3.5's 28.5s, and Aion 3.5 ranks 75th fastest of the 75 models that scored 9/9. For reference, the cheapest and fastest model to clear the suite as of 2026-10-02 is Solar Mini 4, at $0.03 per 1,000 tasks and 2s, with 0 reasoning tokens.
What Doubao Seed Code is, per third parties
Everything in this section is third-party and unverified by us; all sources read 2026-10-02.
- The Doubao name. KrASIA reported that ByteDance's cloud unit Volcano Engine launched a coding model branded Doubao-Seed-Code, integrated with TRAE, ByteDance's AI coding IDE. Doubao is ByteDance's consumer brand; Seed is its model team.
- Seed 2.0 Code. OpenRouter's model page describes
bytedance-seed/seed-2.0-codeas a ByteDance Seed model optimised for agentic coding — frontend work, multilingual programming and coding-agent tools such as Claude Code. A HowAIWorks profile places it in the Seed 2.0 family and describes it as closed-weights; the OpenRouter catalogue lists no Hugging Face ID for it. - Reasoning control. The OpenRouter catalogue lists reasoning as optional, with low, medium and high effort settings and medium as the default.
- Lifespan. The same catalogue entry carries an expiration date, and OpenRouter's listing describes the model as being phased out.
One caution on naming. We tested the OpenRouter endpoint. We cannot confirm that it serves the same weights or settings as the Doubao-Seed-Code model inside TRAE or on Volcano Engine's own API, so read this as a review of seed-2.0-code on OpenRouter, not of every product sold under the Doubao name.
What an incomplete run does tell you
- It is competent when it finishes. Eight attempts, eight passes against hidden asserts. Nothing in the run suggests the model writes bad code.
- Its token spend is the risk, not its accuracy. If you call it with a tight
max_tokens, expect occasional empty responses; if you raise the cap, expect the bill and the wait to grow with it. This is the pattern we covered in our agent model comparison: in a loop, per-call reasoning multiplies by every turn. - The list price will mislead you. $0.5 in and $3 out (2026-10-02) reads as mid-market. What you pay depends on how many tokens it emits, and here it emitted a lot. Prices also move.
- Its natural home may be its own IDE. A model tuned for long agentic sessions in TRAE may be configured there very differently from a 4,000-token single-shot call. Our harness is not that environment.
Our suite is nine self-contained Python functions. It cannot separate a frontier coding model from a competent small one — Devstral, Qwen3 Coder and Solar Mini 4 all clear it. What it can show is whether a model finishes cheaply and predictably on small work, and here the answer is no.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure — including an empty body with finish_reason=length — is recorded separately from a wrong answer; we learned why that separation matters in the content filter post. Cost is derived from measured token counts at the list price captured 2026-10-02, not a billing statement. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page.
What we did not measure
- A re-run of
token_bucketat a higher cap. We did that fairness test for Step 3.5 Flash and not here, so we cannot say whether Seed 2.0 Code would have answered correctly given more room, or how many tokens it would have needed. - Low reasoning effort. The catalogue offers it; we ran at the default, and a lower setting could change everything on this page.
- Per-call variance. We report reasoning tokens per call as one figure, so the claim that the hardest call ran far above it is inference, not measurement.
- Agentic, multi-file or frontend work — the jobs the model is marketed for. Nine single-turn functions say nothing about them.
- The TRAE and Volcano Engine endpoints. Only OpenRouter, only once.
- Repeat runs. One scored attempt per task. A second run might finish, or might lose a different task.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab