Model Reviews

Schematron v2: Our First Genuine Zero (And Why It Is Not a Bug Report)

Schematron v2 Turbo scored 0 out of 9 on our executed Python benchmark. Not a crash, not a timeout, not an empty body: it answered every one of the nine tasks, spending 1,430 output tokens across the suite at a 5.3-second mean, and every single answer failed the hidden asserts. Of the 82 entries in our set this is the first clean, complete, fully-scored zero we have recorded. It is also, at a derived $0.03 per 1,000 tasks, cheaper than the cheapest model that clears the suite. Both of those facts are true and neither one is a reason to use it or avoid it, because we measured a specialist on work it was never shipped to do.

DataLLM Lab article cover: Schematron v2: Our First Genuine Zero (And Why It Is Not a Bug Report)

Most bad benchmark results are infrastructure stories: a rate limit, a truncated body, a provider returning nothing. This one is not. Schematron v2 Turbo showed up, wrote code for all nine problems, and was wrong nine times.

The result

MetricSchematron v2 Turbo
Score0/9
Derived cost / 1,000 tasks$0.03 (list price on 2026-09-15)
Mean latency5.3s
Reasoning tokens per call0
Tokens across the suite867 in / 1,430 out
List price in / out$0.03 / $0.15 per 1M (2026-09-15)
Context window128,000
Measured2026-09-16

The tasks it missed are the whole suite: two_sum, valid_parentheses, merge_intervals, roman_to_int, lcs_len, flatten, top_k_words, token_bucket, parse_csv_line. Nine for nine.

What a real zero looks like

We have been burned by fake zeros before. A provider content filter used to hand our harness an empty completion, and the grader scored that empty string as a wrong answer instead of an infrastructure failure — the write-up is in the post on the content-filter bug. The standing rule since then is that any model which does not score full marks gets re-checked before we publish anything about it.

Schematron passes that check, which is exactly why the zero stands:

So the failure is semantic, not operational. The code came back, ran, and did the wrong thing.

Specialist, not defective

Schematron is not a general coding model and was never sold as one. The name is the product: it is built to take messy input — HTML, free text — and emit JSON conforming to a schema the caller supplies. That is a narrow, extremely useful transformation, and models trained hard on it tend to behave exactly as this one did: fast, no visible chain of thought, cheap, and completely uninterested in the open-ended instruction you actually sent.

Our suite hands a model a function signature plus a prose spec and asks it to invent the body. There is no schema to fill, no source document to parse. A 0/9 here is a measurement of scope, not of quality. If we ran a suite of nine HTML-to-JSON extractions we would expect the ranking to move, possibly hard, and we have not run one — so we cannot tell you where it would land.

This is the honest limit of a general benchmark: it can tell you whether a model is in scope for the work, and it is nearly useless for grading models that are out of scope. For the extraction-shaped side of this, our classification buyer's guide and the notes on schema-constrained output are closer to the right question than this article is.

Where it sits against models that passed

ModelScoreDerived cost / 1k tasksMean latencyOutput tokens, whole suitePriced on
Schematron v2 Turbo0/9$0.035.3s1,4302026-09-15
Ling 3.0 Flash VL9/9$0.074.4s3,2992026-09-15
DeepSeek V3.29/9$0.087.1s1,2652026-07-30
GPT-5.4 Mini9/9$0.532.3s9692026-07-30
GPT-6 Astra9/9$8.195.6s1,3532026-09-15

Read the output-token column before the score column. Schematron wrote 1,430 tokens of code; Astra wrote 1,353 and Claude Fable 5.1 wrote 1,273, both for a perfect score. Length told us nothing. Neither did latency, and neither did price.

Output tokens across the nine-task suiteSchematron v2 Turbo wrote more code than four models that solved every task. It solved none.GPT-5.4 Mini969 · 9/9DeepSeek V3.21,265 · 9/9Claude Fable 5.11,273 · 9/9GPT-6 Astra1,353 · 9/9Schematron v2 Turbo1,430 · 0/9Ling 3.0 Flash VL3,299 · 9/9One scale throughout: 0.14 px per output token, so 1,430 tokens draws 200.2 px. Counts are suite totals for the same nine tasks.
Volume of code produced, and what it scored. The blue bar is the only one that solved nothing.

Three cents buys nothing

At a derived $0.03 per 1,000 tasks (list price on 2026-09-15) Schematron is cheaper than Ling 3.0 Flash VL at $0.07, which is the cheapest model in our set that clears the suite. That comparison is a trap, and we want to be the ones to say so: cost per thousand tasks is only a meaningful number when the tasks get solved. Divide three cents by zero correct answers and you do not get a bargain, you get an undefined quantity.

The same arithmetic runs in the other direction. GPT-6 Astra's $8.19 is exactly 273 times Schematron's $0.03 ($0.03 × 273 = $8.19, both priced 2026-09-15), and on this suite the expensive one is the only one of the two that returned working code. Price tells you nothing about fit until fit is established. That ordering — capability first, then cost — is the whole argument behind our cheapest-API roundup, and it is why we publish cost next to a score rather than on its own.

One pricing note worth keeping: Schematron's output is priced at exactly five times its input ($0.15 ÷ $0.03 = 5, on 2026-09-15), so an extraction workload that reads a large document and emits a small JSON object gets a genuinely favourable shape from that ratio. That is a real advantage of the model. It is just not one our suite exercises, and list prices move.

A zero is a result. An incomplete run is not.

We hold these apart deliberately. In the same 2026-09-16 run, IBM's Granite 4.2 8B is marked excluded in our data: two of its nine tasks failed at the API layer after retries, leaving seven scored. Seven scored tasks out of a nine-task suite is not a score, so we publish none for it, and you will not see it ranked anywhere on this site.

Schematron is the opposite case. Nine tasks attempted, nine tasks graded, nine tasks wrong. That is a complete, valid, publishable measurement — which is precisely why it deserves a fair reading rather than a headline. An excluded run means we do not know. A 0/9 means we know something narrow and specific: this model, on this kind of open-ended Python generation, at temperature 0, one attempt each, produced nothing that passed. More on how we draw that line in our notes on evaluating models honestly.

It is worth saying plainly what the top of this suite cannot do, too. Of 82 entries, 56 scored 9/9, spanning $0.07 to $57.06 per 1,000 tasks. Nine self-contained Python functions cannot separate a frontier model from a competent small one — they tie at the ceiling, and the price spread between them is the only thing left to look at. The suite has real discriminating power in exactly one place: telling you whether a model is in scope at all. Schematron is the clearest data point we have ever collected on that, because it is the only one that landed on the wrong side of the line.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. API-layer failures are recorded separately from wrong answers, which is why a broken run is excluded rather than scored zero. Cost is derived from measured token counts at list price on the date stated beside each figure — it is not a billing statement. Runs go through OpenRouter. Full method on the methodology page.

What we did not measure

FAQ

Is Schematron v2 a bad model? Our data does not say that. It says it scored 0 of 9 on open-ended Python generation, which is not the task it is built for.

Did it error out or time out? No. It returned answers for all nine tasks, 1,430 output tokens in total, at a 5.3-second mean, and all nine were graded wrong.

How much does Schematron v2 Turbo cost? $0.03 per million input tokens and $0.15 output, as of 2026-09-15.

What should I use for Python generation at this price? Ling 3.0 Flash VL scored 9/9 at a derived $0.07 per 1,000 tasks (2026-09-15), and DeepSeek V3.2 scored 9/9 at $0.08 (2026-07-30).

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.