Schematron v2: Our First Genuine Zero (And Why It Is Not a Bug Report)
Schematron v2 Turbo scored 0 out of 9 on our executed Python benchmark. Not a crash, not a timeout, not an empty body: it answered every one of the nine tasks, spending 1,430 output tokens across the suite at a 5.3-second mean, and every single answer failed the hidden asserts. Of the 82 entries in our set this is the first clean, complete, fully-scored zero we have recorded. It is also, at a derived $0.03 per 1,000 tasks, cheaper than the cheapest model that clears the suite. Both of those facts are true and neither one is a reason to use it or avoid it, because we measured a specialist on work it was never shipped to do.
Most bad benchmark results are infrastructure stories: a rate limit, a truncated body, a provider returning nothing. This one is not. Schematron v2 Turbo showed up, wrote code for all nine problems, and was wrong nine times.
The result
| Metric | Schematron v2 Turbo |
|---|---|
| Score | 0/9 |
| Derived cost / 1,000 tasks | $0.03 (list price on 2026-09-15) |
| Mean latency | 5.3s |
| Reasoning tokens per call | 0 |
| Tokens across the suite | 867 in / 1,430 out |
| List price in / out | $0.03 / $0.15 per 1M (2026-09-15) |
| Context window | 128,000 |
| Measured | 2026-09-16 |
The tasks it missed are the whole suite: two_sum, valid_parentheses, merge_intervals, roman_to_int, lcs_len, flatten, top_k_words, token_bucket, parse_csv_line. Nine for nine.
What a real zero looks like
We have been burned by fake zeros before. A provider content filter used to hand our harness an empty completion, and the grader scored that empty string as a wrong answer instead of an infrastructure failure — the write-up is in the post on the content-filter bug. The standing rule since then is that any model which does not score full marks gets re-checked before we publish anything about it.
Schematron passes that check, which is exactly why the zero stands:
- It emitted output on the suite — 1,430 output tokens across nine calls. Empty-body failures produce nothing to grade; this produced something to grade every time.
- It did not run out of room. Our cap is 4,000 tokens per call and its total across all nine tasks is 1,430, so no answer was cut off mid-function.
- It did not stall. A 5.3-second mean is faster than GPT-6 Astra at 5.6s, which scored 9/9.
- It did not think itself into a hole. 0 reasoning tokens per call. It answered directly, the way a model tuned for a narrow transformation does.
So the failure is semantic, not operational. The code came back, ran, and did the wrong thing.
Specialist, not defective
Schematron is not a general coding model and was never sold as one. The name is the product: it is built to take messy input — HTML, free text — and emit JSON conforming to a schema the caller supplies. That is a narrow, extremely useful transformation, and models trained hard on it tend to behave exactly as this one did: fast, no visible chain of thought, cheap, and completely uninterested in the open-ended instruction you actually sent.
Our suite hands a model a function signature plus a prose spec and asks it to invent the body. There is no schema to fill, no source document to parse. A 0/9 here is a measurement of scope, not of quality. If we ran a suite of nine HTML-to-JSON extractions we would expect the ranking to move, possibly hard, and we have not run one — so we cannot tell you where it would land.
This is the honest limit of a general benchmark: it can tell you whether a model is in scope for the work, and it is nearly useless for grading models that are out of scope. For the extraction-shaped side of this, our classification buyer's guide and the notes on schema-constrained output are closer to the right question than this article is.
Where it sits against models that passed
| Model | Score | Derived cost / 1k tasks | Mean latency | Output tokens, whole suite | Priced on |
|---|---|---|---|---|---|
| Schematron v2 Turbo | 0/9 | $0.03 | 5.3s | 1,430 | 2026-09-15 |
| Ling 3.0 Flash VL | 9/9 | $0.07 | 4.4s | 3,299 | 2026-09-15 |
| DeepSeek V3.2 | 9/9 | $0.08 | 7.1s | 1,265 | 2026-07-30 |
| GPT-5.4 Mini | 9/9 | $0.53 | 2.3s | 969 | 2026-07-30 |
| GPT-6 Astra | 9/9 | $8.19 | 5.6s | 1,353 | 2026-09-15 |
Read the output-token column before the score column. Schematron wrote 1,430 tokens of code; Astra wrote 1,353 and Claude Fable 5.1 wrote 1,273, both for a perfect score. Length told us nothing. Neither did latency, and neither did price.
Three cents buys nothing
At a derived $0.03 per 1,000 tasks (list price on 2026-09-15) Schematron is cheaper than Ling 3.0 Flash VL at $0.07, which is the cheapest model in our set that clears the suite. That comparison is a trap, and we want to be the ones to say so: cost per thousand tasks is only a meaningful number when the tasks get solved. Divide three cents by zero correct answers and you do not get a bargain, you get an undefined quantity.
The same arithmetic runs in the other direction. GPT-6 Astra's $8.19 is exactly 273 times Schematron's $0.03 ($0.03 × 273 = $8.19, both priced 2026-09-15), and on this suite the expensive one is the only one of the two that returned working code. Price tells you nothing about fit until fit is established. That ordering — capability first, then cost — is the whole argument behind our cheapest-API roundup, and it is why we publish cost next to a score rather than on its own.
One pricing note worth keeping: Schematron's output is priced at exactly five times its input ($0.15 ÷ $0.03 = 5, on 2026-09-15), so an extraction workload that reads a large document and emits a small JSON object gets a genuinely favourable shape from that ratio. That is a real advantage of the model. It is just not one our suite exercises, and list prices move.
A zero is a result. An incomplete run is not.
We hold these apart deliberately. In the same 2026-09-16 run, IBM's Granite 4.2 8B is marked excluded in our data: two of its nine tasks failed at the API layer after retries, leaving seven scored. Seven scored tasks out of a nine-task suite is not a score, so we publish none for it, and you will not see it ranked anywhere on this site.
Schematron is the opposite case. Nine tasks attempted, nine tasks graded, nine tasks wrong. That is a complete, valid, publishable measurement — which is precisely why it deserves a fair reading rather than a headline. An excluded run means we do not know. A 0/9 means we know something narrow and specific: this model, on this kind of open-ended Python generation, at temperature 0, one attempt each, produced nothing that passed. More on how we draw that line in our notes on evaluating models honestly.
It is worth saying plainly what the top of this suite cannot do, too. Of 82 entries, 56 scored 9/9, spanning $0.07 to $57.06 per 1,000 tasks. Nine self-contained Python functions cannot separate a frontier model from a competent small one — they tie at the ceiling, and the price spread between them is the only thing left to look at. The suite has real discriminating power in exactly one place: telling you whether a model is in scope at all. Schematron is the clearest data point we have ever collected on that, because it is the only one that landed on the wrong side of the line.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. API-layer failures are recorded separately from wrong answers, which is why a broken run is excluded rather than scored zero. Cost is derived from measured token counts at list price on the date stated beside each figure — it is not a billing statement. Runs go through OpenRouter. Full method on the methodology page.
What we did not measure
- Structured extraction, the one job this model exists to do. We have no HTML-to-JSON suite, so we have no number for the thing you would actually buy it for. This is the single largest gap in this article and we are not going to dress it up.
- Schema-constrained decoding. We sent plain prompts with no response format and no grammar. A model tuned for constrained output may behave very differently when it is actually constrained, and we did not try.
- Prompt shape. One phrasing, temperature 0, one attempt. A specialist is more sensitive to prompt format than a generalist, and a system prompt written for this model might well have changed the shape of the failures — though nine wrong answers is a deep enough hole that we doubt formatting alone closes it.
- What the wrong answers were wrong about. We record pass or fail against hidden asserts, not a diff of failure modes. We cannot tell you whether it misread the specs, emitted JSON where code was wanted, or wrote plausible code with broken logic.
- The 128,000-token context. Our prompts are short, and long-document extraction is where a context window like that earns its keep.
- Repeat runs. One scored attempt per task, same as every other model in the set.
FAQ
Is Schematron v2 a bad model? Our data does not say that. It says it scored 0 of 9 on open-ended Python generation, which is not the task it is built for.
Did it error out or time out? No. It returned answers for all nine tasks, 1,430 output tokens in total, at a 5.3-second mean, and all nine were graded wrong.
How much does Schematron v2 Turbo cost? $0.03 per million input tokens and $0.15 output, as of 2026-09-15.
What should I use for Python generation at this price? Ling 3.0 Flash VL scored 9/9 at a derived $0.07 per 1,000 tasks (2026-09-15), and DeepSeek V3.2 scored 9/9 at $0.08 (2026-07-30).
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab