Domain Specific LLM: A Three-Cent Specialist Scored 0 of 9
The cheapest of the six models we chart below is a domain specific LLM, and it answered none of the nine tasks correctly. Schematron V2 Turbo, an HTML-to-JSON extraction model, came in at a derived $0.03 per 1,000 tasks at list price on 2026-09-15 and scored 0/9. In the next row down sits Ling 3.0 Flash VL, a vision-language model, at $0.07 and 9/9 — the cheapest of the 56 models in our set that clear the suite. Two specialists, four cents apart, and the one built furthest from coding is the one that codes.
A vertical model is sold on fit: narrower training, narrower job, better results inside the lane. Our suite runs outside every one of those lanes. What it caught is that the penalty for being outside your lane is not a gentle slope — it is a cliff for one model and, on another, nothing at all.
The three-cent model solves nothing
Schematron V2 Turbo is an HTML-to-JSON extraction model. Its documented interface expects the extraction instructions to arrive as a JSON schema in response_format. Our harness does not send one — it sends a Python function signature and a spec, the same prompt every other model gets. The result:
| Model | Built for | Score on our Python suite | Derived cost / 1,000 tasks | Mean latency |
|---|---|---|---|---|
| Schematron V2 Turbo | HTML-to-JSON extraction | 0/9 | $0.03 (priced 2026-09-15) | 5.3s |
| Ling 3.0 Flash VL | Vision-language | 9/9 | $0.07 (priced 2026-09-15) | 4.4s |
| Granite 4.2 8B | Small open-weight line | no valid score | $0.43 (priced 2026-09-15) | 12.8s |
| GPT-6 Astra | General frontier | 9/9 | $8.19 (priced 2026-09-15) | 5.6s |
It missed all nine: two_sum, valid_parentheses, merge_intervals, roman_to_int, lcs_len, flatten, top_k_words, token_bucket, parse_csv_line. Worth naming, because it rules out the failure mode we have been burned by before. A model that returns nothing — blocked by a content filter, or out of budget mid-thought — used to score in our harness as a wrong answer, and that bug cost us a retraction. Schematron emitted 1,430 output tokens across the suite against 867 in. Non-zero output, nine wrong answers. It answered; the answers did not run.
And it did so on a 5.3-second mean, quicker than every other model in the table above except Ling VL — which is the whole problem with a cost column read on its own.
The vision model that won on cost
Ling 3.0 Flash VL is built for native visual perception. Our suite never shows it an image. It scored 9/9, at $0.07 per 1,000 tasks derived from list price on 2026-09-15, on a 4.4-second mean, using 234 reasoning tokens per call and 741 in / 3,299 out across the suite. Of the 56 models that clear the suite it ranks 1st on cost and 10th on latency.
Set that against the top of the market. GPT-6 Astra scored the same 9 of 9 at $8.19, also priced 2026-09-15. The arithmetic is exact: $0.07 × 117 = $8.19. On this workload, a vision model nobody markets for coding delivers the identical result for one one-hundred-and-seventeenth of the derived cost. The honest reading is not that Astra is overpriced — it is that this suite cannot see whatever Astra is for. We come back to that below.
One base, four domains, one score
InclusionAI does not ship one Ling 3.0 Flash. It ships a base text model and three siblings pointed at different verticals, all served from the same catalog we benchmark through. Here is the split, and here is how much of it we can actually speak to:
| Variant | Domain it is tuned for | What we have |
|---|---|---|
| Ling 3.0 Flash | General text | No valid score — two runs, neither completed |
| Ling 3.0 Flash Fin | Finance and investment | Not measured |
| Ling 3.0 Flash Sante | Health and medicine | Not measured |
| Ling 3.0 Flash VL | Vision-language | 9/9 at $0.07 per 1,000 tasks, priced 2026-09-15 |
One of four rows carries a number. The general-purpose base — the variant a general benchmark is actually designed to measure — is the one we could not score; it failed to finish the suite in either of two attempts, and that write-up is here. The finance and health variants we have never run at all, because we have no finance eval and no medical eval, and running them on nine Python functions would tell you nothing about either domain while producing a number that looks like it did.
That is the trap in one table. A leaderboard row for Ling 3.0 Flash Fin: 9/9 would be true, publishable, and almost completely uninformative about whether it is good at finance.
When a score is not a score
Granite 4.2 8B is in our data with 7 of 9 tasks scored and marked excluded. Two tasks failed at the API layer after retries, so seven is the number of tasks that ran, not a result out of nine. We do not publish it as a score and we will not put it in a ranking next to a 9/9, because an incomplete run and a complete one are not the same measurement — the two missing tasks might have been the two it could not do. Its derived cost, $0.43 per 1,000 tasks at list price on 2026-09-15, and its 12.8-second mean are real; its accuracy is unknown. That is why it sits in the chart above as a dashed outline rather than a filled bar.
How to read a general benchmark on a specialist
Of the 70 usable entries in our set, 56 scored 9/9. That ratio is the most important thing on this page. Nine self-contained Python functions cannot separate a frontier model from a competent small one — at the top, this suite has almost no discriminating power left, and any article that ranks the frontier on it is selling you noise. What it still separates cleanly is three things: derived cost, latency, and catastrophic misfit.
So the rule we use, and the one we would suggest: use a general benchmark to rule out, never to rule in. A 0/9 is real information — it says this model will not survive contact with a prompt shaped like ours. A 9/9 says only that the model is not broken on easy, self-contained work. It is not evidence of domain quality, and for the fin and sante variants above it would not even be evidence about the domain the vendor built them for.
The specialist question is a distribution question, not a capability question. Harvey's Tenet is the clearest case we have written up: a vertical legal model whose economics came from how it is served, not from what it was trained on. If you are buying vertical, the eval that matters is the one you build on the shape of work you actually have — the same caution that applies to vision and classification picks.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. The generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why Granite is excluded rather than scored. Cost is derived from measured token counts at the captured list price — it is not a billing statement, and list prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page, and the cost-per-task framing is laid out in the cheap coding roundup.
What we did not measure
- Schematron in its own interface. It expects a JSON schema in
response_format; we sent it a coding prompt. The 0/9 measures misfit against our harness, not extraction quality. We have not built an HTML-to-JSON eval, so we cannot tell you whether it is good at its actual job — only that it is not a general model at three cents. - Ling 3.0 Flash Fin and Ling 3.0 Flash Sante. Never run. We have no finance or medical eval, and we declined to publish a Python score for them for the reason given above.
- Vision, on the vision model. Ling 3.0 Flash VL earned its 9/9 on the one modality it was not primarily built for.
- Why Granite's two calls failed. Capacity, filtering, or something else — we logged the failure, not the cause.
- Repeat runs. One scored attempt per task, across all of the above. A four-cent gap between two models is inside single-run noise; a 0/9 versus a 9/9 is not.
- Long context. Our prompts are short, so Schematron's 128,000-token window and Ling VL's 131,072 are both untested here — and long HTML pages are exactly where an extraction model would be under load.
FAQ
What is a domain specific LLM? A model post-trained or served for one narrow job — extraction, finance, medicine, vision — usually on top of a general base. InclusionAI ships finance, health and vision-language variants of one Ling 3.0 Flash base.
Are domain specific LLMs cheaper? They can be. Schematron V2 Turbo is the cheapest of the six models charted here, at a derived $0.03 per 1,000 tasks on 2026-09-15 — and scored 0/9, so price per task is not price per solved task.
Should I use a general benchmark to pick one? Only to eliminate. A 9/9 on nine self-contained Python functions says a model is not broken; it says nothing about your vertical.
Which model here would you actually run for coding? Ling 3.0 Flash VL: 9/9 at $0.07 per 1,000 tasks and a 4.4-second mean, cheapest, as of 2026-09-16, of the 56 that clear the suite. It is, uncomfortably, a vision model.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab