Benchmarks

Domain Specific LLM: A Three-Cent Specialist Scored 0 of 9

The cheapest of the six models we chart below is a domain specific LLM, and it answered none of the nine tasks correctly. Schematron V2 Turbo, an HTML-to-JSON extraction model, came in at a derived $0.03 per 1,000 tasks at list price on 2026-09-15 and scored 0/9. In the next row down sits Ling 3.0 Flash VL, a vision-language model, at $0.07 and 9/9 — the cheapest of the 56 models in our set that clear the suite. Two specialists, four cents apart, and the one built furthest from coding is the one that codes.

DataLLM Lab article cover: Domain Specific LLM: A Three-Cent Specialist Scored 0 of 9

A vertical model is sold on fit: narrower training, narrower job, better results inside the lane. Our suite runs outside every one of those lanes. What it caught is that the penalty for being outside your lane is not a gentle slope — it is a cliff for one model and, on another, nothing at all.

The three-cent model solves nothing

Schematron V2 Turbo is an HTML-to-JSON extraction model. Its documented interface expects the extraction instructions to arrive as a JSON schema in response_format. Our harness does not send one — it sends a Python function signature and a spec, the same prompt every other model gets. The result:

ModelBuilt forScore on our Python suiteDerived cost / 1,000 tasksMean latency
Schematron V2 TurboHTML-to-JSON extraction0/9$0.03 (priced 2026-09-15)5.3s
Ling 3.0 Flash VLVision-language9/9$0.07 (priced 2026-09-15)4.4s
Granite 4.2 8BSmall open-weight lineno valid score$0.43 (priced 2026-09-15)12.8s
GPT-6 AstraGeneral frontier9/9$8.19 (priced 2026-09-15)5.6s

It missed all nine: two_sum, valid_parentheses, merge_intervals, roman_to_int, lcs_len, flatten, top_k_words, token_bucket, parse_csv_line. Worth naming, because it rules out the failure mode we have been burned by before. A model that returns nothing — blocked by a content filter, or out of budget mid-thought — used to score in our harness as a wrong answer, and that bug cost us a retraction. Schematron emitted 1,430 output tokens across the suite against 867 in. Non-zero output, nine wrong answers. It answered; the answers did not run.

And it did so on a 5.3-second mean, quicker than every other model in the table above except Ling VL — which is the whole problem with a cost column read on its own.

Cheapest is not a scoreDerived cost per 1,000 tasks. The shortest bar is the one that solved nothing.Schematron V2 Turbo · 0/9$0.03Ling 3.0 Flash VL · 9/9$0.07DeepSeek V3.2 · 9/9$0.08Qwen3 Coder Next · 9/9$0.10GLM 5.3 Flash · 9/9$0.34Granite 4.2 8B · excluded$0.43One scale throughout: 1,200 px per dollar. Costs derived from measured tokens at each model’s own run-date list price.
Six models, one linear scale. The red bar is the cheapest and the only one with a zero.

The vision model that won on cost

Ling 3.0 Flash VL is built for native visual perception. Our suite never shows it an image. It scored 9/9, at $0.07 per 1,000 tasks derived from list price on 2026-09-15, on a 4.4-second mean, using 234 reasoning tokens per call and 741 in / 3,299 out across the suite. Of the 56 models that clear the suite it ranks 1st on cost and 10th on latency.

Set that against the top of the market. GPT-6 Astra scored the same 9 of 9 at $8.19, also priced 2026-09-15. The arithmetic is exact: $0.07 × 117 = $8.19. On this workload, a vision model nobody markets for coding delivers the identical result for one one-hundred-and-seventeenth of the derived cost. The honest reading is not that Astra is overpriced — it is that this suite cannot see whatever Astra is for. We come back to that below.

One base, four domains, one score

InclusionAI does not ship one Ling 3.0 Flash. It ships a base text model and three siblings pointed at different verticals, all served from the same catalog we benchmark through. Here is the split, and here is how much of it we can actually speak to:

VariantDomain it is tuned forWhat we have
Ling 3.0 FlashGeneral textNo valid score — two runs, neither completed
Ling 3.0 Flash FinFinance and investmentNot measured
Ling 3.0 Flash SanteHealth and medicineNot measured
Ling 3.0 Flash VLVision-language9/9 at $0.07 per 1,000 tasks, priced 2026-09-15

One of four rows carries a number. The general-purpose base — the variant a general benchmark is actually designed to measure — is the one we could not score; it failed to finish the suite in either of two attempts, and that write-up is here. The finance and health variants we have never run at all, because we have no finance eval and no medical eval, and running them on nine Python functions would tell you nothing about either domain while producing a number that looks like it did.

That is the trap in one table. A leaderboard row for Ling 3.0 Flash Fin: 9/9 would be true, publishable, and almost completely uninformative about whether it is good at finance.

When a score is not a score

Granite 4.2 8B is in our data with 7 of 9 tasks scored and marked excluded. Two tasks failed at the API layer after retries, so seven is the number of tasks that ran, not a result out of nine. We do not publish it as a score and we will not put it in a ranking next to a 9/9, because an incomplete run and a complete one are not the same measurement — the two missing tasks might have been the two it could not do. Its derived cost, $0.43 per 1,000 tasks at list price on 2026-09-15, and its 12.8-second mean are real; its accuracy is unknown. That is why it sits in the chart above as a dashed outline rather than a filled bar.

How to read a general benchmark on a specialist

Of the 70 usable entries in our set, 56 scored 9/9. That ratio is the most important thing on this page. Nine self-contained Python functions cannot separate a frontier model from a competent small one — at the top, this suite has almost no discriminating power left, and any article that ranks the frontier on it is selling you noise. What it still separates cleanly is three things: derived cost, latency, and catastrophic misfit.

So the rule we use, and the one we would suggest: use a general benchmark to rule out, never to rule in. A 0/9 is real information — it says this model will not survive contact with a prompt shaped like ours. A 9/9 says only that the model is not broken on easy, self-contained work. It is not evidence of domain quality, and for the fin and sante variants above it would not even be evidence about the domain the vendor built them for.

The specialist question is a distribution question, not a capability question. Harvey's Tenet is the clearest case we have written up: a vertical legal model whose economics came from how it is served, not from what it was trained on. If you are buying vertical, the eval that matters is the one you build on the shape of work you actually have — the same caution that applies to vision and classification picks.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. The generated code is executed against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer, which is why Granite is excluded rather than scored. Cost is derived from measured token counts at the captured list price — it is not a billing statement, and list prices move. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page, and the cost-per-task framing is laid out in the cheap coding roundup.

What we did not measure

FAQ

What is a domain specific LLM? A model post-trained or served for one narrow job — extraction, finance, medicine, vision — usually on top of a general base. InclusionAI ships finance, health and vision-language variants of one Ling 3.0 Flash base.

Are domain specific LLMs cheaper? They can be. Schematron V2 Turbo is the cheapest of the six models charted here, at a derived $0.03 per 1,000 tasks on 2026-09-15 — and scored 0/9, so price per task is not price per solved task.

Should I use a general benchmark to pick one? Only to eliminate. A 9/9 on nine self-contained Python functions says a model is not broken; it says nothing about your vertical.

Which model here would you actually run for coding? Ling 3.0 Flash VL: 9/9 at $0.07 per 1,000 tasks and a 4.4-second mean, cheapest, as of 2026-09-16, of the 56 that clear the suite. It is, uncomfortably, a vision model.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.