Decision Model LLMs: Jev, pplx-decider and Clef (and the One Number We Measured)
A decision model LLM does not write. You hand it some state — a ticket, a JSON payload, an image — plus a question with a fixed set of answers, and it returns one of those answers with a probability attached. No prose, no rationale, and on all three products here, a price for input tokens only. TypeSafe's Jev opened the category; Cloudflare shipped open-weight rivals on 2026-10-01, and Perplexity, per AI Weekly, did the same that day. This page compares the three from their own documentation, read 2026-10-03, and is plain about the limit of our evidence: we have not measured any decision model on a decision task. Our one first-party number is a router that runs on Jev, and it scored 9 out of 9 at $1.01 per 1,000 tasks.
Most of what an application asks a language model is not a request for writing. Is this email spam? Which of six queues does this ticket belong in? Should this agent call the tool or stop? A generative model answers those by producing text that your code then has to parse. A decision model skips the text.
What a decision model is
TypeSafe calls these System One models; Sanity's glossary entry on Jev (read 2026-10-03) traces the name to the fast, intuitive mode of thought in Daniel Kahneman's Thinking, Fast and Slow. TypeSafe's homepage (read 2026-10-03) describes them as a new class of model built for decisions inside software, returning typed decisions with calibrated probabilities rather than strings.
The mechanics differ by vendor but the contract is the same. You define the question and the allowed answers in code. The model scores each allowed answer and hands back a distribution. Perplexity's Decisions API documentation (read 2026-10-03) offers three question shapes: a yes/no question returning a probability between 0 and 1, a choice among your options with a probability per option, and a score on an ordered rubric returning a probability-weighted average. Cloudflare's Clef model card on Hugging Face (read 2026-10-03) is blunter about the internals: one logit per allowed option per question, then a softmax.
Two consequences follow. The output can never be off-schema, because the only things the model can say are the options you listed. And output costs nothing, because there is barely any: Jev, pplx-decider and Clef all list a price per input token and none for output.
The three products, side by side
Every cell comes from the vendor's own page unless marked. Where a vendor did not say, neither do we.
| TypeSafe Jev 1.13 | Perplexity pplx-decider-v1-27b | Cloudflare Clef / Clef-flash | |
|---|---|---|---|
| Weights | Hosted API; we found no published weights | Open, Apache 2.0, on Hugging Face | Open, Apache 2.0, on Hugging Face |
| Base model | Not disclosed; TypeSafe cites a new architecture, sampler and training method it calls RLCD | Fine-tuned from Qwen3.8-27B; model card lists 26B parameters | Clef from Qwen3.8-27B, Clef-flash from Qwen3.5-9B |
| Context | 32,000 tokens | Under 262,144 input tokens per request | 64k tokens |
| Price | $0.042 per 1M input, $0 output | $0.04 per 1M input, output free | $0.24 per 1M input for Clef, $0.09 per 1M input for Clef-flash, on the Workers AI pricing page; no output rate listed |
| Answer types | Typed decisions with calibrated probabilities | Yes/no, choice, rubric score | Logit per allowed option, softmax to probabilities |
| Inputs | Not confirmed on the vendor page we read | Text, JSON, arrays, images | Text and images; the model card also lists JSON and video |
| Where to run it | TypeSafe API; also listed on OpenRouter | Perplexity Decisions API, or self-host (the card asks for about 49 GiB of GPU memory for weights) | Workers AI as @cf/cloudflare/clef, or self-host |
| Announced | Early access per TypeSafe's homepage | 2026-10-01, per AI Weekly | 2026-10-01, Cloudflare blog |
Sources, all read 2026-10-03: TypeSafe's homepage, which lists Jev at $42 per billion input tokens — the same rate as the $0.042 per 1M in the OpenRouter catalogue entry we recorded on 2026-10-02; Perplexity's Decisions API docs and its pplx-decider-v1-27b model card; Cloudflare's launch post by Michelle Chen, its Workers AI changelog and pricing page, and the Clef model card. The Jev context figure is from the OpenRouter catalogue, and Cloudflare's post independently gives Jev as 32k.
One detail stands out: Cloudflare describes Clef as fully Jev-API compatible. The newcomers are not just competing with Jev; at least one is built to be swapped in for it.
Confirmed, attributed, unsourced
Three tiers, because a vendor saying something and it being true are different things.
| Claim | Tier | Where it comes from |
|---|---|---|
| pplx-decider scored 85.71% against Jev's 84.51% and Qwen3.8-27B's 74.76% across 11 benchmarks | Confirmed as a vendor claim | Perplexity's model card. Vendor-run, not independently reproduced. |
| Clef median latency 209.3 ms, Clef-flash 38.8 ms, Jev 524.1 ms | Confirmed as a vendor claim | Cloudflare's launch post. Cloudflare ran both sides. |
| Clef beat Jev in 3 of 4 areas of TypeSafe's own eval suite | Confirmed as a vendor claim | Cloudflare's launch post |
| Jev is 193.6x faster and 244.6x cheaper than LLMs on its task workflows, with zero hallucinations | Confirmed as a vendor claim | TypeSafe's homepage |
| The 11-benchmark panel covered 7,210 samples | Attributed | AI Weekly; we did not find the figure on Perplexity's own pages |
| Jev still scored higher than pplx-decider on 6 of the 11 benchmarks | Attributed | KuCoin News flash |
| Jev was released on 15 September 2026 | Attributed | Sanity's glossary entry. TypeSafe's homepage shows a later page date and says early access, so we do not state a launch day. |
| Clef-flash is built on Qwen | Confirmed | Cloudflare's post names Qwen3.5-9B |
| Perplexity and Cloudflare are the start of a decision-model wave | Attributed | The AGTP account on X; a framing, not an announcement |
The unsourced tier is empty: we found no material claim in circulation without a name on it, which says more about the category's age than its rigour. Note what the confirmed tier actually confirms. Every head-to-head number we found was produced by one of the competitors. Perplexity's headline is a 1.2-point aggregate lead (85.71 − 84.51) that, if the KuCoin report is right, hides Jev winning more of the individual tests. Cloudflare's latency table puts Jev at 13.5 times Clef-flash (524.1 ÷ 38.8) in Cloudflare's own test. And “zero hallucinations” is a statement about format, not truth: a calibrated probability placed on the wrong option is still a wrong answer.
Their numbers and ours
Our benchmark scores code generation: nine Python functions, executed against hidden asserts. A decision model cannot enter it, because there is nothing to execute — it returns a probability over options you supply, not a function body. So we have no first-party accuracy, latency or cost figure for Jev, pplx-decider or Clef as decision makers.
What we do have is the router. typesafe/jev-router uses Jev to choose which generative model answers each request, and OpenRouter describes it that way on its model page (read 2026-10-02). We sent it the nine tasks on 2026-10-02 alongside four other routers we probed on 2026-10-03:
| Router | Score | Cost / 1,000 tasks (API-reported) | Median latency | Models it served |
|---|---|---|---|---|
| OpenRouter Auto Beta | 9/9 | $0.26 | 6.321s | GPT-5.6 Luna for all nine |
| Jev Router | 9/9 | $1.01 | 4.414s | GPT-6 Luna 3, DeepSeek V4.1 Flash 3, GPT-6.1 Sol 2, Claude Opus 5.5 1 |
| NVIDIA Switchyard | 9/9 | $2.48 | 5.971s | DeepSeek V4.1 Flash 4, Claude Opus 5.5 5 |
| OpenRouter Fusion | 9/9 | $7.08 | 3.598s | Claude Opus 5.5 for all nine |
| OpenRouter Pareto Code | 9/9 | $8.46 | 4.439s | Claude Fable 5.1 for all nine |
| OpenRouter Free | 8/9 | free, not ranked | 7.722s | Six different :free models |
Among the five paid routers in that probe, Jev Router was the second cheapest and the most willing to mix models. It is not evidence that Jev decides well. Every router passed, and 75 models on our board score 9 out of 9 on their own — including GPT-6 Luna at $0.16 per 1,000 tasks (priced 2026-10-02), which Jev Router itself picked three times. Nine self-contained Python functions cannot separate a frontier model from a competent small one: as of 2026-10-02 the cheapest and fastest model to score 9 out of 9 is Solar Mini 4, at $0.03 per 1,000 tasks and a 2s mean. A router choosing among models that all pass has no hard decision to get right. The full per-task breakdown is in our Jev Router test.
The free router supplied the one accidental lesson. It sent roman_to_int to nvidia/nemotron-3.5-content-safety, a content-safety classification model, and the answer failed. A model built to judge is not a model built to write, and a router that cannot tell them apart will hand a generative task to a classifier.
Where they fit against generative LLMs and routers
- Against a generative LLM doing classification. You can already get a label from any chat model, and structured output keeps it on-schema. What a decision model adds is a probability per option, near-zero output cost and, if the vendor latency claims hold, sub-second answers. What it gives up is the explanation. For the generative route and its costs, see our classification guide.
- Against a router. A router is one product built on a decision: which model should answer this? Jev Router is that decision exposed as a service. Jev 1.13 is the decision engine exposed raw, for your own questions.
- Inside an agent. Stop or continue, call the tool or not, escalate or not — the many small gates in an agent loop are where an input-only price matters most, because they fire on every step. Our note on where agent spend goes covers that arithmetic.
On list price alone, one decision over a 1,000-token state costs $0.000042 on Jev ($0.042 per 1M, as listed 2026-10-02) and $0.00004 on pplx-decider ($0.04 per 1M, as listed 2026-10-03); on Workers AI it is $0.00024 on Clef and $0.00009 on Clef-flash ($0.24 and $0.09 per 1M, as listed 2026-10-03). That is arithmetic on a rate card, not a measurement, and it ignores whatever tokens a vendor adds around your state.
How our numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer. Router costs are the usage.cost value the API returned per call, summed — not reconciled against an invoice. Single-model costs, such as GPT-6 Luna's $0.16 and Solar Mini 4's $0.03, are derived from measured tokens at list price, so the two kinds of figure are produced by different methods. Runs go through OpenRouter, not the DataLLM Lab gateway. Full method on the methodology page, and prices move.
What we did not measure
- Any decision model on a decision task. Every accuracy and latency figure for Jev, pplx-decider and Clef on this page is the vendor's or a named outlet's.
- Calibration. The central promise — that a 0.9 means right nine times in ten — needs a labelled set and many calls. We have run neither.
- Clef's real cost per decision. Workers AI lists an input rate, but we have not called Clef, so we do not know how many tokens it adds around your state.
- Routing quality. Our suite gives a router nothing hard to decide, so a 9 out of 9 says the picks were adequate, not well chosen.
If you are evaluating one of these, the useful test is yours: take a few hundred decisions your system already makes, with known right answers, and check whether the probabilities each model returns match how often it is right.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
Primary sources, checked October 3, 2026: Cloudflare: Clef decision models. Dated measurements above may differ from the current documentation.
DataLLM Lab