Model Reviews

Jev Router: Nine of Nine Across Four Models (and One Task Was 56% of the Bill)

Jev Router scored 9 out of 9 on our executed Python benchmark on 2026-10-02 — and to get there it used four different models for nine tasks: GPT-6 Luna three times, DeepSeek V4.1 Flash three times, GPT-6.1 Sol twice and Claude Opus 5.5 once. The summed cost the API reported was $1.01 per 1,000 tasks. The single task it handed to Opus 5.5 was 56% of that bill — on a task that GPT-6 Luna, one of the router's own two most-used picks, had already passed on its own run that same day for $0.16 per 1,000 across all nine.

DataLLM Lab article cover: Jev Router: Nine of Nine Across Four Models (and One Task Was 56% of the Bill)

A router is a bet that someone else can pick the model better than you can. The only way to check the bet is to look at each pick and what it cost, so that is what this page is: nine requests, nine routing decisions, nine line items.

Every task, and where it went

We sent typesafe/jev-router the same nine prompts every other model on this site gets. All nine passed their hidden asserts. The table below is the complete run — no task omitted. Costs are the per-call figure OpenRouter returned in each response's usage.cost field, printed exactly as the API gave them.

TaskPassServed byCost per call (API-reported)Latency
two_sumyesGPT-6 Luna$0.0000463.3s
valid_parenthesesyesDeepSeek V4.1 Flash$0.00017341.8s
merge_intervalsyesGPT-6.1 Sol$0.0009284.4s
roman_to_intyesDeepSeek V4.1 Flash$0.00019592.3s
lcs_lenyesGPT-6 Luna$0.00006895.9s
flattenyesGPT-6 Luna$0.00009344.7s
top_k_wordsyesDeepSeek V4.1 Flash$0.00037322.0s
token_bucketyesGPT-6.1 Sol$0.002097.7s
parse_csv_lineyesClaude Opus 5.5$0.0051284.9s
Total9/94 models$0.009097 = $1.01 per 1,000 tasks—

The routing is not random-looking. The two tasks with the most state to manage — merge_intervals and token_bucket — went to GPT-6.1 Sol. The one task with real edge-case risk, parse_csv_line, went to the priciest of its four picks. That is a sensible instinct: among the models with complete runs we added to the board on 2026-10-02, every one that dropped a task dropped parse_csv_line — Codestral 2508, Qwen3 Coder Plus, Qwen3 Coder Flash and GLM-5.3 Prime all missed it.

One detail we did not expect. In our direct run, DeepSeek V4.1 Flash had a 15.6s mean and spent 478 reasoning tokens per call. Through the router its three tasks came back in 1.8s, 2.0s and 2.3s. OpenRouter says the router picks a reasoning effort as well as a model, so a low effort setting is the obvious explanation — but we did not record which effort it chose, so treat that as a hypothesis, not a finding.

Who paid for what

The router landed between its cheapest and priciest picksCost per 1,000 tasks, all 9/9 on the same nine executed Python tasks.Jev Router$1.01 · API-reported, 2026-10-02GPT-6 Luna alone$0.16DeepSeek V4.1 Flash alone$0.36 · priced 2026-09-15GPT-6.1 Sol alone$1.61Claude Opus 5.5 alone$4.03THE NINE ROUTER CALLS, STACKED · TOTAL $0.009097token_bucket $0.00209parse_csv_line · Opus 5.5 · 56%$0.009097Light blue GPT-6 Luna · mid grey DeepSeek V4.1 Flash · dark grey GPT-6.1 Sol · blue Claude Opus 5.5Scales: top 140 px per dollar per 1,000 tasks; bottom 66,000 px per dollar (66 px per $0.001). Every width = value × scale.
The router beat always-Sol and always-Opus. It did not beat always-Luna, which also scored 9/9.

The bill is lopsided. parse_csv_line alone cost $0.005128 of the $0.009097 total — 56%. The other eight tasks together cost $0.009097 − $0.005128 = $0.003969. Add token_bucket at $0.00209 and two tasks out of nine account for most of the spend.

That is the honest shape of routing economics: a router does not shave a little off every call. It sends most calls somewhere cheap and occasionally sends one somewhere expensive, and the expensive ones dominate. Whether that one Opus call was worth it is a question our suite can partly answer — see the next section.

Against always using one of its picks

Every model the router chose has its own full nine-task run on our board. All four scored 9 out of 9 on their own:

OptionScoreCost / 1,000 tasksPriced onMean latency (direct models)
Always GPT-6 Luna9/9$0.162026-10-024.9s
Always DeepSeek V4.1 Flash9/9$0.362026-09-1515.6s
Jev Router9/9$1.012026-10-02 (API-reported)—
Always GPT-6.1 Sol9/9$1.612026-10-025.9s
Always Claude Opus 5.59/9$4.032026-10-024.8s

Read plainly: on this suite the router cost about a quarter of always-Opus ($1.01 against $4.03) and less than always-Sol ($1.01 against $1.61). But it cost 6.3 times what always-Luna cost ($1.01 ÷ $0.16), and GPT-6 Luna passed all nine tasks by itself — including the parse_csv_line the router escalated to Opus. If you already knew your workload looked like ours, the router bought you nothing that one of its own picks did not deliver alone.

For the floor of the whole board: Solar Mini 4 is the cheapest and fastest model to score 9 out of 9 as of 2026-10-02, at $0.03 per 1,000 tasks and a 2s mean. The router is about 34 times that ($1.01 ÷ $0.03).

Two caveats keep this comparison fair. First, the router figure is the API-reported usage.cost, while the four single-model figures are derived from measured tokens at list price — two methods, not one. Second, DeepSeek V4.1 Flash has repriced since its run: it was priced at $0.15 in and $0.6 out per 1M on 2026-09-15 and lists at $0.015 in and $1.2 out today, so the router's DeepSeek calls were charged at a rate our $0.36 does not reflect. Neither caveat changes the order of the table; both mean the gaps are approximate.

The fair defence of routing is that nobody's workload is nine identical-difficulty functions. A router earns its keep when you cannot know in advance which requests are hard. Our suite is the case where you can, which is the worst case for a router. Our cheap coding roundup covers the single-model alternatives in depth.

The minus-one-million price

The OpenRouter catalogue lists typesafe/jev-router at −1,000,000 per 1M tokens, in and out. That is not a price; it is a placeholder meaning the cost depends on whichever model the router picks. openrouter/pareto-code, another router, carries the same placeholder.

This matters if you compute costs the way we normally do. Our pipeline multiplies measured tokens by the catalogue rate. Run that against a negative sentinel and you get a large negative number — a router that apparently pays you. That is why this is the one page on the site where the cost column comes from usage.cost rather than tokens times list price. If your own spend dashboard reads prices from the catalogue, check how it handles a negative rate before you point it at a router.

Jev Router is not Jev 1.13

Two TypeSafe entries sit side by side in the catalogue and they do different jobs. According to OpenRouter's model page and launch announcement (third-party, read 2026-10-02), Jev Router picks a model and a reasoning effort per request, and runs on Jev, which TypeSafe describes as its first System One model.

typesafe/jev-1.13 is that underlying decision model exposed on its own. Per the OpenRouter catalogue (read 2026-10-02), it returns a typed choice rather than free text, lists at $0.042 per 1M input and $0 output, and has a 32,000-token context. It does not write code, so it cannot be scored on our harness at all — there is nothing to execute. If you want Jev to make decisions inside your own pipeline, that is the model; if you want Jev to choose a coding model for you, that is the router.

The cache caveat we could not test

The most important cost claim about Jev Router is one our benchmark is structurally unable to check. OpenRouter's own documentation (third-party, read 2026-10-02) warns that switching models mid-conversation can make the new model re-read the full conversation at full price, because the previous model's prompt cache does not carry over. The router is described as cache-aware: it is meant to stay on a model that is working and switch only when the expected gain outweighs that loss.

Our nine tasks are nine independent single-turn requests. There is no conversation to cache, so there is no cache to lose, and every routing decision we observed was free of that trade-off. In a long agent session the arithmetic is different: one switch on a large context can cost more than many cheap turns combined. We explain the mechanics in our prompt caching guide, and why the long-session bill behaves differently in reasoning tokens decide the agent bill. Until someone measures a multi-turn session through the router, treat our $1.01 as the single-turn number only.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code executes against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; the router had neither. For the four single-model runs, cost is derived from measured token counts at the list price captured on the date shown, not a billing statement. For the router, cost is the usage.cost value the API returned per call, summed — we did not reconcile it against an invoice. Runs go through OpenRouter, not through the DataLLM Lab gateway. Full method on the methodology page, and remember that prices move.

What we did not measure

If you are deciding whether to use it, the useful test is not ours. Run a week of your real traffic through it, then price the same requests on its most-used pick alone. Our run says that comparison can go against the router; yours might not.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.