LLM Router Benchmark: Six Routers, Nine Tasks (and a Cheap Model Beat All of Them)
This is an LLM router benchmark run on the same nine executed Python tasks every model on this site gets. We sent them through six routers: typesafe/jev-router, nvidia/switchyard, openrouter/auto-beta, openrouter/fusion, openrouter/pareto-code and openrouter/free. Five scored 9 out of 9, at API-reported costs from $0.26 to $8.46 per 1,000 tasks. The free router scored 8, after sending one coding task to a content-safety classifier. And none of the paid routers beat simply calling Solar Mini 4, which scored 9 out of 9 on its own at $0.03 per 1,000 tasks (derived, priced 2026-10-02).
A router sells one thing: it picks the model so you do not have to. To check that, you have to see each pick and what it cost. We already did this for one router in our Jev Router test. This page runs five more through the same harness and puts all six side by side.
All six routers in one table
The Jev Router run is from 2026-10-02. The other five ran on 2026-10-03. Cost here is the usage.cost value the API returned on each call, summed over the nine tasks and scaled to 1,000. It is not derived from list price, because routers have no list price of their own (more on that below). “Served by” is the model each response reported.
| Router | Pass | Cost / 1,000 tasks (API-reported) | Median latency | Served by (tasks) |
|---|---|---|---|---|
openrouter/auto-beta | 9/9 | $0.26 | 6.321s | GPT-5.6 Luna (9) |
typesafe/jev-router | 9/9 | $1.01 | 4.414s | GPT-6 Luna (3), DeepSeek V4.1 Flash (3), GPT-6.1 Sol (2), Claude Opus 5.5 (1) |
nvidia/switchyard | 9/9 | $2.48 | 5.971s | Claude Opus 5.5 (5), DeepSeek V4.1 Flash (4) |
openrouter/fusion | 9/9 | $7.08 | 3.598s | Claude Opus 5.5 (9) |
openrouter/pareto-code | 9/9 | $8.46 | 4.439s | Claude Fable 5.1 (9) |
openrouter/free | 8/9 | $0 (free; not ranked on cost) | 7.722s | Six different :free models, one of them a safety classifier |
Among these six, Fusion had the lowest median latency and Auto Beta the lowest cost of the five paid routers. Every paid router passed every task, so on this suite the only real differences between them are price and speed.
Switchyard's split matches how OpenRouter describes it. OpenRouter's router benchmark announcement (read 2026-10-03) says Switchyard pairs a cheaper model with a stronger one and switches between them. Here that meant DeepSeek V4.1 Flash on four tasks and Claude Opus 5.5 on five. The split is lopsided in cost. Its cheapest Opus call, $0.003052, cost more than 13 times its priciest DeepSeek call, $0.0002275.
Against one cheap model
As of 2026-10-02, Solar Mini 4 is the cheapest and the fastest of the 75 models that score 9 out of 9 on our board: $0.03 per 1,000 tasks, 2s mean, 0 reasoning tokens. Its cost is derived from measured tokens at list price, while the router figures are API-reported. Latency also differs in statistic: the router figures are medians and Solar Mini 4 is a mean, so this is not a controlled speed ranking. They come from two methods, so the comparison is approximate. The gaps are large enough that the method does not change the order.
Pareto Code cost 282 times Solar Mini 4 ($8.46 ÷ $0.03). Fusion cost 236 times ($7.08 ÷ $0.03). Even Auto Beta, the cheapest paid router here, cost $0.26 against $0.03. If you already know your workload looks like ours, a router adds cost and adds no pass. Our Solar Mini 4 review has the single-model run, and the cheapest-LLM leaderboard shows the whole 9/9 field.
The fair defence of routing is that real workloads are not nine tasks of equal difficulty. A router earns its fee when you cannot tell in advance which requests are hard. Our suite is the case where you can, which makes it the worst case for a router.
Pareto Code defaulted to a $10 / $50 model
Pareto Code sent all nine tasks to anthropic/claude-fable-5.1. Fable 5.1 lists at $10 in and $50 out per 1M (priced 2026-10-03), 2.5 times the $4 / $20 of Claude Opus 5.5, a model Fusion, Switchyard and Jev Router all used. That is the documented default, not a fault. OpenRouter's Pareto Router docs (read 2026-10-03) say the router takes a min_coding_score and picks the cheapest model in that tier. With no score set, it goes to the top tier, the strongest coders available. Our harness sends a plain request with no routing parameters, so it got the top tier.
For reference, Fable 5.1 called directly scored 9 out of 9 on 2026-10-03 at $8.03 per 1,000 tasks, derived. Pareto's API-reported $8.46 is close to that. The lesson is simple: if you call Pareto Code without setting min_coding_score, you are buying a frontier model. The catalogue does not show this either. It lists openrouter/pareto-code, like Jev Router and Switchyard, at a placeholder of −1,000,000 per 1M tokens.
The free router and the safety classifier
openrouter/free routes only to :free models. On our nine tasks it used six different ones. It sent roman_to_int, a short coding task, to nvidia/nemotron-3.5-content-safety:free. That is a content-safety classification model, built to label text rather than write Python, and the answer failed. Every other task passed.
It was also the slowest router here. Its median was 7.722s, and individual calls took 16.5s, 16.8s and 24.0s. Free is a fine price for experiments. Just keep in mind that this router does not appear to check whether the model it picks can do the job at all. Our free LLM API guide covers the alternatives. We do not rank it on cost.
Auto Beta and the previous Luna
openrouter/auto-beta sent all nine tasks to openai/gpt-5.6-luna. By version number, that is the generation before openai/gpt-6-luna, which scored 9 out of 9 when called directly at $0.16 per 1,000 tasks (derived, priced 2026-10-02). Auto Beta's API-reported cost was $0.26.
OpenRouter's August 10, 2026 post on the new Auto Router (read 2026-10-03) says it chooses models based on what OpenRouter users have been using for the same kind of task over the past seven days. It also says openrouter/auto-beta is the channel that gets new improvements first. A popularity-weighted router would pick a model people already use over a newer one. That fits what we saw, but it is our reading, not a confirmed explanation. As with -latest aliases, the name you call tells you less than the model field in the response.
Fusion: a bill we cannot see inside
Every Fusion response reported anthropic/claude-opus-5.5. Opus 5.5 called directly costs $4.03 per 1,000 tasks (derived, priced 2026-10-02). Fusion reported $7.08, which is 1.76 times that ($7.08 ÷ $4.03), and each of its nine calls cost between $0.005631 and $0.008861. That is more than one Opus completion should cost.
OpenRouter's Fusion Router docs (read 2026-10-03) explain why. Fusion can send the prompt to a panel of models (three by default, up to eight), then to an analyst model, and then the outer model writes the final answer. The docs say the model decides whether a task needs that deliberation. They put the cost at roughly 4–5 times a single completion with the default panel. The response's model field shows only the model that handled the request.
So the served field we logged shows the final model only. We cannot see which panel models ran, or on which tasks. Our 1.76 multiple is below the documented 4–5 times, which could mean the panel ran on only some tasks, but our data cannot confirm that. Fusion had the lowest router median latency of the six at 3.598s. The separate direct Opus 5.5 run recorded a 4.8s mean; these are different statistics and do not establish that routing is faster.
What OpenRouter's own router benchmark says
On 2026-10-02, OpenRouter published Model Router Benchmarks on its blog (read 2026-10-03). It has a leaderboard at openrouter.ai/benchmarks/routers. All of the following is confirmed because it comes from OpenRouter's own announcement:
| Claim | Tier |
|---|---|
| Routers are run through a set of benchmarks from varied domains and scored on a 0–10 Router Index | Confirmed (OpenRouter blog) |
| Default weighting: quality 60%, time per task 20%, cost 20%, adjustable with sliders | Confirmed |
| Compared: OpenRouter's Auto Router including Auto Beta, Unbiased's Pareto, Sakana's Fugu Max and Fugu Ultra v2, Typesafe's Jev Router, NVIDIA's Switchyard | Confirmed |
Fusion, Pareto Code and -latest model slugs were not benchmarked, as they serve specialised purposes | Confirmed |
| The Free Models Router is mentioned as one of OpenRouter's first routers but is not in the comparison | Confirmed |
| Results represent general tasks, not your own work | Confirmed |
The two approaches cover different ground. OpenRouter's index weights quality at 60%, and on our suite every paid router tied on quality. That is why our test comes down to cost. We also tested three routers its index does not include: Fusion and Pareto Code, which it set aside as specialised, and Free, which it names but did not benchmark. Two of our findings came from those three. Free routed a coding task to a safety classifier, and Pareto Code defaulted to a $10 / $50 model.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec, with no example tests. Generated code runs against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. API-layer failures are recorded separately from wrong answers. The free router's miss was a wrong answer.
Router cost is the usage.cost the API returned per call, summed, and we did not reconcile it against an invoice. Single-model costs (Solar Mini 4, GPT-6 Luna, Opus 5.5, Fable 5.1) are derived from measured tokens at the list price on the date shown. All runs go through OpenRouter, not through the DataLLM Lab gateway. Full method is on the methodology page, and keep in mind that prices move.
What this cannot tell you
- Whether a router can tell hard from easy. Nine self-contained Python functions cannot separate a frontier model from a competent small one: 75 models on our board score 9 out of 9. A router has nothing to discriminate on here.
- Cache loss on model switches. Every task is a single turn. OpenRouter's own announcement names cache rebuilds as a cost of switching, and our suite never triggers one.
- Routing stability. We made one pass with one decision per task. The same prompt may route somewhere else next time.
- Routing parameters. We sent no
min_coding_score, cost tier or panel settings. Every router ran on its defaults. - Fusion's panel. We logged only the final model.
If you are choosing a router, run a week of your own traffic through it, then price the same requests on its most-used pick alone. On our suite, that comparison went against every router.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
DataLLM Lab