Benchmarks

LLM Router Benchmark: Six Routers, Nine Tasks (and a Cheap Model Beat All of Them)

This is an LLM router benchmark run on the same nine executed Python tasks every model on this site gets. We sent them through six routers: typesafe/jev-router, nvidia/switchyard, openrouter/auto-beta, openrouter/fusion, openrouter/pareto-code and openrouter/free. Five scored 9 out of 9, at API-reported costs from $0.26 to $8.46 per 1,000 tasks. The free router scored 8, after sending one coding task to a content-safety classifier. And none of the paid routers beat simply calling Solar Mini 4, which scored 9 out of 9 on its own at $0.03 per 1,000 tasks (derived, priced 2026-10-02).

DataLLM Lab article cover: LLM Router Benchmark: Six Routers, Nine Tasks (and a Cheap Model Beat All of Them)

A router sells one thing: it picks the model so you do not have to. To check that, you have to see each pick and what it cost. We already did this for one router in our Jev Router test. This page runs five more through the same harness and puts all six side by side.

All six routers in one table

The Jev Router run is from 2026-10-02. The other five ran on 2026-10-03. Cost here is the usage.cost value the API returned on each call, summed over the nine tasks and scaled to 1,000. It is not derived from list price, because routers have no list price of their own (more on that below). “Served by” is the model each response reported.

RouterPassCost / 1,000 tasks (API-reported)Median latencyServed by (tasks)
openrouter/auto-beta9/9$0.266.321sGPT-5.6 Luna (9)
typesafe/jev-router9/9$1.014.414sGPT-6 Luna (3), DeepSeek V4.1 Flash (3), GPT-6.1 Sol (2), Claude Opus 5.5 (1)
nvidia/switchyard9/9$2.485.971sClaude Opus 5.5 (5), DeepSeek V4.1 Flash (4)
openrouter/fusion9/9$7.083.598sClaude Opus 5.5 (9)
openrouter/pareto-code9/9$8.464.439sClaude Fable 5.1 (9)
openrouter/free8/9$0 (free; not ranked on cost)7.722sSix different :free models, one of them a safety classifier

Among these six, Fusion had the lowest median latency and Auto Beta the lowest cost of the five paid routers. Every paid router passed every task, so on this suite the only real differences between them are price and speed.

Switchyard's split matches how OpenRouter describes it. OpenRouter's router benchmark announcement (read 2026-10-03) says Switchyard pairs a cheaper model with a stronger one and switches between them. Here that meant DeepSeek V4.1 Flash on four tasks and Claude Opus 5.5 on five. The split is lopsided in cost. Its cheapest Opus call, $0.003052, cost more than 13 times its priciest DeepSeek call, $0.0002275.

Against one cheap model

As of 2026-10-02, Solar Mini 4 is the cheapest and the fastest of the 75 models that score 9 out of 9 on our board: $0.03 per 1,000 tasks, 2s mean, 0 reasoning tokens. Its cost is derived from measured tokens at list price, while the router figures are API-reported. Latency also differs in statistic: the router figures are medians and Solar Mini 4 is a mean, so this is not a controlled speed ranking. They come from two methods, so the comparison is approximate. The gaps are large enough that the method does not change the order.

Every paid router cost more than one cheap model called directlyCost per 1,000 tasks. All six bars scored 9/9 on the same nine executed Python tasks.Solar Mini 4, direct$0.03 · derived, priced 2026-10-02openrouter/auto-beta$0.26typesafe/jev-router$1.01 · 2026-10-02nvidia/switchyard$2.48openrouter/fusion$7.08openrouter/pareto-code$8.46Scale: width = value × 55 px. Routers API-reported, 2026-10-03 unless marked. openrouter/free omitted (8/9, free).
The Solar Mini 4 bar is 1.65 px wide. That is the point.

Pareto Code cost 282 times Solar Mini 4 ($8.46 ÷ $0.03). Fusion cost 236 times ($7.08 ÷ $0.03). Even Auto Beta, the cheapest paid router here, cost $0.26 against $0.03. If you already know your workload looks like ours, a router adds cost and adds no pass. Our Solar Mini 4 review has the single-model run, and the cheapest-LLM leaderboard shows the whole 9/9 field.

The fair defence of routing is that real workloads are not nine tasks of equal difficulty. A router earns its fee when you cannot tell in advance which requests are hard. Our suite is the case where you can, which makes it the worst case for a router.

Pareto Code defaulted to a $10 / $50 model

Pareto Code sent all nine tasks to anthropic/claude-fable-5.1. Fable 5.1 lists at $10 in and $50 out per 1M (priced 2026-10-03), 2.5 times the $4 / $20 of Claude Opus 5.5, a model Fusion, Switchyard and Jev Router all used. That is the documented default, not a fault. OpenRouter's Pareto Router docs (read 2026-10-03) say the router takes a min_coding_score and picks the cheapest model in that tier. With no score set, it goes to the top tier, the strongest coders available. Our harness sends a plain request with no routing parameters, so it got the top tier.

For reference, Fable 5.1 called directly scored 9 out of 9 on 2026-10-03 at $8.03 per 1,000 tasks, derived. Pareto's API-reported $8.46 is close to that. The lesson is simple: if you call Pareto Code without setting min_coding_score, you are buying a frontier model. The catalogue does not show this either. It lists openrouter/pareto-code, like Jev Router and Switchyard, at a placeholder of −1,000,000 per 1M tokens.

The free router and the safety classifier

openrouter/free routes only to :free models. On our nine tasks it used six different ones. It sent roman_to_int, a short coding task, to nvidia/nemotron-3.5-content-safety:free. That is a content-safety classification model, built to label text rather than write Python, and the answer failed. Every other task passed.

It was also the slowest router here. Its median was 7.722s, and individual calls took 16.5s, 16.8s and 24.0s. Free is a fine price for experiments. Just keep in mind that this router does not appear to check whether the model it picks can do the job at all. Our free LLM API guide covers the alternatives. We do not rank it on cost.

Auto Beta and the previous Luna

openrouter/auto-beta sent all nine tasks to openai/gpt-5.6-luna. By version number, that is the generation before openai/gpt-6-luna, which scored 9 out of 9 when called directly at $0.16 per 1,000 tasks (derived, priced 2026-10-02). Auto Beta's API-reported cost was $0.26.

OpenRouter's August 10, 2026 post on the new Auto Router (read 2026-10-03) says it chooses models based on what OpenRouter users have been using for the same kind of task over the past seven days. It also says openrouter/auto-beta is the channel that gets new improvements first. A popularity-weighted router would pick a model people already use over a newer one. That fits what we saw, but it is our reading, not a confirmed explanation. As with -latest aliases, the name you call tells you less than the model field in the response.

Fusion: a bill we cannot see inside

Every Fusion response reported anthropic/claude-opus-5.5. Opus 5.5 called directly costs $4.03 per 1,000 tasks (derived, priced 2026-10-02). Fusion reported $7.08, which is 1.76 times that ($7.08 ÷ $4.03), and each of its nine calls cost between $0.005631 and $0.008861. That is more than one Opus completion should cost.

OpenRouter's Fusion Router docs (read 2026-10-03) explain why. Fusion can send the prompt to a panel of models (three by default, up to eight), then to an analyst model, and then the outer model writes the final answer. The docs say the model decides whether a task needs that deliberation. They put the cost at roughly 4–5 times a single completion with the default panel. The response's model field shows only the model that handled the request.

So the served field we logged shows the final model only. We cannot see which panel models ran, or on which tasks. Our 1.76 multiple is below the documented 4–5 times, which could mean the panel ran on only some tasks, but our data cannot confirm that. Fusion had the lowest router median latency of the six at 3.598s. The separate direct Opus 5.5 run recorded a 4.8s mean; these are different statistics and do not establish that routing is faster.

What OpenRouter's own router benchmark says

On 2026-10-02, OpenRouter published Model Router Benchmarks on its blog (read 2026-10-03). It has a leaderboard at openrouter.ai/benchmarks/routers. All of the following is confirmed because it comes from OpenRouter's own announcement:

ClaimTier
Routers are run through a set of benchmarks from varied domains and scored on a 0–10 Router IndexConfirmed (OpenRouter blog)
Default weighting: quality 60%, time per task 20%, cost 20%, adjustable with slidersConfirmed
Compared: OpenRouter's Auto Router including Auto Beta, Unbiased's Pareto, Sakana's Fugu Max and Fugu Ultra v2, Typesafe's Jev Router, NVIDIA's SwitchyardConfirmed
Fusion, Pareto Code and -latest model slugs were not benchmarked, as they serve specialised purposesConfirmed
The Free Models Router is mentioned as one of OpenRouter's first routers but is not in the comparisonConfirmed
Results represent general tasks, not your own workConfirmed

The two approaches cover different ground. OpenRouter's index weights quality at 60%, and on our suite every paid router tied on quality. That is why our test comes down to cost. We also tested three routers its index does not include: Fusion and Pareto Code, which it set aside as specialised, and Free, which it names but did not benchmark. Two of our findings came from those three. Free routed a coding task to a safety classifier, and Pareto Code defaulted to a $10 / $50 model.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec, with no example tests. Generated code runs against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. API-layer failures are recorded separately from wrong answers. The free router's miss was a wrong answer.

Router cost is the usage.cost the API returned per call, summed, and we did not reconcile it against an invoice. Single-model costs (Solar Mini 4, GPT-6 Luna, Opus 5.5, Fable 5.1) are derived from measured tokens at the list price on the date shown. All runs go through OpenRouter, not through the DataLLM Lab gateway. Full method is on the methodology page, and keep in mind that prices move.

What this cannot tell you

If you are choosing a router, run a week of your own traffic through it, then price the same requests on its most-used pick alone. On our suite, that comparison went against every router.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.