NVIDIA Switchyard: Nine of Nine on Two Models (and Claude Opus 5.5 Was 97% of the Bill)
NVIDIA Switchyard scored 9 out of 9 on our executed Python benchmark on 2026-10-03, called as nvidia/switchyard on OpenRouter. It used exactly two models: DeepSeek V4.1 Flash for four tasks and Claude Opus 5.5 for five. The summed cost the API reported was $2.48 per 1,000 tasks, with a median latency of 5.971 seconds. Almost all of that money went one way: the five Opus calls were 97% of the bill. On the same nine tasks, Jev Router cost $1.01 and passed everything too.
On OpenRouter, Switchyard's default setup is a two-tier router: a cheap model, an expensive one, and a policy for choosing between them. That makes it easy to audit. Every call either went cheap or went expensive, and the only question is whether the expensive calls were needed.
Every task, and where it went
We sent nvidia/switchyard the same nine prompts every model on this site gets. All nine passed their hidden asserts. The table is the complete run. Costs are the per-call usage.cost figure OpenRouter returned in each response, printed as the API gave them. The last two columns show where Jev Router sent the same task in its run on 2026-10-02.
| Task | Pass | Switchyard sent it to | Cost per call (API-reported) | Latency | Jev Router sent it to | Jev cost |
|---|---|---|---|---|---|---|
| two_sum | yes | DeepSeek V4.1 Flash | $0.0000578 | 3.6s | GPT-6 Luna | $0.0000460 |
| valid_parentheses | yes | Claude Opus 5.5 | $0.0030520 | 6.0s | DeepSeek V4.1 Flash | $0.0001734 |
| merge_intervals | yes | Claude Opus 5.5 | $0.0034560 | 4.9s | GPT-6.1 Sol | $0.0009280 |
| roman_to_int | yes | DeepSeek V4.1 Flash | $0.0001968 | 6.8s | DeepSeek V4.1 Flash | $0.0001959 |
| lcs_len | yes | Claude Opus 5.5 | $0.0037200 | 4.8s | GPT-6 Luna | $0.0000689 |
| flatten | yes | DeepSeek V4.1 Flash | $0.0002275 | 8.6s | GPT-6 Luna | $0.0000934 |
| top_k_words | yes | DeepSeek V4.1 Flash | $0.0001415 | 8.0s | DeepSeek V4.1 Flash | $0.0003732 |
| token_bucket | yes | Claude Opus 5.5 | $0.0062960 | 6.7s | GPT-6.1 Sol | $0.0020900 |
| parse_csv_line | yes | Claude Opus 5.5 | $0.0051280 | 5.3s | Claude Opus 5.5 | $0.0051280 |
| Total | 9/9 | 2 models | $0.0222757 = $2.48 per 1,000 tasks | 5.971s median | 4 models, 9/9 | $0.0090968 = $1.01 |
The split is not the one a human would draw. valid_parentheses — a stack and a lookup table — went to Opus, while flatten and top_k_words went to DeepSeek. The one task where both routers reached for Opus was parse_csv_line, and the two calls cost the identical $0.0051280. On the four other tasks Switchyard escalated, Jev Router used something cheaper and still passed.
The DeepSeek calls were quick for that model. Its own run on our board had a 15.6s mean; through Switchyard its four tasks returned in 3.6s, 6.8s, 8.6s and 8.0s. We did not log the reasoning setting the router used, so we cannot say why.
Where the money went
Add up the five Opus calls in the table and you get $0.021652. The four DeepSeek calls come to $0.0006236. Against the reported total of $0.0222757, Opus is 97.2% of the bill ($0.021652 ÷ $0.0222757). The per-call figures sum to $0.0222756, one ten-millionth of a dollar under the total, which is rounding in the per-call display. token_bucket alone is 28% of the spend.
So the cheap tier did its job — four tasks for well under a tenth of a cent combined — and was irrelevant to the total. Any two-tier router's bill is set almost entirely by how often it escalates. Here it escalated five times out of nine.
Against Jev Router and against either model alone
| Option | Score | Cost / 1,000 tasks | Cost method | Latency (statistic) |
|---|---|---|---|---|
| Always DeepSeek V4.1 Flash | 9/9 | $0.36 | Derived, list price on 2026-09-15 | 15.6s mean |
| Jev Router | 9/9 | $1.01 | API-reported, 2026-10-02 | 4.414s median |
| NVIDIA Switchyard | 9/9 | $2.48 | API-reported, 2026-10-03 | 5.971s median |
| Always Claude Opus 5.5 | 9/9 | $4.03 | Derived, list price on 2026-10-02 | 4.8s mean |
Single-model figures are means; router figures are medians. Do not read across those statistics as a like-for-like latency comparison.
Against Jev Router this is a like-for-like comparison: both figures are summed usage.cost on the same nine tasks, one day apart. Switchyard cost about 2.5 times as much ($2.48 ÷ $1.01), or $0.0131789 more per nine tasks, and was slower at the median. Both scored 9/9. The gap is the four extra Opus calls: on lcs_len, Switchyard's Opus call cost 54 times Jev's GPT-6 Luna call for the same passing result.
Against always-Opus, Switchyard cost about 62% ($2.48 ÷ $4.03). That is a real saving, but it compares two methods: the router figure is what the API reported, the single-model figure is derived from measured tokens times list price. Opus 5.5 listed at $4 in and $20 out per 1M on 2026-10-02, and is unchanged today.
Against always-DeepSeek, Switchyard cost about 6.9 times as much ($2.48 ÷ $0.36), and DeepSeek V4.1 Flash passed all nine tasks on its own — including the five Switchyard escalated. One caveat cuts against DeepSeek: its $0.36 was priced at $0.15 in and $0.6 out per 1M on 2026-09-15, and it lists at $0.3 in and $1.2 out today. Both rates doubled, so the same tokens would cost twice as much now. It would still be far below the router. The trade is latency: DeepSeek alone had a 15.6s mean.
For the floor of the board: Solar Mini 4 is the cheapest and the fastest model to score 9 out of 9 among the 75 that do, at $0.03 per 1,000 tasks and a 2s mean, priced 2026-10-02. Details for the two models Switchyard used are in our DeepSeek V4.1 Flash review and our Claude Opus 5.5 review.
What Switchyard is: the ledger
Everything in this section is third-party, read on 2026-10-03. Confirmed means the vendor's own page; attributed means a named outlet reporting it.
| Claim | Tier | Source |
|---|---|---|
| An open-source library that helps an agent choose which model handles each request; Apache 2.0; OpenAI and Anthropic API compatibility | Confirmed | NVIDIA's NVIDIA-NeMo/Switchyard GitHub README |
| Announced together with Nemotron 3.5 Lightning | Confirmed | NVIDIA's own post on X |
| Announcement date 2026-08-11 | Attributed | SiliconANGLE and VentureBeat, both dated 2026-08-11 |
Listed as nvidia/switchyard, published 2026-09-21, 1,000,000-token context, no routing fee; billed at the rate of whichever model answered | Confirmed | OpenRouter's model page and router docs (OpenRouter operates the hosted endpoint) |
With no models list, picks two candidates from the 20 models with the most OpenRouter spend over the previous seven days: an efficient tier from the cheapest fifth and a capable tier from the 60th to 80th percentile; refreshed hourly | Confirmed | OpenRouter router docs |
Default algorithm is stage, which scores existing tool results and calls a judge when indecisive; judge calls run on google/gemini-2.5-flash-lite and are billed as separate generations | Confirmed | OpenRouter router docs |
| In NVIDIA's internal tests, a Switchyard mix of open models and Opus 4.8 cut task cost to roughly a third of Opus 4.8 alone | Attributed | VentureBeat, 2026-08-11, reporting NVIDIA's own tests. The README shows a cost chart against Opus 4.8 and GLM 5.2 baselines and warns that results depend on benchmark, model pool, serving stack and configuration. |
| Kong, LiteLLM and OpenRouter built Switchyard support into their gateways | Attributed | VentureBeat. The README itself lists LiteLLM support as experimental. |
| A native Rust server installable with Cargo | Attributed | KDnuggets tutorial, 2026-09-04 |
The pair we got, a cheap DeepSeek and an expensive Opus, fits that documented two-tier design. We did not record the candidate list the router chose on the day, and the docs say the ranking refreshes hourly, so your pair may differ from ours.
NVIDIA's one-third claim, and why our suite cannot test it
We measured 62% of always-Opus; the reported NVIDIA figure is about a third. These numbers are not in conflict, because they measure different things. Switchyard is built for agents working through multi-step tasks, and its default stage algorithm reads tool results. Our nine tasks are single-turn requests with no tool calls, so that algorithm had nothing to read. Our reading of the docs is that the judge then made each call, which may explain choices like Opus for valid_parentheses; we did not log it.
That raises a billing point we could not close. The docs say judge calls are billed as ordinary generations on your account. Our $2.48 is the sum of the usage.cost returned on the nine answers. If judge calls happened, they are probably not in that figure. We did not reconcile against account activity, so treat $2.48 as the cost of the answers, not necessarily the whole cost.
The bigger limitation is our suite. Nine self-contained Python functions cannot separate a frontier model from a competent small one: 75 models on our board score 9 out of 9. A router that escalates here is paying for headroom this test never uses. A fair verdict on Switchyard needs long agent sessions, where escalation might rescue a failing run and where switching models costs a prompt cache rebuild.
How these numbers were produced
Nine Python tasks, each a function signature plus a spec and no example tests. Generated code runs against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; Switchyard had neither. Router cost is the per-call usage.cost the API returned, summed; single-model cost is derived from measured tokens times the list price on the date shown. The OpenRouter catalogue lists the router at a negative placeholder rate of −1,000,000 per 1M tokens, so the tokens-times-price method cannot price it. Runs go through OpenRouter, not the DataLLM Lab gateway. Full method on the methodology page; prices move, as DeepSeek's did.
What we did not measure
- Judge calls. Whether any were made, and what they cost.
- Agent workloads. No tool calls, no multi-turn sessions, no cache loss on switches — the conditions Switchyard is designed for.
- Routing stability. One pass, one decision per task. The candidate pair refreshes hourly; a rerun could route differently.
- Custom configuration. We tested the hosted default, not a hand-picked
modelspair or another algorithm.
If you run agents, the useful test is yours: send a week of real sessions through Switchyard, then price the same sessions on its efficient tier alone. On our suite, the cheap tier alone would have won.
Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.
Primary sources, checked October 3, 2026: NVIDIA Switchyard source; OpenRouter hosted Switchyard. Dated measurements above may differ from the current documentation.
DataLLM Lab