Model Reviews

NVIDIA Switchyard: Nine of Nine on Two Models (and Claude Opus 5.5 Was 97% of the Bill)

NVIDIA Switchyard scored 9 out of 9 on our executed Python benchmark on 2026-10-03, called as nvidia/switchyard on OpenRouter. It used exactly two models: DeepSeek V4.1 Flash for four tasks and Claude Opus 5.5 for five. The summed cost the API reported was $2.48 per 1,000 tasks, with a median latency of 5.971 seconds. Almost all of that money went one way: the five Opus calls were 97% of the bill. On the same nine tasks, Jev Router cost $1.01 and passed everything too.

DataLLM Lab article cover: NVIDIA Switchyard: Nine of Nine on Two Models (and Claude Opus 5.5 Was 97% of the Bill)

On OpenRouter, Switchyard's default setup is a two-tier router: a cheap model, an expensive one, and a policy for choosing between them. That makes it easy to audit. Every call either went cheap or went expensive, and the only question is whether the expensive calls were needed.

Every task, and where it went

We sent nvidia/switchyard the same nine prompts every model on this site gets. All nine passed their hidden asserts. The table is the complete run. Costs are the per-call usage.cost figure OpenRouter returned in each response, printed as the API gave them. The last two columns show where Jev Router sent the same task in its run on 2026-10-02.

TaskPassSwitchyard sent it toCost per call (API-reported)LatencyJev Router sent it toJev cost
two_sumyesDeepSeek V4.1 Flash$0.00005783.6sGPT-6 Luna$0.0000460
valid_parenthesesyesClaude Opus 5.5$0.00305206.0sDeepSeek V4.1 Flash$0.0001734
merge_intervalsyesClaude Opus 5.5$0.00345604.9sGPT-6.1 Sol$0.0009280
roman_to_intyesDeepSeek V4.1 Flash$0.00019686.8sDeepSeek V4.1 Flash$0.0001959
lcs_lenyesClaude Opus 5.5$0.00372004.8sGPT-6 Luna$0.0000689
flattenyesDeepSeek V4.1 Flash$0.00022758.6sGPT-6 Luna$0.0000934
top_k_wordsyesDeepSeek V4.1 Flash$0.00014158.0sDeepSeek V4.1 Flash$0.0003732
token_bucketyesClaude Opus 5.5$0.00629606.7sGPT-6.1 Sol$0.0020900
parse_csv_lineyesClaude Opus 5.5$0.00512805.3sClaude Opus 5.5$0.0051280
Total9/92 models$0.0222757 = $2.48 per 1,000 tasks5.971s median4 models, 9/9$0.0090968 = $1.01

The split is not the one a human would draw. valid_parentheses — a stack and a lookup table — went to Opus, while flatten and top_k_words went to DeepSeek. The one task where both routers reached for Opus was parse_csv_line, and the two calls cost the identical $0.0051280. On the four other tasks Switchyard escalated, Jev Router used something cheaper and still passed.

The DeepSeek calls were quick for that model. Its own run on our board had a 15.6s mean; through Switchyard its four tasks returned in 3.6s, 6.8s, 8.6s and 8.0s. We did not log the reasoning setting the router used, so we cannot say why.

Where the money went

Switchyard landed between its two models, closer to the expensive oneCost per 1,000 tasks, all 9/9 on the same nine executed Python tasks.Claude Opus 5.5 alone$4.03NVIDIA Switchyard$2.48 · API-reported, 2026-10-03Jev Router$1.01 · API-reported, 2026-10-02DeepSeek V4.1 Flash alone$0.36 · priced 2026-09-15Solar Mini 4 alone$0.03SWITCHYARD'S NINE CALLS, STACKED · TOTAL $0.0222757token_bucketOpus calls 97%JEV ROUTER'S NINE CALLS, SAME SCALE · TOTAL $0.0090968one Opus call, parse_csv_lineBlue Claude Opus 5.5 · mid grey DeepSeek V4.1 Flash · light blue GPT-6 Luna · dark grey GPT-6.1 SolScales: top 140 px per dollar per 1,000 tasks; bottom 25,000 px per dollar (25 px per $0.001). Every width = value × scale.
Same scale for both stacked bars. The difference between them is mostly four Opus calls Jev Router did not make.

Add up the five Opus calls in the table and you get $0.021652. The four DeepSeek calls come to $0.0006236. Against the reported total of $0.0222757, Opus is 97.2% of the bill ($0.021652 ÷ $0.0222757). The per-call figures sum to $0.0222756, one ten-millionth of a dollar under the total, which is rounding in the per-call display. token_bucket alone is 28% of the spend.

So the cheap tier did its job — four tasks for well under a tenth of a cent combined — and was irrelevant to the total. Any two-tier router's bill is set almost entirely by how often it escalates. Here it escalated five times out of nine.

Against Jev Router and against either model alone

OptionScoreCost / 1,000 tasksCost methodLatency (statistic)
Always DeepSeek V4.1 Flash9/9$0.36Derived, list price on 2026-09-1515.6s mean
Jev Router9/9$1.01API-reported, 2026-10-024.414s median
NVIDIA Switchyard9/9$2.48API-reported, 2026-10-035.971s median
Always Claude Opus 5.59/9$4.03Derived, list price on 2026-10-024.8s mean

Single-model figures are means; router figures are medians. Do not read across those statistics as a like-for-like latency comparison.

Against Jev Router this is a like-for-like comparison: both figures are summed usage.cost on the same nine tasks, one day apart. Switchyard cost about 2.5 times as much ($2.48 ÷ $1.01), or $0.0131789 more per nine tasks, and was slower at the median. Both scored 9/9. The gap is the four extra Opus calls: on lcs_len, Switchyard's Opus call cost 54 times Jev's GPT-6 Luna call for the same passing result.

Against always-Opus, Switchyard cost about 62% ($2.48 ÷ $4.03). That is a real saving, but it compares two methods: the router figure is what the API reported, the single-model figure is derived from measured tokens times list price. Opus 5.5 listed at $4 in and $20 out per 1M on 2026-10-02, and is unchanged today.

Against always-DeepSeek, Switchyard cost about 6.9 times as much ($2.48 ÷ $0.36), and DeepSeek V4.1 Flash passed all nine tasks on its own — including the five Switchyard escalated. One caveat cuts against DeepSeek: its $0.36 was priced at $0.15 in and $0.6 out per 1M on 2026-09-15, and it lists at $0.3 in and $1.2 out today. Both rates doubled, so the same tokens would cost twice as much now. It would still be far below the router. The trade is latency: DeepSeek alone had a 15.6s mean.

For the floor of the board: Solar Mini 4 is the cheapest and the fastest model to score 9 out of 9 among the 75 that do, at $0.03 per 1,000 tasks and a 2s mean, priced 2026-10-02. Details for the two models Switchyard used are in our DeepSeek V4.1 Flash review and our Claude Opus 5.5 review.

What Switchyard is: the ledger

Everything in this section is third-party, read on 2026-10-03. Confirmed means the vendor's own page; attributed means a named outlet reporting it.

ClaimTierSource
An open-source library that helps an agent choose which model handles each request; Apache 2.0; OpenAI and Anthropic API compatibilityConfirmedNVIDIA's NVIDIA-NeMo/Switchyard GitHub README
Announced together with Nemotron 3.5 LightningConfirmedNVIDIA's own post on X
Announcement date 2026-08-11AttributedSiliconANGLE and VentureBeat, both dated 2026-08-11
Listed as nvidia/switchyard, published 2026-09-21, 1,000,000-token context, no routing fee; billed at the rate of whichever model answeredConfirmedOpenRouter's model page and router docs (OpenRouter operates the hosted endpoint)
With no models list, picks two candidates from the 20 models with the most OpenRouter spend over the previous seven days: an efficient tier from the cheapest fifth and a capable tier from the 60th to 80th percentile; refreshed hourlyConfirmedOpenRouter router docs
Default algorithm is stage, which scores existing tool results and calls a judge when indecisive; judge calls run on google/gemini-2.5-flash-lite and are billed as separate generationsConfirmedOpenRouter router docs
In NVIDIA's internal tests, a Switchyard mix of open models and Opus 4.8 cut task cost to roughly a third of Opus 4.8 aloneAttributedVentureBeat, 2026-08-11, reporting NVIDIA's own tests. The README shows a cost chart against Opus 4.8 and GLM 5.2 baselines and warns that results depend on benchmark, model pool, serving stack and configuration.
Kong, LiteLLM and OpenRouter built Switchyard support into their gatewaysAttributedVentureBeat. The README itself lists LiteLLM support as experimental.
A native Rust server installable with CargoAttributedKDnuggets tutorial, 2026-09-04

The pair we got, a cheap DeepSeek and an expensive Opus, fits that documented two-tier design. We did not record the candidate list the router chose on the day, and the docs say the ranking refreshes hourly, so your pair may differ from ours.

NVIDIA's one-third claim, and why our suite cannot test it

We measured 62% of always-Opus; the reported NVIDIA figure is about a third. These numbers are not in conflict, because they measure different things. Switchyard is built for agents working through multi-step tasks, and its default stage algorithm reads tool results. Our nine tasks are single-turn requests with no tool calls, so that algorithm had nothing to read. Our reading of the docs is that the judge then made each call, which may explain choices like Opus for valid_parentheses; we did not log it.

That raises a billing point we could not close. The docs say judge calls are billed as ordinary generations on your account. Our $2.48 is the sum of the usage.cost returned on the nine answers. If judge calls happened, they are probably not in that figure. We did not reconcile against account activity, so treat $2.48 as the cost of the answers, not necessarily the whole cost.

The bigger limitation is our suite. Nine self-contained Python functions cannot separate a frontier model from a competent small one: 75 models on our board score 9 out of 9. A router that escalates here is paying for headroom this test never uses. A fair verdict on Switchyard needs long agent sessions, where escalation might rescue a failing run and where switching models costs a prompt cache rebuild.

How these numbers were produced

Nine Python tasks, each a function signature plus a spec and no example tests. Generated code runs against hidden asserts in an isolated python3 -I subprocess with a 12-second timeout. Temperature 0, max_tokens 4000, one scored attempt per task. An API-layer failure is recorded separately from a wrong answer; Switchyard had neither. Router cost is the per-call usage.cost the API returned, summed; single-model cost is derived from measured tokens times the list price on the date shown. The OpenRouter catalogue lists the router at a negative placeholder rate of −1,000,000 per 1M tokens, so the tokens-times-price method cannot price it. Runs go through OpenRouter, not the DataLLM Lab gateway. Full method on the methodology page; prices move, as DeepSeek's did.

What we did not measure

If you run agents, the useful test is yours: send a week of real sessions through Switchyard, then price the same sessions on its efficient tier alone. On our suite, the cheap tier alone would have won.

Evidence: dated measurements and catalogue fields; method, latency statistics and limitations. Measurements describe these runs, not all providers or future versions.

Primary sources, checked October 3, 2026: NVIDIA Switchyard source; OpenRouter hosted Switchyard. Dated measurements above may differ from the current documentation.

Written by

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show. Articles are drafted with AI assistance and published under his name; every first-party number comes from an executed run.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.