Nvidia released Nemotron 3.5 Lightning and open-source NeMo Switchyard on August 11, 2026, saying its internal benchmark cut AI agent task costs by about 60%. The library selects a model for each workflow step using signals such as the task, price, latency, model state and infrastructure conditions. Nvidia is pitching the pair to businesses that want frontier models for difficult reasoning without paying frontier prices for routine calls.
What Changed
- Nvidia released the Nemotron 3.5 Lightning open-weights model and the Apache 2.0 NeMo Switchyard routing library on August 11, 2026.
- Nvidia’s internal benchmark cut a routed agent workload from about $180 to about $72, or roughly 60%, compared with using Opus 4.8 for every step.
- Launch partners reported lower costs, but the tests used different workloads and one accepted a six-point accuracy loss.
- Switchyard remains pre-alpha software, and the launch data does not establish its production performance outside Nvidia’s partner group.
AI-generated summary, reviewed by an editor. More on our AI guidelines.
The first substitution did most of the work
In its benchmark, Nvidia compared several routed pools with using Anthropic's Opus 4.8 at every step. Adding Lightning to Opus lowered the run's cost from about $180 to about $95, while task completion stayed almost flat. Adding two more open models brought the bill to about $72, but improved completion by less than one percentage point.
That headline result was not independently audited. Nick Patience, an AI platforms analyst, wrote in an analysis that it also uses an expensive all-frontier setup as its baseline, an arrangement a cost-conscious team would be unlikely to choose. Patience said the same data shows that almost all the savings came from the first substitution, not from assembling a large pool.
Switchyard can route a full request or reconsider the choice at each step. It can weigh task content, earlier errors and model responses alongside price and system load, then preserve state across later turns when a policy requires it.
Partners found different tradeoffs
For the August 11 launch, LangChain tested 145 multi-turn tasks across five runs. Its router sent 7% of calls to Opus 4.8 and cut costs 74%, while accuracy fell about six points.
Using its internal SWE-Bench, Ramp lowered costs 58% and runtime 33% while matching a frontier model's performance. On FrontierCode Main, Cognition scored 50.6% at a mean cost of $3.11, within 2.8 percentage points of Opus 5 and about 28% cheaper.
Cost cut, by launch partner
Four partners, four workloads, four baselines. The cost axis is the only thing they share.
LangChain−74%
Gave back 6 points of accuracy
Ramp−58%
Matched frontier performance, and cut runtime 33%
CodeRabbit−50.4%
Accuracy give-back not reported
Cognition−28%
Gave back 2.8 points against Opus 5
All figures vendor-reported at the August 11, 2026 launch. Nvidia and its partners ran the tests; none were independently audited.
In its August 11 launch-partner test, CodeRabbit used a fixed set of 1,000 tasks. After training, its router chose the intended model more often than a GPT baseline. When run at full capacity, CodeRabbit estimated that the cost fell from $2.34 to $1.16, or 50.4%. The experiment took less than three hours and cost less than $100. CodeRabbit ran it with Nvidia and Baseten, so it remains launch-partner evidence rather than independent production evidence.
Get Implicator.ai in your inbox
Strategic AI news from San Francisco. No hype, no "AI will change everything" throat clearing. Just what moved, who won, and why it matters. Daily at 6am PST.
No spam. Unsubscribe anytime.
Lightning is a mixture-of-experts model
Nemotron 3.5 Lightning is an open-weights mixture-of-experts model. It contains 31.6 billion parameters but activates 3.6 billion for each token, letting it draw on a larger model while running only part of it at once.
Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.
Artificial Analysis measured an Intelligence Index score of 24 on August 11, up from Nemotron 3 Nano's 15. That trailed Qwen3.6 35B A3B at 32, Muse Glimmer high at 35 and proprietary Gemini 3.5 Flash-Lite at 37. On a pre-release DeepInfra endpoint serving the final NVFP4 weights, Lightning produced nearly 670 output tokens per second.
The weights use the OpenMDW-1.1 license. Switchyard, by contrast, is open-source software under Apache 2.0.
The production bill is still unknown
The launch data does not independently establish how Switchyard performs in production outside Nvidia's hand-selected partner cohort. Nick Patience wrote in his August 11 analysis that routed systems also require teams to record which model handled each step, then evaluate outputs, debug failures and satisfy compliance rules. He said Nvidia's benchmark does not price that work.
As of August 11, Switchyard's GitHub repository labeled the software pre-alpha and warned that its interfaces and algorithms were expected to change significantly before version 1.0.
Frequently Asked Questions
What is Nvidia NeMo Switchyard?
NeMo Switchyard is an open-source routing library that selects a model for each step of an AI agent workflow using signals such as task type, cost, latency and model state.
How much did Switchyard reduce AI agent costs?
Nvidia’s internal benchmark reduced the cost of a routed workload from about $180 to about $72, or roughly 60%, against an Opus 4.8-only baseline.
What is Nemotron 3.5 Lightning?
Nemotron 3.5 Lightning is an open-weights mixture-of-experts model with 31.6 billion total parameters and 3.6 billion active parameters per token.
Were Nvidia’s cost claims independently verified?
No. The headline benchmark was internal, while the other launch results came from selected partners using different workloads and baselines.
Is NeMo Switchyard production-ready?
Its GitHub repository labeled it pre-alpha as of August 11 and warned that interfaces and algorithms could change before version 1.0.
AI-generated summary, reviewed by an editor. More on our AI guidelines.



IMPLICATOR