Nvidia released Nemotron 3.5 Lightning and open-source NeMo Switchyard on August 11, 2026, saying its internal benchmark cut AI agent task costs by about 60%. The library selects a model for each workflow step using signals such as the task, price, latency, model state and infrastructure conditions. Nvidia is pitching the pair to businesses that want frontier models for difficult reasoning without paying frontier prices for routine calls.

What Changed

AI-generated summary, reviewed by an editor. More on our AI guidelines.

The first substitution did most of the work

In its benchmark, Nvidia compared several routed pools with using Anthropic's Opus 4.8 at every step. Adding Lightning to Opus lowered the run's cost from about $180 to about $95, while task completion stayed almost flat. Adding two more open models brought the bill to about $72, but improved completion by less than one percentage point.

That headline result was not independently audited. Nick Patience, an AI platforms analyst, wrote in an analysis that it also uses an expensive all-frontier setup as its baseline, an arrangement a cost-conscious team would be unlikely to choose. Patience said the same data shows that almost all the savings came from the first substitution, not from assembling a large pool.

Switchyard can route a full request or reconsider the choice at each step. It can weigh task content, earlier errors and model responses alongside price and system load, then preserve state across later turns when a policy requires it.

Partners found different tradeoffs

For the August 11 launch, LangChain tested 145 multi-turn tasks across five runs. Its router sent 7% of calls to Opus 4.8 and cut costs 74%, while accuracy fell about six points.

Using its internal SWE-Bench, Ramp lowered costs 58% and runtime 33% while matching a frontier model's performance. On FrontierCode Main, Cognition scored 50.6% at a mean cost of $3.11, within 2.8 percentage points of Opus 5 and about 28% cheaper.

Cost cut, by launch partner

Four partners, four workloads, four baselines. The cost axis is the only thing they share.

LangChain−74%

145 multi-turn tasks, five runs · 7% of calls kept on Opus 4.8

Gave back 6 points of accuracy

Ramp−58%

Internal SWE-Bench · baseline: a frontier model

Matched frontier performance, and cut runtime 33%

CodeRabbit−50.4%

1,000 fixed tasks at full capacity · $2.34 to $1.16 per task

Accuracy give-back not reported

Cognition−28%

FrontierCode Main · scored 50.6% at $3.11 mean cost

Gave back 2.8 points against Opus 5

All figures vendor-reported at the August 11, 2026 launch. Nvidia and its partners ran the tests; none were independently audited.

In its August 11 launch-partner test, CodeRabbit used a fixed set of 1,000 tasks. After training, its router chose the intended model more often than a GPT baseline. When run at full capacity, CodeRabbit estimated that the cost fell from $2.34 to $1.16, or 50.4%. The experiment took less than three hours and cost less than $100. CodeRabbit ran it with Nvidia and Baseten, so it remains launch-partner evidence rather than independent production evidence.

Lightning is a mixture-of-experts model

Nemotron 3.5 Lightning is an open-weights mixture-of-experts model. It contains 31.6 billion parameters but activates 3.6 billion for each token, letting it draw on a larger model while running only part of it at once.

Know someone who'd find this useful? ✉️ Email it to a friend in one click, or they can subscribe free here.

Artificial Analysis measured an Intelligence Index score of 24 on August 11, up from Nemotron 3 Nano's 15. That trailed Qwen3.6 35B A3B at 32, Muse Glimmer high at 35 and proprietary Gemini 3.5 Flash-Lite at 37. On a pre-release DeepInfra endpoint serving the final NVFP4 weights, Lightning produced nearly 670 output tokens per second.

The weights use the OpenMDW-1.1 license. Switchyard, by contrast, is open-source software under Apache 2.0.

The production bill is still unknown

The launch data does not independently establish how Switchyard performs in production outside Nvidia's hand-selected partner cohort. Nick Patience wrote in his August 11 analysis that routed systems also require teams to record which model handled each step, then evaluate outputs, debug failures and satisfy compliance rules. He said Nvidia's benchmark does not price that work.

As of August 11, Switchyard's GitHub repository labeled the software pre-alpha and warned that its interfaces and algorithms were expected to change significantly before version 1.0.

Frequently Asked Questions

What is Nvidia NeMo Switchyard?

NeMo Switchyard is an open-source routing library that selects a model for each step of an AI agent workflow using signals such as task type, cost, latency and model state.

How much did Switchyard reduce AI agent costs?

Nvidia’s internal benchmark reduced the cost of a routed workload from about $180 to about $72, or roughly 60%, against an Opus 4.8-only baseline.

What is Nemotron 3.5 Lightning?

Nemotron 3.5 Lightning is an open-weights mixture-of-experts model with 31.6 billion total parameters and 3.6 billion active parameters per token.

Were Nvidia’s cost claims independently verified?

No. The headline benchmark was internal, while the other launch results came from selected partners using different workloads and baselines.

Is NeMo Switchyard production-ready?

Its GitHub repository labeled it pre-alpha as of August 11 and warned that interfaces and algorithms could change before version 1.0.

AI-generated summary, reviewed by an editor. More on our AI guidelines.

Five GitHub Projects Show Where Agents Are Heading
San Francisco | Thursday, June 25, 2026 Repo Radar leads today because the agent story has moved from demo clips to the plumbing teams can actually run. The five projects worth watching this week po
This Week's Hottest Repos All Exist to Keep Agents in Check
San Francisco | Thursday, June 18, 2026 Coding agents now move fast enough to strain GitHub itself, which leaned on Amazon's cloud this week to absorb the traffic. The projects climbing beside the ou
Repo Radar: 5 GitHub Projects Worth Your Week
GitHub spent the week leaning on Amazon's cloud to absorb agentic-development traffic, Business Insider reported June 16, after a run of AI-driven outages. As agents run at machine speed, this week's
AI News

San Francisco

Editor-in-Chief and founder of Implicator.ai. Former ARD correspondent and senior broadcast journalist with 10+ years covering tech. Writes daily briefings on policy and market developments. Based in San Francisco. E-mail: editor@implicator.ai