Nvidia unveils AI routing platform to slash enterprise model costs
Nvidia introduced NeMo Switchyard, a software layer that directs AI requests to different models, aiming to reduce enterprise AI spending.
Nvidia's latest offering, NeMo Switchyard, functions as a proxy that intelligently routes AI prompts to the most appropriate model, balancing cost, latency and output fidelity. By diverting simpler tasks to smaller, cheaper or locally hosted models, the company claims enterprises can achieve cost reductions of roughly 74 % relative to relying solely on a premium model such as Claude Opus 4.8, with an estimated six-point drop in accuracy.
The platform is paired with the newly announced Nemotron 3.5-30B-A3B-Lightning, a 30-billion-parameter mixture-of-experts model designed for low-latency, general-purpose use. Nvidia also highlighted specialized models like Nemotron Parse, a one-billion-parameter tool for extracting information from PDFs. While model routing is not new—OpenAI and AT&T have employed similar techniques—the Switchyard aims to automate the decision-making process, potentially allowing AI agents to learn when to employ smaller models versus larger frontier models. Nvidia views this capability as a step toward broader corporate adoption of AI.
Why it matters
It offers businesses a way to manage soaring AI costs while maintaining performance.
In this story