NVIDIA built a model that talks four times faster. The work arrives 30 percent sooner.
The 4x is NVIDIA's own number, measured at the token tap. On NVIDIA's own chart, the agent's finish line moves 30 percent. Nothing independent confirms either.
"NVIDIA's Nemotron 3.5 Lightning delivers up to 4x the output speed of similar-sized models, putting it on the accuracy-speed Pareto frontier for agent workloads." [SOURCE ↗]
THE MOVE: LAB NOT FIELD, it works in the test conditions, not the deployed ones

THE CLAIM. Nemotron 3.5 Lightning, NVIDIA's new 30B agent model with 3B active parameters, generates output up to 4x faster than similar-sized models and wins the accuracy-speed Pareto frontier.
THE CHECK. three sentences after the 4x, NVIDIA's own blog concedes that agent efficiency comes down to completed work, not token speed, and its own PinchBench chart shows tasks finishing 30% faster than Qwen3.6 35B. The 4x is 'up to', measured by the vendor, and Artificial Analysis's page for the model lists no output speed at all. A 4x engine revs; a 30% car arrives. Both numbers are NVIDIA grading NVIDIA.
On August 11 NVIDIA released Nemotron 3.5 Lightning, a 30B-parameter open Mixture-of-Experts model that keeps 3B parameters active per token, shipped under the permissive OpenMDW license with weights, training data, and recipes. It is built for the grunt-work layer of agent systems: the tool calls,
🔒 THE FULL AUTOPSY · FREE WITH AN ACCOUNTYou just read the free check. Sign in free, a code by email, no passwords, and the rest unlocks: the evidence trail, the steelman and the rebuttal, all 4 sources with quotes and screenshots, and our on-record call.
Couldn't verify your access. This looks like our error, not yours.