Nvidia released Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts open model with about 3 billion active parameters per token for high-volume agent steps. The company says its hybrid Mamba-Transformer design and speculative decoding can lift output speed up to 4x and cut agent task time by about 30% versus peers. Weights are on Hugging Face and ModelScope under OpenMDW 1.1, with NeMo Switchyard as an open router across mixed models. Why it matters: Nvidia is productizing the cheap execution layer beside frontier planners. Caveat: vendor speed claims still need independent, workload-matched checks.