NVIDIA's Nemotron 3.5 Lightning model is now available on Amazon SageMaker JumpStart, targeting persistent agent workloads and enterprise automation. The model uses a hybrid Mixture-of-Experts architecture with 30B total parameters but only 3B active per forward pass, enabling up to 4x throughput compared to similar models. It supports up to 1M token context via DFlash speculative decoding and is distilled from the larger Nemotron 3 Ultra.
- Hybrid MoE architecture with 3B active parameters delivers ~410 tokens/sec throughput
- 30% faster task completion for domains like financial processing and cybersecurity
- Supports 1M token context window using DFlash speculative decoding
- Optimized for persistent agents and high-throughput enterprise automation
- Distilled from Nemotron 3 Ultra for efficient performance