Nemotron 3.5 Lightning 30B
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total. Your use is subject to NVIDIA's Terms & Privacy Policies.
ReasoningTool UseImplicit CachingFree
import { streamText } from 'ai'
const result = streamText({ model: 'nvidia/nemotron-3.5-lightning', prompt: 'Why is the sky blue?'})Copy link to headingThroughput24 hours
P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.