Skip to content
Dashboard

Nemotron 3.5 Lightning 30B

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total. Your use is subject to NVIDIA's Terms & Privacy Policies.

ReasoningTool UseImplicit CachingFree
import { streamText } from 'ai'
const result = streamText({
model: 'nvidia/nemotron-3.5-lightning',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingMore models by NVIDIA

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Release Date
1M
0.3s
141tps
$0.50/M
$2.40/M
Read:$0.12/M
Write:
baseten logo
deepinfra logo
togetherai logo
06/04/2026
256K
0.2s
160tps
$0.15/M
$0.65/M
bedrock logo
03/11/2026
262K
0.3s
221tps
$0.05/M
$0.24/M
deepinfra logo
12/15/2025
131K
0.1s
132tps
$0.20/M
$0.60/M
bedrock logo
deepinfra logo
10/28/2025
131K
0.1s
228tps
$0.06/M
$0.23/M
bedrock logo
deepinfra logo
08/18/2025