Skip to content
Dashboard

Nemotron 3 Nano 30B A3B

Nemotron 3 Nano 30B A3B is a sparse hybrid Mamba-Transformer mixture-of-experts (MoE) model with 30B total parameters but only 3B active per token. It supports a context window of 262.1K tokens with throughput closer to a 3B dense model than a 30B one.

Input and output price
Input $0.05, Output $0.24, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'nvidia/nemotron-3-nano-30b-a3b',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingMore models by NVIDIA

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
262K0.2 s65 tps
$0.05/M
$0.15/M
Read$0.01/M
deepinfra logo
fireworks logo
runinfra logo
08/11/2026
1M0.3 s137 tps
$0.50/M
$2.40/M
Read$0.12/M
baseten logo
deepinfra logo
fireworks logo
+1
06/04/2026
256K0.2 s164 tps
$0.15/M
$0.65/M
bedrock logo
03/11/2026
131K0.1 s54 tps
$0.20/M
$0.60/M
bedrock logo
deepinfra logo
10/28/2025
131K0.2 s202 tps
$0.06/M
$0.23/M
bedrock logo
deepinfra logo
08/18/2025

Your use is subject to NVIDIA's Terms & Privacy Policies.