[NVIDIA](/ai-gateway/models/labs/nvidia)

# NVIDIA Nemotron 3 Super 120B A12B

NVIDIA Nemotron 3 Super 120B A12B is NVIDIA's 120B total, 12B active-parameter hybrid Mamba-Transformer MoE built for complex multi-agent applications, featuring latent MoE and multi-token prediction. Your use is subject to NVIDIA's [Terms](https://www.nvidia.com/en-us/about-nvidia/terms-of-service/) & [Privacy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) Policies.

ReasoningTool Use

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'nvidia/nemotron-3-super-120b-a12b',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/nemotron-3-super-120b-a12b) [API](/ai-gateway/models/nemotron-3-super-120b-a12b/api) [About](/ai-gateway/models/nemotron-3-super-120b-a12b/about) [Providers](/ai-gateway/models/nemotron-3-super-120b-a12b/providers) [Throughput](/ai-gateway/models/nemotron-3-super-120b-a12b/throughput) [Latency](/ai-gateway/models/nemotron-3-super-120b-a12b/latency) [Uptime](/ai-gateway/models/nemotron-3-super-120b-a12b/uptime) [Status](/ai-gateway/models/nemotron-3-super-120b-a12b/status) [Similar](/ai-gateway/models/nemotron-3-super-120b-a12b/similar) [FAQ](/ai-gateway/models/nemotron-3-super-120b-a12b/faq)

## [Copy link to heading](#playground)Playground

Try out NVIDIA Nemotron 3 Super 120B A12B by NVIDIA. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png&w=48&q=75)NVIDIA Nemotron 3 Super 120B A12B

![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png&w=96&q=75)

NVIDIA Nemotron 3 Super 120B A12B

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) [Bedrock](/ai-gateway/models/providers/bedrock) Legal:[Terms](https://aws.amazon.com/service-terms/)•[Privacy](https://aws.amazon.com/privacy/) | 256K | 32K | 0.3s | 151tps | $0.15/M | $0.65/M |  | — |  |  |  | 03/11/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-nvidia)More models by NVIDIA

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png&w=48&q=75) [nvidia/nemotron-3.5-lightning](/ai-gateway/models/nemotron-3.5-lightning) | 262K | 0.2s | 391tps | $0.05/M | $0.20/M | Read:$0.01/M Write:— | — |  | ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) |  |  | 08/11/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png&w=48&q=75) [nvidia/nemotron-3-ultra-550b-a55b](/ai-gateway/models/nemotron-3-ultra-550b-a55b) | 1M | 0.3s | 182tps | $0.50/M | $2.40/M | Read:$0.12/M Write:— | — |  | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![togetherai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftogetherai.png&w=48&q=75) |  |  | 06/04/2026 |  |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png&w=48&q=75) [nvidia/nemotron-3-nano\-30b-a3b](/ai-gateway/models/nemotron-3-nano-30b-a3b) | 262K | 0.3s | 188tps | $0.05/M | $0.24/M |  | — |  | ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) |  |  | 12/15/2025 |  |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png&w=48&q=75) [nvidia/nemotron-nano-12b-v2-vl](/ai-gateway/models/nemotron-nano-12b-v2-vl) | 131K | 0.2s | 121tps | $0.20/M | $0.60/M |  | — |  | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) |  |  | 10/28/2025 |  |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png&w=48&q=75) [nvidia/nemotron-nano-9b-v2](/ai-gateway/models/nemotron-nano-9b-v2) | 131K | 0.2s | 205tps | $0.06/M | $0.23/M |  | — |  | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) |  |  | 08/18/2025 |  |

## [Copy link to heading](#about-nvidia-nemotron-3-super-120b-a12b)About NVIDIA Nemotron 3 Super 120B A12B

NVIDIA released NVIDIA Nemotron 3 Super 120B A12B on March 11, 2026 as the second model in the Nemotron 3 family, following Nano. It has 120B total parameters and 12B active parameters per token. The hybrid Mamba-Transformer MoE backbone interleaves Mamba-2 layers for long-sequence processing, Transformer attention layers for precise recall, and MoE layers for compute efficiency. NVIDIA Nemotron 3 Super 120B A12B delivers higher throughput than the previous Nemotron Super generation.

Two architectural innovations distinguish Super from Nano. First, latent MoE: before routing, token embeddings compress into a low-rank latent space. This lets the model consult 4x as many expert specialists at the same inference cost. Finer-grained routing allows distinct experts to activate for different subtasks (Python syntax, SQL logic, multi-hop reasoning) without paying the compute cost of running them all. Second, multi-token prediction (MTP): the model predicts multiple future tokens in a single forward pass. MTP strengthens reasoning during training and provides built-in speculative decoding at inference, yielding up to 3x speedups on structured generation tasks like code and tool calls.

On PinchBench (a benchmark evaluating LLMs as the planning brain of an OpenClaw agent), NVIDIA Nemotron 3 Super 120B A12B scores 85.6%. Full announcement: https://docs.aws.amazon.com/en\_us/bedrock/latest/userguide/model-card-nvidia-nemotron-super-3-120b.html.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: NVIDIA Nemotron 3 Super 120B A12B's multi-agent orientation means it works best as the planning and reasoning backbone in a pipeline where lighter models handle individual steps. Evaluate your task decomposition before choosing a tier. Compare $0.15 and $0.65.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-nvidia-nemotron-3-super-120b-a12b)When to Use NVIDIA Nemotron 3 Super 120B A12B

### Best for

- Complex multi-agent applications: Software development pipelines or cybersecurity triaging that require deep planning across long contexts
- Context explosion workloads: Multi-agent systems with up to 15x the token volume of standard chats that cause goal drift with smaller models
- Dense technical problem-solving: Tasks where higher parameter count provides reasoning headroom
- Super plus nano pattern: Agentic pipelines pairing Super for complex decisions with Nano for efficient individual steps
- Fully open model requirement: Teams that need weights and recipes for enterprise customization, data control, or reproducibility

### Consider alternatives when

- Simpler task steps: Nemotron 3 Nano is more throughput-efficient for lighter workloads
- Vision-language inputs: Super is text-only; Nemotron Nano 12B v2 VL supports multimodal inputs
- Cost-first constraints: A lighter model may deliver acceptable quality at lower cost per token

## [Copy link to heading](#conclusion)Conclusion

NVIDIA Nemotron 3 Super 120B A12B combines latent MoE for expert specialization and multi-token prediction for inference speedups. Route requests through AI Gateway as the planning and reasoning backbone for complex multi-agent applications at scale.