[NVIDIA](/ai-gateway/models/labs/nvidia)

# Nemotron 3 Ultra

Nemotron 3 Ultra is NVIDIA's largest open reasoning model, a hybrid Mamba-Transformer MoE with 550B total and 55B active parameters, latent MoE routing, multi-token prediction, and a context window of 1M tokens for long-running agent workflows. Your use is subject to NVIDIA's [Terms](https://assets.ngc.nvidia.com/products/api-catalog/legal/NVIDIA_Technology_Access_TOU.pdf) & [Privacy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) Policies.

ReasoningTool UseImplicit Caching

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'nvidia/nemotron-3-ultra-550b-a55b',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/nemotron-3-ultra-550b-a55b) [API](/ai-gateway/models/nemotron-3-ultra-550b-a55b/api) [About](/ai-gateway/models/nemotron-3-ultra-550b-a55b/about) [Providers](/ai-gateway/models/nemotron-3-ultra-550b-a55b/providers) [Throughput](/ai-gateway/models/nemotron-3-ultra-550b-a55b/throughput) [Latency](/ai-gateway/models/nemotron-3-ultra-550b-a55b/latency) [Uptime](/ai-gateway/models/nemotron-3-ultra-550b-a55b/uptime) [Status](/ai-gateway/models/nemotron-3-ultra-550b-a55b/status) [Similar](/ai-gateway/models/nemotron-3-ultra-550b-a55b/similar) [FAQ](/ai-gateway/models/nemotron-3-ultra-550b-a55b/faq)

## [Copy link to heading](#playground)Playground

Try out Nemotron 3 Ultra by NVIDIA. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75)Nemotron 3 Ultra

![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=96&q=75)

Nemotron 3 Ultra

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Regional Inference | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![togetherai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftogetherai.png%3Fv%3D1787852239935&w=48&q=75) [Together AI](/ai-gateway/models/providers/togetherai) Legal:[Terms](https://www.together.ai/terms-of-service)•[Privacy](https://www.together.ai/privacy) | 1M | 65K |  |  | $0.60/M | $3.60/M | Read:$0.20/M Write:— | — |  |  |  |  |  | 06/04/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) [DeepInfra](/ai-gateway/models/providers/deepinfra) Legal:[Terms](https://deepinfra.com/terms)•[Privacy](https://deepinfra.com/privacy) | 262K | 65K | 1.0 s | 53 tps | $0.50/M | $2.50/M | Read:$0.15/M Write:— | — |  |  |  |  |  | 06/04/2026 |  |
| ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) [Baseten](/ai-gateway/models/providers/baseten) Legal:[Terms](https://www.baseten.co/terms-and-conditions/)•[Privacy](https://www.baseten.co/privacy-policy/) | 1M | 65K | 0.5 s | 145 tps | $0.60/M | $2.40/M | Read:$0.12/M Write:— | — |  |  |  | US |  | 06/04/2026 |  |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-nvidia)More models by NVIDIA

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75) [nvidia/nemotron-3.5-lightning](/ai-gateway/models/nemotron-3.5-lightning) | 262K | 0.2 s | 452 tps | $0.05/M | $0.15/M | Read:$0.01/M Write:— | — |  | ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![runinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fruninfra.png%3Fv%3D1786910657446&w=48&q=75) |  |  |  | 08/11/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75) [nvidia/nemotron-3-super-120b-a12b](/ai-gateway/models/nemotron-3-super-120b-a12b) | 256K | 0.3 s | 158 tps | $0.15/M | $0.65/M |  | — |  | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) |  |  |  | 03/11/2026 |  |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75) [nvidia/nemotron-3-nano\-30b-a3b](/ai-gateway/models/nemotron-3-nano-30b-a3b) | 262K | 4.5 s | 83 tps | $0.05/M | $0.24/M |  | — |  | ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) |  |  |  | 12/15/2025 |  |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75) [nvidia/nemotron-nano-12b-v2-vl](/ai-gateway/models/nemotron-nano-12b-v2-vl) | 131K | 0.2 s | 88 tps | $0.20/M | $0.60/M |  | — |  | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) |  |  |  | 10/28/2025 |  |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75) [nvidia/nemotron-nano-9b-v2](/ai-gateway/models/nemotron-nano-9b-v2) | 131K | 0.2 s | 141 tps | $0.06/M | $0.23/M |  | — |  | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) |  |  |  | 08/18/2025 |  |

## [Copy link to heading](#about-nemotron-3-ultra)About Nemotron 3 Ultra

NVIDIA released Nemotron 3 Ultra on June 4, 2026 as the largest model in the Nemotron 3 family, completing the tier above Nano and Super. It carries 550B total parameters with 55B active per token, and NVIDIA positions it as the reasoning and orchestration layer for long-running agent workflows: the model that handles planning, synthesis, and verification while lighter models execute routine steps.

The architecture interleaves three layer types. Mamba layers process long sequences with linear-time complexity, which keeps a context window of 1M tokens practical. Transformer attention layers appear at select depths to preserve precise recall from large contexts. Latent mixture-of-experts (MoE) routing compresses token embeddings into a smaller latent space before selecting experts, so distinct specialists activate for reasoning, coding, and tool calls without dense compute. Multi-token prediction (MTP) layers predict several future tokens per forward pass, providing built-in speculative decoding for long outputs.

Nemotron 3 Ultra scores 91% on PinchBench, 82% on IFBench, and 95% on Ruler at 1M tokens. Weights, data, and recipes are released under the Linux Foundation's permissive OpenMDW-1.1 license. Full details: https://www.together.ai/models/nvidia-nemotron-3-ultra.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Long-running agent sessions accumulate tokens quickly, and a context window of 1M tokens makes it easy to carry everything forward. Budget for that before you scale. Compare $0.5 and $2.4, and use prompt caching at $0.12 for repeated prefixes like system prompts and tool definitions.
- Configuration: Output is capped at 65K tokens per request, so plan chunking for very long generations. Nemotron 3 Ultra is the flagship tier of the Nemotron 3 family. Reserve it for the planning and verification calls that need the depth, and route routine steps to smaller Nemotron 3 models.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-nemotron-3-ultra)When to Use Nemotron 3 Ultra

### Best for

- Agent Orchestration Backbones: Planning, synthesis, and verification steps in long-running multi-agent pipelines
- Long-Horizon Coding Agents: Multi-step software tasks that span large codebases and extended tool-call sequences
- Deep Research Workflows: Gathering, cross-checking, and synthesizing evidence across many sources in one context
- Full-Context Session Handling: Keeping complete agent histories, codebases, or document sets in a single pass
- Open-Model Requirements: Teams that need open weights and permissive licensing for governance or reproducibility

### Consider alternatives when

- Lightweight Task Execution: Nemotron 3 Nano handles routine pipeline steps at far lower compute
- Mid-Tier Agent Planning: Nemotron 3 Super covers complex multi-agent decisions at a smaller footprint
- Vision or Multimodal Inputs: Nemotron 3 Ultra is a text reasoning model, so image and video tasks need a vision-language model
- Cost-First Workloads: A smaller model may deliver acceptable quality at lower per-token rates

## [Copy link to heading](#conclusion)Conclusion

Nemotron 3 Ultra closes out the Nemotron 3 family as its reasoning and orchestration tier, pairing latent MoE efficiency with a context window of 1M tokens. Route it through AI Gateway with unified auth and billing, and call it with the AI SDK or through Chat Completions, Responses, Messages, and other API formats.