[NVIDIA](/ai-gateway/models/labs/nvidia)

# Nvidia Nemotron Nano 9B V2

Nvidia Nemotron Nano 9B V2 is a dense hybrid Mamba-Transformer reasoning model that matches or exceeds Qwen3-8B accuracy at up to 6x the throughput, with built-in thinking budget control. Your use is subject to NVIDIA's [Terms](https://assets.ngc.nvidia.com/products/api-catalog/legal/NVIDIA_Technology_Access_TOU.pdf) & [Privacy](https://www.nvidia.com/en-us/about-nvidia/privacy-policy/) Policies.

ReasoningTool Use

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'nvidia/nemotron-nano-9b-v2',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/nemotron-nano-9b-v2) [API](/ai-gateway/models/nemotron-nano-9b-v2/api) [About](/ai-gateway/models/nemotron-nano-9b-v2/about) [Providers](/ai-gateway/models/nemotron-nano-9b-v2/providers) [Throughput](/ai-gateway/models/nemotron-nano-9b-v2/throughput) [Latency](/ai-gateway/models/nemotron-nano-9b-v2/latency) [Uptime](/ai-gateway/models/nemotron-nano-9b-v2/uptime) [Status](/ai-gateway/models/nemotron-nano-9b-v2/status) [Similar](/ai-gateway/models/nemotron-nano-9b-v2/similar) [FAQ](/ai-gateway/models/nemotron-nano-9b-v2/faq)

## [Copy link to heading](#playground)Playground

Try out Nvidia Nemotron Nano 9B V2 by NVIDIA. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75)Nvidia Nemotron Nano 9B V2

![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=96&q=75)

Nvidia Nemotron Nano 9B V2

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) [Bedrock](/ai-gateway/models/providers/bedrock) Legal:[Terms](https://aws.amazon.com/service-terms/)•[Privacy](https://aws.amazon.com/privacy/) | 131K | 131K | 0.1s | 140tps | $0.06/M | $0.23/M |  | — |  |  |  |  | 08/18/2025 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) [DeepInfra](/ai-gateway/models/providers/deepinfra) Legal:[Terms](https://deepinfra.com/terms)•[Privacy](https://deepinfra.com/privacy) | 131K | 131K | 0.3s | 18tps | $0.04/M | $0.16/M |  | — |  |  |  |  | 08/18/2025 |  |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-nvidia)More models by NVIDIA

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75) [nvidia/nemotron-3.5-lightning](/ai-gateway/models/nemotron-3.5-lightning) | 262K | 0.3s | 294tps | $0.05/M | $0.15/M | Read:$0.01/M Write:— | — |  | ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![runinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fruninfra.png%3Fv%3D1786910657446&w=48&q=75) |  |  |  | 08/11/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75) [nvidia/nemotron-3-ultra-550b-a55b](/ai-gateway/models/nemotron-3-ultra-550b-a55b) | 1M | 0.4s | 175tps | $0.50/M | $2.40/M | Read:$0.12/M Write:— | — |  | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![togetherai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftogetherai.png%3Fv%3D1787852239935&w=48&q=75) |  |  |  | 06/04/2026 |  |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75) [nvidia/nemotron-3-super-120b-a12b](/ai-gateway/models/nemotron-3-super-120b-a12b) | 256K | 0.3s | 142tps | $0.15/M | $0.65/M |  | — |  | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) |  |  |  | 03/11/2026 |  |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75) [nvidia/nemotron-3-nano\-30b-a3b](/ai-gateway/models/nemotron-3-nano-30b-a3b) | 262K | 0.4s | 17tps | $0.05/M | $0.24/M |  | — |  | ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) |  |  |  | 12/15/2025 |  |
| ![nvidia logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnvidia.png%3Fv%3D1787340069568&w=48&q=75) [nvidia/nemotron-nano-12b-v2-vl](/ai-gateway/models/nemotron-nano-12b-v2-vl) | 131K | 0.1s | 15tps | $0.20/M | $0.60/M |  | — |  | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) |  |  |  | 10/28/2025 |  |

## [Copy link to heading](#about-nvidia-nemotron-nano-9b-v2)About Nvidia Nemotron Nano 9B V2

NVIDIA released Nvidia Nemotron Nano 9B V2 on August 18, 2025 as the compressed reasoning variant of the Nemotron Nano 2 family. It is a 9B-parameter model with a context window of 131.1K tokens.

Nvidia Nemotron Nano 9B V2 matches or exceeds Qwen3-8B on complex reasoning tasks at up to 6x the throughput. The hybrid Mamba-Transformer architecture contributes to this efficiency. Mamba layers handle sequence processing with sub-quadratic memory scaling, while Transformer attention layers maintain precision on retrieval-heavy tasks within the context window.

Nvidia Nemotron Nano 9B V2 also supports thinking budget control. You can prompt it to reason briefly for simple tasks (faster, cheaper) or thoroughly for hard problems (slower, more accurate). Adjust the latency-accuracy tradeoff at inference time without switching models. Technical report and assets: https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Nvidia Nemotron Nano 9B V2 is a compact dense reasoning model. Evaluate whether its capability tier fits your workload before committing at production scale. Compare $0.06 and $0.23.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-nvidia-nemotron-nano-9b-v2)When to Use Nvidia Nemotron Nano 9B V2

### Best for

- High-throughput reasoning: Workloads where 6x speed over comparable models matters
- Thinking budget control: Applications that vary reasoning depth per request
- Cost-sensitive production: Compact reasoning models that reduce infrastructure spend

### Consider alternatives when

- 1M-token context: Nemotron 3 Nano (30B/3B active) supports that scale
- Vision or multimodal: Nemotron Nano 12B v2 VL is the right choice
- Multi-agent orchestration: The sparse MoE design of Nemotron 3 Nano is better suited to that pattern

## [Copy link to heading](#conclusion)Conclusion

Nvidia Nemotron Nano 9B V2 is a dense reasoning model. It delivers high throughput and accuracy with thinking budget control for tuning the speed-accuracy tradeoff per request. Route it through AI Gateway.