# NVIDIA models

[Get API key](https://vercel.com/d?to=%2F%5Bteam%5D%2F~%2Fai-gateway%3FshowCreateKeyModal%26utm_source%3Dgateway-labs-page%26utm_campaign%3Dai-gateway-labs&title=Get%20API%20key) [Read the docs](https://vercel.com/docs/ai-gateway)

Browse and compare every NVIDIA model available on Vercel AI Gateway. Compare pricing, context windows, and capabilities, then call any of them through a single API.

| Model | Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Regional Inference | Free Tier | Released |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| [nvidia/nemotron-3.5-lightning](/ai-gateway/models/nemotron-3.5-lightning) | 262K | 0.3 s | 581 tps | $0.05/M | $0.15/M | Read:$0.01/M Write:— | — |  |  |  |  |  |  | 08/11/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [nvidia/nemotron-3-ultra-550b-a55b](/ai-gateway/models/nemotron-3-ultra-550b-a55b) | 1M | 0.4 s | 137 tps | $0.50/M | $2.40/M | Read:$0.12/M Write:— | — |  |  |  |  | US |  | 06/04/2026 |  |
| [nvidia/nemotron-3-super-120b-a12b](/ai-gateway/models/nemotron-3-super-120b-a12b) | 256K | 0.2 s | 148 tps | $0.15/M | $0.65/M |  | — |  |  |  |  |  |  | 03/11/2026 |  |
| [nvidia/nemotron-3-nano\-30b-a3b](/ai-gateway/models/nemotron-3-nano-30b-a3b) | 262K | 0.2 s | 128 tps | $0.05/M | $0.24/M |  | — |  |  |  |  |  |  | 12/15/2025 |  |
| [nvidia/nemotron-nano-12b-v2-vl](/ai-gateway/models/nemotron-nano-12b-v2-vl) | 131K | 0.2 s | 127 tps | $0.20/M | $0.60/M |  | — |  |  |  |  |  |  | 10/28/2025 |  |
| [nvidia/nemotron-nano-9b-v2](/ai-gateway/models/nemotron-nano-9b-v2) | 131K | 0.1 s | 144 tps | $0.06/M | $0.23/M |  | — |  |  |  |  |  |  | 08/18/2025 |  |