[Meta](/ai-gateway/models/labs/meta)

# Llama 4 Maverick 17B 128E Instruct FP8

Llama 4 Maverick 17B 128E Instruct FP8 is Meta's natively multimodal Mixture of Experts (MoE) model with 17B active parameters across 128 experts. Published benchmarks span image and text tasks, and the MoE activates a fraction of the parameters that comparable dense models use. Your use is subject to Meta's [Terms](https://www.facebook.com/policies_center/) & [Privacy](https://www.facebook.com/privacy/policy/) Policies.

Tool UseVision (Image)

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'meta/llama-4-maverick',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/llama-4-maverick) [API](/ai-gateway/models/llama-4-maverick/api) [About](/ai-gateway/models/llama-4-maverick/about) [Providers](/ai-gateway/models/llama-4-maverick/providers) [Throughput](/ai-gateway/models/llama-4-maverick/throughput) [Latency](/ai-gateway/models/llama-4-maverick/latency) [Uptime](/ai-gateway/models/llama-4-maverick/uptime) [Status](/ai-gateway/models/llama-4-maverick/status) [Similar](/ai-gateway/models/llama-4-maverick/similar) [FAQ](/ai-gateway/models/llama-4-maverick/faq)

## [Copy link to heading](#playground)Playground

Try out Llama 4 Maverick 17B 128E Instruct FP8 by Meta. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=48&q=75)Llama 4 Maverick 17B 128E Instruct FP8

![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=96&q=75)

Llama 4 Maverick 17B 128E Instruct FP8

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) [DeepInfra](/ai-gateway/models/providers/deepinfra) Legal:[Terms](https://deepinfra.com/terms)•[Privacy](https://deepinfra.com/privacy) | 131K | 8K | 0.2s | 73tps | $0.20/M | $0.80/M |  | — |  |  |  |  | 04/05/2025 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) [Bedrock](/ai-gateway/models/providers/bedrock) Legal:[Terms](https://aws.amazon.com/service-terms/)•[Privacy](https://aws.amazon.com/privacy/) | 128K | 8K | 0.2s |  | $0.24/M | $0.97/M |  | — |  |  |  |  | 04/05/2025 |  |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-meta)More models by Meta

All

Text

Code

Image

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=48&q=75) [meta/muse-image-1.0](/ai-gateway/models/muse-image-1.0) |  |  |  |  | $0.01/img |  | — |  | ![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=48&q=75) |  |  |  | 08/26/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=48&q=75) [meta/muse-glimmer-30b](/ai-gateway/models/muse-glimmer-30b) | 131K | 0.2s | 111tps | $0.30/M | $1.10/M | Read:$0.04/M Write:— | — | +1 | ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![parasail logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fparasail.png&w=48&q=75) ![togetherai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftogetherai.png%3Fv%3D1787852239935&w=48&q=75) |  |  |  | 08/10/2026 |  |
| ![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=48&q=75) [meta/muse-spark-1.2-contributor](/ai-gateway/models/muse-spark-1.2-contributor) | 1M | 3.7s | 103tps | $0.10/M | $0.20/M | Read:$0.002/M Write:— | — | +2 | ![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=48&q=75) |  |  |  | 08/05/2026 |  |
| ![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=48&q=75) [meta/muse-spark-1.2](/ai-gateway/models/muse-spark-1.2) | 1M | 1.7s | 71tps | $1.25/M | $4.25/M | Read:$0.15/M Write:— | — | +2 | ![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=48&q=75) |  |  |  | 08/05/2026 |  |
| ![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=48&q=75) [meta/muse-spark-1.1](/ai-gateway/models/muse-spark-1.1) | 1M | 5.2s | 202tps | $1.25/M | $4.25/M | Read:$0.15/M Write:— | — | +2 | ![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=48&q=75) |  |  |  | 07/09/2026 |  |
| ![meta logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmeta.png&w=48&q=75) [meta/llama-3.1-8b](/ai-gateway/models/llama-3.1-8b) | 131K | 0.2s | 169tps | $0.02/M | $0.05/M |  | — |  | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![novita logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnovita.png&w=48&q=75) |  |  |  | 07/23/2024 |  |

## [Copy link to heading](#about-llama-4-maverick-17b-128e-instruct-fp8)About Llama 4 Maverick 17B 128E Instruct FP8

Meta released Llama 4 Maverick 17B 128E Instruct FP8 on April 5, 2025 as one of the first two models in the Llama 4 generation. The collection is built around two architectural advances: native multimodality through early fusion, and Mixture of Experts (MoE). Llama 4 Maverick 17B 128E Instruct FP8 is the larger and more capable of the two initial releases, with 17 billion active parameters, 128 routed experts plus one shared expert, and 400 billion total parameters. Each token activates only 17B of those 400B parameters (the shared expert plus one routed expert). This makes inference substantially more efficient than a dense 400B model while preserving the quality benefits of the larger total parameter budget.

Llama 4's native multimodality represents a different architectural approach from the adapter-based vision in Llama 3.2. Rather than adding image understanding to an existing text backbone, Llama 4 treats text and vision tokens together from the beginning in a unified backbone. This enables more coherent cross-modal reasoning.

On the LMArena leaderboard, an experimental chat version of Llama 4 Maverick 17B 128E Instruct FP8 scored an Elo of 1417. Llama 4 Maverick 17B 128E Instruct FP8 exceeds comparable frontier models on coding, reasoning, multilingual, long-context, and image benchmarks. It achieves results comparable to other open-weight models on reasoning and coding at less than half the active parameters.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: For workloads that mix images and long text, Llama 4 Maverick 17B 128E Instruct FP8's efficiency advantage over dense models shows most at scale. Validate throughput at your expected concurrency level before you pick a provider tier. Compare $0.24 and $0.97.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-llama-4-maverick-17b-128e-instruct-fp8)When to Use Llama 4 Maverick 17B 128E Instruct FP8

### Best for

- Production multimodal applications: Pairing image understanding with long-form text generation for product catalog processing and document analysis with mixed visual and textual content
- Creative and coding workloads: Multilingual applications where the MoE architecture reaches dense-model scores on published benchmarks at lower active-parameter cost
- Cost-per-quality sensitive workloads: Comparable-capability dense models are significantly more expensive to serve
- Long-context multimodal tasks: Image and text reasoning must be maintained coherently across extended conversations
- General assistant and chat: Meta designates Llama 4 Maverick 17B 128E Instruct FP8 as the intended product workhorse

### Consider alternatives when

- Extreme long documents: Llama 4 Scout's 10M token context window is purpose-built for that use case
- Text-only workload: The MoE overhead of loading all experts into memory is not offset by quality gains over a dense model at similar cost
- Maximum reasoning depth: Llama 4 Behemoth (when available) or other frontier reasoning models may be appropriate

## [Copy link to heading](#conclusion)Conclusion

Llama 4 Maverick 17B 128E Instruct FP8 combines native multimodality, a 128-expert MoE architecture, and strong benchmark results on image and text tasks at a fraction of the active-parameter cost of dense alternatives. For teams building multimodal production applications on open models, Llama 4 Maverick 17B 128E Instruct FP8 is the more capable of the two initial Llama 4 releases.