[Alibaba Cloud](/ai-gateway/models/labs/alibaba)

# Qwen3 Max

Qwen3 Max is Alibaba Cloud's trillion-parameter MoE language model with a context window of 262.1K tokens, delivering competitive performance on coding, mathematics, and enterprise tool-use tasks. Your use is subject to Alibaba Cloud's [Terms](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-product-terms-of-service-v-3-8-0) & [Privacy](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-privacy-policy) Policies.

Tool UseImplicit Caching

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'alibaba/qwen3-max',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/qwen3-max) [API](/ai-gateway/models/qwen3-max/api) [About](/ai-gateway/models/qwen3-max/about) [Providers](/ai-gateway/models/qwen3-max/providers) [Throughput](/ai-gateway/models/qwen3-max/throughput) [Latency](/ai-gateway/models/qwen3-max/latency) [Uptime](/ai-gateway/models/qwen3-max/uptime) [Status](/ai-gateway/models/qwen3-max/status) [Similar](/ai-gateway/models/qwen3-max/similar) [FAQ](/ai-gateway/models/qwen3-max/faq)

## [Copy link to heading](#playground)Playground

Try out Qwen3 Max by Alibaba Cloud. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75)Qwen3 Max

![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=96&q=75)

Qwen3 Max

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) [Alibaba Cloud](/ai-gateway/models/providers/alibaba) Legal:[Terms](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-product-terms-of-service-v-3-8-0)•[Privacy](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-privacy-policy) | 262K | 33K | 2.1s | 46tps | $1.20/M+2 more | $6/M+2 more | Read: $0.24/M+2 more Write: — | — |  |  |  |  | 09/23/2025 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![novita logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnovita.png&w=48&q=75) [Novita AI](/ai-gateway/models/providers/novita) Legal:[Terms](https://novita.ai/legal/terms-of-service)•[Privacy](https://novita.ai/legal/privacy-policy) | 262K | 66K | 1.2s | 62tps | $0.85/M+2 more | $3.38/M+2 more |  | — |  |  |  |  | 09/23/2025 |  |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-alibaba-cloud)More models by Alibaba Cloud

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-flash](/ai-gateway/models/qwen3.8-flash) | 991K | 1.8s | 99tps | $0.16/M | $0.47/M | Read:$0.02/M Write:$0.20/M | — | +2 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) |  |  |  | 08/26/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-27b](/ai-gateway/models/qwen3.8-27b) | 1M | 0.5s | 122tps | $0.10/M | $0.40/M | Read:$0.01/M Write:$0.63/M | — | +1 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![morph logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmorph.png&w=48&q=75) +2 |  |  |  | 08/14/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-2.4t-a95b](/ai-gateway/models/qwen3.8-2.4t-a95b) | 262K | 0.4s | 166tps | $1.65/M | $4.95/M | Read:$0.12/M Write:$2.50/M | — |  | ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![gmicloud logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgmicloud.png%3Fv%3D1785888489595&w=48&q=75) +4 |  |  |  | 08/03/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-max](/ai-gateway/models/qwen3.8-max) | 1M | 3.1s | 157tps | $2/M | $6/M | Read:$0.25/M Write:$2.50/M | — | +1 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) |  |  |  | 08/02/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.7-flash](/ai-gateway/models/qwen3.7-flash) | 991K | 1.2s | 172tps | $0.03/M+2 more | $0.13/M+2 more | Read: $0.006/M+2 more Write: $0.04/M+2 more | — | +2 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) |  |  |  | 07/28/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.7-plus](/ai-gateway/models/qwen3.7-plus) | 1M | 1.7s | 58tps | $0.40/M+1 more | $1.60/M+1 more | Read: $0.08/M+1 more Write: $0.50/M+1 more | — | +2 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) |  |  |  | 06/02/2026 |  |

## [Copy link to heading](#about-qwen3-max)About Qwen3 Max

Qwen3 Max is the largest model in Alibaba Cloud's Qwen3 line, built on a mixture-of-experts (MoE) architecture with over one trillion total parameters. The MoE design allocates computation selectively, enabling performance without activating the full parameter count on every token.

The context window of 262.1K tokens makes it practical for tasks that earlier-generation models had to split across multiple calls: ingesting entire codebases, indexing long legal or financial documents, or tracking dependencies across extended multi-turn conversations. Context caching further reduces the cost of repeatedly processing the same long prefix.

Qwen3 Max performs strongly on structured-output and tool-use benchmarks, recording 74.8 on Tau2-Bench and 79.3% accuracy on LiveBench. On software engineering tasks measured by SWE-bench Verified, Qwen3 Max scored 69.6. These results reflect a consistent emphasis on reliability for enterprise tasks: JSON generation, HTML/CSS formatting, API function calling, and multi-step agentic workflows where predictable output structure matters.

Alibaba Cloud positions Qwen3 Max with native bilingual strength in Chinese and English, alongside broad multilingual support. The model is available via API only. Weights aren't publicly released.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: For regulated industries requiring data-residency guarantees, cross-reference the geographic deployment region of your chosen provider against applicable compliance frameworks before routing production traffic.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-qwen3-max)When to Use Qwen3 Max

### Best for

- Structured enterprise automation: High-volume workloads that require reliable JSON, XML, or formatted report output
- Long-document analysis: Contracts, scientific papers, and codebases where the full context must remain in-window
- Multi-step function calling: Complex agentic workflows that chain multiple tool invocations
- Professional-grade quantitative work: Mathematical reasoning and quantitative problem-solving at expert difficulty
- Bilingual Chinese-English applications: Products where both languages need equal-quality handling

### Consider alternatives when

- Visible chain-of-thought needed: Consider Qwen3-Max-Thinking when you need extended reasoning with visible step traces
- Creative and conversational writing: Open-ended storytelling or conversational warmth is the primary requirement
- Strict token budgets: A smaller open-weight model may meet your quality bar at lower cost per token
- Latency-critical workloads: Response latency is more important than depth of reasoning

## [Copy link to heading](#conclusion)Conclusion

Qwen3 Max brings trillion-parameter scale to tasks that benefit most from it: long-context document work, structured enterprise output, and complex tool use. Its context window of 262.1K tokens and strong benchmark results make it a credible choice for production deployments where reliability and breadth of capability take precedence over speed.