[Alibaba Cloud](/ai-gateway/models/labs/alibaba)

# Qwen 3 32B

Qwen 3 32B is a dense 32-billion-parameter model from Alibaba Cloud with context of 128K tokens and hybrid thinking modes, reaching performance levels previously associated with much larger models. Your use is subject to Alibaba Cloud's [Terms](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-product-terms-of-service-v-3-8-0) & [Privacy](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-privacy-policy) Policies.

ReasoningTool Use

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'alibaba/qwen-3-32b',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/qwen-3-32b) [API](/ai-gateway/models/qwen-3-32b/api) [About](/ai-gateway/models/qwen-3-32b/about) [Providers](/ai-gateway/models/qwen-3-32b/providers) [Throughput](/ai-gateway/models/qwen-3-32b/throughput) [Latency](/ai-gateway/models/qwen-3-32b/latency) [Uptime](/ai-gateway/models/qwen-3-32b/uptime) [Status](/ai-gateway/models/qwen-3-32b/status) [Similar](/ai-gateway/models/qwen-3-32b/similar) [FAQ](/ai-gateway/models/qwen-3-32b/faq)

## [Copy link to heading](#playground)Playground

Try out Qwen 3 32B by Alibaba Cloud. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75)Qwen 3 32B

![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=96&q=75)

Qwen 3 32B

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) [Bedrock](/ai-gateway/models/providers/bedrock) Legal:[Terms](https://aws.amazon.com/service-terms/)•[Privacy](https://aws.amazon.com/privacy/) | 128K | 8K | 0.2s |  | $0.15/M | $0.60/M |  | — |  |  |  | 04/28/2025 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) [Alibaba Cloud](/ai-gateway/models/providers/alibaba) Legal:[Terms](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-product-terms-of-service-v-3-8-0)•[Privacy](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-privacy-policy) | 128K | 8K | 0.4s | 111tps | $0.16/M | $0.64/M |  | — |  |  |  | 04/28/2025 |  |
| ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) [DeepInfra](/ai-gateway/models/providers/deepinfra) Legal:[Terms](https://deepinfra.com/terms)•[Privacy](https://deepinfra.com/privacy) | 41K | 16K | 0.3s |  | $0.10/M | $0.30/M |  | — |  |  |  | 04/28/2025 |  |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-alibaba-cloud)More models by Alibaba Cloud

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-27b](/ai-gateway/models/qwen3.8-27b) | 1M | 0.3s | 123tps | $0.10/M | $0.40/M | Read:$0.01/M Write:— | — | +1 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![runinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fruninfra.png%3Fv%3D1786910657446&w=48&q=75) |  |  | 08/14/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-2.4t-a95b](/ai-gateway/models/qwen3.8-2.4t-a95b) | 262K | 0.4s | 72tps | $1.65/M | $4.95/M | Read:$0.12/M Write:$2.5/M | — |  | ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![gmicloud logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgmicloud.png%3Fv%3D1785888489595&w=48&q=75) +3 |  |  | 08/03/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-max](/ai-gateway/models/qwen3.8-max) | 1M | 4.1s | 122tps | $2/M | $6/M | Read:$0.25/M Write:$2.5/M | — | +1 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) |  |  | 08/02/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.7-flash](/ai-gateway/models/qwen3.7-flash) | 991K | 3.2s | 132tps | $0.03/M+2 more | $0.13/M+2 more | Read: $0.01/M+2 more Write: $0.04/M+2 more | — | +2 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) |  |  | 07/28/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.7-plus](/ai-gateway/models/qwen3.7-plus) | 1M | 1.3s | 230tps | $0.40/M+1 more | $1.60/M+1 more | Read: $0.08/M+1 more Write: $0.5/M+1 more | — | +2 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) |  |  | 06/02/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3-vl-235b-a22b-instruct](/ai-gateway/models/qwen3-vl-235b-a22b-instruct) | 262K | 0.8s | 52tps | $0.20/M | $0.88/M | Read:$0.11/M Write:— | — |  | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) |  |  | 09/23/2025 |  |

## [Copy link to heading](#about-qwen-3-32b)About Qwen 3 32B

Qwen 3 32B is a fully dense model with no expert routing or sparse activation. All 32 billion parameters participate in generating each token. This architecture has a predictable operational profile: memory requirements are fixed, throughput is predictable, and there's no MoE infrastructure complexity to manage.

Alibaba Cloud positions Qwen 3 32B as reaching capability levels that Qwen2.5 required 72 billion parameters to achieve, a meaningful efficiency gain at the same parameter count from the third-generation architecture refinements across 64 transformer layers.

Hybrid thinking mode is available here as in the rest of the Qwen3 family. Activating thinking mode enables Qwen 3 32B to reason step-by-step before producing its answer, improving quality on problems requiring multi-step logic or structured derivation. Non-thinking mode bypasses the reasoning trace for applications where response speed takes priority. The budget control mechanism lets you set a token ceiling on the thinking phase, giving fine-grained control over the latency-quality tradeoff per request.

The model supports tool calling, agentic task scenarios, and MCP. The context window of 128K tokens accommodates long documents, multi-turn conversations, and retrieval-augmented generation (RAG) patterns where large amounts of source material need to fit in a single context.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: If your organization has compliance requirements tied to specific cloud infrastructure, reviewing the provider list and their data handling commitments is worthwhile before deploying at scale.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-qwen-3-32b)When to Use Qwen 3 32B

### Best for

- Long-document processing and analysis: The context window of 128K tokens, combined with dense 32B capacity, handles tasks like full-document summarization, cross-document comparison, and extended conversation history without chunking
- Complex instruction following: Dense models at this parameter scale reliably handle nuanced, multi-constraint instructions. Tasks that require careful attention to several simultaneous requirements (format, tone, content constraints, citation style) are well-served here
- Agentic workflows requiring sustained coherence: The window of 128K tokens helps Qwen 3 32B maintain context across extended multi-step interactions without losing track of earlier steps or decisions
- Coding tasks and technical writing: Strong benchmark performance in coding, combined with a context window large enough to hold substantial codebases or specifications, makes Qwen 3 32B useful for technical assistance workflows

### Consider alternatives when

- Serving cost at high volume dominates: The Qwen3-30B-A3B MoE activates only 3B parameters per inference, which can be substantially cheaper to serve for equivalent throughput. If cost efficiency dominates, the MoE variant is worth evaluating
- You need a higher quality ceiling: The Qwen3-235B-A22B MoE reaches higher benchmark performance on the hardest tasks, making it a better fit where capability headroom outweighs per-token cost
- Tasks are simple and short: For basic question-answering, short-form classification, or simple text formatting, the smaller Qwen3-14B will provide adequate quality at lower cost per token

## [Copy link to heading](#conclusion)Conclusion

Qwen 3 32B delivers strong dense-model performance in the Qwen3 family, reaching capability benchmarks that required a 72B-parameter model in the previous generation. It's a solid choice for long-context tasks, complex instruction following, and teams that want a simple dense model deployment without MoE infrastructure considerations. AI Gateway's provider pool gives it reliable availability through Bedrock, Alibaba Cloud, DeepInfra with a single integration.