[Alibaba Cloud](/ai-gateway/models/labs/alibaba)

# Qwen3-14B

Qwen3-14B is a 14-billion-parameter dense language model from Alibaba Cloud that combines hybrid thinking modes with context of 41.0K tokens, delivering Qwen2.5-32B-class capability at a fraction of the parameter count. Your use is subject to Alibaba Cloud's [Terms](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-product-terms-of-service-v-3-8-0) & [Privacy](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-privacy-policy) Policies.

ReasoningTool Use

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'alibaba/qwen-3-14b',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/qwen-3-14b) [API](/ai-gateway/models/qwen-3-14b/api) [About](/ai-gateway/models/qwen-3-14b/about) [Providers](/ai-gateway/models/qwen-3-14b/providers) [Latency](/ai-gateway/models/qwen-3-14b/latency) [Uptime](/ai-gateway/models/qwen-3-14b/uptime) [Status](/ai-gateway/models/qwen-3-14b/status) [Similar](/ai-gateway/models/qwen-3-14b/similar) [FAQ](/ai-gateway/models/qwen-3-14b/faq)

## [Copy link to heading](#playground)Playground

Try out Qwen3-14B by Alibaba Cloud. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75)Qwen3-14B

![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=96&q=75)

Qwen3-14B

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) [DeepInfra](/ai-gateway/models/providers/deepinfra) Legal:[Terms](https://deepinfra.com/terms)•[Privacy](https://deepinfra.com/privacy) | 41K | 16K | 0.2s |  | $0.12/M | $0.24/M |  | — |  |  |  |  | 04/28/2025 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-alibaba-cloud)More models by Alibaba Cloud

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-flash](/ai-gateway/models/qwen3.8-flash) | 991K | 4.0s | 102tps | $0.16/M | $0.47/M | Read:$0.02/M Write:$0.20/M | — | +2 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) |  |  |  | 08/26/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-27b](/ai-gateway/models/qwen3.8-27b) | 1M | 0.6s | 131tps | $0.10/M | $0.40/M | Read:$0.01/M Write:$0.63/M | — | +1 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![morph logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmorph.png&w=48&q=75) +2 |  |  |  | 08/14/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-2.4t-a95b](/ai-gateway/models/qwen3.8-2.4t-a95b) | 262K | 0.4s | 106tps | $1.65/M | $4.95/M | Read:$0.12/M Write:$2.50/M | — |  | ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![gmicloud logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgmicloud.png%3Fv%3D1785888489595&w=48&q=75) +4 |  |  |  | 08/03/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-max](/ai-gateway/models/qwen3.8-max) | 1M | 3.6s | 128tps | $2/M | $6/M | Read:$0.25/M Write:$2.50/M | — | +1 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) |  |  |  | 08/02/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.7-flash](/ai-gateway/models/qwen3.7-flash) | 991K | 1.9s | 158tps | $0.03/M+2 more | $0.13/M+2 more | Read: $0.006/M+2 more Write: $0.04/M+2 more | — | +2 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) |  |  |  | 07/28/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.7-plus](/ai-gateway/models/qwen3.7-plus) | 1M | 2.2s | 58tps | $0.40/M+1 more | $1.60/M+1 more | Read: $0.08/M+1 more Write: $0.50/M+1 more | — | +2 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) |  |  |  | 06/02/2026 |  |

## [Copy link to heading](#about-qwen3-14b)About Qwen3-14B

Qwen3-14B is a dense transformer model with no sparse routing or mixture-of-experts. Every inference call activates all 14 billion parameters. This architecture trades raw efficiency for predictability: memory requirements and compute costs stay consistent across request types, which simplifies capacity planning.

The model includes Alibaba Cloud's hybrid thinking system. In thinking mode, Qwen3-14B works through a chain-of-thought before producing its final answer, allocating more compute to harder problems. In non-thinking mode, it responds immediately without the intermediate reasoning trace. The `enable_thinking` parameter controls which mode activates. You can adjust the thinking budget per request to match how much latency you're willing to accept.

Within the Qwen3 family, the 14B sits at a practical inflection point. Alibaba Cloud's benchmarks show Qwen3-14B matches Qwen2.5-32B-Base. You get the previous generation's mid-tier performance from a model less than half the size. That translates directly to lower hosting costs for teams running inference at scale.

The model covers 119 languages and dialects across Indo-European, Sino-Tibetan, Afro-Asiatic, Austronesian, and other language families. The result is strong coverage across coding, mathematics, and general instruction-following tasks.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Your choice of provider may matter for latency-sensitive applications or where data residency requirements constrain which infrastructure regions are acceptable.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-qwen3-14b)When to Use Qwen3-14B

### Best for

- Reasoning-intensive tasks on a budget: The hybrid thinking mode lets you activate deep reasoning selectively without committing to a larger model full-time. Use thinking mode for complex derivations and non-thinking mode for fast follow-up queries in the same session
- Multilingual applications: With 119 languages covered, Qwen3-14B suits applications that need to handle user input from diverse linguistic backgrounds, customer support platforms, global content tools, or localization pipelines
- Code generation and review: The model handles code completion, explanation, and debugging across common programming languages
- Balanced latency and quality: When you need better output quality than the smallest models but can't justify the compute cost of the 32B or larger variants, the 14B sits in a useful middle ground

### Consider alternatives when

- Higher reasoning headroom is needed: For the hardest mathematical proofs, complex multi-step logic, or the most demanding coding challenges, Qwen3-32B or the MoE variants offer stronger ceiling performance
- Throughput at the lowest possible cost per token: The Qwen3-30B-A3B MoE model activates only 3B parameters per inference, which can be significantly cheaper to serve despite its larger total parameter count
- Vision or multimodal inputs are required: Qwen3-14B handles text only; a multimodal model would be needed for image or audio processing tasks

## [Copy link to heading](#conclusion)Conclusion

Qwen3-14B gives teams a dense, fully-activating model with the flexibility to run reasoning-heavy or latency-optimized inference from the same checkpoint. Accessing it through AI Gateway removes the overhead of managing multiple provider accounts while keeping automatic failover and consolidated billing in place.