[Tencent Cloud](/ai-gateway/models/labs/tencent)

# Hy3

Hy3 is an Apache 2.0 Mixture-of-Experts model from the Tencent Cloud Hunyuan team, with 295B total parameters, 21B active per token, and three selectable reasoning levels. It supports a context window of 262.1K tokens and a max output of 262.1K tokens per request. Your use is subject to Tencent Cloud's [Terms](https://www.tencentcloud.com/document/product/301/78869) & [Privacy](https://www.tencentcloud.com/document/product/1300/78952) Policies.

Implicit CachingReasoningTool Use

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'tencent/hy3',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/hy3) [API](/ai-gateway/models/hy3/api) [About](/ai-gateway/models/hy3/about) [Providers](/ai-gateway/models/hy3/providers) [Throughput](/ai-gateway/models/hy3/throughput) [Latency](/ai-gateway/models/hy3/latency) [Uptime](/ai-gateway/models/hy3/uptime) [Status](/ai-gateway/models/hy3/status) [Similar](/ai-gateway/models/hy3/similar) [FAQ](/ai-gateway/models/hy3/faq)

## [Copy link to heading](#playground)Playground

Try out Hy3 by Tencent Cloud. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![tencent logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftencent.png%3Fv%3D1787259795229&w=48&q=75)Hy3

![tencent logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftencent.png%3Fv%3D1787259795229&w=96&q=75)

Hy3

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) [DeepInfra](/ai-gateway/models/providers/deepinfra) Legal:[Terms](https://deepinfra.com/terms)•[Privacy](https://deepinfra.com/privacy) | 262K | 262K | 0.2 s | 53 tps | $0.14/M | $0.58/M | Read:$0.04/M Write:— | — |  |  |  |  | 07/06/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![novita logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnovita.png&w=48&q=75) [Novita AI](/ai-gateway/models/providers/novita) Legal:[Terms](https://novita.ai/legal/terms-of-service)•[Privacy](https://novita.ai/legal/privacy-policy) | 262K | 262K | 1.3 s | 139 tps | $0.14/M | $0.58/M | Read:$0.04/M Write:— | — |  |  |  |  | 07/06/2026 |  |
| ![gmicloud logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgmicloud.png%3Fv%3D1785888489595&w=48&q=75) [GMICloud](/ai-gateway/models/providers/gmicloud) Legal:[Terms](https://www.gmicloud.ai/en/terms-and-conditions)•[Privacy](https://www.gmicloud.ai/en/privacy-policy) | 262K | 262K | 1.7 s | 109 tps | $0.13/M | $0.52/M | Read:$0.03/M Write:— | — |  |  |  |  | 07/06/2026 |  |
| ![tencent logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftencent.png%3Fv%3D1787259795229&w=48&q=75) [Tencent Cloud](/ai-gateway/models/providers/tencent) Legal:[Terms](https://www.tencentcloud.com/document/product/301/78869)•[Privacy](https://www.tencentcloud.com/document/product/1300/78952) | 256K | 128K | 1.7 s | 175 tps | $0.13/M | $0.53/M | Read:$0.03/M Write:— | — |  |  |  |  | 07/06/2026 |  |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-tencent-cloud)More models by Tencent Cloud

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![tencent logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftencent.png%3Fv%3D1787259795229&w=48&q=75) [tencent/hy4\-preview](/ai-gateway/models/hy4-preview) | 1M | 3.3 s | 76 tps | $0.83/M | $2.50/M | Read:$0.04/M Write:— | — |  | ![tencent logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftencent.png%3Fv%3D1787259795229&w=48&q=75) |  |  |  | 08/28/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![tencent logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftencent.png%3Fv%3D1787259795229&w=48&q=75) [tencent/hy-mt2-plus](/ai-gateway/models/hy-mt2-plus) | 8K | 1.5 s | 140 tps | $0.07/M | $0.30/M |  | — |  | ![tencent logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftencent.png%3Fv%3D1787259795229&w=48&q=75) |  |  |  | 06/12/2026 |  |
| ![tencent logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftencent.png%3Fv%3D1787259795229&w=48&q=75) [tencent/hy-mt2-lite](/ai-gateway/models/hy-mt2-lite) | 8K | 1.5 s |  | $0.04/M | $0.18/M |  | — |  | ![tencent logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftencent.png%3Fv%3D1787259795229&w=48&q=75) |  |  |  | 06/12/2026 |  |
| ![tencent logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftencent.png%3Fv%3D1787259795229&w=48&q=75) [tencent/hy\-mt2-pro](/ai-gateway/models/hy-mt2-pro) | 8K | 1.5 s |  | $0.07/M | $0.30/M |  | — |  | ![tencent logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftencent.png%3Fv%3D1787259795229&w=48&q=75) |  |  |  | 05/21/2026 |  |

## [Copy link to heading](#about-hy3)About Hy3

Tencent Cloud released Hy3 on July 6, 2026 under the Apache 2.0 license, publishing BF16 and FP8 weights on Hugging Face, ModelScope, GitCode, and CNB. Hy3 comes out of the Hunyuan line and follows the Hy3 Preview release from late April 2026.

The architecture is a Mixture-of-Experts (MoE) with 295B total parameters and 21B active per token, spread across 80 layers and 192 routed experts with top-8 routing. Attention is grouped-query with 64 heads and eight key-value heads. A separate 3.8B multi-token prediction layer drafts more than one token per forward pass, which serving stacks use for speculative decoding. The context window is 262.1K tokens, and one request can return up to 262.1K tokens.

Reasoning depth is a per-request setting. The `reasoning_effort` field accepts `no_think` for a direct answer, `low` for shallow reasoning, and `high` for deep chain-of-thought on math, coding, and analysis. `no_think` is the default, so you opt into reasoning tokens rather than paying for them on every call.

Tencent Cloud focused on tool-call and output-format stability, drawing on feedback from more than 50 of its own products. On SWE-Bench Verified, Hy3 holds accuracy variance within 4% across agent scaffoldings including CodeBuddy, Cline, and KiloCode, so a result from one harness carries to another. Internal evaluations built on real-world scenarios put the hallucination rate at 5.4%, down from 12.5%, and the commonsense error rate at 12.7%, down from 25.4%. On a multi-turn suite covering coreference resolution, ellipsis recovery, and constraint inheritance, the issue rate fell from 17.4% to 7.9%.

Tencent Cloud also ran a blind evaluation in which 270 experts scored tasks drawn from their own work. Hy3 averaged 2.67 out of four against GLM-5.1 at 2.51, with the widest margins on frontend development, data and storage, and CI/CD tasks.

Call Hy3 with one API key and get provider routing, automatic failover, and built-in observability. Integrate through the AI SDK, the Chat Completions API, the Responses API, the Messages API, or other supported API formats. Pay $0.126 per million input tokens, $0.522 per million output tokens, and $0.0315 per million cached input tokens at current list rates.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Hy3 defaults to `no_think`, so a plain request returns a direct answer with no reasoning tokens. Set `reasoning_effort` to `high` on math, coding, and analysis work, and expect longer outputs and higher output-token spend when you do. Harness support for the field varies, so confirm your client forwards it.
- Configuration: Tencent Cloud recommends sampling at `temperature=0.9` and `top_p=1.0`. Start there before you tune, because the published quality figures assume those settings.
- Configuration: Hy3 takes text in and returns text out. Route requests that carry images, audio, or video to a multimodal model.
- Configuration: The hallucination, multi-turn, and expert-panel figures come from Tencent Cloud's own evaluations rather than a neutral third party. Validate on your own workload before you shift production traffic. For current throughput and latency, see live metrics on this page.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-hy3)When to Use Hy3

### Best for

- Tool-Heavy Agent Loops: Multi-step runs where 21B active parameters keep per-request compute low
- Long-Context Analysis: Repositories, contracts, and report sets that fill the 262.1K tokens window
- Mixed Reasoning Traffic: Requests that switch between direct answers and deep chain-of-thought
- Multi-Turn Assistants: Long dialogues that must track references, intent, and inherited constraints
- Grounded Question Answering: Work where a fabricated detail costs more than a missing answer

### Consider alternatives when

- Multimodal Inputs: Hy3 handles text only, so send images, audio, or video to a vision-capable model
- Frontier Coding Ceilings: Larger flagship models still lead on the hardest coding suites
- Fixed Client Parameters: A client that cannot pass `reasoning_effort` leaves Hy3 on direct answers
- Short Single-Turn Chat: Reasoning depth adds output overhead that brief exchanges never recover

## [Copy link to heading](#conclusion)Conclusion

Hy3 pairs a 295B MoE with a 21B active footprint, a context window of 262.1K tokens, and reasoning you turn up per request. The reliability work on tool calls, grounding, and multi-turn tracking makes Hy3 a practical agent default rather than a benchmark showcase. Route it through AI Gateway for failover, observability, and one key across the catalog.