[Thinking Machines](/ai-gateway/models/labs/thinkingmachines)

# Inkling

Inkling is an open-weights multimodal Mixture-of-Experts model that reasons over text, images, and audio. It supports controllable thinking effort and a context window of 262.1K tokens. Call Inkling on AI Gateway with `thinkingmachines/inkling`. Your use is subject to Thinking Machines's [Terms](https://thinkingmachines.ai/legal/terms/) & [Privacy](https://thinkingmachines.ai/legal/privacy/) Policies.

ReasoningTool UseVision (Image)File Input

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'thinkingmachines/inkling',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/inkling) [API](/ai-gateway/models/inkling/api) [About](/ai-gateway/models/inkling/about) [Providers](/ai-gateway/models/inkling/providers) [Throughput](/ai-gateway/models/inkling/throughput) [Latency](/ai-gateway/models/inkling/latency) [Uptime](/ai-gateway/models/inkling/uptime) [Status](/ai-gateway/models/inkling/status) [Similar](/ai-gateway/models/inkling/similar) [FAQ](/ai-gateway/models/inkling/faq)

## [Copy link to heading](#playground)Playground

Try out Inkling by Thinking Machines. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![thinkingmachines logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fthinkingmachines.png%3Fv%3D1787369193129&w=48&q=75)Inkling

![thinkingmachines logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fthinkingmachines.png%3Fv%3D1787369193129&w=96&q=75)

Inkling

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Regional Inference | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) [Baseten](/ai-gateway/models/providers/baseten) Legal:[Terms](https://www.baseten.co/terms-and-conditions/)•[Privacy](https://www.baseten.co/privacy-policy/) | 256K | 256K | 0.5s | 250tps | $1/M | $4.05/M | Read:$0.17/M Write:— | — | +1 |  |  | US | 07/15/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![togetherai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftogetherai.png&w=48&q=75) [Together AI](/ai-gateway/models/providers/togetherai) Legal:[Terms](https://www.together.ai/terms-of-service)•[Privacy](https://www.together.ai/privacy) | 256K | 256K | 0.5s | 232tps | $1.20/M | $4.05/M | Read:$0.20/M Write:— | — | +1 |  |  |  | 07/15/2026 |  |
| ![modal logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmodal-256-black.png%3Fv%3D1786301176874&w=48&q=75) [Modal](/ai-gateway/models/providers/modal) Legal:[Terms](https://modal.com/legal/terms)•[Privacy](https://modal.com/legal/privacy-policy) | 262K | 262K | 0.5s | 192tps | $1.20/M | $5/M | Read:$0.27/M Write:— | — | +1 |  |  |  | 07/15/2026 |  |
| ![thinkingmachines logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fthinkingmachines.png%3Fv%3D1787369193129&w=48&q=75) [Thinking Machines](/ai-gateway/models/providers/thinkingmachines) Legal:[Terms](https://thinkingmachines.ai/legal/terms/)•[Privacy](https://thinkingmachines.ai/legal/privacy/) | 256K | 256K | 1.7s | 137tps | $1/M | $4.05/M | Read:$0.17/M Write:— | — | +1 |  |  |  | 07/15/2026 |  |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-thinking-machines)More models by Thinking Machines

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![thinkingmachines logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fthinkingmachines.png%3Fv%3D1787369193129&w=48&q=75) [thinkingmachines/inkling-small](/ai-gateway/models/inkling-small) | 1M | 0.5s | 317tps | $0.30/M | $1.20/M | Read:$0.06/M Write:— | — | +2 | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![thinkingmachines logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fthinkingmachines.png%3Fv%3D1787369193129&w=48&q=75) +1 |  |  | 07/30/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#about-inkling)About Inkling

Inkling became available on AI Gateway on July 15, 2026. Inkling is a decoder-only transformer with a sparse Mixture-of-Experts (MoE) backbone: 66 layers, 975 billion total parameters, and 41 billion active per token, with each token routed to 6 of 256 experts plus 2 shared experts. Attention mixes local and global layers, and the context window is 262.1K tokens. Thinking Machines released the weights under the Apache 2.0 license.

Multimodality is native rather than bolted on. Images enter through a hierarchical patch encoder and audio through discrete token encoding, and both are processed jointly with text by the same decoder. Inkling accepts pixel-based images with each dimension between 40px and 4096px, and WAV audio sampled at 16kHz. Inkling transcribes speech, follows spoken instructions, and answers questions about recordings, scoring 91.4% on VoiceBench, 77.2% on MMAU, and 56.6% on Audio MC. On vision, Inkling scores 73.5% on MMMU Pro and 78.1% on CharXiv reasoning questions, rising to 82.0% when it uses a Python tool to zoom into and crop the image.

On agentic and reasoning evaluations at maximum effort, Inkling scores 77.6% on SWE-bench Verified, 54.3% on SWE-bench Pro (public), 63.8% on Terminal-Bench 2.1, 76.0% on MCP Atlas, and 45.5% on Toolathlon Verified. Reasoning results include 87.2% on GPQA Diamond, 97.1% on AIME 2026, and 29.7% on Humanity's Last Exam text-only, which rises to 46.0% with tools. Inkling scores 79.8% on IFBench for instruction following.

Controllable thinking effort is the setting you tune most. Raising effort spends more thinking tokens for higher scores, and lowering it returns answers sooner for less. Inkling reaches a given score at fewer thinking tokens than the open-weights models Thinking Machines compared it against, so sweep the effort setting across a representative slice of your traffic before you fix a value.

Inkling also aims for calibrated confidence. It hedges or says it doesn't know rather than guessing, which helps in forecasting and in any workflow where a confident wrong answer costs more than an uncertain one.

Set the model to `thinkingmachines/inkling` in the AI SDK, Chat Completions API, Responses API, Messages API, or other API formats, from TypeScript or Python. AI Gateway serves Inkling through Baseten, Together AI, Modal, Thinking Machines, with retries and failover, and mirrors provider pricing with no markup and no platform fee on inference, including on Bring Your Own Key (BYOK) requests.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Thinking Machines states plainly that Inkling is not the strongest overall model available, open or closed. Its case is breadth: one model that accepts text, images, and audio, follows instructions closely, and exposes an effort dial. Factual recall is the weakest area, with 43.9% on SimpleQA Verified, so pair Inkling with retrieval or search when answers depend on specific facts.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-inkling)When to Use Inkling

### Best for

- Multimodal Agent Backends: Text, image, and audio input handled by a single model rather than three
- Speech and Audio Reasoning: Transcription, spoken instructions, and questions over longer recordings
- Chart and Document Vision: Charts, diagrams, and visual math, with a Python tool for zooming and cropping
- Tool-Heavy Agent Workflows: Broad tool use across harnesses, with 76.0% on MCP Atlas
- Mixed Workload Consolidation: One generalist covering reasoning, coding, chat, and multimodal input
- Effort-Tuned Cost Control: Thinking effort set per request to balance answer quality against latency

### Consider alternatives when

- Peak Coding Scores: Dedicated coding models post higher SWE-bench and Terminal-Bench 2.1 results
- Factual Recall Workloads: 43.9% on SimpleQA Verified means knowledge-heavy answers need a retrieval layer
- Lower Cost Per Task: Inkling Small matches or beats Inkling on many evaluations at a quarter of the size
- Text-Only Pipelines: A text-focused model fits better when image and audio input never apply

## [Copy link to heading](#conclusion)Conclusion

Inkling is a broad, balanced model rather than a leader on any single benchmark family. Use Inkling when one model needs to read images, listen to audio, call tools, and write code, and when you want an effort dial to control what each request spends.