[Thinkingmachines](/ai-gateway/models/labs/thinkingmachines)

# Inkling Small

Inkling Small is the smaller model in the Inkling family at 276 billion total parameters and 12 billion active. It matches or beats Inkling on many coding, reasoning, and tool-use evaluations, reasons natively over images and audio, and supports a context window of 1M tokens. Your use is subject to Thinkingmachines's [Terms](https://thinkingmachines.ai/legal/terms/) & [Privacy](https://thinkingmachines.ai/legal/privacy/) Policies.

ReasoningTool UseVision (Image)File InputImplicit Caching

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'thinkingmachines/inkling-small',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/inkling-small) [API](/ai-gateway/models/inkling-small/api) [About](/ai-gateway/models/inkling-small/about) [Providers](/ai-gateway/models/inkling-small/providers) [Throughput](/ai-gateway/models/inkling-small/throughput) [Latency](/ai-gateway/models/inkling-small/latency) [Uptime](/ai-gateway/models/inkling-small/uptime) [Status](/ai-gateway/models/inkling-small/status) [Similar](/ai-gateway/models/inkling-small/similar) [FAQ](/ai-gateway/models/inkling-small/faq)

## [Copy link to heading](#playground)Playground

Try out Inkling Small by Thinkingmachines. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![thinkingmachines logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fthinkingmachines.png&w=48&q=75)Inkling Small

![thinkingmachines logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fthinkingmachines.png&w=96&q=75)

Inkling Small

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) [Baseten](/ai-gateway/models/providers/baseten) Legal:[Terms](https://www.baseten.co/terms-and-conditions/)•[Privacy](https://www.baseten.co/privacy-policy/) | 1M | 1M | 0.5s | 302tps | $0.50/M | $1.20/M | Read:$0.1/M Write:— | — | +2 |  |  | 07/30/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) [DeepInfra](/ai-gateway/models/providers/deepinfra) Legal:[Terms](https://deepinfra.com/terms)•[Privacy](https://deepinfra.com/privacy) | 1M | 1M | 0.5s | 114tps | $0.58/M | $1.44/M | Read:$0.12/M Write:— | — | +2 |  |  | 07/30/2026 |  |
| ![togetherai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftogetherai.png&w=48&q=75) [Together AI](/ai-gateway/models/providers/togetherai) Legal:[Terms](https://www.together.ai/terms-of-service)•[Privacy](https://www.together.ai/privacy) | 1M | 1M | 1.7s | 56tps | $0.50/M | $1.20/M | Read:$0.1/M Write:— | — | +2 |  |  | 07/30/2026 |  |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-thinkingmachines)More models by Thinkingmachines

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![thinkingmachines logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fthinkingmachines.png&w=48&q=75) [thinkingmachines/inkling](/ai-gateway/models/inkling) | 262K | 0.4s | 247tps | $1/M | $4.05/M | Read:$0.17/M Write:— | — | +1 | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![modal logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmodal-256-black.png%3Fv%3D1786301176874&w=48&q=75) ![togetherai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ftogetherai.png&w=48&q=75) |  |  | 07/15/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#about-inkling-small)About Inkling Small

Inkling Small became available on AI Gateway on July 30, 2026. Inkling Small is a 42-layer decoder-only transformer with a sparse Mixture-of-Experts (MoE) backbone: 276 billion total parameters, 12 billion active per token, and each token routed to 6 of 256 experts plus 2 shared experts. Attention mixes local and global layers, images enter through a hierarchical patch encoder, audio through discrete token encoding, and the context window is 1M tokens. Thinkingmachines released the weights under the Apache 2.0 license.

On agentic coding, Inkling Small scores 80.2% on SWE-bench Verified, 55.9% on SWE-bench Pro (public), 64.7% on Terminal-Bench 2.1, and 48.7% on SciCode. Each of those sits at or above Inkling, which scores 77.6%, 54.3%, 63.8%, and 46.1%. Tool use follows the same pattern: 54.4% on Toolathlon Verified against Inkling's 45.5%, and 79.6% on the public split of MCP Atlas. On reasoning it scores 89.5% on GPQA Diamond, 90.2% on HMMT February 2026, 31.6% on Humanity's Last Exam text-only, and 40.1% on ARC-AGI-2.

The gap is knowledge. Inkling Small scores 20.6% on SimpleQA Verified where Inkling scores 43.9%, and it trails on the AA Omniscience index and on Tau 3 Banking, a multi-turn domain agent evaluation. Audio results sit close behind Inkling at 90.1% on VoiceBench, 77.0% on MMAU, and 54.9% on Audio MC. Treat Inkling Small as a capable reasoner with a smaller store of memorized facts, and give it a retrieval path when questions turn factual.

Vision is a practical strength. Inkling Small scores 74.0% on MMMU Pro and 77.4% on CharXiv reasoning questions, rising to 81.3% when it crops, zooms, and inspects images programmatically. That helps most on documents, forms, and charts where the detail that answers the question is small.

Controllable thinking effort runs from minimal to maximum, so you can trade answer quality against cost and latency per request. Sweep the setting across representative traffic rather than fixing it up front.

Inkling Small is compatible with Zero Data Retention on AI Gateway. Turn it on team-wide from the dashboard, or per request with `zeroDataRetention: true`, and AI Gateway routes only to providers that delete prompts and responses after each request. Set the model to `thinkingmachines/inkling-small` in the AI SDK, Chat Completions API, Responses API, Messages API, or other API formats, from TypeScript or Python. To use Inkling Small in a coding agent, run `vercel ai-gateway coding-agents setup`, then select `thinkingmachines/inkling-small` in the agent's model configuration.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Size shows up in world knowledge, not in reasoning or coding. Inkling Small scores 20.6% on SimpleQA Verified against Inkling's 43.9%, so anything that depends on recalling specific facts needs retrieval or web search alongside the model. Reasoning, coding, and tool-use scores hold up, and several of them land above Inkling's.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-inkling-small)When to Use Inkling Small

### Best for

- High-Volume Coding Agents: Agentic coding and tool use, with 80.2% on SWE-bench Verified
- Tool Orchestration Pipelines: 54.4% on Toolathlon Verified and 79.6% on the public MCP Atlas split
- Document and Chart Analysis: Programmatic cropping and zooming to read small visual detail
- Compact Multimodal Apps: Native text, image, and audio input in a smaller model than Inkling
- Effort-Tuned Latency Budgets: Thinking effort set from minimal to maximum on each request
- Zero Data Retention Routing: Requests routed only to providers that delete prompts and responses

### Consider alternatives when

- World Knowledge Questions: 20.6% on SimpleQA Verified against Inkling's 43.9% is a real recall gap
- Strongest Audio Results: Inkling scores higher on VoiceBench, MMAU, and Audio MC
- Multi-Turn Domain Agents: Inkling scores higher on Tau 3 Banking and on broader factual evaluations
- Frontier Coding Ceiling: Closed frontier models still lead on SWE-bench Verified and Terminal-Bench 2.1

## [Copy link to heading](#conclusion)Conclusion

Inkling Small keeps Inkling's coding, reasoning, tool use, and multimodal input in a model a quarter the size, and gives up world knowledge to get there. Use Inkling Small for coding agents, tool pipelines, and document work, and add retrieval when the questions turn factual.