Skip to content
Dashboard

Inkling Small

Inkling-Small is a lighter-weight model with 12B active parameters, trained with a similar recipe, to Inkling that achieves strong performance with even lower cost and latency. Your use subject to Thinkingmachines's Terms & Privacy Policies.

ReasoningTool UseVision (Image)File InputImplicit Caching
index.ts
import { streamText } from 'ai'
const result = streamText({
model: 'thinkingmachines/inkling-small',
prompt: 'Why is the sky blue?'
})

Playground

Try out Inkling Small by Thinkingmachines. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

thinkingmachines logo
thinkingmachines logo

Inkling Small

Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.

Provider
Context
Max Output
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
ZDR
No Training
Release Date
1M1M
0.2s
528tps
$0.50/M
$1.20/M
Read:$0.1/M
Write:
+2
07/30/2026
Throughput

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.

Latency

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.

Uptime

Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.

More models by Thinkingmachines

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Release Date
256K
0.7s
201tps
$1/M
$4.05/M
Read:$0.17/M
Write:
+1
baseten logo
togetherai logo
07/15/2026