Skip to content
Dashboard

o3-mini

o3-mini is a cost-efficient reasoning model in the o3 family, delivering strong chain-of-thought performance on math, code, and science at a fraction of full o3's cost, with configurable reasoning effort for flexible cost-quality tradeoffs.

Input and output price
Prices from: Input $1.10, Output $4.40, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'openai/o3-mini',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingProviders

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.

Provider
Context
Max Output
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
ZDR
No Training
Free Tier
Release Date
200K100K2.1 s
$1.10/M
$4.40/M
Read$0.55/M
$14/K
01/31/2025
200K100K1.7 s
$1.10/M
$4.40/M
Read$0.55/M
01/31/2025

Copy link to headingPlayground

Try out o3-mini by OpenAI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

openai logo
openai logo

o3-mini

Copy link to headingUptime

Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.

Copy link to headingThroughput

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.

Copy link to headingLatency

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.

Copy link to headingMore models by OpenAI

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
1.1M3.5 s41 tps
$10/M+2 more
$50/M+2 more
Read$1/M
Write$12.50/M
$10/K
+4
azure logo
openai logo
09/04/2026
1.1M1.6 s165 tps
$0.20/M+2 more
$1.20/M+2 more
Read$0.02/M
Write$0.25/M
$10/K
+4
azure logo
bedrock logo
openai logo
07/09/2026
1.1M4.1 s97 tps
$2/M+2 more
$10/M+2 more
Read$0.20/M
Write$2.50/M
$10/K
+4
azure logo
bedrock logo
openai logo
07/09/2026
1.1M2.2 s83 tps
$2/M+2 more
$12/M+2 more
Read$0.20/M
Write$2.50/M
$10/K
+4
azure logo
bedrock logo
openai logo
07/09/2026
1.1M0.8 s77 tps
$2.50/M+2 more
$15/M+2 more
Read$0.25/M
$10/K
+4
azure logo
openai logo
03/05/2026
400K5.8 s121 tps
$0.05/M
$0.40/M
Read$0.005/M
$14/K
+3
azure logo
openai logo
08/07/2025

o3-mini was released on January 31, 2025 as the cost-efficient tier of the o3 reasoning model family. It continues the pattern established by o1-mini: delivering strong chain-of-thought reasoning on structured domains (mathematics, coding, science) at a fraction of the full model's cost.

The model supports the reasoning_effort parameter, letting you control reasoning depth per request. Low effort for straightforward technical queries conserves tokens and reduces cost; high effort for competition-level problems applies the full reasoning capability. This flexibility lets you use o3-mini as the default for all technical queries rather than maintaining a routing layer.

With a context window of 200K tokens and support for the standard API features, o3-mini handles the same types of requests as full o3. The tradeoff is concentrated in reasoning depth: on the hardest problems, full o3 will produce more thorough analysis.

Copy link to headingWhat To Consider When Choosing a Provider

  • Configuration: o3-mini makes chain-of-thought reasoning affordable enough to run on every request rather than reserving it for the hardest problems. The reasoning_effort parameter enables further cost optimization.
  • Configuration: Like o1-mini before it, o3-mini concentrates its reasoning capability on structured problem domains rather than broad general knowledge.
  • Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the documentation for details.
  • Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.

Copy link to headingWhen to Use o3-mini

Best for

  • Math and science reasoning: Competition-level problems, derivations, and quantitative analysis at accessible cost
  • Code reasoning: Algorithm analysis, debugging, and optimization with step-by-step deliberation
  • High-frequency reasoning pipelines: Per-request chain-of-thought on technical workloads at scale
  • Education platforms: Tutoring and problem-solving assistance with visible reasoning steps
  • Cost-optimized reasoning: Tasks that benefit from deliberation but don't justify full o3 pricing

Consider alternatives when

  • Maximum reasoning quality: Full o3 for the hardest problems where every increment of accuracy matters
  • Broader knowledge needed: Full o3 or GPT-5 for tasks requiring wide-ranging factual recall
  • Fastest reasoning: O4-mini for a newer cost-efficient reasoning option with vision support
  • General-purpose tasks: GPT-5 mini for workloads that don't benefit from chain-of-thought

o3-mini makes chain-of-thought reasoning broadly accessible by bringing o3-family performance to a cost tier that scales. For technical workloads on AI Gateway where per-request reasoning is desirable but full o3 pricing is not, it provides the right balance.

Your use is subject to OpenAI's Terms & Privacy Policies.