Skip to content
Dashboard

Gemini 2.5 Flash Lite

Gemini 2.5 Flash Lite is the fastest and most affordable model in the Gemini 2.5 family, with configurable thinking, a context window of 1.0M tokens, and benchmark improvements over 2.0 Flash-Lite across coding, math, and science, at a price designed for high-throughput agentic pipelines.

Input and output price
Prices from: Input $0.10, Output $0.40, Per 1M tokens
24h uptime
Loading AI Gateway uptime
import { streamText } from 'ai'
const result = streamText({
model: 'google/gemini-2.5-flash-lite',
prompt: 'Why is the sky blue?'
})
Read docs

Copy link to headingProviders

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.

Provider
Context
Max Output
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
ZDR
No Training
Regional Inference
Free Tier
Release Date
1M66K0.5 s414 tps
$0.10/M+2 more
$0.40/M+2 more
Read$0.01/M
$35/K+1 more
+3
US
EU
06/17/2025
1M66K0.2 s
$0.10/M+2 more
$0.40/M+2 more
Read$0.01/M
$35/K+1 more
+3
06/17/2025

Copy link to headingPlayground

Try out Gemini 2.5 Flash Lite by Google. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

google logo
google logo

Gemini 2.5 Flash Lite

Copy link to headingUptime

Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.

Copy link to headingThroughput

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.

Copy link to headingLatency

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.

Copy link to headingMore models by Google

Model
Context
Latency
Throughput
Input
Output
Cache
Web Search
Capabilities
Providers
ZDR
No Training
Free Tier
Release Date
1M1.3 s399 tps
$0.75/M
$3.75/M
Read$0.08/M
$14/K+1 more
+3
google logo
vertex logo
09/02/2026
1M0.8 s269 tps
$0.75/M
$3.75/M
Read$0.08/M
$14/K+1 more
+3
google logo
vertex logo
08/13/2026
1M0.4 s247 tps
$0.30/M
$2.50/M
Read$0.03/M
$14/K+1 more
+3
google logo
vertex logo
07/21/2026
1M1.5 s159 tps
$1.50/M
$9/M
Read$0.15/M
$14/K+1 more
+3
google logo
vertex logo
05/19/2026
1M0.4 s225 tps
$0.25/M
$1.50/M
Read$0.03/M
$14/K+1 more
+3
google logo
vertex logo
05/07/2026
1M0.6 s181 tps
$0.50/M+1 more
$3/M+1 more
Read$0.05/M
$14/K+1 more
+3
google logo
vertex logo
12/17/2025

Copy link to headingAbout Gemini 2.5 Flash Lite

Gemini 2.5 Flash Lite is the efficiency tier of the Gemini 2.5 family, released June 17, 2025 alongside 2.5 Flash and 2.5 Pro going to general availability. It runs faster and costs less than any other 2.5 model while outperforming 2.0 Flash-Lite on benchmarks that matter for real-world developer tasks: coding, mathematics, scientific reasoning, and instruction following.

Configurable thinking is the feature that most distinguishes Gemini 2.5 Flash Lite from 2.0 Flash-Lite. At inference time, you set a thinking level (minimal, low, medium, or high) to allocate more deliberation to harder problems without switching endpoints. This is the same thinking mechanism available in 2.5 Flash and 2.5 Pro, scaled down to the lite budget. For tasks that occasionally need more reasoning depth than a pure speed-first model provides, the thinking toggle avoids the cost jump of routing to a full reasoning model.

For teams running 2.0 Flash-Lite in production and evaluating a 2.5 upgrade path, Gemini 2.5 Flash Lite is the migration-friendly choice: better benchmark performance, thinking capability, and latency that matches or beats the previous generation.

Copy link to headingWhat To Consider When Choosing a Provider

  • Configuration: Applications using the thinking feature should benchmark total token cost under realistic thinking budgets, as thinking tokens contribute to output costs.
  • Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the documentation for details.
  • Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.

Copy link to headingWhen to Use Gemini 2.5 Flash Lite

Best for

  • High-volume agentic pipelines needing occasional reasoning: The thinking toggle allows selective deliberation on harder steps without paying full 2.5 Flash prices for every call in the pipeline
  • Migrating from 2.0 Flash-Lite: Benchmark improvements across coding and math mean the upgrade delivers measurable quality gains on common developer tasks at comparable cost
  • Latency-sensitive applications within the 2.5 family: When 2.5 Flash or 2.5 Pro latency is too high for the user experience, Flash-Lite provides 2.5-generation quality at the fastest 2.5 response times
  • Translation, classification, and data extraction at scale: Strong instruction following and fast response make it a reliable workhorse for structured-output production tasks

Consider alternatives when

  • Maximum reasoning depth is required: 2.5 Flash or 2.5 Pro with uncapped thinking budget is more appropriate for the most complex multi-step problems
  • Image generation is needed: Gemini 2.5 Flash Lite does not generate images. Gemini models with native image output are available in the 2.5 Flash Image and 3.x families
  • Your workload is pure annotation/extraction without reasoning: For text-output-only extraction at maximum cost efficiency, 2.0 Flash-Lite's lower price floor may be preferable

Gemini 2.5 Flash Lite closes the gap between 2.0 Flash-Lite and the full 2.5 Flash tier. It delivers better benchmark performance and thinking capability at the same latency profile teams already depend on. For 2.0 Flash-Lite users, it's the natural upgrade.

Your use is subject to Google's Terms & Privacy Policies.