[Google](/ai-gateway/models/labs/google)

# Gemini 3 Flash

Gemini 3 Flash delivers Gemini 3's pro-grade reasoning at flash-level latency and cost, outperforming Gemini 2.5 Pro across most benchmarks with meaningful gains in token efficiency. Your use is subject to Google's [Terms](https://policies.google.com/terms/generative-ai) & [Privacy](https://policies.google.com/privacy) Policies.

ReasoningTool UseFile InputVision (Image)Web Searchtiered-costImplicit Caching

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'google/gemini-3-flash',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/gemini-3-flash) [API](/ai-gateway/models/gemini-3-flash/api) [About](/ai-gateway/models/gemini-3-flash/about) [Providers](/ai-gateway/models/gemini-3-flash/providers) [Throughput](/ai-gateway/models/gemini-3-flash/throughput) [Latency](/ai-gateway/models/gemini-3-flash/latency) [Uptime](/ai-gateway/models/gemini-3-flash/uptime) [Status](/ai-gateway/models/gemini-3-flash/status) [Similar](/ai-gateway/models/gemini-3-flash/similar) [FAQ](/ai-gateway/models/gemini-3-flash/faq)

## [Copy link to heading](#playground)Playground

Try out Gemini 3 Flash by Google. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75)Gemini 3 Flash

![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=96&q=75)

Gemini 3 Flash

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![vertex logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fvertex%2520ai.png&w=48&q=75) [Google Vertex AI](/ai-gateway/models/providers/vertex) Legal:[Terms](https://cloud.google.com/terms/service-terms)•[Privacy](https://cloud.google.com/privacy) | 1M | 65K | 0.8s | 143tps | $0.50/M+3 more | $3/M+3 more | Read:$0.05/M+3 more Write:— | $14/K+1 more \+ input costs | +3 |  |  |  | 12/17/2025 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) [Google](/ai-gateway/models/providers/google) Legal:[Terms](https://policies.google.com/terms/generative-ai)•[Privacy](https://policies.google.com/privacy) | 1M | 64K | 0.6s | 191tps | $0.50/M+3 more | $3/M+3 more | Read:$0.05/M+3 more Write:— | $14/K+1 more \+ input costs | +3 |  |  |  | 12/17/2025 |  |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-google)More models by Google

All

Text

Code

Video Input

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) [google/gemini-3.7-flash](/ai-gateway/models/gemini-3.7-flash) | 1M | 2.4s | 305tps | $1.50/M$0.75/M | $7.50/M$3.75/M | Read:$0.15/M$0.08/M Write:— | — | +3 | ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) ![vertex logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fvertex%2520ai.png&w=48&q=75) |  |  |  | 08/13/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) [google/gemini-3.5-flash-lite](/ai-gateway/models/gemini-3.5-flash-lite) | 1M | 0.5s | 455tps | $0.30/M | $2.50/M | Read:$0.03/M Write:— | $14/K+1 more \+ input costs | +3 | ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) ![vertex logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fvertex%2520ai.png&w=48&q=75) |  |  |  | 07/21/2026 |  |
| ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) [google/gemini-3.6-flash](/ai-gateway/models/gemini-3.6-flash) | 1M | 1.6s | 177tps | $1.50/M$0.75/M | $7.50/M$3.75/M | Read:$0.15/M$0.08/M Write:— | $14/K+1 more \+ input costs | +3 | ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) ![vertex logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fvertex%2520ai.png&w=48&q=75) |  |  |  | 07/21/2026 |  |
| ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) [google/gemini-3.5-flash](/ai-gateway/models/gemini-3.5-flash) | 1M | 3.0s | 175tps | $1.50/M | $9/M | Read:$0.15/M Write:— | $14/K+1 more \+ input costs | +3 | ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) ![vertex logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fvertex%2520ai.png&w=48&q=75) |  |  |  | 05/19/2026 |  |
| ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) [google/gemini-3.1-flash-lite](/ai-gateway/models/gemini-3.1-flash-lite) | 1M | 1.4s | 340tps | $0.25/M | $1.50/M | Read:$0.03/M Write:— | $14/K+1 more \+ input costs | +3 | ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) ![vertex logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fvertex%2520ai.png&w=48&q=75) |  |  |  | 05/07/2026 |  |
| ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) [google/gemini-2.5-flash-lite](/ai-gateway/models/gemini-2.5-flash-lite) | 1M | 0.2s | 427tps | $0.10/M | $0.40/M | Read:$0.01/M Write:— | $35/K+1 more \+ input costs | +3 | ![google logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgoogle.png&w=48&q=75) ![vertex logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fvertex%2520ai.png&w=48&q=75) |  |  |  | 06/17/2025 |  |

## [Copy link to heading](#about-gemini-3-flash)About Gemini 3 Flash

Gemini 3 Flash is Google's speed-optimized model in the Gemini 3 generation, combining Gemini 3's reasoning depth with the efficiency profile of the Flash tier. It outperforms Gemini 2.5 Pro across most benchmarks, meaning a speed-tier model now surpasses a previous-generation flagship. Gemini 3 Flash achieves this with meaningful gains in token efficiency over the 2.5 generation. See live metrics on this page for current throughput.

Thinking is first-class in Gemini 3 Flash. The `thinkingLevel` and `includeThoughts` provider options let you surface intermediate reasoning steps. This helps when debugging multi-step pipelines, constructing chain-of-thought datasets, or validating that Gemini 3 Flash reasons through a problem correctly. Set `thinkingLevel` to `high` when the task demands deeper inference and your latency budget allows it.

Because Gemini 3 Flash sits at the intersection of quality and throughput, it fits a wide range of real-world traffic patterns, from low-latency chat interfaces to batch document processing pipelines. Accessing it through AI Gateway adds observability, automatic retries, and provider failover without requiring a Google Cloud account.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Gemini 3 Flash supports configurable thinking levels (`high` included) via `providerOptions`, giving you direct control over how much reasoning compute the model applies per request.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-gemini-3-flash)When to Use Gemini 3 Flash

### Best for

- Real-time chat and assistants: Interfaces that require pro-level reasoning without high latency
- High-volume agentic pipelines: Per-token cost directly affects operating expenses
- Step-by-step analysis: Tasks where surfacing intermediate reasoning (`includeThoughts`) adds value
- Throughput-bottlenecked apps: Applications previously constrained by Gemini 2.5 Pro throughput limits
- Cost-sensitive production workloads: Production traffic where per-token cost matters but quality still has to stay benchmark-competitive

### Consider alternatives when

- Maximum reasoning depth: Your task requires the deepest reasoning regardless of cost or speed (consider `google/gemini-3-pro-preview` or `google/gemini-3.1-pro-preview`)
- Native image generation needed: You require image output alongside text (consider `google/gemini-3-pro-image` or `google/gemini-3.1-flash-image-preview`)
- Budget and latency dominate: Task quality requirements are low (consider `google/gemini-3.1-flash-lite-preview`)

## [Copy link to heading](#conclusion)Conclusion

Gemini 3 Flash resets expectations for what a speed-tier model can deliver, matching or exceeding previous-generation Pro quality at a fraction of the cost and latency. For teams that need scalable intelligence rather than raw capability, it represents a cost- and latency-efficient entry point into the Gemini 3 generation on AI Gateway.