[Cohere](/ai-gateway/models/labs/cohere)

# Cohere Rerank 4 Fast

Cohere Rerank 4 Fast is a multilingual reranking model from Cohere tuned for low-latency, high-throughput retrieval over English and non-English documents and semi-structured JSON. Your use is subject to Cohere's [Terms](https://cohere.com/terms-of-use) & [Privacy](https://cohere.com/privacy) Policies.

Rerank

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

```
1import { rerank } from 'ai';
2

3const result = await rerank({
4  model: 'cohere/rerank-v4-fast',
5  query: 'What is the capital of France?',
6  documents: [
7    'Paris is the capital of France.',
8    'Berlin is the capital of Germany.',
9    'Madrid is the capital of Spain.',
10  ],
11})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/rerank-v4-fast) [About](/ai-gateway/models/rerank-v4-fast/about) [Providers](/ai-gateway/models/rerank-v4-fast/providers) [Similar](/ai-gateway/models/rerank-v4-fast/similar) [FAQ](/ai-gateway/models/rerank-v4-fast/faq)

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Input | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- |

| ![cohere logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fcohere.png&w=48&q=75) [Cohere](/ai-gateway/models/providers/cohere) Legal:[Terms](https://cohere.com/terms-of-use)•[Privacy](https://cohere.com/privacy) | 32K | $2/K |  |  |  | 12/11/2025 |  |
| --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#more-models-by-cohere)More models by Cohere

All

Text

Code

Embed

Rerank

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Free Tier | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![cohere logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fcohere.png&w=48&q=75) [cohere/rerank-v4-pro](/ai-gateway/models/rerank-v4-pro) | 32K |  |  | $2.50/K |  |  | — |  | ![cohere logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fcohere.png&w=48&q=75) |  |  |  | 12/11/2025 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![cohere logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fcohere.png&w=48&q=75) [cohere/embed-v4.0](/ai-gateway/models/embed-v4.0) | 128K |  |  | $0.12/M |  |  | — |  | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![cohere logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fcohere.png&w=48&q=75) |  |  |  | 04/15/2025 |  |
| ![cohere logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fcohere.png&w=48&q=75) [cohere/command-a](/ai-gateway/models/command-a) | 256K | 0.2s | 76tps | $2.50/M | $10/M |  | — |  | ![cohere logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fcohere.png&w=48&q=75) |  |  |  | 03/13/2025 |  |
| ![cohere logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fcohere.png&w=48&q=75) [cohere/rerank-v3.5](/ai-gateway/models/rerank-v3.5) | 4K |  |  | $2/K |  |  | — |  | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) |  |  |  | 12/02/2024 |  |

## [Copy link to heading](#about-cohere-rerank-4-fast)About Cohere Rerank 4 Fast

Cohere Rerank 4 Fast is the fast tier of Cohere's Rerank 4 generation, released December 11, 2025 alongside `rerank-v4-pro`. It shares the v4 family's multilingual coverage and JSON support but is tuned for lower per-query latency and higher throughput.

Like its siblings, Cohere Rerank 4 Fast is a cross-encoder. It reads the query and each candidate document together through attention, scoring relevance directly rather than relying on independent vector similarity. That structure picks up signal on complex, multi-part, or ambiguous queries that bi-encoder embeddings flatten.

Cohere Rerank 4 Fast covers more than 100 languages, including the same multilingual coverage as Cohere's embed-multilingual family. A French query can match Japanese documents and vice versa under one model. It handles long-form text, tables, code, and semi-structured JSON records with the same per-document context window of 32K tokens.

The fast variant fits at the front of high-traffic search systems and agentic retrieval steps where the reranker runs on every user turn. It pairs naturally with a multilingual embedding retriever for first-pass candidate selection and trades some quality versus `rerank-v4-pro` for response time under load.

See https://cohere.com/blog/rerank-4 for the request format, including `top_n`, `return_documents`, and per-document tokenization rules. Reranking is billed per search query rather than per token.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Rerankers run after an initial retrieval step. Pair Cohere Rerank 4 Fast with an embedding model, keyword search, or hybrid retriever so it has a candidate pool to score. The per-document context of 32K tokens covers query and document tokens together, which is wide enough to score long passages and chunked enterprise documents without truncation in most cases.
- Zero Data Retention: Zero Data Retention is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-cohere-rerank-4-fast)When to Use Cohere Rerank 4 Fast

### Best for

- Multilingual reranking: One model covers 100+ languages for cross-lingual retrieval
- High-throughput search: Lower latency keeps reranker overhead small under load
- Agentic retrieval loops: Tool calls that retrieve then rerank on every turn stay within tight time budgets
- Mixed-format corpora: Long text, JSON, tables, and code share one reranking stage
- Cost-sensitive RAG: A lower per-query price than the Pro variant at scale

### Consider alternatives when

- Maximum quality: `rerank-v4-pro` targets state-of-the-art relevance on complex queries
- English-only corpora: `rerank-v3.5` covers English RAG at a lower price point
- No second stage needed: First-pass retrieval alone meets the accuracy bar
- Image or multimodal retrieval: Use a multimodal retrieval model instead

## [Copy link to heading](#conclusion)Conclusion

Cohere Rerank 4 Fast is the right reranker when multilingual coverage and latency both matter, and a small quality gap versus the Pro variant is acceptable. Route it through AI Gateway with model id `cohere/rerank-v4-fast` for unified billing across the retrieval and generation stages of the same pipeline.