[Perplexity](/ai-gateway/models/labs/perplexity)

# Embed v1 4b

Embed v1 4b is the higher-accuracy tier of Perplexity's pplx-embed-v1 text embedding family. It returns 2560-dimensional vectors quantized to INT8 natively, requires no instruction prefix, and accepts inputs up to 32K tokens. Your use is subject to Perplexity's [Terms](https://www.perplexity.ai/terms) & [Privacy](https://www.perplexity.ai/privacy) Policies.

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

```
1import { embed } from 'ai';
2

3const result = await embed({
4  model: 'perplexity/pplx-embed-v1-4b',
5  value: 'Sunny day at the beach',
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/pplx-embed-v1-4b) [About](/ai-gateway/models/pplx-embed-v1-4b/about) [Providers](/ai-gateway/models/pplx-embed-v1-4b/providers) [Similar](/ai-gateway/models/pplx-embed-v1-4b/similar) [FAQ](/ai-gateway/models/pplx-embed-v1-4b/faq)

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Input | Capabilities | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- |

| ![perplexity logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fperplexity.png&w=48&q=75) [Perplexity](/ai-gateway/models/providers/perplexity) Legal:[Terms](https://www.perplexity.ai/terms)•[Privacy](https://www.perplexity.ai/privacy) | 32K | $0.03/M |  |  |  | 02/26/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#more-models-by-perplexity)More models by Perplexity

All

Text

Embed

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![perplexity logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fperplexity.png&w=48&q=75) [perplexity/pplx-embed\-v1-0.6b](/ai-gateway/models/pplx-embed-v1-0.6b) | 32K |  |  | $0.004/M |  |  | — |  | ![perplexity logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fperplexity.png&w=48&q=75) |  |  | 02/26/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![perplexity logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fperplexity.png&w=48&q=75) [perplexity/sonar](/ai-gateway/models/sonar) | 127K | 2.1s | 72tps | $1/M | $1/M |  | $5/K\*+2 more |  | ![perplexity logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fperplexity.png&w=48&q=75) |  |  | 02/19/2025 |  |
| ![perplexity logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fperplexity.png&w=48&q=75) [perplexity/sonar-pro](/ai-gateway/models/sonar-pro) | 200K | 1.2s | 101tps | $3/M | $15/M |  | $6/K\*+2 more |  | ![perplexity logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fperplexity.png&w=48&q=75) |  |  | 02/19/2025 |  |
| ![perplexity logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fperplexity.png&w=48&q=75) [perplexity/sonar-reasoning-pro](/ai-gateway/models/sonar-reasoning-pro) | 127K | 4.0s | 105tps | $2/M | $8/M |  | $6/K\*+2 more |  | ![perplexity logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fperplexity.png&w=48&q=75) |  |  | 02/19/2025 |  |

## [Copy link to heading](#about-embed-v1-4b)About Embed v1 4b

Embed v1 4b maps queries and documents into a shared vector space so retrieval reduces to approximate nearest neighbor search. Perplexity released it on February 26, 2026 as the larger of the two standard pplx-embed-v1 models. Embed v1 4b uses bidirectional attention with mean pooling over all token representations, rather than the causal attention that decoder-derived embedding models inherit.

On MTEB(Multilingual, v2), Embed v1 4b reaches an average nDCG@10 of 69.66% at INT8 precision, matching Qwen3-Embedding-4B at 69.60% and exceeding gemini-embedding-001 at 67.71%. On ToolRet, which measures retrieval over tool and API descriptions, it scores 44.45% average nDCG@10. On Perplexity's internal PPLXQuery2Query benchmark over a 2.4 million document corpus, it reaches 73.5% Recall@10 against 67.9% for Qwen3-Embedding-4B. On PPLXQuery2Doc over a 30 million page corpus, it reaches 91.7% Recall@1000 against 88.6%.

Two design choices shape how you integrate Embed v1 4b. It produces INT8 embeddings natively rather than as a post-hoc compression step, which cuts storage 4x compared with FP32; binary output cuts it 32x, and at this parameter scale Perplexity measures the binary quality drop at under 1.6 percentage points. Embed v1 4b also requires no instruction prefix, so you embed text directly. That removes a common failure mode where the instruction used at indexing time drifts from the one used at query time and quietly degrades recall. Matryoshka representation learning lets you request shorter vectors through the `dimensions` parameter when storage matters more than the last point of accuracy.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Embed v1 4b returns unnormalized embeddings. Compare INT8 vectors with cosine similarity and binary vectors with Hamming distance. If your vector database only supports inner product, convert to float32 and L2-normalize before storing, or similarity scores will be wrong. Most managed vector databases offer cosine similarity, but confirm the setting before you index anything.
- Configuration: Pick your dimension count before you build the index. Embed v1 4b returns 2560 dimensions by default and supports shorter vectors through Matryoshka representation learning, but changing the dimension later means re-embedding the whole corpus.
- Configuration: Perplexity recommends embedding documents and queries with the same model. Mixing Embed v1 4b with `pplx-embed-v1-0.6b` across the two sides makes similarity scores unreliable.
- Zero Data Retention: Zero Data Retention is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-embed-v1-4b)When to Use Embed v1 4b

### Best for

- Web-Scale First-Stage Retrieval: High recall at large depths feeds a downstream reranker
- Large RAG Corpora: Native INT8 and binary output keep vector storage within budget
- Multilingual Semantic Search: Evaluated on MTEB(Multilingual, v2) across many languages
- Tool and API Retrieval: Scores 44.45% average nDCG@10 on the ToolRet benchmark
- Prompt-Free Integration: No instruction prefix can drift between indexing and query time

### Consider alternatives when

- Per-Token Cost Dominates: `pplx-embed-v1-0.6b` returns 1024-dimensional vectors at a lower price
- Code-Only Corpora: `voyage-code-3` is purpose-built for source code retrieval
- Reranking Stage Needed: `rerank-2.5` reorders candidates after first-stage retrieval
- Generated Text Required: This model returns vectors only, not completions

## [Copy link to heading](#conclusion)Conclusion

Embed v1 4b is a practical default once your index grows large enough that storage and recall both constrain the design. Native INT8 output, binary compression, and a context window of 32K tokens cover most retrieval workloads without prompt engineering. Call it through AI Gateway with the AI SDK `embed` and `embedMany` functions, or through the OpenAI-compatible REST API, and keep one integration across embedding providers.