Browse all AI Gateway models
Every model available on Vercel AI Gateway, with API access, pricing, and a playground. 307 models · Page 5 of 6.
Search and filter all models →- Alibaba CloudQwen3 Embedding 8BThe Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B).
- Alibaba CloudQwen3 MaxThe Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios.
- Alibaba CloudQwen3 Max PreviewQwen3-Max-Preview shows substantial gains over the 2.5 series in overall capability, with significant enhancements in Chinese-English text understanding, complex instruction following, handling of subjective open-ended tasks, multilingual ability, and tool invocation; model knowledge hallucinations are reduced.
- Alibaba CloudQwen3 Next 80B A3B InstructA new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507).
- Alibaba CloudQwen3 Next 80B A3B ThinkingA new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507).
- Alibaba CloudQwen3 VL 235B A22B InstructThe Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement.
- Alibaba CloudQwen3 VL 235B A22B ThinkingQwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade.
- Alibaba CloudQwen3-14BQwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support
- Alibaba CloudQwen3-30B-A3BQwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support
- RecraftRecraft V2Recraft V2 is an image generation model released in March 2024 and the first model trained from scratch by Recraft. With 20 billion parameters, it was a breakthrough in human anatomical accuracy and the first to support brand consistency and brand color inputs. It also introduced vector image generation (SVG output), as well as minimalistic icon and illustration styles.
- RecraftRecraft V3V3 introduced major advances in photorealism and text rendering. It was the first Recraft model to generate mid-size text accurately and, as of 2025, is the only model capable of placing text at specific positions in an image.
- RecraftRecraft V4The model delivers strong photorealism, including realistic skin rendering and natural textures, while avoiding common synthetic artifacts. It produces more distinctive lighting, composition, diverse subjects, contemporary styling, and carefully considered scene elements. For illustration, it generates original characters and forms with sophisticated and unexpected color combinations.
- RecraftRecraft V4 ProThe model delivers strong photorealism, including realistic skin rendering and natural textures, while avoiding common synthetic artifacts. It produces more distinctive lighting, composition, diverse subjects, contemporary styling, and carefully considered scene elements. For illustration, it generates original characters and forms with sophisticated and unexpected color combinations.
- RecraftRecraft V4.1V4.1 is built on the same visual aesthetic and the same eye for what looks right, just pushed further across every dimension. The photorealism is more natural. The gradients are dreamier. Illustration styles are now possible that simply weren't before. And this is a model that reads a short prompts easier and creates something worth stopping for.
- RecraftRecraft V4.1 ProV4.1 is built on the same visual aesthetic and the same eye for what looks right, just pushed further across every dimension. The photorealism is more natural. The gradients are dreamier. Illustration styles are now possible that simply weren't before. And this is a model that reads a short prompts easier and creates something worth stopping for. V4.1 Pro generates higher-resolution images for when the idea deserves more room.
- RecraftRecraft V4.1 UtilityV4.1 is built on the same visual aesthetic and the same eye for what looks right, just pushed further across every dimension. The photorealism is more natural. The gradients are dreamier. Illustration styles are now possible that simply weren't before. And this is a model that reads a short prompts easier and creates something worth stopping for. V4.1 Utility is designed for when restraint is the aesthetic choice, with flat lighting, front-facing composition, and simple, controlled scenes.
- RecraftRecraft V4.1 Utility ProV4.1 is built on the same visual aesthetic and the same eye for what looks right, just pushed further across every dimension. The photorealism is more natural. The gradients are dreamier. Illustration styles are now possible that simply weren't before. And this is a model that reads a short prompts easier and creates something worth stopping for. V4.1 Utility is designed for when restraint is the aesthetic choice, with flat lighting, front-facing composition, and simple, controlled scenes.
- ByteDanceSeed 1.6ByteDance's new multimodal deep-thinking model, supporting both text and visual inputs with enhanced reasoning capabilities.
- ByteDanceSeedance 2.0Built with a unified multimodal audio-video joint generation architecture, Seedance 2.0 supports four input modalities: text, image, audio, and video. Compared with Version 1.5, Seedance 2.0 delivers a substantial leap in generation quality. It achieves a higher usability rate for complex interaction and motion scenes, with significant improvements in physical accuracy, visual realism, and controllability, making it well-suited for high-quality creation scenarios.
- ByteDanceSeedance 2.0 FastSeedance 2.0 Fast is a new-generation multimodal video creation model, inheriting the core functions and advantages of Seedance 2.0, with faster speed.
- ByteDanceSeedance v1.0 ProA video generation model that supports multi-shot storytelling. It excels in semantic understanding and instruction following, producing smooth, detailed, and cinematic 1080P HD videos.
- ByteDanceSeedance v1.0 Pro FastSeedance 1.0 Pro Fast delivers top performance at an unbeatable price, balancing quality, speed, and cost. Built on Seedance 1.0 Pro’s core strengths, it’s faster and more cost-efficient for creators.
- ByteDanceSeedance v1.5 ProByteDance's Seedance 1.5 Pro is a professional video model using V2A native generation for integrated, synced audio-visual output, enhancing efficiency of professional video creation.
- ByteDanceSeedream 4.0Seedream 4.0 is a SOTA multimodal image creation model built on leading architecture. It breaks through the boundaries of traditional text-to-image models by natively supporting text, single-image, and multi-image inputs. Users can freely combine text and images to achieve diverse creative modes within a single model—such as multi-image blending, image editing, and sequentially batch image generation, featuring subject consistency, making image creation more free and controllable.
- ByteDanceSeedream 4.5Seedream 4.5 is the latest in-house image generation model developed by ByteDance. Compared with Seedream 4.0, it delivers comprehensive improvements—especially in editing consistency, including better preservation of subject details, lighting, and color tone. It also enhances portrait refinement and small-text rendering. The model’s multi-image composition capabilities have been significantly strengthened, and both reasoning performance and visual aesthetics continue to advance, enabling more accurate and artistically expressive image generation.
- ByteDanceSeedream 5.0 LiteByteDance-Seedream-5.0-lite is the latest image generation model released by BytePlus. For the first time, it introduces web-connected retrieval, enabling the model to fuse real-time online information to significantly improve the timeliness and relevance of generated images. The model’s reasoning and comprehension capabilities are further upgraded, allowing it to accurately interpret complex prompts and visual inputs. In addition, ByteDance-Seedream-5.0-lite delivers notable improvements in global knowledge coverage, reference consistency, and professional-grade scene generation, making it well suited for enterprise-level visual creation workflows.
- ByteDanceSeedream 5.0 ProSeedream-5.0-Pro, ByteDance's newest image generation model, delivers comprehensive upgrades for complex, lifelike image creation and editing, ushering in a new phase of controllable visual production. It stands out with precise editing control, robust commercial applicability and natural rendering results.
- PerplexitySonarPerplexity's lightweight offering with search grounding, quicker and cheaper than Sonar Pro.
- PerplexitySonar ProPerplexity's premier offering with search grounding, supporting advanced queries and follow-ups.
- PerplexitySonar Reasoning ProA premium reasoning-focused model that outputs Chain of Thought (CoT) in responses, providing comprehensive explanations with enhanced search capabilities and multiple search queries per request.
- StepFunStep 3.7 FlashStepFun’s flagship multimodal reasoning model. Powered by a 198B-parameter / 11B-activation sparse MoE architecture, with native support for image and video understanding.
- StepFunStepFun 3.5 FlashStep 3.5 Flash is an open-source reasoning model by StepFun with 196B total parameters (11B active) using Mixture of Experts. It features a 256K context window, deep reasoning, tool calling, and agentic capabilities, achieving 97.3 on AIME 2025 and 74.4% on SWE-bench Verified.
- GoogleText Embedding 005English-focused text embedding model optimized for code and English language tasks.
- GoogleText Multilingual Embedding 002Multilingual text embedding model optimized for cross-lingual tasks across many languages.
- OpenAItext-embedding-3-largeOpenAI's most capable embedding model for both english and non-english tasks.
- OpenAItext-embedding-3-smallOpenAI's improved, more performant version of their ada embedding model.
- OpenAItext-embedding-ada-002OpenAI's legacy text embedding model.
- AmazonTitan Text Embeddings V2Amazon Titan Text Embeddings V2 is a light weight, efficient multilingual embedding model supporting 1024, 512, and 256 dimensions.
- Arcee AITrinity Large ThinkingTrinity-Large-Thinking is a reasoning-optimized variant of Arcee AI's Trinity-Large family — a 398B-parameter sparse Mixture-of-Experts (MoE) model with approximately 13B active parameters per token. Built on Trinity-Large-Base and post-trained with extended chain-of-thought reasoning and agentic RL, Trinity-Large-Thinking delivers state-of-the-art performance on agentic benchmarks while maintaining strong general capabilities.
- Arcee AITrinity MiniTrinity Mini is a 26B-parameter (3B active) sparse mixture-of-experts language model, engineered for efficient inference over long contexts with robust function calling and multi-step agent workflows.
- OpenAITTS-1TTS is a model that converts text to natural sounding spoken text.
- OpenAITTS-1 HDTTS is a model that converts text to natural sounding spoken text. The tts-1-hd model is optimized for high quality text-to-speech use cases.
- GoogleVeo 3.0Veo 3 is designed to handle a range of video generation tasks, from cinematic narratives to dynamic character animations. With Veo 3, you can create more immersive experiences by not only generating stunning visuals, but also audio like dialogue and sound effects.
- GoogleVeo 3.0 Fast GenerateVeo 3 Fast is a quicker and more cost effective version of Veo 3, allowing developers to create videos with sound while maintaining high quality and optimizing for speed and business use cases. Veo 3 Fast offers both text-to-video and image-to-video modalities.
- GoogleVeo 3.1Veo 3.1 is Google's state-of-the-art model for generating high-fidelity, 8-second 720p, 1080p or 4k videos featuring stunning realism and natively generated audio.
- GoogleVeo 3.1 Fast GenerateVeo 3.1 Fast is a specialized, high-speed variant of Google DeepMind’s Veo 3.1 text-to-video model, optimized for rapid generation of 8-second, high-fidelity videos. It is designed to create cinematic, 1080p, or 720p content with improved prompt adherence and native audio, making it ideal for creating quick, high-quality video clips, social media content, and ad creatives.
- Voyage AIVoyage 3.5Voyage AI's embedding model optimized for general-purpose and multilingual retrieval quality.
- Voyage AIVoyage 3.5 LiteVoyage AI's embedding model optimized for latency and cost.
- Voyage AIVoyage 4Optimized for general-purpose and multilingual retrieval quality. All embeddings created with the 4 series are compatible with each other.
- Voyage AIVoyage 4 LargeThe best general-purpose and multilingual retrieval quality. All embeddings created with the 4 series are compatible with each other.
- Voyage AIVoyage 4 LiteOptimized for latency and cost. All embeddings created with the 4 series are compatible with each other.
- Voyage AIVoyage Code 2Voyage AI's embedding model optimized for code retrieval (17% better than alternatives). This is the previous generation of code embeddings models.
- Voyage AIVoyage Code 3Voyage AI's embedding model optimized for code retrieval.
- Voyage AIVoyage Finance 2Voyage AI's embedding model optimized for finance retrieval and RAG.
- Voyage AIVoyage Law 2Voyage AI's embedding model optimized for legal retrieval and RAG.
- Voyage AIVoyage Rerank 2.5A generalist reranker optimized for quality with instruction-following and multilingual support.
- Voyage AIVoyage Rerank 2.5 LiteA generalist reranker optimized for both latency and quality with instruction-following and multilingual support.
- Voyage AIvoyage-3-largeVoyage AI's embedding model with the best general-purpose and multilingual retrieval quality.
- Alibaba CloudWan v2.5 Text-to-Video Preview
- Alibaba CloudWan v2.6 Image-to-Video