Browse all AI Gateway models
Every model available on Vercel AI Gateway, with API access, pricing, and a playground. 307 models · Page 3 of 6.
Search and filter all models →- OpenAIGPT-5.3 ChatThe model powering ChatGPT is gpt-5.3-chat-latest: this is OpenAI's best general-purpose model, part of the GPT-5 flagship model family.
- OpenAIGPT-Realtime miniGPT-Realtime mini is capable of responding to audio and text inputs in realtime over WebRTC, WebSocket, or SIP connections.
- OpenAIGPT-Realtime-1.5GPT-Realtime-1.5 is our flagship audio model for voice agents and customer support.
- OpenAIgpt-realtime-2GPT Realtime 2 is our most capable realtime voice model. It supports speech-to-speech interactions with configurable reasoning effort, stronger instruction following, and more reliable tool use for complex voice-agent workflows.
- OpenAIgpt-realtime-2.1GPT-Realtime-2.1 updates GPT-Realtime-2 with improved alphanumeric recognition, silence and noise handling, and interruption behavior. It supports speech-to-speech interactions with configurable reasoning effort, instruction following, and tool use for complex voice-agent workflows.
- OpenAIgpt-realtime-whisperGPT Realtime Whisper is a streaming speech-to-text model for applications that need low-latency transcript deltas from live audio. It is designed for realtime use cases where developers need to tune latency and accuracy. GPT Realtime Whisper is priced by audio duration rather than text tokens.
- xAIGrok 4.1 Fast Non-Reasoning
- xAIGrok 4.1 Fast Reasoning
- xAIGrok 4.20 Beta Non-ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truthful responses.
- xAIGrok 4.20 Beta ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truthful responses.
- xAIGrok 4.20 Multi Agent BetaMultiple agents collaborate in parallel to perform deep research tasks.
- xAIGrok 4.20 Multi-AgentMultiple agents collaborate in parallel to perform deep research tasks.
- xAIGrok 4.20 Non-ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, delivering consistently precise and truthful responses.
- xAIGrok 4.20 ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, delivering consistently precise and truthful responses.
- xAIGrok 4.3Grok 4.3 is a new model matching the scale of Grok 4.20 with an improved architecture and a December 2025 knowledge cutoff.
- xAIGrok 4.5SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
- xAIGrok Build 0.1xAI's fast coding model trained specifically for agentic coding.
- xAIGrok ImagineState-of-the-art video generation across quality, cost, and latency. Grok Imagine is x.AI's most powerful video-audio generative model yet. Bring an image to life, start from a simple text prompt, or even refine a complex cinematic sequence.
- xAIGrok Imagine ImageGenerate high-quality images from text prompts with xAI's imagine API.
- xAIGrok Imagine Video 1.5
- xAIGrok Imagine Video 1.5 Preview
- xAIGrok STTTranscribe audio to text in 25 languages with batch and streaming modes.
- xAIGrok TTSGenerate speech with 5 expressive voices, speech tags, and telephony codecs.
- xAIGrok Voice Think Fast 1.0Build real-time voice applications powered by Grok. Stream audio and text bidirectionally via WebSocket for voice assistants, phone agents, and interactive voice systems.
- TencentHy3
- GoogleImagen 4Imagen 4: Google's flagship text-to-image model that serves as the go-to choice for a wide variety of high-quality image generation tasks, featuring significant improvements in text rendering over previous models. It now supports up to 2K resolution generation for creating detailed and crisp visuals, making it suitable for everything from marketing assets to artistic compositions.
- GoogleImagen 4 FastImagen 4 Fast is Google’s speed-optimized variant of the Imagen 4 text-to-image model, designed for rapid, high-volume image generation. It’s ideal for workflows like quick drafts, mockups, and iterative creative exploration. Despite emphasizing speed, it still benefits from the broader Imagen 4 family’s improvements in clarity, text rendering, and stylistic flexibility, and supports high-resolution outputs up to 2K.
- GoogleImagen 4 UltraImagen 4 Ultra: Highest quality image generation model for detailed and photorealistic outputs.
- ThinkingmachinesInklingInkling is a multimodal MoE model (975B total, 41B active, 256k context) reasoning over text, image, and audio inputs.
- InterfazeInterfaze BetaInterfaze is an AI model built on a new architecture that merges specialized DNN/CNN models with LLMs for developer tasks that require deterministic output and high consistency like OCR, scraping, classification, STT and more.
- KwaiPilotKat Coder Air V2.5Fast response version of Kat Coder V2.5 optimized for Agent and Claw use cases.
- KwaiPilotKat Coder Pro V2A high-performance edition designed for complex enterprise projects and SaaS integration.
- KwaiPilotKat Coder Pro V2.5KAT-Coder-V2.5 is a coding-focused agentic model trained to act autonomously inside real, executable repositories rather than as a single-turn code generator.
- KwaiPilotKAT-Coder-Pro V1KAT-Coder-Pro V1 is KwaiKAT's most advanced agentic coding model in the KwaiKAT series. Designed specifically for agentic coding tasks, it excels in real-world software engineering scenarios, achieving a remarkable 73.4% solve rate on the SWE-Bench Verified benchmark. KAT-Coder-Pro V1 delivers top-tier coding performance and has been rigorously tested by thousands of in-house engineers. The model has been optimized for tool-use capability, multi-turn interaction, instruction following, generalization and comprehensive capabilities through a multi-stage training process, including mid-training, supervised fine-tuning (SFT), reinforcement fine-tuning (RFT), and scalable agentic RL.
- Moonshot AIKimi K2 InstructKimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.
- Moonshot AIKimi K2 ThinkingKimi K2 Thinking is an advanced open-source thinking model by Moonshot AI. It can execute up to 200 – 300 sequential tool calls without human interference, reasoning coherently across hundreds of steps to solve complex problems. Built as a thinking agent, it reasons step by step while using tools, achieving state-of-the-art performance on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks, with major gains in reasoning, agentic search, coding, writing, and general capabilities.
- Moonshot AIKimi K2.5kimi-k2.5 is Kimi's most versatile model to date, featuring a native multimodal architecture that supports both visual and text input, thinking and non-thinking modes, and dialogue and agent tasks.
- Moonshot AIKimi K2.6Kimi K2.6 demonstrates particularly strong performance in long-horizon coding tasks and produces professional-grade design with code and vision.
- Moonshot AIKimi K2.7 CodeKimi-K2.7-Code is a coding model from Moonshot AI. It has improved coding & agent performance over K2.6, more reasoning efficiency with less overthinking, and improved instruction following for long-horizon coding.
- Moonshot AIKimi K2.7 Code High SpeedKimi K2.7 Code HighSpeed is the high-speed version of Kimi K2.7 Code, the same model as Kimi K2.7 Code, but with an output speed of approximately 180 Tokens/s and up to 260 Tokens/s in short context scenarios, delivering a more extreme coding experience.
- Moonshot AIKimi K3Kimi’s flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window.
- Kling AIKling v2.5 Turbo Image-to-VideoKling 2.5 Turbo is a major update to the AI video generation model focused on significantly improving speed, video quality, temporal stability, and creative control for creators, making professional-grade AI-generated video faster, more coherent, and easier to direct from text prompts.
- Kling AIKling v2.5 Turbo Text-to-VideoKling 2.5 Turbo is a major update to the AI video generation model focused on significantly improving speed, video quality, temporal stability, and creative control for creators, making professional-grade AI-generated video faster, more coherent, and easier to direct from text prompts.
- Kling AIKling v2.6 Image-to-VideoKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.
- Kling AIKling v2.6 Motion ControlKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.
- Kling AIKling v2.6 Text-to-VideoKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.
- Kling AIKling v3.0 Image-to-VideoBuild upon an All-in-One product framework, the Kling 3.0 model series supports full multimodal input and output spanning text, images, audio, and video, bringing the understanding, generation, and editing of video together in one streamlined AI workflow. The models integrate multiple tasks, including text-to-video, image-to-video, reference-to-video, and in-video editing, into a single, native multimodal architecture, enabling the models to follow complex narrative logic, deliver precise shot control, and maintain strong prompt adherence.
- Kling AIKling v3.0 Motion ControlKling 3.0 delivers a major leap in character fidelity for motion-driven generation, with stable facial features across multi-angle and long-duration motion, accurate complex emotions from multi-image face references, identity preservation through partial occlusions (hats, hands, fans), and steady clarity as the camera zooms, pans, or tracks.
- Kling AIKling v3.0 Text-to-VideoBuild upon an All-in-One product framework, the Kling 3.0 model series supports full multimodal input and output spanning text, images, audio, and video, bringing the understanding, generation, and editing of video together in one streamlined AI workflow. The models integrate multiple tasks, including text-to-video, image-to-video, reference-to-video, and in-video editing, into a single, native multimodal architecture, enabling the models to follow complex narrative logic, deliver precise shot control, and maintain strong prompt adherence.
- PoolsideLaguna S 2.1Laguna S 2.1 is Poolside's new open-weight model for agentic coding and long-horizon work.
- PoolsideLaguna S 2.1 FreeLaguna S 2.1 is Poolside's new open-weight model for agentic coding and long-horizon work.
- InclusionaiLing 3.0 FlashLing 3.0 Flash is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.
- MetaLlama 3.1 70B InstructAn update to Meta Llama 3 70B Instruct that includes an expanded 128K context length, multilinguality and improved reasoning capabilities.
- MetaLlama 3.1 8B InstantLlama 3.1 8B with 128K context window support, making it ideal for real-time conversational interfaces and data analysis while offering significant cost savings compared to larger models. Served by Groq with their custom Language Processing Units (LPUs) hardware to provide fast and efficient inference.
- MetaLlama 3.3 70B InstructWhere performance meets efficiency. This model supports high-performance conversational AI designed for content creation, enterprise applications, and research, offering advanced language understanding capabilities, including text summarization, classification, sentiment analysis, and code generation.
- MetaLlama 4 Maverick 17B 128E Instruct FP8The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Maverick, a 17 billion parameter model with 128 experts. Served by DeepInfra.
- MetaLlama 4 Scout 17B 16E InstructThe Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Scout, a 17 billion parameter model with 16 experts. Served by DeepInfra.
- MistralMagistral Medium 2509Complex thinking, backed by deep understanding, with transparent reasoning you can follow and verify. The model excels in maintaining high-fidelity reasoning across numerous languages, even when switching between languages mid-task.
- MistralMagistral Small 2509Complex thinking, backed by deep understanding, with transparent reasoning you can follow and verify. The model excels in maintaining high-fidelity reasoning across numerous languages, even when switching between languages mid-task.
- InceptionMercury 2A diffusion-based reasoning LLM that generates text via parallel refinement (not token-by-token), delivering real-time latency with ~1k tokens/sec plus 128K context and built-in tool/JSON support.