# Browse all AI Gateway models

Every model available on Vercel AI Gateway, with API access, pricing, and a playground. 345 models · Page 3 of 6.

[Search and filter all models →](/ai-gateway/models)

- [OpenAIGPT-5GPT-5 is OpenAI's flagship language model that excels at complex reasoning, broad real-world knowledge, code-intensive, and multi-step agentic tasks.](/ai-gateway/models/gpt-5)
- [OpenAIGPT-5 miniGPT-5 mini is a cost optimized model that excels at reasoning/chat tasks. It offers an optimal balance between speed, cost, and capability.](/ai-gateway/models/gpt-5-mini)
- [OpenAIGPT-5 nanoGPT-5 nano is a high throughput model that excels at simple instruction or classification tasks.](/ai-gateway/models/gpt-5-nano)
- [OpenAIGPT-5 proGPT-5 pro uses more compute to think harder and provide consistently better answers. Since GPT-5 pro is designed to tackle tough problems, some requests may take several minutes to finish.](/ai-gateway/models/gpt-5-pro)
- [OpenAIGPT-5-CodexGPT-5-Codex is a version of GPT-5 optimized for agentic coding tasks in Codex or similar environments.](/ai-gateway/models/gpt-5-codex)
- [OpenAIGPT-5.1-CodexGPT-5.1-Codex is a version of GPT-5.1 optimized for agentic coding tasks in Codex or similar environments.](/ai-gateway/models/gpt-5.1-codex)
- [OpenAIGPT-Realtime miniGPT-Realtime mini is capable of responding to audio and text inputs in realtime over WebRTC, WebSocket, or SIP connections.](/ai-gateway/models/gpt-realtime-mini)
- [OpenAIGPT-Realtime-1.5GPT-Realtime-1.5 is our flagship audio model for voice agents and customer support.](/ai-gateway/models/gpt-realtime-1.5)
- [OpenAIgpt-realtime-2GPT Realtime 2 is our most capable realtime voice model. It supports speech-to-speech interactions with configurable reasoning effort, stronger instruction following, and more reliable tool use for complex voice-agent workflows.](/ai-gateway/models/gpt-realtime-2)
- [OpenAIgpt-realtime-2.1GPT-Realtime-2.1 updates GPT-Realtime-2 with improved alphanumeric recognition, silence and noise handling, and interruption behavior. It supports speech-to-speech interactions with configurable reasoning effort, instruction following, and tool use for complex voice-agent workflows.](/ai-gateway/models/gpt-realtime-2.1)
- [OpenAIgpt-realtime-whisperGPT Realtime Whisper is a streaming speech-to-text model for applications that need low-latency transcript deltas from live audio. It is designed for realtime use cases where developers need to tune latency and accuracy. GPT Realtime Whisper is priced by audio duration rather than text tokens.](/ai-gateway/models/gpt-realtime-whisper)
- [SpaceXAIGrok 4.1 Fast Non-Reasoning](/ai-gateway/models/grok-4.1-fast-non-reasoning)
- [SpaceXAIGrok 4.1 Fast Reasoning](/ai-gateway/models/grok-4.1-fast-reasoning)
- [SpaceXAIGrok 4.20 Beta Non-ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truthful responses.](/ai-gateway/models/grok-4.20-non-reasoning-beta)
- [SpaceXAIGrok 4.20 Beta ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truthful responses.](/ai-gateway/models/grok-4.20-reasoning-beta)
- [SpaceXAIGrok 4.20 Multi Agent BetaMultiple agents collaborate in parallel to perform deep research tasks.](/ai-gateway/models/grok-4.20-multi-agent-beta)
- [SpaceXAIGrok 4.20 Multi-AgentMultiple agents collaborate in parallel to perform deep research tasks.](/ai-gateway/models/grok-4.20-multi-agent)
- [SpaceXAIGrok 4.20 Non-ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, delivering consistently precise and truthful responses.](/ai-gateway/models/grok-4.20-non-reasoning)
- [SpaceXAIGrok 4.20 ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, delivering consistently precise and truthful responses.](/ai-gateway/models/grok-4.20-reasoning)
- [SpaceXAIGrok 4.3Grok 4.3 is a new model matching the scale of Grok 4.20 with an improved architecture and a December 2025 knowledge cutoff.](/ai-gateway/models/grok-4.3)
- [SpaceXAIGrok 4.5SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.](/ai-gateway/models/grok-4.5)
- [SpaceXAIGrok 4.6Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. It stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact.](/ai-gateway/models/grok-4.6)
- [SpaceXAIGrok Build 0.1xAI's fast coding model trained specifically for agentic coding.](/ai-gateway/models/grok-build-0.1)
- [SpaceXAIGrok ImagineState-of-the-art video generation across quality, cost, and latency. Grok Imagine is x.AI's most powerful video-audio generative model yet. Bring an image to life, start from a simple text prompt, or even refine a complex cinematic sequence.](/ai-gateway/models/grok-imagine-video)
- [SpaceXAIGrok Imagine ImageGenerate high-quality images from text prompts with xAI's imagine API.](/ai-gateway/models/grok-imagine-image)
- [SpaceXAIGrok Imagine Image 2.0](/ai-gateway/models/grok-imagine-image-2.0)
- [SpaceXAIGrok Imagine Video 1.5](/ai-gateway/models/grok-imagine-video-1.5)
- [SpaceXAIGrok Imagine Video 1.5 Preview](/ai-gateway/models/grok-imagine-video-1.5-preview)
- [SpaceXAIGrok STTTranscribe audio to text in 25 languages with batch and streaming modes.](/ai-gateway/models/grok-stt)
- [SpaceXAIGrok TTSGenerate speech with 5 expressive voices, speech tags, and telephony codecs.](/ai-gateway/models/grok-tts)
- [SpaceXAIGrok Voice Think Fast 1.0Build real-time voice applications powered by Grok. Stream audio and text bidirectionally via WebSocket for voice assistants, phone agents, and interactive voice systems.](/ai-gateway/models/grok-voice-think-fast-1.0)
- [SpaceXAIGrok Voice Think Fast 2.0Grok Voice Think Fast 2.0, our next-generation voice model with improved intelligence, transcription accuracy, and conversational capabilities.](/ai-gateway/models/grok-voice-think-fast-2.0)
- [Tencent CloudHy3](/ai-gateway/models/hy3)
- [Thinking MachinesInklingInkling is a multimodal MoE model (975B total, 41B active, 256k context) reasoning over text, image, and audio inputs.](/ai-gateway/models/inkling)
- [Thinking MachinesInkling SmallInkling-Small is a lighter-weight model with 12B active parameters, trained with a similar recipe, to Inkling that achieves strong performance with even lower cost and latency.](/ai-gateway/models/inkling-small)
- [InterfazeInterfaze BetaInterfaze is an AI model built on a new architecture that merges specialized DNN/CNN models with LLMs for developer tasks that require deterministic output and high consistency like OCR, scraping, classification, STT and more.](/ai-gateway/models/interfaze-beta)
- [KwaiPilotKat Coder Air V2.5Fast response version of Kat Coder V2.5 optimized for Agent and Claw use cases.](/ai-gateway/models/kat-coder-air-v2.5)
- [KwaiPilotKat Coder Pro V2A high-performance edition designed for complex enterprise projects and SaaS integration.](/ai-gateway/models/kat-coder-pro-v2)
- [KwaiPilotKat Coder Pro V2.5KAT-Coder-V2.5 is a coding-focused agentic model trained to act autonomously inside real, executable repositories rather than as a single-turn code generator.](/ai-gateway/models/kat-coder-pro-v2.5)
- [KwaiPilotKAT-Coder-Pro V1KAT-Coder-Pro V1 is KwaiKAT's most advanced agentic coding model in the KwaiKAT series. Designed specifically for agentic coding tasks, it excels in real-world software engineering scenarios, achieving a remarkable 73.4% solve rate on the SWE-Bench Verified benchmark. KAT-Coder-Pro V1 delivers top-tier coding performance and has been rigorously tested by thousands of in-house engineers. The model has been optimized for tool-use capability, multi-turn interaction, instruction following, generalization and comprehensive capabilities through a multi-stage training process, including mid-training, supervised fine-tuning (SFT), reinforcement fine-tuning (RFT), and scalable agentic RL.](/ai-gateway/models/kat-coder-pro-v1)
- [Moonshot AIKimi K2 InstructKimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.](/ai-gateway/models/kimi-k2)
- [Moonshot AIKimi K2 ThinkingKimi K2 Thinking is an advanced open-source thinking model by Moonshot AI. It can execute up to 200 – 300 sequential tool calls without human interference, reasoning coherently across hundreds of steps to solve complex problems. Built as a thinking agent, it reasons step by step while using tools, achieving state-of-the-art performance on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks, with major gains in reasoning, agentic search, coding, writing, and general capabilities.](/ai-gateway/models/kimi-k2-thinking)
- [Moonshot AIKimi K2.5kimi-k2.5 is Kimi's most versatile model to date, featuring a native multimodal architecture that supports both visual and text input, thinking and non-thinking modes, and dialogue and agent tasks.](/ai-gateway/models/kimi-k2.5)
- [Moonshot AIKimi K2.6Kimi K2.6 demonstrates particularly strong performance in long-horizon coding tasks and produces professional-grade design with code and vision.](/ai-gateway/models/kimi-k2.6)
- [Moonshot AIKimi K2.7 CodeKimi-K2.7-Code is a coding model from Moonshot AI. It has improved coding & agent performance over K2.6, more reasoning efficiency with less overthinking, and improved instruction following for long-horizon coding.](/ai-gateway/models/kimi-k2.7-code)
- [Moonshot AIKimi K2.7 Code High SpeedKimi K2.7 Code HighSpeed is the high-speed version of Kimi K2.7 Code, the same model as Kimi K2.7 Code, but with an output speed of approximately 180 Tokens/s and up to 260 Tokens/s in short context scenarios, delivering a more extreme coding experience.](/ai-gateway/models/kimi-k2.7-code-highspeed)
- [Moonshot AIKimi K3Kimi’s flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window.](/ai-gateway/models/kimi-k3)
- [Moonshot AIKimi K3 FastFast version of Kimi’s flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window.](/ai-gateway/models/kimi-k3-fast)
- [Kling AIKling v2.5 Turbo Image-to-VideoKling 2.5 Turbo is a major update to the AI video generation model focused on significantly improving speed, video quality, temporal stability, and creative control for creators, making professional-grade AI-generated video faster, more coherent, and easier to direct from text prompts.](/ai-gateway/models/kling-v2.5-turbo-i2v)
- [Kling AIKling v2.5 Turbo Text-to-VideoKling 2.5 Turbo is a major update to the AI video generation model focused on significantly improving speed, video quality, temporal stability, and creative control for creators, making professional-grade AI-generated video faster, more coherent, and easier to direct from text prompts.](/ai-gateway/models/kling-v2.5-turbo-t2v)
- [Kling AIKling v2.6 Image-to-VideoKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.](/ai-gateway/models/kling-v2.6-i2v)
- [Kling AIKling v2.6 Motion ControlKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.](/ai-gateway/models/kling-v2.6-motion-control)
- [Kling AIKling v2.6 Text-to-VideoKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.](/ai-gateway/models/kling-v2.6-t2v)
- [Kling AIKling v3.0 Image-to-VideoBuild upon an All-in-One product framework, the Kling 3.0 model series supports full multimodal input and output spanning text, images, audio, and video, bringing the understanding, generation, and editing of video together in one streamlined AI workflow. The models integrate multiple tasks, including text-to-video, image-to-video, reference-to-video, and in-video editing, into a single, native multimodal architecture, enabling the models to follow complex narrative logic, deliver precise shot control, and maintain strong prompt adherence.](/ai-gateway/models/kling-v3.0-i2v)
- [Kling AIKling v3.0 Motion ControlKling 3.0 delivers a major leap in character fidelity for motion-driven generation, with stable facial features across multi-angle and long-duration motion, accurate complex emotions from multi-image face references, identity preservation through partial occlusions (hats, hands, fans), and steady clarity as the camera zooms, pans, or tracks.](/ai-gateway/models/kling-v3.0-motion-control)
- [Kling AIKling v3.0 Text-to-VideoBuild upon an All-in-One product framework, the Kling 3.0 model series supports full multimodal input and output spanning text, images, audio, and video, bringing the understanding, generation, and editing of video together in one streamlined AI workflow. The models integrate multiple tasks, including text-to-video, image-to-video, reference-to-video, and in-video editing, into a single, native multimodal architecture, enabling the models to follow complex narrative logic, deliver precise shot control, and maintain strong prompt adherence.](/ai-gateway/models/kling-v3.0-t2v)
- [PoolsideLaguna S 2.1Laguna S 2.1 is Poolside's new open-weight model for agentic coding and long-horizon work.](/ai-gateway/models/laguna-s-2.1)
- [PoolsideLaguna S 2.1 FreeLaguna S 2.1 is Poolside's new open-weight model for agentic coding and long-horizon work.](/ai-gateway/models/laguna-s-2.1-free)
- [InclusionaiLing 3.0 FlashLing 3.0 Flash is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.](/ai-gateway/models/ling-3.0-flash)
- [InclusionaiLing 3.0 Flash FinLing 3.0 Flash Fin is InclusionAI’s finance-enhanced MoE language model, combining 124 billion total parameters with approximately 5.1 billion active parameters for efficient financial reasoning. Its 256K context window, function calling, and support for complex, multi-step investment workflows make it ideal for financial research, analysis, long-horizon planning, and execution, while retaining strong capabilities in coding and mathematics.](/ai-gateway/models/ling-3.0-flash-fin)