[Sakana AI](/ai-gateway/models/labs/sakana)

# Fugu Ultra

Fugu Ultra is Sakana AI's quality-maximizing multi-agent system released June 21, 2026. Behind a single model API, Fugu Ultra coordinates a pool of expert agents to solve hard, multi-step problems. Available through AI Gateway with unified billing, observability, and provider routing. Your use is subject to Sakana AI's [Terms](https://console.sakana.ai/terms-of-service) & [Privacy](https://console.sakana.ai/privacy-policy) Policies.

Vision (Image)Tool UseReasoning

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'sakana/fugu-ultra',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/fugu-ultra) [API](/ai-gateway/models/fugu-ultra/api) [About](/ai-gateway/models/fugu-ultra/about) [Providers](/ai-gateway/models/fugu-ultra/providers) [Latency](/ai-gateway/models/fugu-ultra/latency) [Uptime](/ai-gateway/models/fugu-ultra/uptime) [Status](/ai-gateway/models/fugu-ultra/status) [Similar](/ai-gateway/models/fugu-ultra/similar) [FAQ](/ai-gateway/models/fugu-ultra/faq)

## [Copy link to heading](#playground)Playground

Try out Fugu Ultra by Sakana AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![sakana logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fsakana.png&w=48&q=75)Fugu Ultra

![sakana logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fsakana.png&w=96&q=75)

Fugu Ultra

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![sakana logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fsakana.png&w=48&q=75) [Sakana AI](/ai-gateway/models/providers/sakana) Legal:[Terms](https://console.sakana.ai/terms-of-service)•[Privacy](https://console.sakana.ai/privacy-policy) | 1M | 1M | 3.8s |  | $5/M+1 more | $30/M+1 more | Read: $0.5/M+1 more Write: — | — |  |  |  | 06/21/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-sakana-ai)More models by Sakana AI

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![sakana logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fsakana.png&w=48&q=75) [sakana/namazu](/ai-gateway/models/namazu) | 256K | 0.8s | 264tps | $0.95/M | $4/M | Read:$0.15/M Write:— | — | +3 | ![sakana logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fsakana.png&w=48&q=75) |  |  | 08/03/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#about-fugu-ultra)About Fugu Ultra

Fugu Ultra was released June 21, 2026 as Sakana AI's premium tier in the Fugu line. Instead of answering with a single network, Fugu Ultra runs a multi-agent system behind one model API. A lightweight coordinator assigns Thinker, Worker, and Verifier roles across a pool of frontier models, then routes between one and three agents depending on the problem. The approach builds on Sakana AI's published TRINITY and Conductor research.

Sakana AI tuned Fugu Ultra for maximum answer quality rather than low latency, drawing on a deeper, fixed agent pool and spending more compute per request. Early users applied Fugu Ultra to AI research, paper reproduction, cybersecurity analysis, and literature and patent investigations. In Sakana AI's published evaluations, Fugu Ultra scores 73.7 on SWE-Bench Pro, 95.5 on GPQA-Diamond, and 50.0 on Humanity's Last Exam.

Fugu Ultra supports a context window of 1M tokens with output up to 1M tokens, plus prompt caching for repeated context. Through AI Gateway, you call Fugu Ultra with a unified API key, automatic retries, and built-in observability, using the AI SDK, Chat Completions API, Responses API, Messages API, or other API formats.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Fugu Ultra optimizes for answer quality, not speed. Each request can fan out across multiple coordinated agents, so responses take longer than single-model calls. See live metrics on this page, and reserve Fugu Ultra for problems where quality justifies the wait.
- Configuration: Pricing is tiered by context length. Requests beyond 272K tokens bill at higher rates for input, output, and cached input. Track actual spend with AI Gateway's built-in observability, and trim prompts where you can.
- Configuration: The agent pool behind Fugu Ultra is fixed, and per-request routing details aren't exposed. Teams with strict model-provenance requirements should confirm the orchestrated approach fits their compliance needs before deploying. Fugu Ultra is available through Sakana AI.
- Zero Data Retention: Zero Data Retention is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-fugu-ultra)When to Use Fugu Ultra

### Best for

- Deep Research Workflows: Literature reviews, patent investigations, and paper reproduction benefit from coordinated multi-agent analysis
- Hard Multi-Step Problems: Engineering and scientific questions where answer quality outweighs latency and cost per request
- Cybersecurity Analysis: Careful multi-step reasoning over code, configurations, and threat reports fits the verifier-backed workflow
- Long-Context Investigations: The context window of 1M tokens fits large codebases, document sets, and extended research threads
- Quality-Critical Agent Steps: Route your hardest pipeline steps to Fugu Ultra while cheaper models handle routine work

### Consider alternatives when

- Latency-Sensitive Applications: Multi-agent coordination adds response time, so single-model options such as GPT-5.4 or Gemini 3.5 Flash fit interactive chat better
- High-Volume Simple Tasks: Classification, extraction, and summarization run cheaper on efficiency models such as DeepSeek V4 Flash or Kimi K2.6
- Predictable Single-Model Behavior: Teams that must know exactly which model produced each answer should pick a conventional model, since Fugu Ultra routes internally
- Budget-Constrained Workloads: Premium per-token rates make Fugu Ultra expensive for routine generation at high volume

## [Copy link to heading](#conclusion)Conclusion

Fugu Ultra packages Sakana AI's multi-agent orchestration research into a single endpoint tuned for maximum answer quality on hard problems. For research, security analysis, and quality-critical agent steps, Fugu Ultra is available through AI Gateway with unified billing, observability, and provider routing.