Skip to content
Dashboard

What are confidence-based decision fallbacks in AI Gateway?

Content Engineer

Confidence-based decision fallbacks let AI Gateway ask another model to reconsider a decision when the first answer meets a condition you define. Choice and Score questions can escalate on low confidence; Boolean questions use a probability range. The feature is in beta, and a triggered fallback bills both stages.

Release workflows can classify changes with a decision model, then ask a language model to reassess ambiguous descriptions. Your application receives the final result and decides whether to accept the classification or send it for review.

Copy link to headingHow does this differ from an error fallback?

An error fallback handles a model that cannot complete its request. Conditional decision fallback also handles a successful response whose answer meets your escalation rule. An HTTP success can contain an uncertain judgment, so these triggers address different problems.

In providerOptions.gateway.models, plain model strings retain error-fallback behavior. Conditional objects name a model and a when condition. The decision fallback configuration permits one conditional object, placed first, followed by optional string fallbacks.

If the primary model fails before returning answers, Gateway tries the conditional object's model as an error fallback. It cannot assess confidence without an answer. If the primary succeeds and the condition matches, Gateway reruns the whole decision and returns the fallback's answers.

Copy link to headingWhich condition should you use?

Choice and Score confidence measures how concentrated the answer distribution is. The selected option's probability is a separate value.

Choose the condition according to the question's answer type:

Question

Condition

Example meaning

Choice or Score

confidenceBelow: 0.6

Reassess an answer with confidence below 0.6

Boolean

probabilityBetween: [0.4, 0.6]

Reassess when P(true) is between 0.4 and 0.6, including both endpoints

These numbers illustrate the syntax, not recommended production thresholds. For Boolean questions, a probability near zero favors false; treating every low probability as uncertainty would escalate clear negative answers.

Name a question to target one judgment. Omit question to trigger when any answer of the matching type meets the condition. Conditions can also combine signals with any, all, or atLeast.

For Choice and Score, missing or non-finite confidence also triggers confidenceBelow. Plan for this when comparing providers: a missing confidence field can cause escalation even when the selected category looks plausible.

Copy link to headingHow can you reassess an ambiguous release change?

Consider a release description that adds an optional filter but removes an existing request field. The optional addition is compatible by itself; the removal may require client changes. The classification should account for the full description before assigning a release label.

This example asks Jev to classify the change and configures GPT-6 Astra as the fallback when the answer's confidence is below 0.6. It prints the final answer and routing details so you can inspect which model answered. Your release workflow can use that result to request approval before publishing.

Use Node.js 22.18 or later and install AI SDK 7 and the Gateway provider with npm install ai @ai-sdk/gateway. Follow the decision quickstart to configure authentication with AI_GATEWAY_API_KEY or Vercel OIDC.

import { gateway } from '@ai-sdk/gateway';
import { experimental_decide as decide } from 'ai';
const result = await decide({
model: gateway.decisionModel('typesafe-ai/jev'),
state: {
change: 'Add an optional region filter. If omitted, include all regions. ' +
'Remove the legacy zone request field; clients sending it receive an error.',
},
questions: {
compatibility: {
type: 'choice',
instructions:
'Classify the described API change. Choose review if evidence is insufficient.',
criteria: {
breaking: 'Previously valid client requests can fail or require changes',
compatible: 'Existing valid requests retain their behavior',
review: 'The description does not establish compatibility',
},
},
},
providerOptions: {
gateway: {
models: [{
model: 'openai/gpt-6-astra',
when: { question: 'compatibility', confidenceBelow: 0.6 },
}],
},
},
});
console.log(JSON.stringify({
model: result.response.modelId,
answer: result.answers.compatibility,
routing: result.providerMetadata?.gateway?.routing,
}, null, 2));

With an API key exported in your shell, run the example as node release-decision.mjs. For local OIDC, first run vercel link and vercel env pull .env.local in your project directory. Then run node --env-file=.env.local release-decision.mjs to load the downloaded token.

Under this rubric, the expected label is breaking because an existing request field stops working. The fallback runs only if the primary answer meets the condition or the primary request fails.

Astra handles this Gateway fallback through structured output. This is separate from OpenAI's native Decisions API, whose current supported model is Luna. The OpenAI Decisions API explainer describes that endpoint and its answer types.

Copy link to headingWhat happens to the first model's answers?

Gateway reruns every original question against the original state and returns the complete fallback result. It doesn't merge selected answers from the two models or recheck the condition against the fallback. One request supports at most two successful decision stages.

This matters when one question controls escalation in a larger request. If you ask both compatibility and documentation-impact questions, low compatibility confidence also causes the documentation-impact question to run again. Your application should treat the returned answers as one result from the final model.

The final answer can also have a different shape within the SDK's supported contract. Language-model fallbacks omit Choice and Score probability distributions. Code that expects a selected-option probability must handle its absence instead of treating the fallback label as permission to act.

For the release workflow, you could retain human approval when the final result lacks the evidence required by your publishing policy. The condition determines when to ask another model; the application still owns the acceptance rule.

Copy link to headingHow do you know why a fallback ran?

Read the final model from result.response.modelId and inspect result.providerMetadata?.gateway?.routing. The routing attempts identify matched conditions through triggeredBy, including the question and reason.

Keep that record alongside the final classification during testing. Separate escalations caused by low confidence from those caused by unavailable confidence. If a provider omits confidence for most inputs, your policy may send much more work to the fallback than a trial with another provider suggested.

If the fallback and any configured error fallbacks fail, the request fails. Gateway doesn't return the earlier answer as a successful result. Catch that failure at the application's boundary and preserve the change for review or retry.

Copy link to headingHow should you choose a threshold?

Build a labeled set of release descriptions that includes additive changes, removed fields, and ambiguous wording. Compare primary-only decisions with the complete fallback policy. Record which mistakes the second stage corrects and which new mistakes it introduces.

Measure the share of requests that escalate as well as the accuracy of the final decisions. Both calls contribute to the cost of a triggered request, and the stages run sequentially. Frequent escalation is useful only if the improvement in final decisions justifies that extra work.

Hold back separate examples for the final assessment after choosing the threshold. Include cases where the primary is confidently wrong: low-confidence escalation cannot catch every incorrect answer. Retain a review path for incomplete change descriptions regardless of which model produced the final label.

Copy link to headingFrequently asked questions

Copy link to headingCan a successful decision request trigger a fallback?

Yes. Conditional fallbacks can run after the primary model returns a valid answer when that answer meets the configured confidence or probability condition. Ordinary execution failures can also use the configured fallback model.

Copy link to headingCan you use confidenceBelow for a Boolean question?

No. Boolean answers express the probability that a condition is true, so use probabilityBetween to select an uncertainty range. Low P(true) can represent a clear false answer.

Copy link to headingDoes the fallback rerun only the uncertain question?

Gateway sends the original state and all original questions to the fallback model. Its complete result replaces the primary result, and the same condition does not run again against the fallback answers.

Copy link to headingDoes escalation guarantee a correct answer?

No. Another model can repeat the mistake or introduce a different one. Evaluate the full policy on labeled cases and keep application review rules separate from the fallback trigger.

More Decision models articles

Ready to deploy?