Decision models are AI models that evaluate supplied context against defined questions and return structured answers, often with probabilities. They suit tasks such as assigning a support request to a team or rating an incident's urgency. Your application defines the allowed outcomes and how to use the prediction.
Imagine a customer writes, "I was charged twice. Please return the extra payment." The application requires a billing category and a refund-request flag before it can receive a written reply. You can use a decision model for those judgments and a generative model to draft the response.
Copy link to headingHow does a decision model work?
You supply state, the evidence the model should consider, along with questions and their answer definitions. State could contain a customer message, a document excerpt, or a record of an agent's actions. Each question should ask for one judgment that you can check against that evidence.
The model returns values your application can use in a branch, a queue, or a review process. Your code can use the billing label to select a support team and its associated probability to decide whether to route automatically or request review.
Native decision models can score the allowed outcomes without generating an answer token by token. Their implementations differ. Laya uses an encoder with a decision head, while Typical reads shared state into a cache and scores question-specific options. The category describes a task and interface; it doesn't imply that every provider uses the same architecture.
Several independent questions can share the same evidence. For example, Jev evaluates questions separately against one state. If a later question needs an earlier answer, your application must pass that result into a subsequent step.
Copy link to headingWhat kinds of answers can you ask for?
Three common question types cover different judgments. Liquid d1 and Jev support these types, and AI Gateway exposes them as Choice, Score, and Boolean.
TypeSafe-compatible APIs call the yes/no type Noul. Values near zero mean the model favors no; they don't signal uncertainty. An urgency score has a different meaning again: it describes the assessed urgency, rather than confidence in that assessment.
Define categories so that they separate the outcomes you care about. Include a review or insufficient-evidence option when a forced choice would hide missing information. If multiple labels can apply, ask separate yes/no questions. The same ticket can describe both a login problem and a refund request.
Copy link to headingHow do decision models differ from structured outputs and classifiers?
An LLM can already return a category in valid JSON. Structured Outputs constrains generated responses to a schema, including allowed enum values. The reason to consider a native decision model is its focused prediction interface and probability information, not the assumption that every LLM answer requires scraping prose.
Caller-defined labels also exist in zero-shot classifiers such as GLiClass. Classification and probability estimation have a longer history than the current decision-model products. Evaluate the actual question types, output semantics, and task accuracy instead of treating the product label as a technical guarantee.
Similarly, reward models learn to score outputs according to preferences. They overlap with decision models in evaluation work, but a preference score doesn't automatically represent the probability that a business condition is true.
An SDK can hide some of these differences. AI SDK's evaluation adapters support both native Jev evaluations and structured-output LLMs. Its OpenAI, Anthropic, and Google adapters omit Choice and Score distributions; their Boolean probabilities are prompted estimates. An identical function call therefore doesn't establish identical model behavior.
Copy link to headingWhich decision models and APIs are available?
The category includes hosted APIs and models whose weights you can run yourself. These examples illustrate different implementations and access methods.
OpenAI introduced Decisions API at DevDay on September 29, 2026. Its announced finite-answer interface doesn't establish support for every question type or probability field offered by the other products.
For open-weight models, inspect the checkpoint's training scope before choosing it. Laya's model card, for example, distinguishes its base models from a checkpoint fine-tuned for typed-decision tasks. Benchmark results apply to the tested checkpoint and shouldn't be attributed to the whole family.
Copy link to headingWhen should you use a decision model?
Start with a repeated judgment whose answer space you can define and whose mistakes you can measure. Useful candidates include:
Routing a form submission to sales, support, or a review queue based on the message's intent.
Assessing an incident against an urgency rubric so responders can prioritize investigation.
Checking an agent trace for a specific outcome, such as whether the agent asked the customer for missing information.
Choosing whether a request needs a reasoning model or can follow an existing workflow.
Keep exact checks in code. If a payment record already has a trusted refunded field, reading that field is more direct than asking a model to infer it. Distinguishing a promised refund from a completed one in a conversation requires interpretation, which is where a model can help.
Existing decision tables can remain responsible for business policy. You can combine a model's interpretation of a refund request with explicit checks of the purchase date and the user's permissions. The request classification alone doesn't authorize a refund.
Use a generative model when you need a written explanation, a new plan, or source code. For a task that requires several deductions, test whether decomposing it into bounded questions preserves the necessary context. Multiple questions in one call don't automatically form a reasoning chain.
Copy link to headingHow can you try a decision model with Liquid d1?
Start by printing predictions for representative examples before connecting them to actions. Liquid d1 is available on AI Gateway under liquid/d1, so you can ask a routing question and a refund question about the same message.
Use AI SDK 7 or later, with AI_GATEWAY_API_KEY set in your server environment. The evaluation quickstart covers installation and credentials. The evaluation API is experimental.
import { experimental_evaluate as evaluate } from 'ai';
const result = await evaluate({ model: 'liquid/d1', state: 'I was charged twice for one order. Please return the extra payment.', questions: { queue: { type: 'choice', instructions: 'Which queue should handle this message?', criteria: { billing: 'Charges, payments, or refunds', delivery: 'Shipping or delivery problems', technical: 'Problems using the application', review: 'Insufficient information or none of the other queues fits', }, }, requestsRefund: { type: 'boolean', instructions: 'Is the customer asking for money to be returned?', }, },});
console.log(result.answers.queue);console.log(result.answers.requestsRefund.probability);The first answer contains the selected queue and its option probabilities. The second value is the estimated probability that a refund was requested. Neither establishes that the customer is entitled to a refund or that a payment has been returned.
The native AI SDK response uses probability for a Boolean answer. Liquid's TypeSafe-compatible interface calls that field noul, and its direct API uses d1:free rather than the Gateway model ID. Preserve the request and response contract for the integration you choose.
Copy link to headingHow should you evaluate probabilities before automating a decision?
To use probabilities in application logic, first check how well they correspond to outcomes on your data. With a well-calibrated binary classifier, roughly 80% of cases assigned a probability near 0.8 should belong to the positive class. Fields called confidence can have a different meaning. Jev's confidence value summarizes the answer distribution; it isn't interchangeable with the selected option's probability.
Build a labeled test set containing routine cases, ambiguous requests, and missing evidence. Compare the model with your current approach using the same inputs. For support routing, measure both incorrect assignments and the share of tickets sent for review. Even a threshold with few mistakes may not improve the workflow if it sends almost everything to a person. Include fallback requests and review work when assessing the total cost and time to resolve a case.
Choose thresholds using those results and the consequences of each error. Mistaking a refund request for a shipping question has different consequences from letting a model approve a payment. Keep an explicit review outcome even when its probability is high: certainty that a case needs review is a reason to send it there.
Treat timeouts and invalid responses as failed evaluations, with their own fallback. After changing a model, question, or category definition, rerun the labeled examples. The decisions can change even when the schema stays the same.
Copy link to headingFrequently asked questions
Copy link to headingAre decision models a replacement for LLMs?
No. Decision models fit judgments with defined outcomes, while generative models can produce explanations and other open-ended content. An application can use a decision model to select a workflow and an LLM to carry out the writing within it.
Copy link to headingIs a decision model the same as a classifier?
The tasks overlap, especially when choosing labels for text. Decision-model APIs often combine classification with ordered scoring and yes/no questions over shared context. Some zero-shot classifiers also accept labels at request time, so that capability alone doesn't distinguish the categories.
Copy link to headingDoes a high probability guarantee a correct answer?
No. Missing evidence or inputs that differ from the training data can lead a model to favor the wrong outcome with high probability. Measure calibration and errors on representative examples before using a threshold to trigger actions.
Copy link to headingDo all decision models support images?
Input support depends on the model and endpoint. OpenAI announced text and image context for Decisions API, while many typed-decision integrations use text-based state. Check the selected provider's input contract before passing media.
Copy link to headingCan you run a decision model yourself?
Yes. Open-weight projects such as Laya and Typical provide models you can run on your own infrastructure. Check the specific checkpoint's license, hardware requirements, and task coverage, then evaluate it on your application's examples.