OpenAI Decisions API and TypeSafe AI's Jev both target decisions with defined answers. Decisions API includes image input and is in limited preview. Jev accepts text-based state and documents choices, rubric scores, and yes-or-no probabilities. For an existing Jev application, the comparison depends on which of those outputs its code uses.
Two models can select the same category without providing the same information for deciding whether to act on it. Before replacing a decision service, inspect the data your application reads after the answer arrives.
Copy link to headingWhich capabilities can you compare now?
The comparison covers the capabilities confirmed for the Decisions API preview and Jev's documented interface. OpenAI's announcement does not specify whether Decisions API returns option probabilities, rubric scores, or yes-or-no probabilities.
Sources: OpenAI's announcement, Jev input formats, and TypeSafe's question types.
If the application only needs a selected category, define that category consistently for your evaluation. If it also uses a probability distribution to decide whether to request review, the response contract becomes part of the migration requirements.
Copy link to headingWhy do probabilities and confidence matter for a Jev migration?
A Jev Choice answer includes the selected option, probabilities for the available options, and a confidence value. Those fields can support different decisions in your code. The selected option might identify a support queue, while the distribution helps you detect a request that spans two teams.
TypeSafe's confidence value summarizes the shape of that distribution. It differs from the probability assigned to the selected option. Probability spread across several answers produces lower confidence than probability concentrated on one answer.
Suppose your application sends a ticket for review when Billing and Account access both receive substantial probability. A replacement that returns only one category would leave that rule without its inputs. A replacement that also returns uncertainty information would still require checking what each number means and whether the rule works on the new model's answers.
Do not carry a numeric threshold across providers because both outputs fall between zero and one. Test the acceptance rule again against labeled examples, including the cases it sends for review. A threshold is useful only when the resulting routing behavior meets your requirements.
Jev's other question types also need separate mappings. Score rates input against ordered rubric levels. Noul estimates the probability of a yes answer. Turning either into a fixed list of labels changes what the application receives unless the replacement explicitly preserves the original meaning.
Copy link to headingWhen does image input change the comparison?
Decisions API's announced image support makes it a candidate when the evidence is visual. Jev's state format is text-only, so a Jev workflow needs relevant visual information converted into text before evaluation.
Consider a user who reports that an export failed and attaches a screenshot. The message may omit the error visible in the image. A proposed Decisions API workflow could evaluate that screenshot when choosing the next support step. A Jev workflow could first extract the error text and supply it with the message.
Compare those complete workflows using the same original inputs. Check whether text extraction loses information needed for the decision, and include that extraction step when measuring elapsed time or reviewing failures. Direct image input is a useful capability to test; it doesn't establish that one workflow will make better assignments.
If your application already has the error as structured text, use that record in the comparison. There is no need to add a screenshot-processing step to a decision whose evidence is available in the application.
Copy link to headingDoes AI SDK make the two interfaces interchangeable?
AI SDK's current OpenAI evaluation adapter uses the Responses API with structured output. Its Choice and Score results omit probability distributions, and its Boolean results contain prompted probability estimates. This describes that adapter, not Decisions API.
Calling openai.evaluationModel('gpt-6-luna') uses that Responses-based integration. If your Jev routing logic depends on probabilities for each category, this adapter's selected label does not supply enough information to preserve the same acceptance rule.
Vercel's guide to OpenAI evaluation with AI SDK explains the existing Responses-based approach. Use it as a separate baseline if that approach already fits your application.
Copy link to headingWhat would justify changing an existing Jev workflow?
A provider change needs a benefit in the task your application performs. Start with one decision, such as assigning the first support team, and preserve its category definitions. Run the same reviewed cases through each candidate when access permits.
Inspect errors by category. Confusing Billing with Technical support may create extra handling time; misclassifying a request that needs review may send an unresolved problem into an automated path. An overall accuracy score can hide that difference. Record how often each workflow defers a case and whether the remaining automated assignments are acceptable.
Keep the downstream action separate during evaluation. You can compare proposed queue assignments without moving live tickets. After selecting a provider, test failures and timeouts as well as successful answers so the application preserves requests when a decision call cannot finish.
Neither a valid label nor an uncertainty value guarantees the underlying judgment. If the task expands into investigating the customer's problem and writing a response, evaluate that work separately. Vercel's Jev and GPT-6 Astra comparison covers where a general-purpose model fits alongside a decision step.
Copy link to headingFrequently asked questions
Copy link to headingIs Decisions API a drop-in replacement for Jev?
OpenAI's announcement does not establish compatibility with Jev's request and response formats. For applications that use Jev's option probabilities, returning the same category is insufficient to preserve existing routing rules.
Copy link to headingIs Jev's confidence the probability that its chosen answer is correct?
TypeSafe defines confidence as a summary of the distribution across answer options. The selected option also has its own probability, so code should distinguish the two fields and validate any acceptance rule against reviewed cases.
Copy link to headingCan both products evaluate screenshots directly?
OpenAI has announced image context for Decisions API. Jev accepts text-based state, so a screenshot workflow needs to extract the relevant information before sending it to Jev.
Copy link to headingDo decision APIs remove the need for a general-purpose model?
No. A defined choice can handle one step in an application, such as selecting a support queue. Work that also requires a written explanation or further investigation still needs a component capable of that work.