Skip to content
Dashboard

Jev vs. Laya vs. Liquid d1: Which decision model should you use?

Content Engineer

Jev, Laya, and Liquid d1 support typed decisions over supplied context. If you already use Jev, compare alternatives against that working integration. Laya adds a documented self-hosting path with downloadable checkpoints; Liquid d1 offers another hosted implementation. For hosted use, all three can be evaluated through AI Gateway with the same question definitions.

The choice depends on where you need the model to run and whether its predictions meet your application's requirements. Shared question types make a comparison possible, but don't establish which model will handle your data best.

Copy link to headingHow do Jev, Laya, and Liquid d1 compare?

You can compare these decision models using the same category definitions and scoring rubrics. Their access methods and response formats determine how much of your existing integration you can reuse.

Attribute

Jev

Laya

Liquid d1

Access paths covered here

TypeSafe's hosted API; AI Gateway

Downloadable checkpoints and local serving; AI Gateway

Liquid's hosted API; AI Gateway

AI Gateway model ID

typesafe-ai/jev

convaiinnovations/laya

liquid/d1

AI SDK question types

choice, score, boolean

choice, score, boolean

choice, score, boolean

TypeSafe-compatible yes/no type

noul

noul

noul

Direct integration to assess

TypeSafe API and SDK

Laya runtime or its Jev-compatible HTTP server

Liquid endpoint using a TypeSafe client

Through AI Gateway, the application uses one evaluation interface. The providers' direct integrations require their own endpoint and credential configuration. For example, TypeSafe's direct API uses jev-latest, while Liquid's direct examples use d1:free. Those identifiers differ from the Gateway IDs in the table.

Question-type support alone gives you little reason to switch between these models. The useful distinctions are whether the deployment fits your requirements, whether the integration exposes the fields your application needs, and how the resulting decisions compare on your task.

Copy link to headingWhen does Laya's self-hosting option matter?

Laya publishes downloadable checkpoints and a Jev-compatible HTTP server. That gives teams a path to operate the model on their own infrastructure while retaining familiar state-and-question requests. It is relevant when controlling the serving environment is a requirement.

Self-hosting also makes checkpoint selection part of your application design. Laya includes English, multilingual, and typed-decision checkpoints. Record which one you deploy and test it with the languages and input lengths your application receives. Results from one checkpoint don't establish how another will behave.

Running Laya yourself means managing the inference process and its capacity. Calling Laya through AI Gateway uses a hosted provider instead. The Gateway example below selects Boundless; it doesn't configure Laya's local checkpoint router or demonstrate a self-hosted deployment.

Copy link to headingHow can you compare all three on a product-catalog task?

Suppose a store needs to assign incoming product listings to catalog categories. One listing describes a backpack with a laptop compartment and an insulated pouch. Keyword matching could mistake an accessory for the main product, so the question must distinguish what is being sold from what it can hold.

The following example sends that listing to each model through AI Gateway. It uses one Choice question and prints each returned answer, including the selected category's probability when available. It doesn't write to the catalog.

Use Node.js 22.18 or later and AI SDK 7 or later. The Gateway evaluation quickstart covers installing ai and setting AI_GATEWAY_API_KEY in your shell.

You can also authenticate with Vercel OIDC. The AI SDK uses an available OIDC token when AI_GATEWAY_API_KEY isn't set. For local OIDC use, run vercel link to connect your working directory to a Vercel project. Then run vercel env pull .env.local to save the project's environment variables, including its OIDC token, to .env.local. Repeat that command when the token expires.

Run node compare-catalog.mjs if you've exported an API key in your shell. For the Vercel OIDC token, use node --env-file=.env.local compare-catalog.mjs to load the token from .env.local.

compare-catalog.mjs
import { experimental_evaluate as evaluate } from 'ai';
const models = [
'typesafe-ai/jev',
'convaiinnovations/laya',
'liquid/d1',
];
const state = {
title: 'Harbor commuter pack',
description:
'Roll-top backpack with two shoulder straps, a padded laptop ' +
'compartment, and a removable insulated lunch pouch. Sold as one backpack.',
};
const questions = {
category: {
type: 'choice',
instructions:
'Classify the main product being sold, ignoring included accessories. ' +
'Use review if the description is insufficient or no category fits.',
criteria: {
backpacks: 'Bags designed to be worn on the back with shoulder straps',
laptop_sleeves: 'Standalone protective sleeves for laptops',
lunch_bags: 'Standalone insulated bags primarily for carrying food',
review: 'Insufficient evidence or no matching category',
},
},
};
for (const model of models) {
try {
const result = await evaluate({
model,
state,
questions,
...(model === 'convaiinnovations/laya'
? { providerOptions: { gateway: { only: ['boundless'] } } }
: {}),
});
const answer = result.answers.category;
console.log(JSON.stringify({
model,
choice: answer.choice,
selectedProbability: answer.probabilities?.[answer.choice] ?? null,
probabilities: answer.probabilities ?? null,
}, null, 2));
} catch (error) {
console.error(model, 'Evaluation failed:', error.message);
}
}

Under the stated catalog rule, the intended label is backpacks. That is the example's human-defined expectation, not an observed result from any model. The code reports the responses it receives and continues to the next model if a request fails.

Extend the comparison with listings that test the category boundaries:

Listing to test

Intended label

What the case tests

Standalone padded laptop sleeve advertised as fitting inside a backpack

laptop_sleeves

Whether compatibility language overrides the actual product

Insulated lunch tote with a shoulder strap

lunch_bags

Whether the model mistakes one shoulder strap for a backpack design

Travel organizer described only as a multipurpose carry pouch

review

Whether the model recognizes insufficient evidence

These labels follow this example's catalog policy. If your store puts bundles in a separate category, change the rubric and expected labels before comparing models. Otherwise, disagreements can reflect an undefined merchandising rule.

Copy link to headingWhich probability fields can your application rely on?

The AI SDK evaluation contract requires a probability for Boolean answers, but makes Choice and Score distributions optional. The example checks for a distribution and prints null when it is absent. Treat that as missing information if your publishing rule requires a selected-category probability; don't substitute certainty.

Provider-specific confidence is another value to inspect separately. Jev derives its Choice confidence from the leading probability and the number of options. With four options, a leading probability of 0.7 produces a confidence value of 0.6. Those values describe the same answer but would behave differently under a cutoff of 0.65.

Liquid's direct API also returns confidence for Choice and Score answers. Matching field names don't establish that a cutoff will accept the same mix of correct and incorrect listings. Measure the behavior of the field your application uses, through the integration you intend to deploy.

Copy link to headingWhat should you measure before switching models?

For this catalog task, compare incorrect assignments among automatically accepted listings alongside the proportion sent to review. Include categories with similar descriptions and products outside the catalog taxonomy. An aggregate accuracy figure can hide a model that consistently confuses laptop sleeves with backpacks.

Keep a labeled set for selecting each model's thresholds and a separate set for assessing the final policy. Preserve the same category definitions across the initial comparison. If you later tune a question for one model, record that change so the results describe the model and its configuration together.

Test missing distributions and failed requests as well as successful predictions. Both should remain distinguishable from a deliberate review answer. The example only prints results; a catalog importer needs an explicit path for unresolved listings before it can publish categories automatically.

If Jev already meets your acceptance criteria, use it as the baseline. Evaluate Liquid d1 when assessing another hosted decision API. Include Laya when you want to compare its hosted predictions or need to investigate self-hosting. Choose among the candidates using the deployment you can operate and the errors your catalog workflow can tolerate.

Copy link to headingWhat doesn't this comparison establish?

The shared example demonstrates request compatibility, not comparative accuracy, speed, or an overall winner. One clear backpack listing cannot establish how any model handles an entire catalog. Laya's local checkpoints also need their own evaluation; a request to its Gateway model ID doesn't test every checkpoint in the family.

The task uses product text. If your classification depends on photographs, assess an image-capable workflow separately. OpenAI's Decisions API uses Luna to answer questions with predefined answers from text or image context. For a catalog workflow, that makes classifying a product from its photograph another candidate to evaluate.

Our OpenAI Decisions API vs. Jev comparison covers the input and output requirements to check before adapting a Jev application. The Jev vs. Perplexity's Decisions API comparison examines another option for including images in a decision workflow.

Copy link to headingFrequently asked questions

Copy link to headingCan you compare Jev, Laya, and Liquid d1 without rewriting each request?

Yes. AI Gateway lets you send the same supported state and question definitions to these models using different model IDs. Keep provider-specific routing settings explicit, and inspect each response before connecting it to application actions.

Copy link to headingIs Laya through AI Gateway the same as running Laya yourself?

No. Gateway sends the evaluation to a hosted provider, while self-hosting requires you to choose and operate a Laya checkpoint. Test the deployment you plan to use rather than transferring results between them.

Copy link to headingDoes switching from Jev to Liquid d1 preserve decision thresholds?

An existing cutoff needs validation with Liquid d1's predictions on your data. Even when the answer fields match, a different model can assign different probabilities and change which cases your application accepts automatically.

Copy link to headingCan these models assign more than one category to a product?

Choice selects one outcome from the supplied categories. If your catalog allows several independent labels, define a Boolean question for each label and assess the resulting probabilities against your labeling policy.

More Build with AI articles

Ready to deploy?