landscape · pricing · calibration

Decision model APIs compared: Jev, OpenAI Luna, Perplexity, Laya, Clef & D1 (2026)

Six vendors now sell the same product shape - a typed answer with a probability instead of generated text. This page is the category map: what a decision model API is, who ships one, what a decision actually costs, and which vendor fits which workload.

Quick answer

A decision model API returns typed answers with probabilities - no text generation. In October 2026 the hosted field is three vendors with input-only billing: Jev at $0.042 per million input tokens, Perplexity at $0.02 (halved from its $0.04 launch rate), and OpenAI's Luna-based Decisions API at $0.10, in public beta since 2026-10-06. Output is free on all three. If the data cannot leave your infrastructure, the open-weight route - Clef, Liquid D1, Laya, the OpenJev family - carries no API price at all.

Pricing and status on this page were verified against official vendor documentation on 2026-10-07: OpenAI's Decisions guide and pricing page, and Perplexity's Decisions quickstart and pricing page. Beta labels are literal: OpenAI's GA is promised, not shipped, and beta rates can move. Independent benchmark figures are quoted from the named third-party evaluations.

What is a decision model API?

A decision model API is an endpoint that answers questions about your content with typed answers and probabilities instead of generated text. You send a state - a ticket, a document, a JSON record, sometimes an image - plus named questions with predefined answers; it returns one answer per question with a probability attached, in a single forward pass. There is no prose to parse, no JSON schema for the model to break, and no output-token bill, because nothing is generated. The category is also called typed decision APIs or System One models, and both OpenAI and Perplexity ship products named literally "Decisions API" - which is why the generic term matters.

The economics come from what the model does not do. Because it never writes tokens, cost scales with input only, latency stays in the tens-to-hundreds of milliseconds, and the same state can carry many questions in one call. The trade-off is scope: a decision model classifies, routes, and scores - it does not write the reply, reason about novel problems, or explain its answer. Your code does the reasoning with the numbers it returns.

The three-primitive contract - and one naming split

Yes / no - Noul vs predicate

A yes/no question returning the probability of yes. Jev and Perplexity both call this noul; OpenAI's beta calls it predicate - same shape, different name, and the one rename you must map when switching vendors.

One of N - choice

Pick one option from your declared set: ticket queues, intents, categories. All three vendors call this choice and return per-option probabilities plus the most likely option.

Ordered rubric - score

Rate content on an ordered scale (severity 1-5, quality levels). All three call it score and return a probability distribution over the levels; Perplexity documents up to 10 levels per rubric, and OpenAI's beta guide defines score as an ordered rubric with its own response shape.

The landscape: six decision model APIs

Every vendor selling typed decisions as of 2026-10-07, by deployment, status, primitive naming, and vision input. Vendor-published performance claims are omitted - they belong to benchmarks, not a landscape map.

VendorDeploymentStatus (2026-10-07)PrimitivesVision input
Jev (TypeSafe)Hosted API (also via OpenRouter and gateways)GAnoul / choice / scoreNo - text state
OpenAI Decisions (gpt-6-luna)Hosted (api.openai.com)Public beta - GA "in the coming weeks"predicate / choice / score - predicate is OpenAI's name for noulYes - inline base64 data URLs only
Perplexity DecisionsHosted + Apache 2.0 weights on Hugging FaceGA (launch week, 2026-10-05)noul / choice / score - identical names to Jev; v1.1 and v1 models both servedYes - base64 image parts, 32x32 tiles
Clef (Cloudflare)Hosted on Cloudflare + Apache 2.0 open weightsGA (launched 2026-10-01)Typed decision answers; documents a "fully Jev-API compatible" surfaceYes - multimodal
Laya (open source)Open weights (Apache-2.0, 421M) - pip install, CPU-capable; free hosted route via gatewaysGA (open source)choice / score / noulNo - text
Liquid D1Open weights (Apache 2.0) + hosted via OpenRouterGA (launched 2026-09-30; vision update 2026-10-05)Typed decision answers on Liquid's own API surfaceYes - "now with vision"

Two reading notes. OpenAI's row is a beta, not a launch: the endpoint, field names, and the price below can move before GA, and its pricing is published on the developer guide page while the main pricing page still carries no Decisions line - check both. And the two noul-named rows (Jev, Perplexity) are compatible in vocabulary only: same primitive names, different contracts - which is exactly what the head-to-head pages below test.

Head-to-head deep reads

Pricing: what a decision actually costs

All three hosted vendors bill input tokens only - output is free, with no cache fees and no per-request fee. The philosophy is identical; the rates are not.

VendorInput price (per 1M tokens)OutputNotes
Jev (hosted)$0.042FreeGA; published since launch; also reachable via OpenRouter and gateways
Perplexity Decisions$0.02FreeGA; no per-request fee; halved from the $0.04 launch-week rate (verified 2026-10-07)
OpenAI Decisions (gpt-6-luna)$0.10FreePublic beta; no cache-read/cache-write charges; rate published on the guide page - the pricing page has no Decisions line yet
Self-hosted (Clef, D1, Laya, OpenJev)$0 API - your GPUFreeApache 2.0 weights; the marginal cost is electricity and hardware you already own

The same decision, priced three ways

Take a 500-token support ticket routed once. Jev: $0.000021 per decision - about $21 per million decisions. Perplexity: $0.00001 - about $10 per million. Luna at the beta rate: $0.00005 - about $50 per million. A general-purpose LLM at $2.50/M input and $10/M output that writes a 35-token JSON verdict pays $0.0016 for the same call - roughly 75x Jev - and pays again on every schema-validation retry. At decision-model rates the gaps are real but small: the $0.02-vs-$0.042 difference starts to matter around tens of millions of decisions a month.

The line the pricing trackers miss: self-hosting

Third-party pricing content generally stops at hosted rates. The open-weight row changes the question: Clef, Liquid D1, and Perplexity's own decider weights are Apache 2.0, Laya ships as a 421M pip install that runs on a CPU, and the OpenJev family covers fully local routes. On those paths the marginal cost of a decision is the GPU you already own - effectively electricity. You trade the API fee for hardware, latency, and the calibration work the hosted vendor used to do for you - which is exactly why the trust section below exists. As of 2026-10-07, the third-party content ranking for these pricing keywords (OpenAI explainers and hosted-rate trackers) does not cover this line.

Jev pricing in depth - with the interactive cost calculator

Which decision model API should you use?

By use case, not by leaderboard - the honest answer to "best" is conditional on four things: the price floor, existing agreements, calibration requirements, and where the data lives.

Lowest input price, or vision input today

Perplexity Decisions

$0.02/M input is the category floor (half of Jev, a fifth of Luna's beta rate), with output free, GA status, image input, 262k-token context, and Apache 2.0 weights as your exit. Mind the two documented contract quirks: confidence is "the model's own certainty estimate," and identical requests can differ in the second decimal place.

Already running on OpenAI - enterprise agreement, one vendor, one bill

OpenAI Decisions (gpt-6-luna)

First-party integration and procurement are real advantages no rate sheet captures, and image input works today (inline base64). The trade is everything a beta carries: pricing is published on the guide page only, confidence semantics are undocumented, and GA is promised, not shipped. Budget $0.10/M input - 2.4x Jev's rate.

Audits, thresholds, and calibration you can cite

Jev

RLCD-calibrated probabilities, published ECE methodology, per-option criteria in versioned code, multi-question calls at ~70-100ms, plus the audit-trail and fallback stack around it. Read the trust section below first: the independent evaluation also found real limits (last of twelve on LegalBench) - calibrated is not the same as universally accurate.

Data that cannot leave your infrastructure

Clef, Liquid D1, or Laya

All open weights (Apache 2.0), all free of API fees: Clef for a Cloudflare-native multimodal route, D1 for a vision-capable hosted-or-self path, Laya for a CPU-sized encoder you can fine-tune - with its own limits candidly documented. The OpenJev family covers the fully local route. The honest cost is hardware plus calibration work you now own.

Calibration & trust: what independent measurement says

Every vendor publishes an accuracy story; fewer publish the honesty of their confidence. Two independent evaluations from October 2026 cut both ways - we cite both sides.

The endorsement: vals.ai's preregistered evaluation (2026-10-06; twelve systems, 400 human-audited claim-verification items built from SEC filings) scored hosted Jev at 0.975 - statistically tied with GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna - at $0.02 per 1,000 cases, about 1/498th of Astra's cost, with the lowest ECE of the twelve (0.011) and ~0.1s median latency (194x lower than Astra at 32 judgments per call). TypeSafe's launch headline of "193.6x faster, 444.6x cheaper," the evaluators concluded, holds up.

The criticism, from the same evaluation: on a preregistered 12-subtask slice of LegalBench, Jev finished last of twelve on class-balanced accuracy (0.730 vs GPT-5.6 Terra's 0.932, with every other system's lead statistically significant), it collapsed on one contract-NLI subtask (answering yes to 30 of 33 items), and its ECE rose to 0.108 - the worst of the twelve. Calibration measured on one task did not transfer to the other.

Neither finding generalizes automatically - that is the lesson this category keeps re-proving. The 575,000-call Synthpop audit, Red Hat's decision-model-vs-classifier benchmark, and DoubtBench each test a different property of decision-model confidence, and all of them are collected on our calibration page. The rule that survives every audit: a confidence number is a hypothesis until you have measured it on your own labels.

Decision model API FAQ

Is there a "Decisions API"?

Two products carry that exact name. OpenAI's Decisions API runs on gpt-6-luna and entered public beta on 2026-10-06; Perplexity's Decisions API went GA in launch week (2026-10-05) with Apache 2.0 weights. The generic category term is decision model API - and the naming collision is why searches for "decisions api" mix meeting-software docs with AI products. This page maps the AI meaning; each vendor's deep read lives on the head-to-head pages.

Is the Decisions API GA?

Depends which one. Perplexity's Decisions API is GA: any Perplexity API key works, pricing is published, and the docs carry no beta language. OpenAI's Decisions API is in public beta (since 2026-10-06): anyone can call it, but the guide states OpenAI expects to GA "in the coming weeks," the beta pricing ($0.10/M input-only) is published on the guide page while the main pricing page has no Decisions line, and field shapes can still move. Budget for the rate, but pin the version.

What does a decision model API cost?

Hosted, all three bill input tokens only with free output: Jev $0.042 per million input tokens, Perplexity $0.02 (halved from its $0.04 launch rate), OpenAI's Luna beta $0.10. A 500-token decision therefore costs about $0.000021 on Jev, $0.00001 on Perplexity, and $0.00005 on Luna. Self-hosting the open-weight routes (Clef, Liquid D1, Laya, OpenJev) carries no API fee at all - just the hardware.

Can I self-host a decision model?

Yes - the open-weight row is real. Cloudflare's Clef and Liquid AI's D1 ship Apache 2.0 weights; Perplexity publishes Apache 2.0 weights for its decider with official inference code; Laya is a 421M Apache-2.0 encoder that installs with pip and runs on a CPU; and the OpenJev family (Kev, SemIf, Von, OpenJev) covers fully local routes. Jev itself is hosted-only - the local route runs through the open clones. The trade: the calibration and validation work the hosted vendor does for you becomes yours. Our self-hosted matrix and the Clef local install guide are the practical starting points.

Decision model API vs LLM API - when do I need which?

Use a decision model API when the output is a choice your code acts on: routing, classification, gating, scoring against a rubric. You get typed answers with probabilities, input-only billing, and tens-of-milliseconds latency. Use an LLM when the output is language: the customer reply, the summary, the code. The common production shape is both - the LLM writes, a decision model gates whether the reply ships. Forcing an LLM to emit JSON adds a schema-retry tax and uncalibrated confidence to every gate.

Which is the best decision model API?

There is no unconditional best - there is a best for your constraints. Cheapest input with vision: Perplexity. Already inside an OpenAI agreement: Luna's beta. Audit-grade thresholds and published calibration: Jev. Data that cannot leave your hardware: Clef, D1, or Laya. Then decide with data no vendor can give you: run ~100 labeled examples from your own workload through the finalists before anything automates on confidence alone.