Jev recipe / vendor comparison

Jev vs OpenAI Decision API (Luna): Typed Decisions Compared

OpenAI answered Jev with the Decision API on GPT-6 Luna. Compare the contracts: answer sets vs typed Choice/Score/Noul, 150ms vs ~95ms, unpublished vs public pricing, and what each confidence number is worth.

Six checks before you pick a decision API

01Confirm you are comparing the same contract, not a chat model: both products take context plus a question plus a predefined answer set and return one answer with a confidence score. Two vendors converging on that shape within a month is the market validating typed decisions.
02Map your decisions to primitives before picking a vendor: Choice for one-of-N routing, Score for ordered quality, Noul for yes/no gates. The Luna preview covers the answer-set case; an ordered Score question has no Luna equivalent yet.
03Price your decision volume both ways: Jev publishes $0.042 per million input tokens with output free; Decision API pricing per call was still undisclosed at the limited preview. At classification volume, unpublished pricing is a migration risk, not a detail.
04Gate automation on calibration, not on the existence of a confidence number: Jev probabilities are RLCD-calibrated, so a 0.85 threshold is an accuracy contract. Whatever confidence semantics Luna ships, validate them on ~100 of your own labeled examples before any auto-action.
05Check the modality and deployment edges honestly: the Luna context field accepts images today while Jev state is text (visual judgment lives in the OpenJev ecosystem); if decisions must stay on your hardware, the local route is OpenJev clones — Luna has no self-host path.
06Keep vendor optionality by design: wrap the decision call behind one interface so a Jev gate and a Luna call are swappable, and let the confidence-gated fallback chain route the low band from either model to the same review queue.
schema / typed multi-question contract
{
  "routing": {
    "type": "choice",
    "instructions": "Which queue should this ticket go to?",
    "criteria": {
      "billing": "Payment, invoice or refund issue",
      "technical": "Product malfunction or bug",
      "sales": "Buying or upgrade question",
      "abuse": "Safety or abuse report"
    }
  },
  "needs_safety_escalation": {
    "type": "noul",
    "instructions": "Does this ticket require a safety escalation regardless of queue?"
  },
  "urgency": {
    "type": "score",
    "instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)"
  }
}
This schema is the contract depth in one view — run all three questions in the Playground and inspect the calibrated confidence.

Decision API contracts: Jev typed primitives vs OpenAI Decision API (Luna)

METRIC
Jev
OpenAI Decision API (Luna)
Availability (October 2026)
Generally available

self-serve API key, also reachable through OpenRouter and gateway providers

Limited preview announced at DevDay (2026-09-29); broad release promised "in the coming days"

preview shape may still change

Question contract
Typed Choice (one of N options, criteria per option), Score (ordered 1–5), Noul (yes/no)

multiple questions share one call

Question + predefined flat answer set + context

one decision per call; max answer count not yet documented

Latency per decision
~70–100ms single forward pass

no text generated

~150ms (OpenAI-cited, vs 1.6s for GPT-6 Luna chat)

also a single pass, not generation

Pricing transparency
Public: $0.042 per million input tokens, output free

pennies per 100k decisions

Not disclosed at the preview

"more details at broad rollout" (The New Stack, 2026-09-29)

Confidence semantics
RLCD-calibrated probabilities

0.85 means ~85% right, so thresholds act as accuracy contracts

Confidence score shipped, calibration methodology unpublished

a ranking signal until validated on your own labels

Modality & deployment
Text state (visual judgment via the OpenJev ecosystem); local route with OpenJev clones on your own hardware
Context accepts images today (per OpenAI changelog); OpenAI-hosted only

no self-host or on-prem path

Side-by-side code: OpenAI decisions endpoint vs the Jev typed call

python / openai decisions endpoint (limited-preview shape)
import os
import requests

# Limited-preview shape (announced 2026-09-29). Verify against the
# current reference at broad rollout - preview APIs move.
resp = requests.post(
    "https://api.openai.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"},
    json={
        "model": "gpt-6-luna",
        "context": "Ticket: my invoice shows the same charge twice...",
        "question": "Which queue should this ticket go to?",
        "answers": ["billing", "technical", "sales", "abuse"],
    },
    timeout=5,
)
resp.raise_for_status()
decision = resp.json()  # -> {"answer": "billing", "confidence": 0.91}

# One flat answer set per call: no per-option criteria, no second
# question sharing the call, no ordered score primitive.
python / jev typed multi-question gate
import requests

JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_ROUTE_CONFIDENCE = 0.85

QUESTIONS = {
    "routing": {
        "type": "choice",
        "instructions": "Which queue should this ticket go to?",
        "criteria": {
            "billing": "Payment, invoice or refund issue",
            "technical": "Product malfunction or bug",
            "sales": "Buying or upgrade question",
            "abuse": "Safety or abuse report",
        },
    },
    "needs_safety_escalation": {
        "type": "noul",
        "instructions": "Does this ticket require a safety escalation regardless of queue?",
    },
    "urgency": {
        "type": "score",
        "instructions": "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
    },
}

resp = requests.post(
    JEV_ENDPOINT,
    json={"state": {"ticket": "..."}, "questions": QUESTIONS},
    timeout=5,
)
resp.raise_for_status()
data = resp.json()

route = data["routing"]  # criteria-checked answer + calibrated confidence
if (
    route["confidence"] >= AUTO_ROUTE_CONFIDENCE
    and not data["needs_safety_escalation"]["answer"]
):
    lane = f"auto:{route['answer']}"  # ~70-100ms, input tokens only
else:
    lane = "review"  # 0.60-0.85 band or safety flag
typescript / openai decisions endpoint (limited-preview shape)
const resp = await fetch("https://api.openai.com/v1/decisions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "gpt-6-luna",
    context: "Ticket: my invoice shows the same charge twice...",
    question: "Which queue should this ticket go to?",
    answers: ["billing", "technical", "sales", "abuse"],
  }),
});
// Limited-preview shape (announced 2026-09-29) - verify at broad rollout.
const decision = (await resp.json()) as {
  answer: string;
  confidence: number;
};
typescript / jev typed multi-question call
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const AUTO_ROUTE_CONFIDENCE = 0.85;

const QUESTIONS = {
  routing: {
    type: "choice",
    instructions: "Which queue should this ticket go to?",
    criteria: {
      billing: "Payment, invoice or refund issue",
      technical: "Product malfunction or bug",
      sales: "Buying or upgrade question",
      abuse: "Safety or abuse report",
    },
  },
  needs_safety_escalation: {
    type: "noul",
    instructions:
      "Does this ticket require a safety escalation regardless of queue?",
  },
  urgency: {
    type: "score",
    instructions:
      "Rate how urgent a human reply is, 1 (routine) to 5 (business-stopping)",
  },
} as const;

const resp = await fetch(JEV_ENDPOINT, {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ state: { ticket: "..." }, questions: QUESTIONS }),
});
const data = await resp.json();

const lane =
  data.routing.confidence >= AUTO_ROUTE_CONFIDENCE &&
  !data.needs_safety_escalation.answer
    ? `auto:${data.routing.answer}`
    : "review";

Jev vs OpenAI Decision API FAQ

What is OpenAI's Decision API?

A non-chat endpoint announced at OpenAI DevDay on 2026-09-29, built on Luna — the smallest and most affordable model in the GPT-6 family. You supply a question, a predefined answer set, and context; it returns one of the predefined answers plus a confidence score in about 150ms. It shipped as a limited preview with broad release promised "in the coming days", and its per-call pricing was still undisclosed at announcement. The New Stack reported it as likely a reaction to TypeSafe's Jev.

Is the Decision API the same thing as Jev?

Same category, different contracts. Both are non-generative: context plus a question plus a fixed answer set in, one answer plus confidence out — no prose generated. Jev's contract is deeper: Choice questions carry criteria per option (the policy lives in reviewable code), Score covers ordered decisions, multiple questions share one call, probabilities are RLCD-calibrated, input-only pricing is public ($0.042/M), and the OpenJev ecosystem offers a self-hosted route. Luna counters with a simpler request shape, image context today, and first-party OpenAI integration — at preview-stage pricing and calibration transparency.

Which one should I pick for production today?

Default to Jev while Luna is in limited preview: it is generally available, its thresholds are calibrated, its pricing is public, and migration is one adapter function away. Choose Luna first if you are all-in on OpenAI, your decisions need image context, or single-vendor billing matters more than price transparency. Either way, run ~100 labeled examples from your own workload through both gates before automating anything on confidence alone.

Can I use both in one pipeline?

Yes, and the pattern is already documented: put both behind one decision interface in your code, then let the confidence-gated fallback chain own the routing — a Jev gate as the always-on first pass, a Luna call where image context helps, and the same 0.60–0.85 review band catching low confidence from either vendor. Vendor swaps then cost an adapter, not a rewrite.

How is this different from the Jev vs Luna benchmark guide?

This page compares the product and API contracts: request shape, primitives, latency, pricing, calibration, deployment. The Jev vs Luna benchmark guide compares measured task results — a 505-sample third-party evaluation where Jev took 382/505 at about $0.01 versus Luna's $0.06, including the event where Luna won. Read both: the contract tells you what you can build; the benchmark tells you what to expect when you build it.

What happens to my code if the preview shape changes?

OpenAI promised more detail "at broad rollout", so assume the request and response fields may move — which is exactly why the Luna call should live inside one adapter function your application never sees past. The Jev evaluate endpoint has been generally available and unchanged since launch. Keep your ~100-example eval set handy either way: re-validating a swapped gate is an afternoon, not a quarter.

Go deeper on contracts and measurements