Jev recipe / output contract comparison

Jev vs JSON Mode: Typed Decisions vs Constrained Generation

JSON mode guarantees valid JSON, not correct answers. Compare the retry tax, latency, output billing, and calibration of response_format / structured outputs with Jev zero-generation typed calls.

Six checks before you drop JSON mode

01Name what json_mode actually guarantees: syntactic validity. response_format: {type: "json_object"} promises the model never emits a broken JSON literal — it says nothing about your keys, your enum values, or whether the answer inside is right.
02Separate the three guarantees a pipeline needs: valid syntax (json_mode), schema conformance (structured outputs / json_schema), and correct values (no one guarantees this). Constrained decoding closes the first two gaps and leaves the third — the one that breaks routing — wide open.
03Price the retry tax before you ship: a schema-shaped reply with a hallucinated enum value still costs full output generation, and every validation failure re-bills it. At classification volume, retries are not an edge case — they are the cost model.
04Treat confidence honestly: per-token logprobs measure how surely the model was about to type the next character, not how likely the decision is correct. Only RLCD-calibrated probabilities turn a 0.85 threshold into an accuracy contract.
05Check the latency curve against your volume: generation time scales with output length — 500ms to 2s+ for a typical JSON payload, multiplied by retries. A Jev evaluate call returns typed answers in ~70–100ms no matter how many questions you pack into it.
06Know when constrained generation is still the right tool: free-form content — summaries, variable-length lists, nested generative fields — is beyond Jev, whose heads return bounded values only. Route bounded decisions to Jev and generative payloads to a constrained call, in the same pipeline.
schema / typed decision contract
{
  "intent": {
    "type": "choice",
    "instructions": "Classify what this inbound message is about",
    "criteria": {
      "billing": "Payment, invoice or refund issue",
      "technical": "Product malfunction or bug",
      "sales": "Buying, upgrade or pricing question",
      "abuse": "Safety or abuse report"
    }
  },
  "is_actionable": {
    "type": "noul",
    "instructions": "Does this message require an action from the team, or is it a pure FYI?"
  },
  "priority": {
    "type": "score",
    "instructions": "Rate handling priority, 1 (whenever) to 5 (drop everything)"
  }
}
This schema is what a JSON-schema-plus-validation layer has been approximating — run it in the Playground and watch answers with calibrated confidence arrive in one pass.

Output contract: Jev typed decision heads vs json_mode / structured outputs

METRIC
Jev
json_mode / structured outputs
What is actually guaranteed
Typed answers from decision heads

a value-level contract: one of your options, a score, or a yes/no

json_mode: valid JSON syntax only; structured outputs: schema-shaped output

field values are still generated

Failure mode at scale
No parse step

0% format errors by construction

Valid JSON with wrong or missing values: hallucinated enum labels, plausible-but-wrong decisions

each failure triggers a billed retry

Latency profile
~70–100ms flat

no generation, independent of question count

Scales with output length

500ms to 2s+ for a typical payload, multiplied by every retry

Output token cost
$0

no output tokens exist

Full output pricing per attempt; retries re-bill the entire generation
Confidence semantics
RLCD-calibrated decision confidence

0.85 ≈ 85% accuracy on thresholds

Per-token logprobs at best

sampling statistics, not decision calibration

Free-form content in the payload
Not supported

decision heads return bounded values, never prose

Native strength

summaries, variable-length lists, and nested generative fields

Side-by-side code: constrained generation vs the Jev typed call

python / openai json_mode with value-validation retry loop
import json
import time
from openai import OpenAI

client = OpenAI()
VALID_QUEUES = {"billing", "technical", "sales", "abuse"}
MAX_RETRIES = 3

def route_ticket(ticket_text: str) -> dict:
    for attempt in range(MAX_RETRIES):
        resp = client.chat.completions.create(
            model="gpt-4o-mini",
            response_format={"type": "json_object"},
            messages=[
                {"role": "system", "content": 'Reply as JSON: {"queue": "...", "urgent": true|false}'},
                {"role": "user", "content": ticket_text},
            ],
        )
        payload = json.loads(resp.choices[0].message.content)  # json_mode: always parseable
        if payload.get("queue") in VALID_QUEUES:  # value check: json_mode says nothing about this
            return payload
        time.sleep(2 ** attempt)  # hallucinated enum -> re-bill the full generation
    raise ValueError("queue never converged")

# Typical path: 500-1500ms per attempt, 20-60 output tokens billed on every try
python / jev typed multi-question call
import requests

JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_ROUTE_CONFIDENCE = 0.85

QUESTIONS = {
    "intent": {
        "type": "choice",
        "instructions": "Classify what this inbound message is about",
        "criteria": {
            "billing": "Payment, invoice or refund issue",
            "technical": "Product malfunction or bug",
            "sales": "Buying, upgrade or pricing question",
            "abuse": "Safety or abuse report",
        },
    },
    "is_actionable": {
        "type": "noul",
        "instructions": "Does this message require an action from the team?",
    },
    "priority": {
        "type": "score",
        "instructions": "Rate handling priority, 1 (whenever) to 5 (drop everything)",
    },
}

resp = requests.post(
    JEV_ENDPOINT,
    json={"state": {"message": "..."}, "questions": QUESTIONS},
    timeout=5,
)
resp.raise_for_status()
data = resp.json()

if data["intent"]["confidence"] >= AUTO_ROUTE_CONFIDENCE:
    lane = f"auto:{data['intent']['answer']}"  # ~70-100ms, input tokens only, zero retries
else:
    lane = "review"
typescript / openai structured outputs (json_schema)
import OpenAI from "openai";
import { z } from "zod";
import { zodResponseFormat } from "openai/helpers/zod";

const RouteDecision = z.object({
  queue: z.enum(["billing", "technical", "sales", "abuse"]),
  urgent: z.boolean(),
});

const client = new OpenAI();

const resp = await client.chat.completions.create({
  model: "gpt-4o-mini",
  response_format: zodResponseFormat(RouteDecision, "route_decision"),
  messages: [{ role: "user", content: ticketText }],
});

const decision = RouteDecision.parse(
  JSON.parse(resp.choices[0].message.content!),
);
// Shape is guaranteed. Whether "queue" is the right queue - and how sure the
// model is - are not expressible: you get a guess, not a calibrated decision.
typescript / jev typed multi-question call
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const AUTO_ROUTE_CONFIDENCE = 0.85;

const QUESTIONS = {
  intent: {
    type: "choice",
    instructions: "Classify what this inbound message is about",
    criteria: {
      billing: "Payment, invoice or refund issue",
      technical: "Product malfunction or bug",
      sales: "Buying, upgrade or pricing question",
      abuse: "Safety or abuse report",
    },
  },
  is_actionable: {
    type: "noul",
    instructions: "Does this message require an action from the team?",
  },
  priority: {
    type: "score",
    instructions: "Rate handling priority, 1 (whenever) to 5 (drop everything)",
  },
} as const;

const resp = await fetch(JEV_ENDPOINT, {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    state: { message: ticketText },
    questions: QUESTIONS,
  }),
});
const data = await resp.json();

const lane =
  data.intent.confidence >= AUTO_ROUTE_CONFIDENCE
    ? `auto:${data.intent.answer}`
    : "review"; // low band or safety flag -> same review queue

Jev vs JSON Mode FAQ

What is JSON mode?

An OpenAI chat-completions feature — response_format: {"type": "json_object"} — now also offered by most major providers, that constrains the model to emit syntactically valid JSON: no markdown fences, no prose around it. It guarantees parseability only. Your keys, your enum values, and the correctness of the answer inside remain the model's guess, which is why production code still wraps json_mode calls in schema validation and retry loops.

Is JSON mode the same as structured outputs?

No. JSON mode (json_object) guarantees valid JSON syntax; structured outputs (json_schema) goes further and constrains the shape — the model must emit fields matching your schema, via constrained decoding. Both still generate the payload token by token: latency scales with output length, output tokens are billed, and value correctness is never guaranteed. Structured outputs close the shape gap; the decision-quality gap stays open.

Why does valid JSON still break my pipeline?

Because parseability was never the hard part at scale — value correctness is. A schema-conforming reply with a hallucinated enum value, a plausible-but-wrong routing call, or a missing field fails your application even though it parses cleanly. Each of those failures triggers a validation retry that re-bills the full generation. Multiply your measured failure rate by per-call output cost and QPS: that product is the price of guessing without calibration.

Can I get confidence scores out of JSON mode or structured outputs?

Not in a decision-useful form. Logprobs measure how likely each generated token was under the sampling distribution — typing confidence, not decision accuracy. A route label emitted at high token probability can still be the wrong call. Jev's probabilities are RLCD-calibrated against outcomes, so a 0.85 threshold behaves like an accuracy contract; that property is what makes automation on the number safe.

When should I keep using JSON mode instead of Jev?

Whenever the payload must contain free-form generative content: summaries, explanations, variable-length lists, extracted prose, nested objects with text fields. Jev's decision heads return bounded values only — they cannot write a paragraph. The production pattern is to split the work: Jev answers the bounded questions (routing, gating, scoring, yes/no) in one ~95ms call, and a constrained-generation call fills the generative fields.

How does this compare to Instructor and to OpenAI's Decision API (Luna)?

Different layers of the same problem. Instructor is a library that wraps generation with Pydantic validation and retry loops — see the Jev vs Instructor comparison for that axis. json_mode and structured outputs are API-level constraints on generation — this page. OpenAI's Decisions API on GPT-6 Luna is a separate non-generative product returning one answer plus confidence from a predefined set — see Jev vs OpenAI Decision API for the contract-level comparison. The direction of travel is the same everywhere: away from parsing generated text, toward typed decisions.

Go deeper on the structured-output family