Jev recipe / output contract comparison
Jev vs JSON Mode: Typed Decisions vs Constrained Generation
JSON mode guarantees valid JSON, not correct answers. Compare the retry tax, latency, output billing, and calibration of response_format / structured outputs with Jev zero-generation typed calls.
Six checks before you drop JSON mode
{
"intent": {
"type": "choice",
"instructions": "Classify what this inbound message is about",
"criteria": {
"billing": "Payment, invoice or refund issue",
"technical": "Product malfunction or bug",
"sales": "Buying, upgrade or pricing question",
"abuse": "Safety or abuse report"
}
},
"is_actionable": {
"type": "noul",
"instructions": "Does this message require an action from the team, or is it a pure FYI?"
},
"priority": {
"type": "score",
"instructions": "Rate handling priority, 1 (whenever) to 5 (drop everything)"
}
}Output contract: Jev typed decision heads vs json_mode / structured outputs
a value-level contract: one of your options, a score, or a yes/no
field values are still generated
0% format errors by construction
each failure triggers a billed retry
no generation, independent of question count
500ms to 2s+ for a typical payload, multiplied by every retry
no output tokens exist
0.85 ≈ 85% accuracy on thresholds
sampling statistics, not decision calibration
decision heads return bounded values, never prose
summaries, variable-length lists, and nested generative fields
Side-by-side code: constrained generation vs the Jev typed call
import json
import time
from openai import OpenAI
client = OpenAI()
VALID_QUEUES = {"billing", "technical", "sales", "abuse"}
MAX_RETRIES = 3
def route_ticket(ticket_text: str) -> dict:
for attempt in range(MAX_RETRIES):
resp = client.chat.completions.create(
model="gpt-4o-mini",
response_format={"type": "json_object"},
messages=[
{"role": "system", "content": 'Reply as JSON: {"queue": "...", "urgent": true|false}'},
{"role": "user", "content": ticket_text},
],
)
payload = json.loads(resp.choices[0].message.content) # json_mode: always parseable
if payload.get("queue") in VALID_QUEUES: # value check: json_mode says nothing about this
return payload
time.sleep(2 ** attempt) # hallucinated enum -> re-bill the full generation
raise ValueError("queue never converged")
# Typical path: 500-1500ms per attempt, 20-60 output tokens billed on every tryimport requests
JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_ROUTE_CONFIDENCE = 0.85
QUESTIONS = {
"intent": {
"type": "choice",
"instructions": "Classify what this inbound message is about",
"criteria": {
"billing": "Payment, invoice or refund issue",
"technical": "Product malfunction or bug",
"sales": "Buying, upgrade or pricing question",
"abuse": "Safety or abuse report",
},
},
"is_actionable": {
"type": "noul",
"instructions": "Does this message require an action from the team?",
},
"priority": {
"type": "score",
"instructions": "Rate handling priority, 1 (whenever) to 5 (drop everything)",
},
}
resp = requests.post(
JEV_ENDPOINT,
json={"state": {"message": "..."}, "questions": QUESTIONS},
timeout=5,
)
resp.raise_for_status()
data = resp.json()
if data["intent"]["confidence"] >= AUTO_ROUTE_CONFIDENCE:
lane = f"auto:{data['intent']['answer']}" # ~70-100ms, input tokens only, zero retries
else:
lane = "review"import OpenAI from "openai";
import { z } from "zod";
import { zodResponseFormat } from "openai/helpers/zod";
const RouteDecision = z.object({
queue: z.enum(["billing", "technical", "sales", "abuse"]),
urgent: z.boolean(),
});
const client = new OpenAI();
const resp = await client.chat.completions.create({
model: "gpt-4o-mini",
response_format: zodResponseFormat(RouteDecision, "route_decision"),
messages: [{ role: "user", content: ticketText }],
});
const decision = RouteDecision.parse(
JSON.parse(resp.choices[0].message.content!),
);
// Shape is guaranteed. Whether "queue" is the right queue - and how sure the
// model is - are not expressible: you get a guess, not a calibrated decision.const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const AUTO_ROUTE_CONFIDENCE = 0.85;
const QUESTIONS = {
intent: {
type: "choice",
instructions: "Classify what this inbound message is about",
criteria: {
billing: "Payment, invoice or refund issue",
technical: "Product malfunction or bug",
sales: "Buying, upgrade or pricing question",
abuse: "Safety or abuse report",
},
},
is_actionable: {
type: "noul",
instructions: "Does this message require an action from the team?",
},
priority: {
type: "score",
instructions: "Rate handling priority, 1 (whenever) to 5 (drop everything)",
},
} as const;
const resp = await fetch(JEV_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
state: { message: ticketText },
questions: QUESTIONS,
}),
});
const data = await resp.json();
const lane =
data.intent.confidence >= AUTO_ROUTE_CONFIDENCE
? `auto:${data.intent.answer}`
: "review"; // low band or safety flag -> same review queueJev vs JSON Mode FAQ
What is JSON mode?
An OpenAI chat-completions feature — response_format: {"type": "json_object"} — now also offered by most major providers, that constrains the model to emit syntactically valid JSON: no markdown fences, no prose around it. It guarantees parseability only. Your keys, your enum values, and the correctness of the answer inside remain the model's guess, which is why production code still wraps json_mode calls in schema validation and retry loops.
Is JSON mode the same as structured outputs?
No. JSON mode (json_object) guarantees valid JSON syntax; structured outputs (json_schema) goes further and constrains the shape — the model must emit fields matching your schema, via constrained decoding. Both still generate the payload token by token: latency scales with output length, output tokens are billed, and value correctness is never guaranteed. Structured outputs close the shape gap; the decision-quality gap stays open.
Why does valid JSON still break my pipeline?
Because parseability was never the hard part at scale — value correctness is. A schema-conforming reply with a hallucinated enum value, a plausible-but-wrong routing call, or a missing field fails your application even though it parses cleanly. Each of those failures triggers a validation retry that re-bills the full generation. Multiply your measured failure rate by per-call output cost and QPS: that product is the price of guessing without calibration.
Can I get confidence scores out of JSON mode or structured outputs?
Not in a decision-useful form. Logprobs measure how likely each generated token was under the sampling distribution — typing confidence, not decision accuracy. A route label emitted at high token probability can still be the wrong call. Jev's probabilities are RLCD-calibrated against outcomes, so a 0.85 threshold behaves like an accuracy contract; that property is what makes automation on the number safe.
When should I keep using JSON mode instead of Jev?
Whenever the payload must contain free-form generative content: summaries, explanations, variable-length lists, extracted prose, nested objects with text fields. Jev's decision heads return bounded values only — they cannot write a paragraph. The production pattern is to split the work: Jev answers the bounded questions (routing, gating, scoring, yes/no) in one ~95ms call, and a constrained-generation call fills the generative fields.
How does this compare to Instructor and to OpenAI's Decision API (Luna)?
Different layers of the same problem. Instructor is a library that wraps generation with Pydantic validation and retry loops — see the Jev vs Instructor comparison for that axis. json_mode and structured outputs are API-level constraints on generation — this page. OpenAI's Decisions API on GPT-6 Luna is a separate non-generative product returning one answer plus confidence from a predefined set — see Jev vs OpenAI Decision API for the contract-level comparison. The direction of travel is the same everywhere: away from parsing generated text, toward typed decisions.