Jev recipe / abstention & honesty
Decision Model Abstention with Jev: When the Right Answer Is "I Don't Know"
Some questions cannot be answered from the state in front of the model — and no confidence threshold fixes that. Pair a typed Choice with an is_answerable Noul and a missing_information Choice, return a structured "I don't know" that names the missing piece, and gate automation on answerability and confidence together.
Build the abstention contract in six steps
{
"answer": {
"type": "choice",
"instructions": "Answer using only what the state contains",
"criteria": {
"billing": "Payment, invoice or refund",
"technical": "Product malfunction",
"sales": "Buying or upgrade question"
}
},
"is_answerable": {
"type": "noul",
"instructions": "Does the state contain enough information to pick exactly one option — not more data, not a guess?"
},
"missing_information": {
"type": "choice",
"instructions": "If is_answerable is no: what single piece is missing?",
"criteria": {
"none": "The state is sufficient",
"a_date": "A date or deadline",
"an_amount": "An amount or balance",
"an_identity": "Who a person or account is",
"a_policy_rule": "Which policy or version applies",
"external_state": "Live data from another system"
}
}
}Six questions, one answerability gate
Every question below already ran through the evaluate endpoint with the abstention contract: a typed answer, its calibrated confidence, and an is_answerable Noul. Drag the answerability gate and watch the lanes recompute. The dangerous rows are the confident ones: order #9012 shows 0.88 confidence on a date the state cannot contain — confidence measures sureness, not whether the answer exists in the state.
| Question & state | Answer & confidence | is_answerable | Lane |
|---|---|---|---|
Is order #5521 eligible for a refund under policy v3? q_101 · order lines + return window dates + policy event | eligible0.96 | 0.98 | answered |
Which team owns this: billing / technical / sales? q_102 · full ticket text with one clear topic | billing0.97 | 0.95 | answered |
Did the customer change their address before this order? q_103 · profile with two timestamped addresses | true0.94 | 0.91 | answered |
When will order #9012 ship? q_104 · order lines only — no inventory, no carrier SLA | 2026-10-140.88 | 0.31 | abstainedwould have auto-executed on confidence alone |
missing:inventory state + carrier SLA table | |||
Why did account #4183 churn? q_105 · usage metrics — no survey, no conversation | price_sensitivity0.71 | 0.22 | abstained |
missing:any direct customer statement — metrics show what happened, not why | |||
Will contract #77 renew next quarter? q_106 · usage trend — renewal outcome is future information | true0.66 | 0.41 | abstained |
missing:the outcome does not exist yet — only intent signals do | |||
Demo data only — six simulated questions. The lane recompute uses the same answerability rule documented in the recipe copy.
Abstention behavior: typed contract vs prompted "I don't know" vs confidence-only gating
is_answerable returns a probability on every call, whether or not you gate on it
the follow-up question is whatever a human infers from the wording
the two readings gate independently
the workflow collects exactly that piece, and the abstention is logged as a decision
AbstentionBench exists because prompted refusals do not track answerability
Production code: the abstention contract in TypeScript and Python
import os
import requests
resp = requests.post(
"https://api.openai.com/v1/chat/completions",
headers={"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"},
json={
"model": "gpt-4o-mini",
"messages": [{
"role": "user",
"content": (
"When will order #9012 ship? Use only the data below. "
"If the data does not contain the answer, reply exactly: I DON'T KNOW\n"
"Data: {'lines': [...], 'channel': 'web'}"
),
}],
},
timeout=30,
)
text = resp.json()["choices"][0]["message"]["content"]
if "I DON'T KNOW" in text:
...
# The failure you cannot see: rephrase the question ("by which date..."),
# warm up the style, or lengthen the data dump, and the same state returns
# "Ships Oct 14" — fluent, unfalsifiable, and indistinguishable from a
# grounded answer. You gated on prose, and prose has no threshold.import requests
JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
ANSWERABILITY_GATE = 0.60 # is_answerable below this -> abstain
AUTO_CONFIDENCE = 0.85 # answer confidence below this -> review
QUESTIONS = {
"answer": {
"type": "choice",
"instructions": "Answer using only what the state contains",
"criteria": {
"billing": "Payment, invoice or refund",
"technical": "Product malfunction",
"sales": "Buying or upgrade question",
},
},
"is_answerable": {
"type": "noul",
"instructions": "Does the state contain enough information to pick exactly one option — not more data, not a guess?",
},
"missing_information": {
"type": "choice",
"instructions": "If is_answerable is no: what single piece is missing?",
"criteria": {
"none": "The state is sufficient",
"a_date": "A date or deadline",
"an_amount": "An amount or balance",
"an_identity": "Who a person or account is",
"a_policy_rule": "Which policy or version applies",
"external_state": "Live data from another system",
},
},
}
def decide_or_abstain(state: dict) -> dict:
resp = requests.post(
JEV_ENDPOINT,
json={"state": state, "questions": QUESTIONS},
timeout=5,
)
resp.raise_for_status()
data = resp.json()
knows = data["is_answerable"]
if not knows["answer"] or knows["confidence"] < ANSWERABILITY_GATE:
# Typed abstention: the workflow routes on the missing key,
# and the record below is loggable, countable, replayable.
return {
"abstained": True,
"missing": data["missing_information"]["answer"],
"is_answerable": knows["confidence"],
}
answer = data["answer"]
lane = "auto" if answer["confidence"] >= AUTO_CONFIDENCE else "review"
return {
"abstained": False,
"answer": answer["answer"],
"confidence": answer["confidence"],
"lane": lane,
}import OpenAI from "openai";
const openai = new OpenAI();
export async function whenDoesItShip(orderData: object) {
const resp = await openai.chat.completions.create({
model: "gpt-4o-mini",
messages: [
{
role: "user",
content: `When will this order ship? Use only the data below. If the data does not contain the answer, reply exactly: I DON'T KNOW\nData: ${JSON.stringify(orderData)}`,
},
],
});
const text = resp.choices[0].message.content ?? "";
// You gated on prose. A paraphrased question, a warmer style, or a
// longer data dump shifts refusal behavior — and nothing in this
// return type can tell a grounded date from a confident guess.
return { text, refused: text.includes("I DON'T KNOW") };
}const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const ANSWERABILITY_GATE = 0.6; // is_answerable below -> abstain
const AUTO_CONFIDENCE = 0.85; // answer confidence below -> review
const QUESTIONS = {
answer: {
type: "choice",
instructions: "Answer using only what the state contains",
criteria: {
billing: "Payment, invoice or refund",
technical: "Product malfunction",
sales: "Buying or upgrade question",
},
},
is_answerable: {
type: "noul",
instructions: "Does the state contain enough information to pick exactly one option — not more data, not a guess?",
},
missing_information: {
type: "choice",
instructions: "If is_answerable is no: what single piece is missing?",
criteria: {
none: "The state is sufficient",
a_date: "A date or deadline",
an_amount: "An amount or balance",
an_identity: "Who a person or account is",
a_policy_rule: "Which policy or version applies",
external_state: "Live data from another system",
},
},
} as const;
export type Outcome =
| { abstained: true; missing: string; is_answerable: number }
| {
abstained: false;
answer: string;
confidence: number;
lane: "auto" | "review";
};
export async function decideOrAbstain(state: object): Promise<Outcome> {
const resp = await fetch(JEV_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ state, questions: QUESTIONS }),
});
if (!resp.ok) throw new Error("Jev evaluate failed");
const data = await resp.json();
const knows = data.is_answerable as { answer: boolean; confidence: number };
if (!knows.answer || knows.confidence < ANSWERABILITY_GATE) {
return {
abstained: true,
missing: (data.missing_information as { answer: string }).answer,
is_answerable: knows.confidence,
};
}
const answer = data.answer as { answer: string; confidence: number };
return {
abstained: false,
answer: answer.answer,
confidence: answer.confidence,
lane: answer.confidence >= AUTO_CONFIDENCE ? "auto" : "review",
};
}
// One call, three typed readings. The abstention branch carries the
// missing key, so the workflow can go collect exactly that piece.Decision model abstention FAQ
What is decision model abstention?
The model returning an explicit, typed "I don't know" — with the missing piece named — instead of guessing when the input state does not contain the answer. In the contract this page builds, abstention is a lane decided by an is_answerable Noul plus your threshold, and the abstention object carries a missing_information key so the workflow can go collect exactly what is absent. It is the evidence axis of decision quality; confidence thresholds are the certainty axis, and production systems gate on both.
Why not just add an "unknown" option to the Choice?
Three reasons. The option distribution degrades — every real answer now competes with a catch-all, so per-option probabilities stop meaning the same thing. The signal is unlabeled — after the fact you cannot tell "the model lacked evidence" from "the criteria are badly written." And there is no lever — a merged option cannot treat over-abstention and over-confidence differently, while a separate Noul gives each error direction its own measured rate and its own gate.
Doesn't a confidence threshold already handle this?
No — confidence and answerability fail differently. Confidence is calibrated: a 0.62 means the model saw the evidence and is 62% sure. Abstention covers the case where no amount of sureness is warranted: the ship date is not in the state, the churn reason was never recorded, the renewal has not happened yet. The audits and the evaluation cited on this page both show models stay confident on exactly those states — which is why the gate reads is_answerable first and confidence second.
Is abstention built into Jev natively?
Jev gives you the primitives — typed answers, calibrated confidence, the Noul question type, one-pass parallel questions — and the abstention lane is five lines of application code on top: ask is_answerable alongside your decision, gate, and route on missing_information. There is no "IDK mode" to switch on, which is honest in both directions: nothing abstains behind your back, and every abstention your system makes is one you wrote a rule for and can test.
How do I pick the answerability threshold?
The same way you pick a confidence threshold: from measured error costs. Build a labeled set with both answerable and provably-unanswerable states, sweep the is_answerable gate, and read two curves — false abstentions (human work wasted) and hallucinations (the expensive direction). The crossing point of those cost curves is your gate. Re-run the sweep after schema changes; the calibration guide shows the measurement methodology.
Do decision models actually abstain well?
It differs wildly by model, and it is measurable. The five-model evaluation cited on this page (October 2026, 18 messages, exploratory) found hosted Jev abstained on all six unanswerable cases while every small local model answered at least one — Clef Flash 9B only one of six. That is consistent with the broader overconfidence evidence on the limitations page. Test with unanswerable states drawn from your own domain before trusting any vendor number, ours included.