Jev recipe / abstention & honesty

Decision Model Abstention with Jev: When the Right Answer Is "I Don't Know"

Some questions cannot be answered from the state in front of the model — and no confidence threshold fixes that. Pair a typed Choice with an is_answerable Noul and a missing_information Choice, return a structured "I don't know" that names the missing piece, and gate automation on answerability and confidence together.

AD

Build the abstention contract in six steps

01Name the two failure modes before writing any gate. "Not sure" is a calibration problem: the model saw the evidence and is not confident enough — confidence thresholds, calibrated on labeled data, are the fix. "Not answerable" is an evidence problem: the state simply does not contain the answer, and no threshold can save it — ask a model when order #9012 will ship, give it order lines but no inventory or carrier data, and it will produce a confident-looking date anyway. Third-party audits have caught decision models reporting 30%+ confidence on unknowable answers, and an independent five-model evaluation published the day before this page went up found the same pattern: every small local model answered at least one question no state could support. Abstention is the typed answer to the second failure mode: the contract itself says "this question has no answer here."
02Write the abstention contract: keep the answer clean, add two questions. The temptation is adding an "unknown" option to your existing Choice — do not. It pollutes the option distribution (every real answer now competes with a catch-all), the signal is unlabeled (you cannot tell "missing evidence" from "badly written criteria" after the fact), and it gives you no separate lever to tune. Instead keep the answer Choice limited to real outcomes and add is_answerable as a Noul ("Does the state contain enough information to pick exactly one option — not more data, not a guess?") plus missing_information as a Choice over the kinds of things that go absent: a date, an amount, an identity, a policy rule, live state from another system. One evaluate call returns all three: the answer it would give, whether the state supports answering at all, and what to collect when it does not.
03Gate on answerability first, confidence second. The two numbers answer different questions and must never merge into one threshold: confidence below your calibrated line routes to review — a calibration decision; is_answerable below its line produces an abstention — an evidence decision. The dangerous quadrant is high confidence with low answerability, the 0.88 ship-date on a carrier-less state, because a confidence-only gate auto-executes exactly that row. Abstain whenever is_answerable falls below the gate, regardless of how good the answer looks, and return a typed object: abstained true, the missing_information key, the answerability reading. The rule is five lines of code, and it is the difference between an "I don't know" and a hallucination with a confidence score.
04Make the abstention actionable. An abstention that dead-ends is just a polite error. Because the contract names the missing piece, the abstention routes work: a missing date sends the workflow to the order system, a missing policy rule to the compliance wiki, a missing identity to a clarifying email — the same collect-exactly-this pattern the confidence-gated fallback chain applies to uncertain answers, pointed at evidence instead of certainty. In practice teams treat missing_information as a routing key: hash it, count it, and the categories that dominate your abstention rate become your data-model TODO list, ranked by how often they block automation.
05Test the abstention lane with unanswerable states. Build a set of questions your state provably cannot answer — future outcomes, fields you never collect, live values from other systems — alongside your normal labeled set, and measure the two error directions separately: abstaining on answerable questions wastes human time, answering unanswerable ones hallucinates, and they trade off against each other. The independent evaluation this page cites ran 18 customer messages through five decision models: all five handled the twelve clear requests nearly perfectly, and abstention was the entire difference — hosted Jev abstained on all six unanswerable cases, while the small local models guessed through most of them (Clef Flash 9B managed one abstention in six). Tune the answerability threshold exactly like a confidence threshold: from measured error costs on your own states, not vibes.
06Log abstentions as first-class decisions. The moment "I don't know" changes what happens next — a human gets a task, a customer gets an email, a ticket enters a queue — it is a decision, and it belongs in the audit trail with the same fields as any other: the contract version, the answerability and confidence readings, the lane, and the missing_information key. Abstention records age particularly well: when the missing data ships next quarter, replaying this quarter's abstentions under the new contract shows exactly which cases the enrichment fixes. And a rising abstention rate after a schema change is not a regression — it is the contract telling you the new questions are reaching states that cannot support them.
schema / answer, answerability & missing-info contract
{
  "answer": {
    "type": "choice",
    "instructions": "Answer using only what the state contains",
    "criteria": {
      "billing": "Payment, invoice or refund",
      "technical": "Product malfunction",
      "sales": "Buying or upgrade question"
    }
  },
  "is_answerable": {
    "type": "noul",
    "instructions": "Does the state contain enough information to pick exactly one option — not more data, not a guess?"
  },
  "missing_information": {
    "type": "choice",
    "instructions": "If is_answerable is no: what single piece is missing?",
    "criteria": {
      "none": "The state is sufficient",
      "a_date": "A date or deadline",
      "an_amount": "An amount or balance",
      "an_identity": "Who a person or account is",
      "a_policy_rule": "Which policy or version applies",
      "external_state": "Live data from another system"
    }
  }
}
Send a question whose answer is not in the state and watch is_answerable fall while confidence stays high.

Six questions, one answerability gate

Every question below already ran through the evaluate endpoint with the abstention contract: a typed answer, its calibrated confidence, and an is_answerable Noul. Drag the answerability gate and watch the lanes recompute. The dangerous rows are the confident ones: order #9012 shows 0.88 confidence on a date the state cannot contain — confidence measures sureness, not whether the answer exists in the state.

answered · 3abstained · 31 · would have auto-executed on confidence alone
Question & stateAnswer & confidenceis_answerableLane

Is order #5521 eligible for a refund under policy v3?

q_101 · order lines + return window dates + policy event

eligible0.960.98answered

Which team owns this: billing / technical / sales?

q_102 · full ticket text with one clear topic

billing0.970.95answered

Did the customer change their address before this order?

q_103 · profile with two timestamped addresses

true0.940.91answered

When will order #9012 ship?

q_104 · order lines only — no inventory, no carrier SLA

2026-10-140.880.31abstainedwould have auto-executed on confidence alone

missing:inventory state + carrier SLA table

Why did account #4183 churn?

q_105 · usage metrics — no survey, no conversation

price_sensitivity0.710.22abstained

missing:any direct customer statement — metrics show what happened, not why

Will contract #77 renew next quarter?

q_106 · usage trend — renewal outcome is future information

true0.660.41abstained

missing:the outcome does not exist yet — only intent signals do

Demo data only — six simulated questions. The lane recompute uses the same answerability rule documented in the recipe copy.

Abstention behavior: typed contract vs prompted "I don't know" vs confidence-only gating

METRIC
Jev
Prompted "I don't know" / confidence-only gating
Where the "I don't know" lives
A Noul in the contract

is_answerable returns a probability on every call, whether or not you gate on it

A sentence the model may or may not produce; you parse prose and hope the refusal appears
What the abstention carries back
The missing piece, typed: missing_information names the date, amount, identity or policy rule to collect
A refusal with no structure

the follow-up question is whatever a human infers from the wording

Can you tune it?
The answerability threshold is a gate like any other: false-abstain and hallucination rates are measurable on labeled sets
Re-prompting ("only say I don't know when…") shifts refusals unpredictably and re-fails on the next batch
Confident hallucinations
Caught by the answerability axis even when confidence runs high

the two readings gate independently

Nothing catches them: a fluent wrong answer and a fluent right answer look identical without a second signal
What happens next
missing_information is a routing key

the workflow collects exactly that piece, and the abstention is logged as a decision

A dead-end refusal in a chat log: no field to route on, nothing to threshold, nothing to replay
Independent evidence
18/18 overall and 6/6 abstained on unanswerable cases in a five-model independent evaluation (Oct 2026), n=18 exploratory; RLCD-calibrated confidence with published ECE methodology
LLM abstention is miscalibrated in both directions

AbstentionBench exists because prompted refusals do not track answerability

Production code: the abstention contract in TypeScript and Python

python / baseline: prompting for "I don't know"
import os

import requests

resp = requests.post(
    "https://api.openai.com/v1/chat/completions",
    headers={"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"},
    json={
        "model": "gpt-4o-mini",
        "messages": [{
            "role": "user",
            "content": (
                "When will order #9012 ship? Use only the data below. "
                "If the data does not contain the answer, reply exactly: I DON'T KNOW\n"
                "Data: {'lines': [...], 'channel': 'web'}"
            ),
        }],
    },
    timeout=30,
)
text = resp.json()["choices"][0]["message"]["content"]

if "I DON'T KNOW" in text:
    ...

# The failure you cannot see: rephrase the question ("by which date..."),
# warm up the style, or lengthen the data dump, and the same state returns
# "Ships Oct 14" — fluent, unfalsifiable, and indistinguishable from a
# grounded answer. You gated on prose, and prose has no threshold.
python / jev: the abstention contract and gate
import requests

JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
ANSWERABILITY_GATE = 0.60  # is_answerable below this -> abstain
AUTO_CONFIDENCE = 0.85     # answer confidence below this -> review

QUESTIONS = {
    "answer": {
        "type": "choice",
        "instructions": "Answer using only what the state contains",
        "criteria": {
            "billing": "Payment, invoice or refund",
            "technical": "Product malfunction",
            "sales": "Buying or upgrade question",
        },
    },
    "is_answerable": {
        "type": "noul",
        "instructions": "Does the state contain enough information to pick exactly one option — not more data, not a guess?",
    },
    "missing_information": {
        "type": "choice",
        "instructions": "If is_answerable is no: what single piece is missing?",
        "criteria": {
            "none": "The state is sufficient",
            "a_date": "A date or deadline",
            "an_amount": "An amount or balance",
            "an_identity": "Who a person or account is",
            "a_policy_rule": "Which policy or version applies",
            "external_state": "Live data from another system",
        },
    },
}

def decide_or_abstain(state: dict) -> dict:
    resp = requests.post(
        JEV_ENDPOINT,
        json={"state": state, "questions": QUESTIONS},
        timeout=5,
    )
    resp.raise_for_status()
    data = resp.json()

    knows = data["is_answerable"]
    if not knows["answer"] or knows["confidence"] < ANSWERABILITY_GATE:
        # Typed abstention: the workflow routes on the missing key,
        # and the record below is loggable, countable, replayable.
        return {
            "abstained": True,
            "missing": data["missing_information"]["answer"],
            "is_answerable": knows["confidence"],
        }

    answer = data["answer"]
    lane = "auto" if answer["confidence"] >= AUTO_CONFIDENCE else "review"
    return {
        "abstained": False,
        "answer": answer["answer"],
        "confidence": answer["confidence"],
        "lane": lane,
    }
typescript / baseline: parsing prose refusals
import OpenAI from "openai";

const openai = new OpenAI();

export async function whenDoesItShip(orderData: object) {
  const resp = await openai.chat.completions.create({
    model: "gpt-4o-mini",
    messages: [
      {
        role: "user",
        content: `When will this order ship? Use only the data below. If the data does not contain the answer, reply exactly: I DON'T KNOW\nData: ${JSON.stringify(orderData)}`,
      },
    ],
  });
  const text = resp.choices[0].message.content ?? "";

  // You gated on prose. A paraphrased question, a warmer style, or a
  // longer data dump shifts refusal behavior — and nothing in this
  // return type can tell a grounded date from a confident guess.
  return { text, refused: text.includes("I DON'T KNOW") };
}
typescript / jev: typed abstention object
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const ANSWERABILITY_GATE = 0.6; // is_answerable below -> abstain
const AUTO_CONFIDENCE = 0.85; // answer confidence below -> review

const QUESTIONS = {
  answer: {
    type: "choice",
    instructions: "Answer using only what the state contains",
    criteria: {
      billing: "Payment, invoice or refund",
      technical: "Product malfunction",
      sales: "Buying or upgrade question",
    },
  },
  is_answerable: {
    type: "noul",
    instructions: "Does the state contain enough information to pick exactly one option — not more data, not a guess?",
  },
  missing_information: {
    type: "choice",
    instructions: "If is_answerable is no: what single piece is missing?",
    criteria: {
      none: "The state is sufficient",
      a_date: "A date or deadline",
      an_amount: "An amount or balance",
      an_identity: "Who a person or account is",
      a_policy_rule: "Which policy or version applies",
      external_state: "Live data from another system",
    },
  },
} as const;

export type Outcome =
  | { abstained: true; missing: string; is_answerable: number }
  | {
      abstained: false;
      answer: string;
      confidence: number;
      lane: "auto" | "review";
    };

export async function decideOrAbstain(state: object): Promise<Outcome> {
  const resp = await fetch(JEV_ENDPOINT, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ state, questions: QUESTIONS }),
  });
  if (!resp.ok) throw new Error("Jev evaluate failed");
  const data = await resp.json();

  const knows = data.is_answerable as { answer: boolean; confidence: number };
  if (!knows.answer || knows.confidence < ANSWERABILITY_GATE) {
    return {
      abstained: true,
      missing: (data.missing_information as { answer: string }).answer,
      is_answerable: knows.confidence,
    };
  }

  const answer = data.answer as { answer: string; confidence: number };
  return {
    abstained: false,
    answer: answer.answer,
    confidence: answer.confidence,
    lane: answer.confidence >= AUTO_CONFIDENCE ? "auto" : "review",
  };
}
// One call, three typed readings. The abstention branch carries the
// missing key, so the workflow can go collect exactly that piece.

Decision model abstention FAQ

What is decision model abstention?

The model returning an explicit, typed "I don't know" — with the missing piece named — instead of guessing when the input state does not contain the answer. In the contract this page builds, abstention is a lane decided by an is_answerable Noul plus your threshold, and the abstention object carries a missing_information key so the workflow can go collect exactly what is absent. It is the evidence axis of decision quality; confidence thresholds are the certainty axis, and production systems gate on both.

Why not just add an "unknown" option to the Choice?

Three reasons. The option distribution degrades — every real answer now competes with a catch-all, so per-option probabilities stop meaning the same thing. The signal is unlabeled — after the fact you cannot tell "the model lacked evidence" from "the criteria are badly written." And there is no lever — a merged option cannot treat over-abstention and over-confidence differently, while a separate Noul gives each error direction its own measured rate and its own gate.

Doesn't a confidence threshold already handle this?

No — confidence and answerability fail differently. Confidence is calibrated: a 0.62 means the model saw the evidence and is 62% sure. Abstention covers the case where no amount of sureness is warranted: the ship date is not in the state, the churn reason was never recorded, the renewal has not happened yet. The audits and the evaluation cited on this page both show models stay confident on exactly those states — which is why the gate reads is_answerable first and confidence second.

Is abstention built into Jev natively?

Jev gives you the primitives — typed answers, calibrated confidence, the Noul question type, one-pass parallel questions — and the abstention lane is five lines of application code on top: ask is_answerable alongside your decision, gate, and route on missing_information. There is no "IDK mode" to switch on, which is honest in both directions: nothing abstains behind your back, and every abstention your system makes is one you wrote a rule for and can test.

How do I pick the answerability threshold?

The same way you pick a confidence threshold: from measured error costs. Build a labeled set with both answerable and provably-unanswerable states, sweep the is_answerable gate, and read two curves — false abstentions (human work wasted) and hallucinations (the expensive direction). The crossing point of those cost curves is your gate. Re-run the sweep after schema changes; the calibration guide shows the measurement methodology.

Do decision models actually abstain well?

It differs wildly by model, and it is measurable. The five-model evaluation cited on this page (October 2026, 18 messages, exploratory) found hosted Jev abstained on all six unanswerable cases while every small local model answered at least one — Clef Flash 9B only one of six. That is consistent with the broader overconfidence evidence on the limitations page. Test with unanswerable states drawn from your own domain before trusting any vendor number, ours included.

Extend the trustworthy-decision stack