Jev recipe / compliance infrastructure
Decision Audit Trails with Jev: Every Answer Logged, Replayed, Explainable
Turn automated decisions into compliance-grade evidence: every Jev evaluate call returns a typed answer with calibrated confidence, so the audit log records the contract, the answer, and the gate — small enough to keep, deterministic enough to replay, explicit enough for GDPR Art. 22.
Build a decision audit trail in six steps
{
"decision": {
"type": "choice",
"instructions": "Decide the case under the published policy",
"criteria": {
"approve": "Meets every policy condition — no discretionary judgment required",
"refer": "Meets most conditions but a stated edge applies — hand to the queue with a reason",
"deny": "Fails a hard condition — deny with the failed criterion named"
}
},
"reversibility": {
"type": "score",
"instructions": "If this decision is wrong, how costly is it to reverse? 1 (trivially reversible, e.g. a refund) to 4 (irreversible harm to a person, e.g. account closure reported to a bureau)"
},
"needs_human_review": {
"type": "noul",
"instructions": "Given this case and this confidence, should a human confirm before the decision takes effect?"
}
}Six decisions, one audit log
Every decision below already ran through the evaluate endpoint. Move the two gates — the auto-execute confidence threshold and the human-handoff threshold — and watch the lanes recompute. High confidence is not enough on its own: the credit-limit call sits at 0.93 and still routes to review, because reversibility is the second axis.
| Decision | Answer & confidence | Reversibility | Lane |
|---|---|---|---|
Refund auto-approval · ticket #4812 dec_8f21 · audit-contract@1.3.0 · jev-1.13 · 96ms | approve0.97 | rev 1/4 | auto_executed |
Content publish gate · draft #2214 dec_8f22 · audit-contract@1.3.0 · jev-1.13 · 88ms | block0.88 | rev 2/4 | auto_executed |
Credit-limit increase · customer #7741 dec_8f23 · audit-contract@1.3.0 · jev-1.13 · 102ms | refer0.93 | rev 3/4Art. 22 record | review_queue |
Offer targeting · customer #2298 dec_8f24 · audit-contract@1.3.0 · jev-1.13 · 91ms | standard0.71 | rev 1/4 | review_queue |
Ticket escalation · ticket #5190 dec_8f25 · audit-contract@1.3.0 · jev-1.13 · 84ms | false0.55 | rev 2/4 | human_decision |
Account suspension · customer #3306 dec_8f26 · audit-contract@1.3.0 · jev-1.13 · 99ms | suspend0.78 | rev 4/4Art. 22 record | human_decision |
Demo data only — six simulated decisions. The lanes recompute from the same two-axis rule documented in the recipe copy.
Decision audit evidence: Jev typed log vs generative chat log vs no decision log
a fixed-size record, cheap enough to keep forever
the auditor reads the option definitions plus the per-option distribution
a logged 0.85 behaves like an 85% accuracy contract, so gate behavior is provable from the log alone
independent audits have caught decision models at 30%+ reported confidence on answers wrong far more often; rules carry no confidence at all
every historical case is re-runnable as a regression test
calibration loss has an early symptom
Production code: the audit pipeline in TypeScript and Python
import logging
import os
import requests
logger = logging.getLogger("decisions")
def decide_with_chat_model(customer_id: str, case: dict) -> str:
resp = requests.post(
"https://api.openai.com/v1/chat/completions",
headers={"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"},
json={
"model": "gpt-4o-mini",
"messages": [{
"role": "user",
"content": f"Decide this case under policy v3. Reply with the decision and a one-paragraph rationale.\nCase: {case}",
}],
},
timeout=30,
)
decision_text = resp.json()["choices"][0]["message"]["content"]
# What the audit log gets: a paragraph. It cannot be thresholded,
# replayed, or explained — only re-read. Which prompt version made
# this call? Was "likely approve" a 0.8 of confidence? The log
# cannot say, because none of it was recorded.
logger.info("decision for %s: %s", customer_id, decision_text)
return decision_textimport hashlib
import json
import requests
JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_EXECUTE_CONFIDENCE = 0.85 # >= gate AND reversibility <= 2 -> auto
HUMAN_CONFIDENCE = 0.60 # below this, or reversibility 4 -> human
QUESTIONS = {
"decision": {
"type": "choice",
"instructions": "Decide the case under the published policy",
"criteria": {
"approve": "Meets every policy condition — no discretionary judgment",
"refer": "Meets most conditions but a stated edge applies",
"deny": "Fails a hard condition — name the failed criterion",
},
},
"reversibility": {
"type": "score",
"instructions": "If wrong, how costly to reverse? 1 = trivial, 4 = irreversible harm to a person",
},
"needs_human_review": {
"type": "noul",
"instructions": "Should a human confirm before this decision takes effect?",
},
}
def decide_and_record(subject_ref: str, state: dict) -> dict:
payload = {"state": state, "questions": QUESTIONS}
resp = requests.post(JEV_ENDPOINT, json=payload, timeout=5)
resp.raise_for_status()
data = resp.json()
decision = data["decision"]
reversibility = data["reversibility"]["answer"]
if data["needs_human_review"]["answer"] or reversibility == 4:
lane = "human_decision"
elif decision["confidence"] >= AUTO_EXECUTE_CONFIDENCE and reversibility <= 2:
lane = "auto_executed"
else:
lane = "review_queue"
# The record is the audit trail: contract, answer, gate — one
# fixed-size row, replayable under any future schema version.
return {
"subject_ref": subject_ref,
"decision": decision["answer"],
"confidence": decision["confidence"],
"reversibility": reversibility,
"lane": lane,
"logic": {
"schema_version": "audit-contract@1.3.0",
"questions_hash": hashlib.sha256(
json.dumps(payload["questions"], sort_keys=True).encode()
).hexdigest()[:16],
"model_version": "jev-1.13",
},
"human_intervention_path": (
"available_on_request"
if lane == "auto_executed"
else "mandatory_before_effect"
),
}import { z } from "zod";
import OpenAI from "openai";
const Decision = z.object({ decision: z.enum(["approve", "refer", "deny"]) });
const openai = new OpenAI();
export async function decideCase(caseData: unknown) {
const resp = await openai.chat.completions.create({
model: "gpt-4o-mini",
response_format: { type: "json_object" },
messages: [
{ role: "user", content: `Decide under policy v3: ${JSON.stringify(caseData)}` },
],
});
const { decision } = Decision.parse(
JSON.parse(resp.choices[0].message.content!),
);
// The shape is guaranteed; the evidence is not. No confidence, no
// schema version, no distribution — an audit six months from now
// gets the what but never the why or the how-often-right.
return { decision, auditRecord: { decision } };
}const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const AUTO_EXECUTE_CONFIDENCE = 0.85; // gate AND reversibility <= 2
const HUMAN_CONFIDENCE = 0.6; // below, or reversibility 4 -> human
type Lane = "auto_executed" | "review_queue" | "human_decision";
const QUESTIONS = {
decision: {
type: "choice",
instructions: "Decide the case under the published policy",
criteria: {
approve: "Meets every policy condition",
refer: "A stated edge applies — queue with a reason",
deny: "Fails a hard condition — name it",
},
},
reversibility: {
type: "score",
instructions: "If wrong, how costly to reverse? 1 = trivial, 4 = irreversible harm",
},
needs_human_review: {
type: "noul",
instructions: "Should a human confirm before this takes effect?",
},
} as const;
export async function decideAndRecord(subjectRef: string, state: object) {
const resp = await fetch(JEV_ENDPOINT, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ state, questions: QUESTIONS }),
});
if (!resp.ok) throw new Error("Jev evaluate failed");
const data = await resp.json();
const { answer, confidence } = data.decision as {
answer: string;
confidence: number;
};
const reversibility = data.reversibility.answer as number;
const lane: Lane =
data.needs_human_review.answer || reversibility === 4
? "human_decision"
: confidence >= AUTO_EXECUTE_CONFIDENCE && reversibility <= 2
? "auto_executed"
: "review_queue";
return {
subject_ref: subjectRef,
decision: answer,
confidence,
reversibility,
lane,
logic: { schema_version: "audit-contract@1.3.0", model_version: "jev-1.13" },
human_intervention_path:
lane === "auto_executed" ? "available_on_request" : "mandatory_before_effect",
};
}
// Art. 22 record = this object. Replay = re-post the stored state under
// a new schema version and diff the lanes.Decision audit trail FAQ
What is a decision audit trail?
A timestamped, queryable record of every automated decision your system makes: what was decided, with what confidence, under which schema and model versions, through which gate, and with what human-intervention path. With Jev the record comes almost for free from the evaluate response — the typed answer, the calibrated confidence, and the option distribution are already structured, so the log is a fixed-size row instead of a paragraph.
Does GDPR Article 22 require this?
Article 22 gives individuals rights around solely automated decisions with legal or similarly significant effects — including not being subject to them, and obtaining meaningful information about the logic involved when exceptions apply. Whether any specific decision falls under it is a legal determination for your team, not something a logging library decides. What this recipe provides is the evidence layer: if you treat a decision as significant, the Art. 22 record (subject, decision, confidence, logic versions, human path, retention) is one object; if you decide it is not significant, the log is what lets you defend that call consistently.
What should one decision record contain?
Six fields: a subject reference (pseudonymized where possible), the decision with its confidence, the logic block (schema version, questions hash, model version, the gate thresholds in force), the per-option distribution for Choice questions, the lane the gate assigned, and the human-intervention path with the reviewer action if one occurred. Everything except the subject and the reviewer action comes from the single evaluate response.
How is this different from ordinary application logs?
Application logs answer "what did the code do"; an audit trail answers "what did the system decide, how confident was it, and under which rules". The difference is structure: a typed decision record can be thresholded (how many auto-executions happened at confidence below 0.9?), replayed (what would the new schema have decided?), and aggregated (is the confidence distribution drifting?) — none of which works on free-text log lines.
Can I re-decide old cases after upgrading the schema?
Yes — that is replay, and it is deterministic. Because both the state and the versioned questions are stored, upgrading the contract means re-posting every logged state to the new schema and diffing the lanes. Teams use replay as a migration check (did the upgrade change any auto-executed decisions?), as a regression suite (logged states are historical labeled data), and as an explanation engine when someone asks what the rules in force at the time would say today.
Why does calibrated confidence matter for an audit?
Because a confidence number is only evidence if it tracks reality. Jev confidence is RLCD-calibrated, so a logged 0.85 gate corresponds to a measurable accuracy contract — you can compute, from your own log, how often 0.85 decisions were right. Uncalibrated self-reported confidence does not support that: independent audits of decision models have found 30%+ reported confidence on answers that are wrong far more often, which turns every logged threshold into decoration. Calibration is also what makes drift visible — when the distribution shifts, your review rate tells you before your users do.