Jev recipe / compliance infrastructure

Decision Audit Trails with Jev: Every Answer Logged, Replayed, Explainable

Turn automated decisions into compliance-grade evidence: every Jev evaluate call returns a typed answer with calibrated confidence, so the audit log records the contract, the answer, and the gate — small enough to keep, deterministic enough to replay, explicit enough for GDPR Art. 22.

Build a decision audit trail in six steps

01Log the contract, not just the answer. A decision nobody can explain is a liability, so every record carries the schema version, a hash of the questions object, and the model version next to the answer. The questions object is your decision logic in machine-readable form — criteria included — which is exactly what an auditor or a data subject asks for when they ask why. Store it once per schema version and reference it, so a million decisions cost the same as one.
02Store the full typed answer, confidence included. The answer alone says what the model chose; the calibrated confidence says how much that choice is worth. Because Jev confidence is RLCD-calibrated, a logged 0.85 behaves like an 85% accuracy contract — the number means something years later, when a regulator or a customer asks how often decisions like this one were right. Log the per-option distribution when the question is a Choice: the runner-up options are the explanation nobody thought to ask for at build time.
03Gate on two axes: confidence and reversibility. A 0.97 on a refund approval and a 0.97 on a credit-limit increase are not the same decision — the first reverses with a click, the second changes someones financial footprint. Auto-execute only when confidence clears the gate AND reversibility is low; route the middle band to a review queue; and let high reversibility force a human sign-off even at high confidence. The two-axis rule is four lines of code, and it is the difference between a gate and a rubber stamp.
04Write Art. 22-grade records for significant automated decisions. GDPR Article 22 gives people the right not to be subject to solely automated decisions with legal or similarly significant effects — and to obtain an explanation of that decision. The record that honors it: subject reference, the decision with its confidence, the logic block (schema version, questions hash, gate thresholds in force), the human-intervention path, and retention. Whether a given decision counts as significant is a legal call your team makes — the log exists so the answer is provable either way.
05Replay instead of migrate. Because the state and the questions are both stored, a schema upgrade re-decides every historical case under the new contract — deterministically, in one forward pass per case. Replay is three tools in one: a migration check (did the new schema change any lane?), a regression suite (your logged states are your labeled set), and an explanation engine (this decision, under the rules in force that day, said this).
06Monitor the confidence distribution, not just accuracy. Track the auto-executed share, the review-queue rate, and the population drifting through the contested band. Calibration degrades quietly — third-party audits of decision models have caught models reporting 30%+ confidence on answers that are wrong far more often than that — and the earliest symptom is your own distribution moving. A rising review rate is not a bug; it is the sensor working. The same drift discipline the confidence-gated fallback chain applies to live traffic applies here to history.
schema / decision, reversibility & human-review contract
{
  "decision": {
    "type": "choice",
    "instructions": "Decide the case under the published policy",
    "criteria": {
      "approve": "Meets every policy condition — no discretionary judgment required",
      "refer": "Meets most conditions but a stated edge applies — hand to the queue with a reason",
      "deny": "Fails a hard condition — deny with the failed criterion named"
    }
  },
  "reversibility": {
    "type": "score",
    "instructions": "If this decision is wrong, how costly is it to reverse? 1 (trivially reversible, e.g. a refund) to 4 (irreversible harm to a person, e.g. account closure reported to a bureau)"
  },
  "needs_human_review": {
    "type": "noul",
    "instructions": "Given this case and this confidence, should a human confirm before the decision takes effect?"
  }
}
Send a real case as state and inspect decision, reversibility, and needs_human_review with calibrated confidence.

Six decisions, one audit log

Every decision below already ran through the evaluate endpoint. Move the two gates — the auto-execute confidence threshold and the human-handoff threshold — and watch the lanes recompute. High confidence is not enough on its own: the credit-limit call sits at 0.93 and still routes to review, because reversibility is the second axis.

auto_executed · 2review_queue · 2human_decision · 2
DecisionAnswer & confidenceReversibilityLane

Refund auto-approval · ticket #4812

dec_8f21 · audit-contract@1.3.0 · jev-1.13 · 96ms

approve0.97rev 1/4auto_executed

Content publish gate · draft #2214

dec_8f22 · audit-contract@1.3.0 · jev-1.13 · 88ms

block0.88rev 2/4auto_executed

Credit-limit increase · customer #7741

dec_8f23 · audit-contract@1.3.0 · jev-1.13 · 102ms

refer0.93rev 3/4Art. 22 recordreview_queue

Offer targeting · customer #2298

dec_8f24 · audit-contract@1.3.0 · jev-1.13 · 91ms

standard0.71rev 1/4review_queue

Ticket escalation · ticket #5190

dec_8f25 · audit-contract@1.3.0 · jev-1.13 · 84ms

false0.55rev 2/4human_decision

Account suspension · customer #3306

dec_8f26 · audit-contract@1.3.0 · jev-1.13 · 99ms

suspend0.78rev 4/4Art. 22 recordhuman_decision

Demo data only — six simulated decisions. The lanes recompute from the same two-axis rule documented in the recipe copy.

Decision audit evidence: Jev typed log vs generative chat log vs no decision log

METRIC
Jev
Chat log / rules audit
What the log contains
Typed answer + calibrated confidence + schema version + questions hash

a fixed-size record, cheap enough to keep forever

A free-text paragraph (or parsed JSON with no confidence); the prompt version that produced it is usually lost
Explaining a decision
The criteria block IS the explanation

the auditor reads the option definitions plus the per-option distribution

Post-hoc rationalization of generated text; ask the same question twice and the explanation changes
Confidence as evidence
RLCD-calibrated

a logged 0.85 behaves like an 85% accuracy contract, so gate behavior is provable from the log alone

Self-reported and uncalibrated

independent audits have caught decision models at 30%+ reported confidence on answers wrong far more often; rules carry no confidence at all

Replay after upgrades
Logged state + versioned questions re-decide deterministically

every historical case is re-runnable as a regression test

Prompts drift and models deprecate; the original decision is not reproducible, so changes ship blind
Drift detection
Confidence distribution, review rate, and contested-band traffic are measurable time series

calibration loss has an early symptom

Text outputs have no comparable signal; drift surfaces as complaints or incidents
Human-handoff evidence
The escalation lane, the thresholds in force, and the reviewer action all live in the same record
Approvals scattered across tickets and chat threads, unlinked to the decision that triggered them

Production code: the audit pipeline in TypeScript and Python

python / baseline: the chat log you cannot explain
import logging
import os

import requests

logger = logging.getLogger("decisions")

def decide_with_chat_model(customer_id: str, case: dict) -> str:
    resp = requests.post(
        "https://api.openai.com/v1/chat/completions",
        headers={"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"},
        json={
            "model": "gpt-4o-mini",
            "messages": [{
                "role": "user",
                "content": f"Decide this case under policy v3. Reply with the decision and a one-paragraph rationale.\nCase: {case}",
            }],
        },
        timeout=30,
    )
    decision_text = resp.json()["choices"][0]["message"]["content"]

    # What the audit log gets: a paragraph. It cannot be thresholded,
    # replayed, or explained — only re-read. Which prompt version made
    # this call? Was "likely approve" a 0.8 of confidence? The log
    # cannot say, because none of it was recorded.
    logger.info("decision for %s: %s", customer_id, decision_text)
    return decision_text
python / jev: one evaluate call, one audit record
import hashlib
import json

import requests

JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate"
AUTO_EXECUTE_CONFIDENCE = 0.85  # >= gate AND reversibility <= 2 -> auto
HUMAN_CONFIDENCE = 0.60         # below this, or reversibility 4 -> human

QUESTIONS = {
    "decision": {
        "type": "choice",
        "instructions": "Decide the case under the published policy",
        "criteria": {
            "approve": "Meets every policy condition — no discretionary judgment",
            "refer": "Meets most conditions but a stated edge applies",
            "deny": "Fails a hard condition — name the failed criterion",
        },
    },
    "reversibility": {
        "type": "score",
        "instructions": "If wrong, how costly to reverse? 1 = trivial, 4 = irreversible harm to a person",
    },
    "needs_human_review": {
        "type": "noul",
        "instructions": "Should a human confirm before this decision takes effect?",
    },
}

def decide_and_record(subject_ref: str, state: dict) -> dict:
    payload = {"state": state, "questions": QUESTIONS}
    resp = requests.post(JEV_ENDPOINT, json=payload, timeout=5)
    resp.raise_for_status()
    data = resp.json()

    decision = data["decision"]
    reversibility = data["reversibility"]["answer"]
    if data["needs_human_review"]["answer"] or reversibility == 4:
        lane = "human_decision"
    elif decision["confidence"] >= AUTO_EXECUTE_CONFIDENCE and reversibility <= 2:
        lane = "auto_executed"
    else:
        lane = "review_queue"

    # The record is the audit trail: contract, answer, gate — one
    # fixed-size row, replayable under any future schema version.
    return {
        "subject_ref": subject_ref,
        "decision": decision["answer"],
        "confidence": decision["confidence"],
        "reversibility": reversibility,
        "lane": lane,
        "logic": {
            "schema_version": "audit-contract@1.3.0",
            "questions_hash": hashlib.sha256(
                json.dumps(payload["questions"], sort_keys=True).encode()
            ).hexdigest()[:16],
            "model_version": "jev-1.13",
        },
        "human_intervention_path": (
            "available_on_request"
            if lane == "auto_executed"
            else "mandatory_before_effect"
        ),
    }
typescript / baseline: structured output, unstructured evidence
import { z } from "zod";
import OpenAI from "openai";

const Decision = z.object({ decision: z.enum(["approve", "refer", "deny"]) });
const openai = new OpenAI();

export async function decideCase(caseData: unknown) {
  const resp = await openai.chat.completions.create({
    model: "gpt-4o-mini",
    response_format: { type: "json_object" },
    messages: [
      { role: "user", content: `Decide under policy v3: ${JSON.stringify(caseData)}` },
    ],
  });
  const { decision } = Decision.parse(
    JSON.parse(resp.choices[0].message.content!),
  );

  // The shape is guaranteed; the evidence is not. No confidence, no
  // schema version, no distribution — an audit six months from now
  // gets the what but never the why or the how-often-right.
  return { decision, auditRecord: { decision } };
}
typescript / jev: the two-axis gate, typed end to end
const JEV_ENDPOINT = "https://api.typesafe.ai/v1/jev/evaluate";
const AUTO_EXECUTE_CONFIDENCE = 0.85; // gate AND reversibility <= 2
const HUMAN_CONFIDENCE = 0.6; // below, or reversibility 4 -> human

type Lane = "auto_executed" | "review_queue" | "human_decision";

const QUESTIONS = {
  decision: {
    type: "choice",
    instructions: "Decide the case under the published policy",
    criteria: {
      approve: "Meets every policy condition",
      refer: "A stated edge applies — queue with a reason",
      deny: "Fails a hard condition — name it",
    },
  },
  reversibility: {
    type: "score",
    instructions: "If wrong, how costly to reverse? 1 = trivial, 4 = irreversible harm",
  },
  needs_human_review: {
    type: "noul",
    instructions: "Should a human confirm before this takes effect?",
  },
} as const;

export async function decideAndRecord(subjectRef: string, state: object) {
  const resp = await fetch(JEV_ENDPOINT, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({ state, questions: QUESTIONS }),
  });
  if (!resp.ok) throw new Error("Jev evaluate failed");
  const data = await resp.json();

  const { answer, confidence } = data.decision as {
    answer: string;
    confidence: number;
  };
  const reversibility = data.reversibility.answer as number;
  const lane: Lane =
    data.needs_human_review.answer || reversibility === 4
      ? "human_decision"
      : confidence >= AUTO_EXECUTE_CONFIDENCE && reversibility <= 2
        ? "auto_executed"
        : "review_queue";

  return {
    subject_ref: subjectRef,
    decision: answer,
    confidence,
    reversibility,
    lane,
    logic: { schema_version: "audit-contract@1.3.0", model_version: "jev-1.13" },
    human_intervention_path:
      lane === "auto_executed" ? "available_on_request" : "mandatory_before_effect",
  };
}
// Art. 22 record = this object. Replay = re-post the stored state under
// a new schema version and diff the lanes.

Decision audit trail FAQ

What is a decision audit trail?

A timestamped, queryable record of every automated decision your system makes: what was decided, with what confidence, under which schema and model versions, through which gate, and with what human-intervention path. With Jev the record comes almost for free from the evaluate response — the typed answer, the calibrated confidence, and the option distribution are already structured, so the log is a fixed-size row instead of a paragraph.

Does GDPR Article 22 require this?

Article 22 gives individuals rights around solely automated decisions with legal or similarly significant effects — including not being subject to them, and obtaining meaningful information about the logic involved when exceptions apply. Whether any specific decision falls under it is a legal determination for your team, not something a logging library decides. What this recipe provides is the evidence layer: if you treat a decision as significant, the Art. 22 record (subject, decision, confidence, logic versions, human path, retention) is one object; if you decide it is not significant, the log is what lets you defend that call consistently.

What should one decision record contain?

Six fields: a subject reference (pseudonymized where possible), the decision with its confidence, the logic block (schema version, questions hash, model version, the gate thresholds in force), the per-option distribution for Choice questions, the lane the gate assigned, and the human-intervention path with the reviewer action if one occurred. Everything except the subject and the reviewer action comes from the single evaluate response.

How is this different from ordinary application logs?

Application logs answer "what did the code do"; an audit trail answers "what did the system decide, how confident was it, and under which rules". The difference is structure: a typed decision record can be thresholded (how many auto-executions happened at confidence below 0.9?), replayed (what would the new schema have decided?), and aggregated (is the confidence distribution drifting?) — none of which works on free-text log lines.

Can I re-decide old cases after upgrading the schema?

Yes — that is replay, and it is deterministic. Because both the state and the versioned questions are stored, upgrading the contract means re-posting every logged state to the new schema and diffing the lanes. Teams use replay as a migration check (did the upgrade change any auto-executed decisions?), as a regression suite (logged states are historical labeled data), and as an explanation engine when someone asks what the rules in force at the time would say today.

Why does calibrated confidence matter for an audit?

Because a confidence number is only evidence if it tracks reality. Jev confidence is RLCD-calibrated, so a logged 0.85 gate corresponds to a measurable accuracy contract — you can compute, from your own log, how often 0.85 decisions were right. Uncalibrated self-reported confidence does not support that: independent audits of decision models have found 30%+ reported confidence on answers that are wrong far more often, which turns every logged threshold into decoration. Calibration is also what makes drift visible — when the distribution shifts, your review rate tells you before your users do.

Extend the compliance and gating stack