What is decision model calibration?
Calibration is the statistical alignment between a model’s stated confidence and its actual accuracy: a model is calibrated if, across all predictions where it reports 70% confidence, it is correct about 70% of the time. For decision models — which return a typed option plus a probability — calibration is what makes thresholds meaningful. “Automate above 0.85” is a claim about your workload’s future that only holds if 0.85 has actually corresponded to ~85% correctness on your labels.
How is ECE (Expected Calibration Error) computed?
Bucket all predictions into confidence bins (ten is conventional), compute the gap between observed accuracy and mean stated confidence within each bin, then average those gaps weighted by the share of predictions in each bin. A model that is right whenever it says 0.9 and wrong whenever it says 0.4 has an ECE near zero even at mediocre accuracy — ECE measures honesty, not skill. Always read the reliability diagram alongside it: two models with the same ECE can fail in opposite directions, one systematically overconfident, one underconfident.
Does high confidence mean an answer is trustworthy?
Not by itself. The Synthpop audit is the cleanest counterexample: on unanswerable questions where the model was 1% accurate, it still reported 32–36% confidence, and after its knowledge cutoff accuracy fell to a coin flip while confidence rose to 82%. High confidence is a claim, not evidence. It becomes evidence only after you have measured, on your own labeled workload, how often that confidence level has actually been correct — which is what the five-step workflow on this page produces.
How do I calibrate my own decision model?
Freeze a held-out labeled set from your workload, run it through the model, log all probabilities, draw the reliability diagram, compute ECE, then set auto/review/human thresholds only where the curve is honest — and re-verify quarterly, or whenever the model version or traffic mix changes. If the diagram shows a systematic temperature-like distortion, a scaling fit on the calibration set can help. If it shows confident answers on unknowable questions — the audit failure above — no post-hoc recalibration repairs it; only routing those questions to a different process does.
Where does Jev’s confidence come from?
Not from the model commenting on itself. A Jev call returns a typed answer drawn from your declared options plus the parallel scores over all options from a single pass — confidence is a structural output of the decision contract, not generated prose. That makes it inspectable and reproducible under a frozen schema, which is why our benchmarks can publish ECE per fixture. It does not exempt the number from audit: the published Synthpop calibration audit ran against this exact architecture and found confident failures at the knowledge boundary.
What is the difference between calibration and accuracy?
Accuracy is how often the model is right. Calibration is whether its confidence tells you the truth about when it is right. A model can be 95% accurate with terrible calibration (every error comes back stamped 0.99), or 70% accurate with near-perfect calibration (it reliably knows when it is guessing). Automation needs the second property at least as much as the first: thresholds, fallback lanes and audit evidence are all built on the confidence number, not the accuracy number.