Every vendor publishes an accuracy story; fewer publish the honesty of their confidence. Two independent evaluations from October 2026 cut both ways - we cite both sides.
The endorsement: vals.ai's preregistered evaluation (2026-10-06; twelve systems, 400 human-audited claim-verification items built from SEC filings) scored hosted Jev at 0.975 - statistically tied with GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna - at $0.02 per 1,000 cases, about 1/498th of Astra's cost, with the lowest ECE of the twelve (0.011) and ~0.1s median latency (194x lower than Astra at 32 judgments per call). TypeSafe's launch headline of "193.6x faster, 444.6x cheaper," the evaluators concluded, holds up.
The criticism, from the same evaluation: on a preregistered 12-subtask slice of LegalBench, Jev finished last of twelve on class-balanced accuracy (0.730 vs GPT-5.6 Terra's 0.932, with every other system's lead statistically significant), it collapsed on one contract-NLI subtask (answering yes to 30 of 33 items), and its ECE rose to 0.108 - the worst of the twelve. Calibration measured on one task did not transfer to the other.
Neither finding generalizes automatically - that is the lesson this category keeps re-proving. The 575,000-call Synthpop audit, Red Hat's decision-model-vs-classifier benchmark, and DoubtBench each test a different property of decision-model confidence, and all of them are collected on our calibration page. The rule that survives every audit: a confidence number is a hypothesis until you have measured it on your own labels.