Guides / 動画ウォークスルー

AutoTrust JEV-27B Tested Locally: Four Arms, 72 Cases, and One 0.969 Score That Was Wrong

A third-party field test of AutoTrust JEV-27B turned into a step-by-step page: four scoring arms (rules, decision head, ordinary generation, reasoning) on 72 held-out policy cases, a DGX Spark BF16 latency reality check, and the sections most launch posts would cut — the testers own 5-case wording flaw, six confidence thresholds that all failed, and a 0.969-score mistake.

要点

This is RUNTIME.'s independent local evaluation of AutoTrust JEV-27B (Hugging Face autotrust/JEV-27B, Apache-2.0) — a 10:34 field test that treats the model's official claims as things to verify rather than repeat. First, identity, because the 27B Jev-family space now has two players: AutoTrust AI Lab's JEV-27B is a finetune of Alibaba's Qwen3.8-27B that distills decision behavior from TypeSafe's hosted Jev (the HF card calls it a student of TypeSafe Jev 1.13) by freezing the backbone and training roughly 109 million parameters — a decision adapter plus a scoring head; the local release bundles those additions, so it never calls the teacher. It is not the same thing as ZefanCai/Open-Jev-27B-v1.1, a separate community release shipped as a PEFT adapter package (adapter plus head.pt) on the same Qwen3.8-27B base — different publisher, different artifact — and AutoTrust's own JEV-27B-VL is the vision variant of this model, not the one tested here. The official claims on the slide: about 137 ms per decision on a B200, roughly half the hosted TypeSafe Jev API's published 238-301 ms (which includes network time), with coding ability unchanged — 78.0% versus 78.0% on 128 of 164 tasks. The test rig: a DGX Spark (GB10) running the pinned revision in 16-bit BF16 — explicitly not Q4 or Q8 — with two accelerated components unavailable and fallback implementations in their place, so none of this is an optimized throughput benchmark. Four arms see identical inputs (a written policy, the customer message, the same options list — Billing/technical, Security/ask): the rules control is a deliberately small deterministic parser; the decision head is the trained typed-choice path; the ordinary answer runs the same backbone with the adapter off, thinking off and a strict structured format; the reasoning path generates thinking first, then scores. Protocol: 24 development cases exist to test one thing — whether a confidence score could gate when to escalate to reasoning — then 72 held-out cases are scored, and repeats (reordered options, irrelevant details, misleading instructions) check consistency only, never extra independent problems; no LLM judges the answers, the cases were written and checked in advance, and a high score never overrides a wrong action. Case 1 shows where interpretation earns its place: a duplicate charge from yesterday that was refunded, plus an export failure today, routes to technical support — the rules control matched the word 'charge' and picked billing, while quick head, ordinary generation and reasoning all read the change in state correctly. Case 2 shows where ordered conditions bite: retry is authorized, severity is normal, three of four attempts are used, but the diagnostic is 14 minutes old against a 13-minute freshness limit — the required action is collect; rules and reasoning get it, quick head and ordinary generation both retry, and the reasoning answer cost 85.94 s against the quick answer's 0.66 s in that setup (a related critical-severity variant requires escalate first; generation still retries). The 72-case scoreboard: ordinary generation 50, quick head 63, reasoning 71, rules 70 — and moving from quick head to reasoning fixed eight answers without flipping a single correct one wrong in this run, while the tiny rules control finished one point behind reasoning. Then the honesty sections, which are the reason to watch. First, the testers report a flaw in their own test: five routing cases contain a policy exception that sends a billing issue to Security for a human audit while the Security option itself describes unauthorized account activity — the expected route follows the explicit exception, but the wording pulls against it. They keep the original scores, then re-run without the five: quick head 63/67, reasoning 67/67, rules 65/67, generation 50/67 — a sensitivity check, not a new benchmark, on a small synthetic set with an ordinary-generation control that never got a prompt search. Second, the confidence shortcut fails: six thresholds from 0.50 to 0.95 were tried on the 24 development cases and none met the accuracy requirement, so the video contains no working hybrid gate and no held-out result for one. The clean proof is a single development case where the quick head chose retry — the required action was collect — with a score of 0.969: even the highest candidate threshold, 0.950, would have accepted it without reasoning. The score describes the model's preference among the supplied choices; reading 0.97 as a 97% chance of correctness needs evidence that predicted confidence matches observed accuracy across many representative examples, and picking a new threshold after inspecting answers is just a new experiment (new design, separate development examples, freeze, fresh held-out cases). Third, the runtime reality: median end-to-end request time was 0.65 s for the quick head (p95 0.72 s), 2.34 s for ordinary generation (p95 2.60 s) and 102.20 s for reasoning (p95 128.50 s), one request at a time; reasoning hit its 512-token limit on 20 of the 72 cases and was forced to read out an answer — a correct final choice does not mean the reasoning completed; the model files weigh about 50 GiB with the highest sampled GPU process around 52 GiB, and because the Spark uses unified memory those two overlap rather than add. The fast decision head changes how you get the answer; it does not turn a 27-billion-parameter model into a small download. Finally, 96 presentation variants built from 24 parent cases — reversed options, permutations, irrelevant detail, misleading instructions planted in the customer text where they are data, not policy — left the semantic choice unchanged 91/96 times for the quick head, 95/96 for reasoning, 91/96 for generation and 96/96 for rules, which is consistency, not correctness: a system can repeat the same wrong answer every time, and a few misleading messages are not a security guarantee. The video's closing advice is the operational summary: the quick head is interesting when an agent picks from a known set and the wording needs interpretation; keep clear conditions in code so the model interprets the message while the program checks diagnostic age and permission, act only when they agree, and treat disagreements as concrete things to inspect — then try another backend or a better test, and share prompts, settings and results.

動画ソース

RUNTIME.

10:34hYSeZrYUDxA

ステップごとのウォークスルー

  1. 1

    Meet the model: JEV-27B from AutoTrust AI Lab

    The opening slide names the exact pair under test: JEV-27B, an independent release from AutoTrust AI Lab, running on Alibaba's Qwen3.8-27B backbone. The architecture, per the video and the Hugging Face card (autotrust/JEV-27B, Apache-2.0): the Qwen backbone stays frozen while roughly 109 million parameters are trained on top — a decision adapter plus a scoring head — learned by distilling decision behavior from TypeSafe's hosted Jev (the card calls the model a student of TypeSafe Jev 1.13). The local release bundles those trained additions into the weights, so it makes decisions without calling the teacher. One naming caution before you install anything: this is not the only open 27B Jev-family model. ZefanCai/Open-Jev-27B-v1.1 is a separate community release — shipped as a PEFT adapter package (adapter plus head.pt) on the same Qwen3.8-27B base — from a different publisher entirely, and AutoTrust also publishes JEV-27B-VL, the vision variant of this model. Both orgs' repos are live on the Hub; this page follows the video's subject, AutoTrust's integrated JEV-27B.

    Opening slide introducing AutoTrust JEV-27B with a model card reading JEV-27B by AutoTrust AI Lab beside a backbone card reading Qwen3.8-27B, Alibaba language model
    The opener pins down the exact subject: AutoTrust's JEV-27B sitting on its frozen Qwen3.8-27B backbone.タイムスタンプ 0:08 を見る
  2. 2

    The claims under test: 137 ms and half of hosted Jev

    Before running anything, the video puts the official numbers on screen — and labels them honestly: 'their speed claim — not our Spark result'. AutoTrust reports a median single-decision latency of 137 ms measured on an NVIDIA B200; hosted TypeSafe Jev's published measurements run 238-301 ms per decision, and that figure includes network time. So the claim is roughly half of hosted Jev. Two more claims round out the card: coding ability is unchanged — 78.0% for the original Qwen versus 78.0% for JEV-27B generation across 128 of 164 coding tasks — and the decision behavior was learned from TypeSafe's Jev, with the local model carrying the trained additions so it never needs to call the teacher. Everything after this step is the channel finding out what those claims look like on very different hardware.

    Claims slide contrasting a local B200 median single-decision latency of 137 ms against hosted TypeSafe Jev published measurements of 238 to 301 ms including network time
    Author-reported numbers, labeled as claims: the B200 figure is what a local re-test is for, not a result from this rig.タイムスタンプ 0:30 を見る
  3. 3

    Four controls, one scorecard: rules, quick head, ordinary answer, reasoning

    The protocol is the spine of the whole review, so it gets its own slide. Every arm sees identical inputs: a written policy (written rules, current facts, exceptions, the runbook, missing-facts-means-ask), the customer message, and the same options list — Billing/technical, Security/ask — with no correct label supplied. Then four runners: a small deterministic rules parser; the trained decision head emitting a typed choice; an ordinary answer from the same backbone with the adapter off and a strict structured format; and a reasoning path that thinks first, then scores the choices. The ground rules at the bottom matter as much as the arms: count mistakes, invalid answers and actual request time, and a high score does not override the outcome — if the selected action is wrong, it is wrong. On top of the four arms sits a dataset design: 24 development cases exist to test exactly one thing — whether a confidence score could gate when to escalate to reasoning — and the 72 held-out cases carry the scoring; repeats with reordered options and tweaked details check consistency only, and are never counted as extra independent problems. No language model judges the results — the cases were written and checked before the models ran.

    Four-arm test design listing the deterministic rules control, the decision head for typed choices, an ordinary answer on the same backbone and a reasoning path that thinks before scoring
    Same policy, message and options for every arm — and the scorecard never lets a confidence score override a wrong action.タイムスタンプ 2:06 を見る
  4. 4

    Case 1: read the change in state

    The first scored example is a language-understanding case. The customer message: yesterday there were two posted charges, but one was refunded; today the export fails; no login anomaly. The policy says to use current facts — a problem mentioned in a message is not necessarily a problem that still exists. Three facts to separate: the duplicate charge happened yesterday, the refund resolved it, the export failure is happening now — which makes technical support the required route. Results as recorded: the rules control picked billing, matching the word 'charge' but missing that the charge problem was already resolved; the quick decision head picked technical, and so did ordinary generation and the reasoning path. The video's framing is worth keeping: this is where a language model earns its place against a rigid text parser — but one example does not doom the rules approach, which is deliberately small, and a better parser could fix this particular case. Note also that extra reasoning changed nothing here: both model paths were correct, and if you are paying for longer decisions you want to know when that extra work actually changes something.

    Case timeline sequencing a duplicate charge, a refund that resolved it and today export failure, with the required route resolved to technical support
    Same words, different times: the refund closed the billing problem yesterday, so the live export failure owns the route.タイムスタンプ 2:45 を見る
  5. 5

    Case 2: a retry is authorized — but is it allowed yet?

    The second case tests whether a model applies conditions in the right order. The runbook facts: retry is authorized, severity is normal, three of four attempts are used — but the diagnostic must be no more than 13 minutes old, and this one is 14. Pause on the comparison: 14 is greater than 13, and permission to retry does not override the freshness requirement, so the required action is collect. The recorded choices split cleanly: rules and the reasoning path pick collect; the quick head and ordinary generation both pick retry — wrong. The price of correctness is the striking part: the quick answer took about 0.66 seconds, the reasoning answer about 85.94 seconds in this unoptimized setup — extra reasoning corrected a real mistake, with a substantial weight attached. A related variant raises severity to critical, where the runbook checks escalation before diagnostic age: the required action becomes escalate, the quick head and reasoning both get it right, and ordinary generation still retries. A useful test needs these interactions — not just recognizing a number, but applying the right condition in the right order.

    Test facts asking whether an authorized retry is allowed yet with severity normal, three of four attempts used, and a diagnostic age of 14 minutes against a 13 minute limit
    14 against 13 is the whole case: freshness outranks permission, and two of the four arms missed it.タイムスタンプ 3:52 を見る
  6. 6

    The 72-case scoreboard: reasoning 71, rules 70, quick head 63, generation 50

    Across the 72 original held-out cases, the four arms finish in this order: ordinary generation 50 correct, the quick decision head 63, the reasoning path 71, and the small rules control 70. Two readings the video stops to make explicit. First, upgrading from the quick head to reasoning fixed eight answers, and in this run not a single correct quick answer flipped to wrong — the paid-for latency bought strictly more correctness. Second, a deliberately small deterministic parser finished within one point of the reasoning path, which the video offers as a reminder to test the simple solution before assuming a model is needed. Remember the scoreboard's own subtitle: these are synthetic supplied-policy tasks, and a wording caveat follows — which is the next step.

    Bar chart of correct choices on the 72-case held-out test showing generation 50, quick head 63, reasoning 71 and rules 70 on a common zero to 72 scale
    Four arms, 72 held-out cases: the tiny rules control finishes one point behind the reasoning path.タイムスタンプ 5:00 を見る
  7. 7

    The test found its own flaw: five cases with colliding wording

    Here is the section almost no launch post would include, and the reason this video is worth a page. While reviewing results, the testers found a defect in their own test set: five routing cases contain a policy exception that sends a billing issue to Security for a human audit, while the Security option itself is described as escalating unauthorized account activity. Those two descriptions do not line up — the expected route follows the explicit exception, but the wording pulls the model the other way. The handling is the lesson: they kept the original cases and scores exactly as recorded, published the flaw next to them, and then ran a separate sensitivity check with the five cases set aside. The stated principle: a test we design can contain a flaw, just like a model can make a mistake — and both belong on screen.

    Wording collision between a policy exception sending a billing issue to Security for human audit and a Security option described as escalating unauthorized account activity
    Two Security meanings that do not agree — the testers published their own defect beside the scores instead of quietly dropping it.タイムスタンプ 5:15 を見る
  8. 8

    Setting the five aside: reasoning goes 67 for 67

    The sensitivity check re-scores the remaining 67 cases: quick head 63, reasoning a perfect 67 of 67, rules 65, ordinary generation 50. The video is careful about what this is and is not — a post-hoc diagnostic, not a replacement benchmark and not a new claim of perfect reliability. The caveats get their own airtime: 72 examples from a small set of related policy patterns say nothing about production workloads; and the ordinary-generation control has limits of its own — same backbone with the adapter disabled, thinking off, a strict structured answer, and no search for a better prompt or a separate reasoning strategy, so it measures configured paths rather than the ceiling of language models. The strong rules score stays in the frame too: it is the standing reminder that the simple solution deserves a test. What the check does establish is specific: the five flawed cases were genuinely ambiguous, and on unambiguous wording the reasoning path did not miss once in this run.

    Sensitivity bar chart recounting quick head 63 of 67, reasoning a perfect 67 of 67, rules 65 of 67 and generation 50 of 67 after five flawed cases are set aside
    Labeled post-hoc, and honestly: remove five ambiguous cases and the reasoning path is perfect on the remaining 67.タイムスタンプ 5:43 を見る
  9. 9

    The confidence shortcut fails: no threshold qualified

    The whole premise of a fast-head-plus-reasoning hybrid is a gate: take the quick answer, and escalate to reasoning only when its score falls below a threshold. Before the held-out test, the channel tried to build that gate on the 24 development cases, sweeping six candidate thresholds — 0.50, 0.60, 0.70, 0.80, 0.90, 0.95. None met the accuracy requirement. The slide states the consequence plainly: no approved hybrid result — and, importantly, the held-out test therefore contains no result for a successful hybrid system either. This is the negative result the video refuses to hide: the sensible-sounding idea ('use the fast answer most of the time, think only when needed') collapsed at the one step everyone skips — recognizing when the first answer needs help. The next step shows exactly why it collapsed.

    Gate verdict over six candidate thresholds from 0.50 to 0.95 on 24 development cases concluding no approved hybrid result and no working confidence shortcut
    Six gates tried, zero approved: the video never shows the hybrid working — because on this evidence, it does not.タイムスタンプ 6:44 を見る
  10. 10

    Exhibit A: a 0.969 score on the wrong answer

    Why did every threshold fail? One development case is the clean demonstration. A stale-diagnostic case much like Case 2: the quick head chose retry — the required action was collect — and assigned the choice a score of about 0.969. The highest candidate threshold, 0.950, would have accepted that answer without reasoning, so even the most conservative gate waves the mistake through. The video draws the right conclusion about what the number means: the score describes the model's preference among the supplied choices, and nobody has established that 0.97 means a 97% chance of being correct on this work — that reading requires evidence that predicted confidence matches observed accuracy across many representative examples, which is exactly the calibration work that has not been done here. And the escape routes are honest ones: checking diagnostic age directly in code, or changing the threshold — but the latter is a new experiment needing a new design, separate development examples, a frozen approach, and fresh held-out cases; choosing a setting after inspecting test answers proves nothing about new problems.

    Confidence error pairing a quick-head score of 0.969 for a retry that should have been collect with the 0.950 highest threshold that would have skipped reasoning
    The calibration poster child: 0.969 confidence on the wrong action, above every gate the testers tried.タイムスタンプ 6:52 を見る
  11. 11

    What it costs to run: BF16 weights and a 102-second reasoning path

    The runtime card closes the gap between claim and desk. Everything ran locally on a DGX Spark (GB10) in 16-bit BF16 — explicitly not a 4-bit or 8-bit version — at a pinned revision, with the full backbone staying loaded for every decision. Median end-to-end request time across the 72 original cases, one request at a time: 0.65 s for the quick head (p95 0.72 s), 2.34 s for ordinary generation (p95 2.60 s), and 102.20 s for reasoning (p95 128.50 s); the rules control ran directly in Python and is timed separately. Two honesty notes ride along: two accelerated components were unavailable, so fallback implementations were used and this is not an optimized throughput benchmark — a different backend could substantially change the gap; and reasoning hit its 512-token limit on 20 of the 72 cases, where an answer was forced and the cutoff recorded — a correct final choice does not mean the reasoning completed. Memory tells the deployment story: model files about 50 GiB, highest sampled GPU process about 52 GiB, and because the Spark uses unified memory the two overlap — do not add them. The fast decision head changes how you get the answer; it does not turn a 27-billion-parameter model into a small download.

    Latency table tabulating the quick head at 0.65 seconds median and 0.72 seconds p95, generation at 2.34 and 2.60 seconds, and reasoning at 102.20 and 128.50 seconds
    One request at a time on 16-bit weights: the quick head at 0.65 s is the product story; the 102 s reasoning path is the bill.タイムスタンプ 8:04 を見る
  12. 12

    96 perturbations later: consistency is not correctness

    The last experiment changes presentation without changing the required answer: reverse the options, permute them, add irrelevant detail, or insert a misleading instruction inside the customer message — where it should be treated as data, not as a new policy. That yields 96 variants built from 24 parent cases, and the same semantic choice survived 91 of 96 for the quick head, 95 of 96 for reasoning, 91 of 96 for ordinary generation — and all 96 for the rules control. The video refuses to let that last number comfort you: consistency and correctness are different, and a system can repeat the same wrong answer every time, which is why these checks sit beside the accuracy results instead of replacing them. The caveats stay attached too — a few misleading messages are not proof of security against every attack. The closing advice ties the whole test together: the quick head earns its place when an agent picks from a known set and the wording takes interpretation; keep clear conditions in code, so the model interprets the message while the program checks diagnostic age and permission; act only when both agree, treat disagreements as concrete things to inspect — and if you run your own backend or a better test, publish prompts, settings and results.

    Consistency chart tracking the same semantic choice across 96 prompt variants with quick head 91, reasoning 95, generation 91 and rules 96 of 96
    Rules never blinked across 96 rewordings — and the video still reminds you a repeated answer can be consistently wrong.タイムスタンプ 9:26 を見る

よくある質問(FAQ)

What is AutoTrust JEV-27B?

AutoTrust JEV-27B (huggingface.co/autotrust/JEV-27B, Apache-2.0) is an open decision model from AutoTrust AI Lab: a frozen Alibaba Qwen3.8-27B backbone carrying roughly 109 million trained parameters — a decision adapter plus a scoring head — distilled from TypeSafe's hosted Jev (the model card calls it a student of TypeSafe Jev 1.13). The local weights include the trained additions, so decisions run without calling the teacher. AutoTrust reports about 137 ms per decision on a B200 and unchanged coding ability (78.0% vs 78.0%); the video in this page re-tests those claims on a DGX Spark in BF16 and scores the model against three controls on 72 held-out policy cases.

Is AutoTrust JEV-27B the same as Open-Jev-27B-v1.1?

No — they are two different open 27B Jev-family models that search results routinely mix up. The model tested here is AutoTrust AI Lab's integrated release at huggingface.co/autotrust/JEV-27B: a full finetune of Qwen3.8-27B whose frozen backbone carries a trained decision adapter and scoring head distilled from TypeSafe Jev 1.13, published by the AutoTrust org. ZefanCai/Open-Jev-27B-v1.1 is a separate community project by a different publisher, distributed on Hugging Face as a PEFT adapter package (adapter plus head.pt) on the same Qwen3.8-27B base. Different team, different artifact format, different training. If a page or repo path does not start with autotrust/, it is not the model this video tests.

What about JEV-27B-VL — is that the model in this test?

JEV-27B-VL (huggingface.co/autotrust/JEV-27B-VL) is the vision sibling of the model tested here: the same AutoTrust org ships it as an image-text-to-text model on the same Qwen3.8-27B base for multimodal decisions. The video and this page cover the text-only JEV-27B; none of the scores, latencies or memory figures below apply to the VL variant. Mentioning it because the two live in the same HF org and a name glance is all it takes to grab the wrong checkpoint.

Did the confidence-score gate work in the video?

No — and that is the video’s core negative result. Six thresholds (0.50, 0.60, 0.70, 0.80, 0.90, 0.95) were tried on 24 development cases to decide when a low quick-head score should escalate to reasoning; none met the accuracy requirement, so the video contains no approved hybrid gate and no held-out result for one. The reason is a single clean error case: the quick head scored 0.969 on a wrong answer (it chose retry when the required action was collect), which sits above even the highest threshold tried. The score measures preference among supplied choices, not probability of correctness — treating 0.97 as 97% reliability requires calibration evidence that was not shown, and picking a new threshold after seeing the answers would just be a new experiment needing fresh validation.

What hardware does JEV-27B need to run locally?

The video ran everything on a DGX Spark (GB10) with the pinned revision in 16-bit BF16 — not Q4 or Q8 — keeping the whole backbone loaded. The download is about 50 GiB and the highest sampled GPU process memory was about 52 GiB (Spark unified memory, so those overlap rather than add). Median end-to-end request times were 0.65 s for the quick head, 2.34 s for ordinary generation, and 102.20 s for the reasoning path, one request at a time — with two accelerated components unavailable and fallback implementations in place, so treat it as a reference run, not an optimized serving benchmark. The practical takeaway: the fast decision head does not shrink a 27-billion-parameter download.

Can you trust the 71/72 and 67/67 reasoning scores?

Treat them as carefully labeled behaviors to investigate, not a benchmark to quote. The 71/72 comes from 72 synthetic held-out cases built from a small set of related policy patterns; the perfect 67/67 is a post-hoc sensitivity check after the testers set aside five of their own cases with a wording flaw (a billing-to-Security policy exception colliding with the Security option's description) — original scores were kept and the flaw published. Two more limits ride along: reasoning hit its 512-token budget on 20 of the 72 cases and was forced to answer (a correct choice can follow unfinished reasoning), and the ordinary-generation control never received a prompt search, so it measures configured paths rather than a ceiling. The video's own framing is the right one: these are specific behaviors to investigate on your workflow's cases — before production, not after.

関連ガイド

その他の動画ウォークスルー