Guides / illustrated walkthrough

Jev, Hands-On: Where the Official Claims Meet Independent Remeasurement

A 16:55 hands-on from Mandarin channel 01Coder: third-party audits two days after launch re-measured the 20-400x claim at 5-25x and overseas latency at 1.6-3.7s, "no hallucination" does not hold, and calibration remains publicly unverifiable — walked through the Vercel AI Gateway route, a one-file AI SDK playground running six scenes, and the official agent skill.

Quick takeaway

Two days after Jev launched, third-party audits re-checked the official numbers, and this Mandarin video (01Coder, the second entry from this channel on this site) walks the whole reconciliation: the claimed 20-400x speed-and-cost advantage re-measures at 5-25x; official 70-500ms latency lands at 1.6-3.7 seconds from overseas (this site’s own benchmark harness measured 290-410ms P95 across three task suites — latency follows the route, and both sources converge on "measure your own scenario"); "frontier intelligence" oversold an accuracy that sits level with mid-tier models; "no hallucination" does not hold — the official docs say typed output guarantees the interface, not the truth; and calibration, the core selling point, still has no public benchmark to verify it. Access runs through Vercel AI Gateway (typesafe-ai/jev), callable only via AI SDK 7’s experimental_evaluate — OpenAI-compatible endpoints unsupported; the official API was waitlisted at filming time and has since opened up, with a local Ollama route available. The creator built a one-file Next.js playground and ran six scenes: a refund boolean with states swapped to watch the probability move; five parallel questions returning in 245ms/599 tokens, essentially the same latency as one; a contradictory ticket and then gibberish, showing confidence-gated escalation to humans; model routing (a typo goes cheap, an LRU+TTL cache with tests picks strong at probability 1.0); comment moderation with a raw JSON state; and an LLM-output gate that catches a deliberately planted internal-policy leak — establishing the position: the LLM generates, Jev judges behind it. The official skill (typesafe-ai/skills, committed three weeks before launch) teaches coding agents three things: read docs on demand, derive needed judgments from app requirements, and unlearn LLM-era habits. The final boundary: use it when the answer space is definable in advance; skip it when you need explanations, generation, or your option design is shaky; it is a closed-source hosted API with no weights, so weigh privacy yourself — and reconcile probabilities against a few dozen labeled examples before adopting.

Video source

01Coder

16:55tYvu6IpSfiM

Step-by-step walkthrough

  1. 1

    The official numbers on the table: cheap, fast, level with Terra

    The video opens by laying the official claims out in full: $0.042 per million input tokens, free output (there are no output tokens), 70-500ms end-to-end latency, a claimed workflow accuracy of 67.8% against GPT-5.6 Terra’s 67.9% — at 1/76 the cost and 25x the speed — and a 0% structured-output error rate because outputting a wrong type is mathematically impossible. The backdrop: Jev launched September 15, took 1,679 points and 400+ comments on Hacker News the same day, and Vercel announced AI Gateway support the next day. It is TypeSafe’s first model out of stealth, founded by Diogo, co-inventor of RLHF at OpenAI. The narration is the channel host’s own Mandarin voiceover (he introduces himself as Xiaomutou) over flat motion-graphic cards, with no talking head anywhere — which is where every frame on this page comes from.

    The official numbers card prices Jev input at $0.042 per million tokens with free output, quotes 70-500ms end-to-end latency, and shows workflow accuracy of 67.8% against GPT-5.6 Terra’s 67.9%.
    Official claims: $0.042/M input, free output, 70-500ms, 67.8% vs 67.9%.Watch at 4:25
  2. 2

    Two days after launch, third-party audits started line-by-line

    Two days after the numbers dropped, independent developers published audits (the frame credits a jev-exploration remeasure project; figures here are relayed by the channel): the 20-400x speed-and-cost multiplier re-measured at 5-25x, and overseas latency of 1.6-3.7 seconds rather than hundreds of milliseconds — though the channel stresses results vary with the network route. Accuracy sits level with mid-tier models and behind reasoning models, so "frontier intelligence" oversold it. "No hallucination" does not hold either: the official documentation itself states that typed output guarantees the interface, not the truth — the answer will be one of your options, not necessarily the right one. And calibration — whether a stated 80% really means an 80% chance of being right — has no public calibration study or benchmark to check against, despite being the product’s core selling point. For balance: this site’s own benchmarks measured P95 latencies of 290-410ms across three task suites, the same order as the official claim. Both sides point to the same conclusion — latency follows your route, so measure your own.

    A claim-audit table lines official marketing up against third-party remeasurement: the 20-400x multiplier corrected to 5-25x, overseas latency of 1.6-3.7 seconds, accuracy level with mid-tier models, and a no-hallucination row marked as not holding.
    Five audited rows: the multiplier shrinks to 5-25x, latency runs 1.6-3.7s by route, "no hallucination" fails.Watch at 5:52
  3. 3

    The channel’s three verdicts: cheap and fast are real; format safety is not answer correctness

    The channel condenses the whole audit into three verdicts: cheap and fast are real, but the multiplier depends on what you compare against; type safety equals format safety, not answer correctness; and whether answers are right must be measured in your own scenario. That stance matches how this site runs its benchmark reports — automation thresholds, human-fallback tiers, and a reproducible harness all serve the same "measure it yourself" conclusion. It is also why this page exists: the site’s first critical-hands-on review, where critical findings are attributed to the channel and the third-party audit while the official claims are presented alongside them.

    Three verdict banners condense the audit table: cheap and fast are real but the multiplier depends on your baseline, type safety means format safety rather than correct answers, and correctness must be measured in your own scenario.
    Channel verdicts: cheap and fast are real — whether answers are right, measure in your own scenario.Watch at 6:05
  4. 4

    Integration paths: a waitlisted official API and Vercel AI Gateway on the shelf

    There are four integration paths: the official API, Vercel AI Gateway, the AI SDK’s evaluate call, and the official agent skill. When this was filmed (mid-September) the official API was still invite-by-waitlist — the creator was still in the queue himself — so he went through Vercel AI Gateway, where the model ID is typesafe-ai/jev. Two constraints worth memorizing: it is only callable through AI SDK 7’s experimental_evaluate, and OpenAI-compatible endpoints are not supported. Timeliness note: the waitlist story is history now — this site has since confirmed self-serve access and an Ollama-based local route (see run-jev-locally-guide).

    The Vercel AI Gateway access card lists model ID typesafe-ai/jev with a badge restricting calls to AI SDK 7’s experimental_evaluate while the OpenAI-compatible endpoint option is struck through.
    Still queued for the official API? typesafe-ai/jev is already on the Gateway — but only via the AI SDK.Watch at 7:42
  5. 5

    A one-file evaluate route: the minimal AI SDK surface

    The creator hand-wrote a small playground in TypeScript: the server is a single file, app/api/evaluate/route.ts — evaluate takes a model (the Gateway’s typesafe-ai/jev), the state, and the questions, returns answers and usage, and he wraps the call with performance.now to report latency alongside. That one-file shape is the minimal surface for calling Jev through the AI SDK. All six demo scenes run inside this playground: three panes — preset scenes on the left, request in the middle, response on the right — with both request and response switchable to raw JSON.

    A one-file Next.js route, app/api/evaluate/route.ts, imports experimental_evaluate from the ai package, targets gateway typesafe-ai/jev, passes state and questions, and times each call with performance.now.
    One file is enough: evaluate takes model, state, questions — and times itself.Watch at 8:02
  6. 6

    Five parallel questions and conditional relevance: the category decides which answers matter

    Scene two is one compound ticket: cannot log in, overcharged last month, demands a refund, urgent tone. Five questions go out in a single request — a four-option Choice for category, a four-level Score for fault severity, a boolean for reproduction steps, a boolean for refund demanded, a three-level Score for mood — and five answers come back together: 245ms and 599 tokens measured, essentially the same latency as the single-question scene. The usage lesson is conditional relevance: fault severity only matters if the category is a fault; refund only matters if the category is billing. The traditional pattern asks serially; Jev’s guidance is to ask everything in parallel and let code keep only the relevant answers. He then swaps in a self-contradictory ticket — "everything is fine but nothing works. I don’t want a refund, give me my money back" — the category must still pick one, but the probability distribution now shows the uncertainty.

    The five-question triage scene draws a conditional-relevance tree where bug_severity only matters once the category lands on bug and wants_refund only once it lands on billing, beside a serial-asking versus ask-all-at-once comparison.
    Five answers, one round trip: the category decides which answers matter; code keeps the relevant ones.Watch at 9:58
  7. 7

    Gibberish tickets and confidence: the type is always right, the answer not necessarily

    Replace the state with gibberish and the model still picks a category — look at the confidence: this is the direct demo of "the type is always right, the answer is not necessarily right." In code you add one rule: below a confidence threshold, route to a human — exactly the confidence-gated routing the official docs recommend. The skill adds two corrections: confidence only measures how concentrated the probability distribution is, not whether the overall flow is correct; and a Noul near 0.5 means yes and no are evenly split, not medium intensity. The official FAQ adds the inverse advice: if you just want the best option, take the highest-probability one and stop sprinkling thresholds everywhere.

    Feeding a gibberish ticket, the playground response panel still returns a fixed category choice but with visibly depressed confidence numbers — the cue for confidence-gated routing to a human reviewer.
    Even gibberish must pick one — below the confidence threshold, escalate to a human (the official pattern).Watch at 10:20
  8. 8

    The LLM generates, Jev judges: one gate after every generation

    Scene five positions Jev in the system: the state holds two user questions and an LLM-written support draft; three questions ask for a quality Score, a leaks-internal-policy boolean, and a tone Choice (apologetic / neutral / deflecting). The draft deliberately embeds an internal policy line — "complain twice or more and you can request an extra discount coupon" — and the leak boolean catches it; delete the line, re-run, and the number falls back. Scene four adds a friendly detail about state: pass a JSON object directly (author, time, body, report count) instead of flattening it into prose — the model reads the fields itself; the moderation trio (spam? offensiveness? allow / review / delete) runs exactly that way. The conclusion becomes the diagram: Jev does not replace the language model — the LLM generates, Jev judges behind it, one gate after every generation.

    A position diagram routes the user question through an LLM that generates the draft, then through a Jev gate deciding whether it can be sent or should be rewritten, with two escalation branches hanging off the gate.
    Not an LLM replacement: the model generates, Jev stands guard behind it.Watch at 13:10
  9. 9

    The official skill’s three jobs: point at docs, teach decomposition, unlearn LLM habits

    When it is time to wire Jev into a real project you will probably have Claude Code or a similar coding agent write it — so TypeSafe ships a skill: repo typesafe-ai/skills, created August 25, three weeks before the model launched; a single SKILL.md under 150 lines; two commands to install as a Claude Code plugin, npx skills add for other agents. It is deliberately not an API reference (the online docs are the source of truth; append .md to any Mintlify page URL to get Markdown). It does three things: points the agent at the right doc pages on demand; teaches it to decompose requirements — reason backward from what the app must show, choose, or change — and lists six scenario shapes (route and fill parameters, select rather than generate, judge evidence after retrieval, turn judgments into reusable data, validate and escalate, decide the next step as state changes); and corrects LLM-era habits — ask all independent questions at once, including speculative ones, and keep questions and threshold constants in one file for human review.

    The official skill card lays out three step chips — point to the online docs, teach the agent to decompose requirements, correct LLM-era habits — with bullets like asking all independent questions at once and confidence measuring distribution concentration only.
    Not an API manual — a workflow: read docs on demand, derive judgments, unlearn habits.Watch at 15:04
  10. 10

    The use/avoid boundary, and the channel’s way to start

    The closing card draws the boundary. Use it when the answer space can be defined in advance — classification, routing, scoring, validation, and real-time loops making several decisions per second, where the overhead of LLM text generation is real. Avoid it when you need explanations (it cannot give reasons), generation (it cannot write), or when options are poorly designed — the correct answer may not be among them, and the probability still has to land on the remaining options; even TypeSafe admits the hard part shifted from writing prompts to designing decision patterns. Two footnotes: it is a closed-source hosted API with no weights, so privacy-sensitive uses are your call (the community is already fine-tuning Qwen into calibrated open-weight replicas); and the channel’s advice — take a real classification or routing task, run a few dozen labeled examples, and check whether the probabilities match your labels before shipping it into a system.

    The final verdict card splits use cases side by side: answers whose space you can define in advance — classification, routing, scoring, validation, real-time loops — versus explanations, generation, and poorly designed option sets where the right answer may be missing.
    Definable answer space: use it. Explanations, generation, shaky options: skip it.Watch at 15:57

Frequently asked questions

Is Jev worth it?

The channel’s verdict after this hands-on: cheap and fast are real, so the value depends on task shape. If your answers live in a space you can define in advance — classification, routing, scoring, validation, or real-time loops making several decisions per second — the LLM text-generation overhead you remove is real, and it is worth it. If you need explanations, free-form generation, or your option design is shaky, it is not. The practical test: run a few dozen labeled examples from your own task and check whether the probabilities match your labels before adopting it.

What did this Jev hands-on review conclude?

Five rows: the official 20-400x multiplier re-measured at 5-25x; overseas latency of 1.6-3.7s (this site’s own benchmark harness measured 290-410ms P95 — route-dependent); accuracy level with mid-tier models, behind reasoning models; "no hallucination" does not hold — typed output guarantees the interface, not the truth; and calibration remains publicly unverifiable despite being the core selling point. Cheap and fast: real. Type safety: format safety only. Answer correctness: measure it in your own scenario.

How fast is Jev, really?

It depends on whose numbers and which route. Official: 70-500ms end-to-end. The third-party audit relayed in the video: 1.6-3.7s from overseas. This site’s own three benchmark suites: 290-410ms P95. All three are honest — latency follows the network route and integration path, which is why every source ends at the same advice: benchmark your own route. One more data point from the video: five parallel questions returned in essentially the same latency as one (245ms, 599 tokens), so question count is nearly free.

Does Jev hallucinate?

The official "no hallucination" claim does not hold, as the video demonstrates: typed output guarantees the interface, not the truth — the answer is always one of your options, but not guaranteed correct. Feed it a contradictory ticket and the category still picks one while the probability distribution reveals the uncertainty; feed it gibberish and confidence drops. That is the feature to build on: gate automation on confidence and route low-confidence cases to humans (the official confidence-gated routing pattern) — while remembering confidence measures distribution concentration, not overall flow correctness.

How does Jev’s cost compare with calling a raw LLM?

Official pricing: $0.042 per million input tokens, output free — there are no output tokens. Against GPT-5.6 Terra on TypeSafe’s workflow evaluation: parity at 67.8% vs 67.9% accuracy at 1/76 the cost and 25x the speed; the independent remeasure puts the speed-and-cost multiplier at 5-25x depending on baseline. The routing pattern also saves real money downstream: deciding which model answers costs negligible time next to actually calling the big model. And in the LLM-output-gate pattern the two cooperate — the LLM still generates, Jev adds a cheap judgment layer after each generation.

What are the integration paths for Jev?

Four: (1) the official API — still waitlisted when the video was filmed, now self-serve with a local Ollama route available; (2) Vercel AI Gateway, model ID typesafe-ai/jev; (3) the AI SDK 7 experimental_evaluate call — the only supported invocation on the Gateway, OpenAI-compatible endpoints not supported; (4) the official agent skill at typesafe-ai/skills — a sub-150-line SKILL.md that points your coding agent at the docs, teaches requirement decomposition, and corrects LLM-era habits.

Related guides

More video walkthroughs