Guides / illustrated walkthrough
Jev, Hands-On: Where the Official Claims Meet Independent Remeasurement
A 16:55 hands-on from Mandarin channel 01Coder: third-party audits two days after launch re-measured the 20-400x claim at 5-25x and overseas latency at 1.6-3.7s, "no hallucination" does not hold, and calibration remains publicly unverifiable — walked through the Vercel AI Gateway route, a one-file AI SDK playground running six scenes, and the official agent skill.
Quick takeaway
Two days after Jev launched, third-party audits re-checked the official numbers, and this Mandarin video (01Coder, the second entry from this channel on this site) walks the whole reconciliation: the claimed 20-400x speed-and-cost advantage re-measures at 5-25x; official 70-500ms latency lands at 1.6-3.7 seconds from overseas (this site’s own benchmark harness measured 290-410ms P95 across three task suites — latency follows the route, and both sources converge on "measure your own scenario"); "frontier intelligence" oversold an accuracy that sits level with mid-tier models; "no hallucination" does not hold — the official docs say typed output guarantees the interface, not the truth; and calibration, the core selling point, still has no public benchmark to verify it. Access runs through Vercel AI Gateway (typesafe-ai/jev), callable only via AI SDK 7’s experimental_evaluate — OpenAI-compatible endpoints unsupported; the official API was waitlisted at filming time and has since opened up, with a local Ollama route available. The creator built a one-file Next.js playground and ran six scenes: a refund boolean with states swapped to watch the probability move; five parallel questions returning in 245ms/599 tokens, essentially the same latency as one; a contradictory ticket and then gibberish, showing confidence-gated escalation to humans; model routing (a typo goes cheap, an LRU+TTL cache with tests picks strong at probability 1.0); comment moderation with a raw JSON state; and an LLM-output gate that catches a deliberately planted internal-policy leak — establishing the position: the LLM generates, Jev judges behind it. The official skill (typesafe-ai/skills, committed three weeks before launch) teaches coding agents three things: read docs on demand, derive needed judgments from app requirements, and unlearn LLM-era habits. The final boundary: use it when the answer space is definable in advance; skip it when you need explanations, generation, or your option design is shaky; it is a closed-source hosted API with no weights, so weigh privacy yourself — and reconcile probabilities against a few dozen labeled examples before adopting.
Video source
01Coder
Step-by-step walkthrough
- 1
The official numbers on the table: cheap, fast, level with Terra
The video opens by laying the official claims out in full: $0.042 per million input tokens, free output (there are no output tokens), 70-500ms end-to-end latency, a claimed workflow accuracy of 67.8% against GPT-5.6 Terra’s 67.9% — at 1/76 the cost and 25x the speed — and a 0% structured-output error rate because outputting a wrong type is mathematically impossible. The backdrop: Jev launched September 15, took 1,679 points and 400+ comments on Hacker News the same day, and Vercel announced AI Gateway support the next day. It is TypeSafe’s first model out of stealth, founded by Diogo, co-inventor of RLHF at OpenAI. The narration is the channel host’s own Mandarin voiceover (he introduces himself as Xiaomutou) over flat motion-graphic cards, with no talking head anywhere — which is where every frame on this page comes from.

Official claims: $0.042/M input, free output, 70-500ms, 67.8% vs 67.9%.Watch at 4:25 - 2
Two days after launch, third-party audits started line-by-line
Two days after the numbers dropped, independent developers published audits (the frame credits a jev-exploration remeasure project; figures here are relayed by the channel): the 20-400x speed-and-cost multiplier re-measured at 5-25x, and overseas latency of 1.6-3.7 seconds rather than hundreds of milliseconds — though the channel stresses results vary with the network route. Accuracy sits level with mid-tier models and behind reasoning models, so "frontier intelligence" oversold it. "No hallucination" does not hold either: the official documentation itself states that typed output guarantees the interface, not the truth — the answer will be one of your options, not necessarily the right one. And calibration — whether a stated 80% really means an 80% chance of being right — has no public calibration study or benchmark to check against, despite being the product’s core selling point. For balance: this site’s own benchmarks measured P95 latencies of 290-410ms across three task suites, the same order as the official claim. Both sides point to the same conclusion — latency follows your route, so measure your own.

Five audited rows: the multiplier shrinks to 5-25x, latency runs 1.6-3.7s by route, "no hallucination" fails.Watch at 5:52 - 3
The channel’s three verdicts: cheap and fast are real; format safety is not answer correctness
The channel condenses the whole audit into three verdicts: cheap and fast are real, but the multiplier depends on what you compare against; type safety equals format safety, not answer correctness; and whether answers are right must be measured in your own scenario. That stance matches how this site runs its benchmark reports — automation thresholds, human-fallback tiers, and a reproducible harness all serve the same "measure it yourself" conclusion. It is also why this page exists: the site’s first critical-hands-on review, where critical findings are attributed to the channel and the third-party audit while the official claims are presented alongside them.

Channel verdicts: cheap and fast are real — whether answers are right, measure in your own scenario.Watch at 6:05 - 4
Integration paths: a waitlisted official API and Vercel AI Gateway on the shelf
There are four integration paths: the official API, Vercel AI Gateway, the AI SDK’s evaluate call, and the official agent skill. When this was filmed (mid-September) the official API was still invite-by-waitlist — the creator was still in the queue himself — so he went through Vercel AI Gateway, where the model ID is typesafe-ai/jev. Two constraints worth memorizing: it is only callable through AI SDK 7’s experimental_evaluate, and OpenAI-compatible endpoints are not supported. Timeliness note: the waitlist story is history now — this site has since confirmed self-serve access and an Ollama-based local route (see run-jev-locally-guide).

Still queued for the official API? typesafe-ai/jev is already on the Gateway — but only via the AI SDK.Watch at 7:42 - 5
A one-file evaluate route: the minimal AI SDK surface
The creator hand-wrote a small playground in TypeScript: the server is a single file, app/api/evaluate/route.ts — evaluate takes a model (the Gateway’s typesafe-ai/jev), the state, and the questions, returns answers and usage, and he wraps the call with performance.now to report latency alongside. That one-file shape is the minimal surface for calling Jev through the AI SDK. All six demo scenes run inside this playground: three panes — preset scenes on the left, request in the middle, response on the right — with both request and response switchable to raw JSON.

One file is enough: evaluate takes model, state, questions — and times itself.Watch at 8:02 - 6
Five parallel questions and conditional relevance: the category decides which answers matter
Scene two is one compound ticket: cannot log in, overcharged last month, demands a refund, urgent tone. Five questions go out in a single request — a four-option Choice for category, a four-level Score for fault severity, a boolean for reproduction steps, a boolean for refund demanded, a three-level Score for mood — and five answers come back together: 245ms and 599 tokens measured, essentially the same latency as the single-question scene. The usage lesson is conditional relevance: fault severity only matters if the category is a fault; refund only matters if the category is billing. The traditional pattern asks serially; Jev’s guidance is to ask everything in parallel and let code keep only the relevant answers. He then swaps in a self-contradictory ticket — "everything is fine but nothing works. I don’t want a refund, give me my money back" — the category must still pick one, but the probability distribution now shows the uncertainty.

Five answers, one round trip: the category decides which answers matter; code keeps the relevant ones.Watch at 9:58 - 7
Gibberish tickets and confidence: the type is always right, the answer not necessarily
Replace the state with gibberish and the model still picks a category — look at the confidence: this is the direct demo of "the type is always right, the answer is not necessarily right." In code you add one rule: below a confidence threshold, route to a human — exactly the confidence-gated routing the official docs recommend. The skill adds two corrections: confidence only measures how concentrated the probability distribution is, not whether the overall flow is correct; and a Noul near 0.5 means yes and no are evenly split, not medium intensity. The official FAQ adds the inverse advice: if you just want the best option, take the highest-probability one and stop sprinkling thresholds everywhere.

Even gibberish must pick one — below the confidence threshold, escalate to a human (the official pattern).Watch at 10:20 - 8
The LLM generates, Jev judges: one gate after every generation
Scene five positions Jev in the system: the state holds two user questions and an LLM-written support draft; three questions ask for a quality Score, a leaks-internal-policy boolean, and a tone Choice (apologetic / neutral / deflecting). The draft deliberately embeds an internal policy line — "complain twice or more and you can request an extra discount coupon" — and the leak boolean catches it; delete the line, re-run, and the number falls back. Scene four adds a friendly detail about state: pass a JSON object directly (author, time, body, report count) instead of flattening it into prose — the model reads the fields itself; the moderation trio (spam? offensiveness? allow / review / delete) runs exactly that way. The conclusion becomes the diagram: Jev does not replace the language model — the LLM generates, Jev judges behind it, one gate after every generation.

Not an LLM replacement: the model generates, Jev stands guard behind it.Watch at 13:10 - 9
The official skill’s three jobs: point at docs, teach decomposition, unlearn LLM habits
When it is time to wire Jev into a real project you will probably have Claude Code or a similar coding agent write it — so TypeSafe ships a skill: repo typesafe-ai/skills, created August 25, three weeks before the model launched; a single SKILL.md under 150 lines; two commands to install as a Claude Code plugin, npx skills add for other agents. It is deliberately not an API reference (the online docs are the source of truth; append .md to any Mintlify page URL to get Markdown). It does three things: points the agent at the right doc pages on demand; teaches it to decompose requirements — reason backward from what the app must show, choose, or change — and lists six scenario shapes (route and fill parameters, select rather than generate, judge evidence after retrieval, turn judgments into reusable data, validate and escalate, decide the next step as state changes); and corrects LLM-era habits — ask all independent questions at once, including speculative ones, and keep questions and threshold constants in one file for human review.

Not an API manual — a workflow: read docs on demand, derive judgments, unlearn habits.Watch at 15:04 - 10
The use/avoid boundary, and the channel’s way to start
The closing card draws the boundary. Use it when the answer space can be defined in advance — classification, routing, scoring, validation, and real-time loops making several decisions per second, where the overhead of LLM text generation is real. Avoid it when you need explanations (it cannot give reasons), generation (it cannot write), or when options are poorly designed — the correct answer may not be among them, and the probability still has to land on the remaining options; even TypeSafe admits the hard part shifted from writing prompts to designing decision patterns. Two footnotes: it is a closed-source hosted API with no weights, so privacy-sensitive uses are your call (the community is already fine-tuning Qwen into calibrated open-weight replicas); and the channel’s advice — take a real classification or routing task, run a few dozen labeled examples, and check whether the probabilities match your labels before shipping it into a system.

Definable answer space: use it. Explanations, generation, shaky options: skip it.Watch at 15:57
Frequently asked questions
Is Jev worth it?
The channel’s verdict after this hands-on: cheap and fast are real, so the value depends on task shape. If your answers live in a space you can define in advance — classification, routing, scoring, validation, or real-time loops making several decisions per second — the LLM text-generation overhead you remove is real, and it is worth it. If you need explanations, free-form generation, or your option design is shaky, it is not. The practical test: run a few dozen labeled examples from your own task and check whether the probabilities match your labels before adopting it.
What did this Jev hands-on review conclude?
Five rows: the official 20-400x multiplier re-measured at 5-25x; overseas latency of 1.6-3.7s (this site’s own benchmark harness measured 290-410ms P95 — route-dependent); accuracy level with mid-tier models, behind reasoning models; "no hallucination" does not hold — typed output guarantees the interface, not the truth; and calibration remains publicly unverifiable despite being the core selling point. Cheap and fast: real. Type safety: format safety only. Answer correctness: measure it in your own scenario.
How fast is Jev, really?
It depends on whose numbers and which route. Official: 70-500ms end-to-end. The third-party audit relayed in the video: 1.6-3.7s from overseas. This site’s own three benchmark suites: 290-410ms P95. All three are honest — latency follows the network route and integration path, which is why every source ends at the same advice: benchmark your own route. One more data point from the video: five parallel questions returned in essentially the same latency as one (245ms, 599 tokens), so question count is nearly free.
Does Jev hallucinate?
The official "no hallucination" claim does not hold, as the video demonstrates: typed output guarantees the interface, not the truth — the answer is always one of your options, but not guaranteed correct. Feed it a contradictory ticket and the category still picks one while the probability distribution reveals the uncertainty; feed it gibberish and confidence drops. That is the feature to build on: gate automation on confidence and route low-confidence cases to humans (the official confidence-gated routing pattern) — while remembering confidence measures distribution concentration, not overall flow correctness.
How does Jev’s cost compare with calling a raw LLM?
Official pricing: $0.042 per million input tokens, output free — there are no output tokens. Against GPT-5.6 Terra on TypeSafe’s workflow evaluation: parity at 67.8% vs 67.9% accuracy at 1/76 the cost and 25x the speed; the independent remeasure puts the speed-and-cost multiplier at 5-25x depending on baseline. The routing pattern also saves real money downstream: deciding which model answers costs negligible time next to actually calling the big model. And in the LLM-output-gate pattern the two cooperate — the LLM still generates, Jev adds a cheap judgment layer after each generation.
What are the integration paths for Jev?
Four: (1) the official API — still waitlisted when the video was filmed, now self-serve with a local Ollama route available; (2) Vercel AI Gateway, model ID typesafe-ai/jev; (3) the AI SDK 7 experimental_evaluate call — the only supported invocation on the Gateway, OpenAI-compatible endpoints not supported; (4) the official agent skill at typesafe-ai/skills — a sub-150-line SKILL.md that points your coding agent at the docs, teaches requirement decomposition, and corrects LLM-era habits.
Related guides
Jev Getting Started: Your First API Call
This review deliberately skips the beginner concepts (one paragraph at most); the step-by-step first-call tutorial lives there.
ReadJev Playground Walkthrough (our tool)
The video uses its author’s hand-built playground; the walkthrough of our own in-browser playground lives on that page — swap states and watch probabilities without writing code.
ReadOur Jev Benchmark Reports
The counterweight to the third-party audit relayed in the video: ticket-routing Macro F1 78.4% at 410ms P95, spam detection at 320ms, prompt-injection defense at 290ms, with a reproducible harness.
ReadJev vs GPT-5.6 Luna Benchmark
The third-party-benchmark axis: how to read the accuracy-latency-cost triangle fairly.
ReadRun Jev Locally (Ollama route)
The waitlist story in the video is history: self-serve access and the local route, installation and endpoints.
ReadMore video walkthroughs
- Jev Classification Quickstart: OpenRouter API, Primitives & Real Probabilities
- Jev Architecture Explained: Why 70ms Decision Models Beat LLMs for Workflow Automation
- Ultra-Fast Browser Agents with Jev: 178ms DOM Loops & Dual Model Orchestration
- Open Jev Models Are Here: Semif, Nimble, Decider, DiffusionGemma & Laya Hands-On
- Jev vs LLM: Will Jev Replace LLMs? Krish Naik's Whiteboard Explainer
- Jev Trader Tutorial: Build a Subsecond AI Trading Bot on Monad
- Jev Model Router: Build a Privacy-Gated LLM Router with Jev & OpenJev
- Jev Tutorial for Beginners: State, Questions & the TypeScript SDK
- Run Jev Locally: Kev, SemIf & Von on Your Own GPU (OpenJev Guide)
- Jev RAG Reranker: Policy-Steered Reranking for Retrieval-Augmented Generation
- When to Use Jev: An Engineer's Audit of Claims, Gates, and Failure Modes
- LangChain + Jev Integration Tutorial: Routing, Guardrails & Evals
- Jev MCP Server: Connect Jev Decisions to Claude Code & Cursor
- Jev vs Luna: Independent Benchmarks Put "Better, Faster, Cheaper" to the Test
- Jev Agent Harness: Where the Decision Gate Sits in Your LLM Loop
- Jev Playground Walkthrough: The Hotdog Lesson, Criteria, and a Four-Console Token Test
- Jev Text Classification API: Zero-Shot CLI & REST with classifier.dev
- Jev API Examples: First Request, curl & All Three Question Types
- Jev Log Triage with Expanso Edge
- Jev Lead Enrichment with Treg: ICP and Signup Scoring
- Use Jev Decision Nodes in Heym for Model Routing
- Laya Tutorial: Open-Source AI Routing With Calibrated Probabilities (Laya vs Jev Setup)
- Train Your Own Jev: Fine-Tune a Jev-Style Decision Model for $5–$17 (What You Can and Cannot Train)
- Jev Tips: 8 Best Practices for Better Decisions (State, Questions, Criteria & Thresholds)
- Jev Context Compaction: Prune AI Agent Memory Without Generative Summaries
- Jev as an LLM Judge: Confidence-Gated Cascades at 0.36% of the Cost
- TypeSafe Computer Use: Local Desktop Automation with Jev, Step by Step
- Jev Resume Screening: Build an AI Resume Evaluator with the Jev JavaScript SDK
- Jev + Claude Code Guide: Voice-Controlled Browser Automation with Typed Decisions
- Jev vs Ollama: Can Local AI Replace Hosted Jev Without Sending Your Data Away?
- Build Your Own Jev: Train a Free Open-Source Zero-Shot Classifier (That Plays Doom)
- CUA-S1-Forms: a 706K-Parameter Jev-Like Model That Fills GUI Forms on Your CPU
- NOC/SOC Alert Triage with Jev: Rules First, One Typed Question, a Policy Gate
- Ollama Decision Models: Run tev1 and Nimble Locally (Tested on an 8 GB Card)
- Jev Guardrails in Production: A Five-Step Playbook for Decision Automation
- A Session Drift Guard for Pi Agent: Let the Jev Model Propose, Let Code Decide
- 50 Tev1 Use Cases: What a Local Decision Model Can Actually Do (Tested on an 8 GB Laptop)
- Jev Ticket Classification in a Real App: the After-Insert Hook and the Calculated Field