Guides / illustrated walkthrough

Jev + Claude Code Guide: Voice-Controlled Browser Automation with Typed Decisions

A step-by-step walkthrough of Prompt Engineer 48’s 18:25 build: install the TypeSafe agent skill in Claude Code, watch typed Jev decisions route a Stripe ticket, and see a spoken command drive Chrome at P=0.78 with a click in 1176 ms.

Quick takeaway

Prompt Engineer 48’s 18:25 video makes the agent-integration case for Jev in four moves. First the indictment: every “LLM as decision engine” is the same loop — Prompt, Parse, Pray, Babysit — and RLHF’s people-pleasing training leaves you mode dropping, overconfidence, and a human babysitting the automation. Then the contract: Jev System One takes a state plus typed questions and returns typed answers, a full probability distribution, and a confidence number — “LLMs produce words for people. Jev produces typed decisions.” The proof is on screen: a terminal race where uv run typesafe-race gets 28 typed answers in 0.114 s for $0.000081 while an autoregressive OpenAI model is still waiting for its first token; a playground where “Hi, my stripe is failing” wavers between payment (86%, confidence 79%) until one added word — “billing” — collapses the choice to 100% at confidence 99%; and a pricing slide at $0.042 per million input tokens with output free, 238x lower input price than Claude Fable 5.1. Then the Claude Code integration: the TypeSafe agent skill installs with claude plugin marketplace add typesafe-ai/skills + claude plugin install typesafe@typesafe-ai, handing the coding agent the three question types and architectural patterns. The payoff project, voice-chrome, is a pipeline of mic → faster-whisper (local GPU) → Jev (~400 ms) → Command dataclass (typed, validated) → Playwright → real Chrome — no prompt engineering, no JSON parsing, no retry loop — and in the final demo the spoken “Click on videos” is gated at ADDRESSED TO BROWSER P=0.78, picks click_link at 0.62, and clicks in 1176 ms.

Video source

Prompt Engineer 48

18:25rGNZxt8eUaI

Step-by-step walkthrough

  1. 1

    Open with the problem: text models make terrible decision engines

    The video opens on a live demo — the presenter speaks, and Chrome obeys — then spends its first two minutes on why the obvious alternative fails. The slide at 1:35 names the loop every integration team knows: 1. Prompt (“beg the model, in English, to answer in JSON”), 2. Parse (“strip the markdown fence; retry when it apologises instead”), 3. Pray (“no calibrated confidence, so you cannot tell a sure answer from a guess”), 4. Babysit (“so you keep a human in the loop — and the automation never actually ships”). The footnote carries TypeSafe’s thesis: RLHF tuned models to please people, and that training created mode dropping, overconfidence and unreliability — fine for chat, fatal for software.

    Slide from the Jev Claude Code video listing the Prompt, Parse, Pray and Babysit loop of using text models as decision engines with an RLHF overconfidence footnote
    Prompt, parse, pray, babysit — the four-step loop Jev was built to delete.Watch at 1:35
  2. 2

    Jev in one slide: send state and typed questions, get typed decisions

    The core slide defines the whole contract. What you send: a state — the ticket, the email, the row, the page, the app’s current facts — plus a set of typed questions about it. What comes back: typed answers, a full probability distribution, and a confidence number. No prose, no reasoning trace, nothing to parse. The quote beneath it does the positioning in one line — “LLMs produce words for people. Jev produces typed decisions.” — with code owning the workflow while the model supplies programmable common sense. The presenter frames this as backing Diogo Almeida’s TypeSafe bet (chapter 1:57): a model trained for calibrated decisions rather than people-pleasing chat, so the confidence number it prints is one your code can branch on.

    System One slide explaining that Jev takes a state plus typed questions and returns typed answers with a full probability distribution and a confidence number, quoting that LLMs produce words for people
    “LLMs produce words for people. Jev produces typed decisions.” — the slide that frames the video.Watch at 4:43
  3. 3

    Watch the race: 28 typed answers in 0.114 s while the LLM waits for token one

    In the benchmark clip played mid-video, two terminals run uv run typesafe-race side by side. The left pane asks the parallel TypeSafe API and instantly fills with 28 typed answers — noul gates like “Revenue currently impacted?” at 0.85, a choice like “Which primary department?” = technical at confidence 1.0, scores like “Churn likelihood level?” at 1.6 — finishing at a cost of $0.000081 in 0.114 s. The right pane, pointed at the autoregressive OpenAI API (gpt-5.6-terra), is still stuck on “Waiting for first token…”. This one frame is the JSON-output / confidence / latency chapter in full: a whole screen of decisions lands before a chat model emits a single character.

    Split terminal race showing uv run typesafe-race returning 28 typed Jev answers for $0.000081 in 0.114 seconds while the gpt-5.6-terra LLM pane still waits for its first token
    28 decisions, $0.000081, 114 ms — the LLM pane is still waiting for token one.Watch at 2:45
  4. 4

    Learn the entire API surface: Choice, Noul, Score — then grab a free key

    Chapter 4:13 puts the whole API on one slide. Choice returns the winning option, the probability of every option, and a confidence — built for routing, triage, intent. Noul returns a single 0–1 probability that a condition holds; there is no separate confidence because the number is the answer — built for guardrails, flags, checks. Score returns a probability-weighted position on ordered rubric levels you define in words — ranking, quality, risk. The slide’s selection rule: pick by what the answer means, not by what is easiest to prompt — a Noul near 0.5 means “equally likely yes or no”, never “medium”. Chapter 3:06, right before it, covers the on-ramp: typesafe.ai’s free tier and an API key from the console are enough to reproduce every demo in the video.

    Typesafe slide presenting the three Jev primitives Choice, Noul and Score with their use cases and the rule that a Noul near 0.5 means equally likely, never medium
    Three primitives are the entire API surface — pick by what the answer means.Watch at 4:52
  5. 5

    Playground: ask “does this message express urgency?” on a Stripe complaint

    In the console playground at console.typesafe.ai, the author pastes the state — “Hi, my stripe is failing. Please help ASAP. I am losing numerous high value clients.” — and defines three typed questions: urgency (noul, “Does this message express urgency?”), Problem Type (choice over accounts / billing / payment), and Severity (score over Low / Middle / High). The response from jev-latest: urgency 98% true; Problem Type payment at 86% with billing 13% and accounts 1% — overall confidence 79%; Severity 2 of 2 at 100%. The interesting part is the wobble: the distribution visibly hesitates between payment and billing, and that readable uncertainty is exactly what you want from a decision model.

    Typesafe AI playground showing a Stripe complaint state with urgency noul at 98 percent true, Problem Type choice on payment at 86 percent with 79 percent confidence, and Severity 2 of 2
    The choice distribution wavers between payment and billing — you can see the hesitation.Watch at 6:35
  6. 6

    One word flips the route: “stripe billing is failing”

    The rerun adds a single word — “Hi, my stripe billing is failing…” — and the choice collapses: Problem Type billing at 100% with confidence 99%, urgency 96% true, Severity still 2 of 2 at 100%. This is chapter 8:00’s real Stripe-ticket routing point: the same three typed questions sort tickets into queues, and when the state actually carries the signal, an 86/13/1 split sharpens into 100/0/0. In production the confidence number is the threshold — auto-route when it is high, escalate to a human when it is not, and both branches are one if statement.

    Typesafe playground rerun on the Stripe billing ticket showing Problem Type choice collapse to billing at 100 percent with 99 percent confidence and urgency noul at 96 percent true
    Add “billing” and the 86/13 wobble becomes a 100-percent routing decision.Watch at 8:20
  7. 7

    What it costs: $0.042 per million input tokens, output free

    The pricing slide shown at 12:22 lists Jev (TypeSafe AI) at $0.042 per million input tokens — $42 per billion — with output tokens FREE, against a field where GPT-6 Astra and Claude Fable 5.1 list $10.00 input / $50.00 output per million. The slide’s own headline: 238x lower input price than Claude Fable 5.1. (The video’s description rounds the pitch to 193x faster and 244x cheaper; 238x is the on-screen chart’s claim.) Chapter 11:07 pairs the cost talk with honest limitations: this is a decision model, not a writer — it classifies, routes, and scores at ~400 ms, and your code does the acting.

    Jev cost comparison chart listing $0.042 per million input tokens and free output for TypeSafe AI next to GPT-6 Astra and Claude Fable 5.1 at $10 with a 238x claim
    Decision tokens are priced like a utility, not like chat — $42 per billion input.Watch at 12:22
  8. 8

    Install the TypeSafe agent skill in Claude Code

    Chapter 13:01 is the Claude Code integration proper. docs.typesafe.ai/agent-skill ships a drop-in skill for Claude Code, Codex, and other agent environments — it hands your AI coding agent full context on the TypeSafe API: the three question types, the architectural patterns, and best practices for structuring evaluations. Installation is two terminal commands: claude plugin marketplace add typesafe-ai/skills, then claude plugin install typesafe@typesafe-ai (the manual route is copying the skills/typesafe-ai directory into your agent’s skills folder — pick one method to avoid duplicate copies). The doc’s common-issues list is worth reading before your first run: the agent not using the skill, routing not working as expected, confidence thresholds everywhere, and agents inventing request or response fields.

    Typesafe agent skill documentation page for Claude Code showing the claude plugin marketplace add typesafe-ai/skills and claude plugin install typesafe@typesafe-ai commands
    Two commands give Claude Code the full TypeSafe API context.Watch at 13:34
  9. 9

    The architecture: mic to Whisper to Jev to Playwright to Chrome

    The voice-chrome README shown at 15:06 is the whole blueprint in one line: mic → faster-whisper (local GPU) → Jev (~400 ms) → Command dataclass (typed, validated) → Playwright → Chrome. “No prompt engineering. No JSON parsing. No retry loop. Jev returns an enum, a probability distribution and a confidence number, and ordinary if statements do the rest.” The README’s comparison table draws the line against typical LLM tool-calling: instead of free text you must parse, the agent loop receives a typed Command. Whisper transcribes; Jev decides which command the utterance is — and whether it was meant for the browser at all; Python executes against a real Chrome over the DevTools protocol.

    VS Code showing the voice-chrome README with the pipeline mic, faster-whisper, Jev, Command dataclass, Playwright and Chrome plus a table contrasting a typed Command with LLM tool-calling
    A typed Command dataclass replaces prompt parsing at the heart of the loop.Watch at 15:06
  10. 10

    The payoff: “Click on videos” — spoken, gated at P=0.78, clicked in 1176 ms

    In the full demo at 15:25, the author is browsing YouTube and simply says — “Click on videos.” The HUD shows what Jev decided: the utterance was ADDRESSED TO BROWSER at P=0.78 (the gate that keeps the agent from obeying speech meant for a human), with action probabilities click_link 0.62, scroll_to_section 0.33, media 0.03, scroll 0.01, none 0.01 — and the browser clicked through to the “Claude Code + Ollama = FULL LOCAL AI AGENT” video in 1176 ms. The closing chapters (17:01) run the same shape through two more demos — a fitting-room flow and a news filter — before the outro. Every one of them is state in, typed decision with a probability out, code acting on it.

    Voice command HUD over YouTube reading Click on videos addressed to browser at P 0.78 with click_link probability 0.62 and a confirmation the browser clicked in 1176 ms
    The gate, the action distribution, and the click — one spoken sentence, end to end.Watch at 15:55

Frequently asked questions

What is the Jev Claude Code integration?

The official path shown in the video is the TypeSafe agent skill: install it into Claude Code with claude plugin marketplace add typesafe-ai/skills followed by claude plugin install typesafe@typesafe-ai, and your coding agent gains the three Jev question types, architectural patterns, and evaluation best practices — everything it needs to wire typed decisions into whatever it builds, which in this video is a voice-controlled browser agent.

How fast and cheap are Jev decisions?

In the video’s terminal race, 28 typed answers came back in 0.114 s at a cost of $0.000081 while an autoregressive LLM was still waiting for its first token. The pricing slide lists $0.042 per million input tokens with output free — 238x lower input price than Claude Fable 5.1 as charted on screen. All figures are as shown in the video; check typesafe.ai for current pricing.

Why not just let the LLM choose the browser action?

Because a text model puts you back in the Prompt / Parse / Pray / Babysit loop: no calibrated confidence, retries on malformed JSON, and a human babysitting every run. Jev returns a typed Command with a probability distribution and confidence, so ordinary if statements act on it — and a noul gate like ADDRESSED TO BROWSER at P=0.78 stops the agent from executing speech that was meant for a person.

What does P=0.78 in the voice HUD mean?

It is the model’s probability that the spoken sentence was addressed to the browser rather than to another person in the room — the gate checked before any action probabilities (click_link 0.62, scroll_to_section 0.33 in the demo) are considered. The numbers are demo values from the video, but the pattern is the guardrail every voice agent needs: below your threshold, do nothing.

Related guides

More video walkthroughs