Guides / 動画ウォークスルー
TypeSafe Computer Use: Local Desktop Automation with Jev, Step by Step
A sketch-by-sketch guide to Alex Hitt’s 6-minute walkthrough of the typesafe-computer-use repo: Vision OCR plus AX tree parsing on-device, Jev classifying the next action for $0.0002 a step, dual API keys, --act dry runs, and hardware-level emergency stops.
要点
Alex Hitt’s 6-minute whiteboard walkthrough builds the case for the typesafe-computer-use repository: autonomous desktop control used to mean shipping multi-megabyte screenshots to a large multimodal model — more than 5 seconds and up to 32 cents per click. The repo flips the architecture: the desktop state is analyzed locally (Apple Vision OCR plus the accessibility tree on macOS 14+, Python 3.12+, installed with uv), compiled into a hard list of candidate actions with coordinates, and Jev is queried for pure classification of the next action — a decision that costs $0.0002 and completes in under a second. Setup has rules. The dual API configuration separates concerns: your TypeSafe API key handles structural logic and action selection, while the Anthropic key is used solely for text generation such as filling in search fields — and every free-text entry is immediately verified by a Jev boolean proposition, with the action reversed and the field cleared if validation fails, which is what prevents the model from choosing coordinates or elements that do not exist in the local state. macOS automation permissions must be granted explicitly or synthetic clicks are quietly dismissed, and the terminal-cleaning mandate is real: the agent OCRs your whole screen, so leftover automation logs become clickable targets and trap it in a recursive loop. Run the clicker diagnostic first (it prints the intended action sequence without sending hardware events), then add --act to move from read-only planning to execution — and relinquish the mouse, because fighting the agent corrupts its coordinate mapping. Two emergency stops exist: Ctrl-C if the terminal keeps focus, or shoving the physical mouse to the upper-left corner, which breaches a monitored coordinate boundary and terminates the thread instantly. Every run writes a flight-data-recorder folder — JSON states, screen bitmaps, per-phase timings — so you debug from annotated bounding-box screenshots without spending API credits. Known edges: two identical on-screen labels split Jev’s probability (0.4/0.4 against a 0.5 safety threshold) and the loop auto-stops, the agent is stateless between runs, and instructions like “click the submit button in the upper-right corner” beat “click submit” every time.
ステップごとのウォークスルー
- 1
The old economics: 5-second screenshots at 32 cents a click
Autonomous desktop control has been the most expensive way to use a computer: the agent screenshots the whole display (several megabytes), uploads it to a large multimodal model, and waits for generated reasoning — more than 5 seconds and up to 32 cents for a single click. The video’s balance-scale frame makes the contrast concrete before any code appears: raw screenshots on the heavy side of the scale, a sub-second local state on the other.

Shipping pixels to a reasoning model is the slow, expensive side of the scale.タイムスタンプ 0:32 を見る - 2
The new economics: $0.0002 per decision, under a second
typesafe-computer-use abandons deliberative thought chains for a System One decision model. The desktop state is analyzed on your machine, a candidate list is compiled locally, and Jev answers a pure classification question about the next action. The headline numbers: execution cost drops to fractions of a cent — $0.0002 per step — and decision latency falls under one second, moving desktop automation out of research-demo territory into per-step pricing you can actually run a loop on.

One classification per step — the price of a click falls four orders of magnitude.タイムスタンプ 0:48 を見る - 3
The perception stack: Native OCR, Jev classification, AX tree parsing
Instead of sending raw pixel data anywhere, the framework reads the screen locally: Apple Vision OCR (the reason for the macOS 14+ requirement), the accessibility tree parser, and CoreGraphics — wired together with the TypeSafe SDK and Python 3.12+ via uv. Jev’s role in the stack is exactly one box: classify the next action from the parsed local state. No pixels leave the machine for the decision itself.

Perception is local; only the classification question travels to Jev.タイムスタンプ 0:42 を見る - 4
The candidate list: coordinates and targets, compiled up front
The local analysis produces a hard list of candidate actions — an explicit table of X/Y coordinates mapped to named targets like Form Field 1, Button A, Dropdown C — alongside the parsed UI. That structural constraint is the safety rail: the model cannot choose coordinates or select interface elements that do not exist in the local state, because its choices are constrained to the compiled list rather than invented from a screenshot.

The agent picks from a compiled menu of real targets — hallucinated coordinates are structurally impossible.タイムスタンプ 2:02 を見る - 5
System hygiene: clear your terminal before every run
The video’s non-obvious operational rule. The agent captures the visual state of your entire screen — so if old automation logs are still sitting in the terminal, the OCR engine detects the agent’s own previous output as clickable UI, and the agent gets stuck trying to interact with its own history. The mandate: physically clear the active terminal window before starting any execution. The SYSTEM HYGIENE frame is the reminder worth pinning next to your runbook.

Your agent reads the whole screen — yesterday’s logs are today’s phantom buttons.タイムスタンプ 2:45 を見る - 6
Dry-run first, then --act for real execution
Verification comes before autonomy. Running the clicker command with a natural-language target prints the intended sequence of actions without sending hardware events — a read-only diagnostic that shows exactly what the agent plans to do. Only when the plan looks right do you add the --act flag, which moves the script from planning to execution. From that moment you relinquish physical control: moving the mouse or typing mid-loop creates focus conflicts that corrupt the coordinate mapping and derail the run.

Print the plan first; add --act only when the sequence reads correctly.タイムスタンプ 3:38 を見る - 7
Two emergency stops: Ctrl-C and the upper-left corner
Delegating an operating system to an agent is only viable if you can take control back instantly. If your terminal retains focus, a standard Ctrl-C interrupt safely stops the Python loop. If focus is elsewhere, push the physical mouse firmly toward the upper-left corner of the screen — the framework continuously monitors coordinate boundaries and terminates the execution thread the moment the cursor breaches them. Anchoring the kill switch to a physical gesture means software errors cannot block your override.

Ctrl-C for the terminal, the corner shove for everything else — both stop the loop instantly.タイムスタンプ 4:28 を見る - 8
Debug from the flight recorder, not the API
Analysis failures in borderline cases are inevitable, so every run writes a timestamped local folder: JSON states, screen bitmaps, and per-phase time metrics. Developers review screenshots annotated with numeric bounding boxes to confirm which elements Vision OCR and the accessibility tree actually detected — the AX Node highlight in the video — replaying failures offline without consuming a single API credit.

Every run records what the perception layer saw — failures become replayable, not mysterious.タイムスタンプ 5:02 を見る - 9
Probabilistic collisions stop the loop for you
When an interface presents two identical labels, Jev evaluates both valid options and splits the probability evenly — 0.4 and 0.4 in the video’s example, under the 0.5 safety threshold. The vote ties, confidence collapses, and the loop automatically stops rather than guessing. The fix is on you: write highly explicit natural-language targets (“click the submit button in the upper-right corner”, never just “click submit”), remember the agent is stateless between runs, and give every run deterministic rules, strict probability limits, and extreme environmental clarity.

A 0.4/0.4 tie against a 0.5 threshold halts the run — ambiguity is a stop sign, not a coin flip.タイムスタンプ 5:30 を見る
よくある質問(FAQ)
What is typesafe-computer-use?
An open repository for local desktop automation built on TypeSafe’s Jev. Instead of sending multi-megabyte screenshots to a multimodal model, it parses the screen locally with Apple Vision OCR and the accessibility tree, compiles a hard list of candidate actions with coordinates, and asks Jev to classify the next action — about $0.0002 and under one second per decision.
What are the system requirements?
macOS 14 or later (for Apple Vision OCR), Python 3.12 or later, and the uv package manager to install dependencies. You also need to grant macOS automation/accessibility permissions explicitly — without them, the system’s synthetic clicks are quietly dismissed and nothing appears to happen.
Why does it need both a TypeSafe API key and an Anthropic key?
The dual API configuration separates structural logic from language: the TypeSafe key powers Jev’s action selection and verification, while the Anthropic key is used solely for text generation such as filling in search fields. Every free-text entry is verified by a Jev boolean proposition afterward — if validation fails, the action is reversed and the field cleared.
Why does my computer-use agent click on its own logs?
The agent OCRs your entire screen, so leftover automation output in a terminal window is detected as clickable interface text — the recursive-loop failure the video demonstrates. The fix is the terminal-cleaning mandate: clear the active terminal before starting any execution run.
How do I stop the agent in an emergency?
Two ways. If the terminal window keeps focus, Ctrl-C interrupts the Python loop safely. Otherwise, push the physical mouse firmly toward the upper-left corner of the screen — the framework watches coordinate boundaries and terminates the execution thread immediately on the breach. The hardware-anchored stop works even when the software is misbehaving.
How is this different from a browser agent?
Scope and layer. A browser agent automates one application over the network; computer use parses the operating system’s full visual state and synthesizes real clicks and keystrokes. The trade-offs match: OS-level reach comes with permissions, hygiene, focus, and safety-threshold rules that browser agents never face.
関連ガイド
CUA-S1-Forms: Form Filling Without Screenshots
The form-specialized sibling: a 2.8MB scorer and the accessibility tree fill real GUI forms while the GPU stays cold.
読むJev Browser Agent Guide
The browser-scoped sibling: speeding up web agents with Jev decision points instead of full-page model calls.
読むJev MCP Server Guide
Give coding agents typed decision tools over the Model Context Protocol — the keyboard-and-terminal side of agent control.
読むWhat Is Jev? System One Models Explained
Why a non-generative classifier is the right engine for per-step UI decisions.
読むその他の動画ウォークスルー
- Jev Classification Quickstart: OpenRouter API, Primitives & Real Probabilities
- Jev Architecture Explained: Why 70ms Decision Models Beat LLMs for Workflow Automation
- Ultra-Fast Browser Agents with Jev: 178ms DOM Loops & Dual Model Orchestration
- Open Jev Models Are Here: Semif, Nimble, Decider, DiffusionGemma & Laya Hands-On
- Jev vs LLM: Will Jev Replace LLMs? Krish Naik's Whiteboard Explainer
- Jev Trader Tutorial: Build a Subsecond AI Trading Bot on Monad
- Jev Model Router: Build a Privacy-Gated LLM Router with Jev & OpenJev
- Jev Tutorial for Beginners: State, Questions & the TypeScript SDK
- Run Jev Locally: Kev, SemIf & Von on Your Own GPU (OpenJev Guide)
- Jev RAG Reranker: Policy-Steered Reranking for Retrieval-Augmented Generation
- When to Use Jev: An Engineer's Audit of Claims, Gates, and Failure Modes
- LangChain + Jev Integration Tutorial: Routing, Guardrails & Evals
- Jev MCP Server: Connect Jev Decisions to Claude Code & Cursor
- Jev vs Luna: Independent Benchmarks Put "Better, Faster, Cheaper" to the Test
- Jev Agent Harness: Where the Decision Gate Sits in Your LLM Loop
- Jev Playground Walkthrough: The Hotdog Lesson, Criteria, and a Four-Console Token Test
- Jev Text Classification API: Zero-Shot CLI & REST with classifier.dev
- Jev API Examples: First Request, curl & All Three Question Types
- Expanso Edge と Jev でログをトリアージする
- Treg と Jev でリードをエンリッチする: ICP 判定と登録スコアリング
- Heym で Jev Decision Node を使う: モデルルーティングを制御する
- Laya Tutorial: Open-Source AI Routing With Calibrated Probabilities (Laya vs Jev Setup)
- Train Your Own Jev: Fine-Tune a Jev-Style Decision Model for $5–$17 (What You Can and Cannot Train)
- Jev Tips: 8 Best Practices for Better Decisions (State, Questions, Criteria & Thresholds)
- Jev Context Compaction: Prune AI Agent Memory Without Generative Summaries
- Jev as an LLM Judge: Confidence-Gated Cascades at 0.36% of the Cost
- Jev Resume Screening: Build an AI Resume Evaluator with the Jev JavaScript SDK
- Jev + Claude Code Guide: Voice-Controlled Browser Automation with Typed Decisions
- Jev + Codex: Install the TypeSafe Skill and Triage a Real Gmail Inbox
- Jev vs Ollama: Can Local AI Replace Hosted Jev Without Sending Your Data Away?
- Build Your Own Jev: Train a Free Open-Source Zero-Shot Classifier (That Plays Doom)
- CUA-S1-Forms: a 706K-Parameter Jev-Like Model That Fills GUI Forms on Your CPU
- NOC/SOC Alert Triage with Jev: Rules First, One Typed Question, a Policy Gate
- Ollama Decision Models: Run tev1 and Nimble Locally (Tested on an 8 GB Card)
- Jev Guardrails in Production: A Five-Step Playbook for Decision Automation
- A Session Drift Guard for Pi Agent: Let the Jev Model Propose, Let Code Decide
- 50 Tev1 Use Cases: What a Local Decision Model Can Actually Do (Tested on an 8 GB Laptop)
- Jev, Hands-On: Where the Official Claims Meet Independent Remeasurement
- Jev Ticket Classification in a Real App: the After-Insert Hook and the Calculated Field
- OpenJev RLCD: Run the Open-Source Calibrated Decision Model Locally (Full Guide)
- Clef-Flash vs Jev: I Tested Cloudflare's Decision Model Locally in Ollama (Q4_K_M)
- Julia-1 Tutorial: Install the Open-Source Jev Replacement in Pure Python (and Watch It Beat If-Statements 9 to 2)
- OpenJev 0.8B on CPU: I Built a Ticket-Routing Inbox and Calibrated the Thresholds
- Jev n8n Integration: the JevGate Community Node, Step by Step (Plus a Plain-HTTP Fallback)