Guides / illustrated walkthrough

TypeSafe Computer Use: Local Desktop Automation with Jev, Step by Step

A sketch-by-sketch guide to Alex Hitt’s 6-minute walkthrough of the typesafe-computer-use repo: Vision OCR plus AX tree parsing on-device, Jev classifying the next action for $0.0002 a step, dual API keys, --act dry runs, and hardware-level emergency stops.

Quick takeaway

Alex Hitt’s 6-minute whiteboard walkthrough builds the case for the typesafe-computer-use repository: autonomous desktop control used to mean shipping multi-megabyte screenshots to a large multimodal model — more than 5 seconds and up to 32 cents per click. The repo flips the architecture: the desktop state is analyzed locally (Apple Vision OCR plus the accessibility tree on macOS 14+, Python 3.12+, installed with uv), compiled into a hard list of candidate actions with coordinates, and Jev is queried for pure classification of the next action — a decision that costs $0.0002 and completes in under a second. Setup has rules. The dual API configuration separates concerns: your TypeSafe API key handles structural logic and action selection, while the Anthropic key is used solely for text generation such as filling in search fields — and every free-text entry is immediately verified by a Jev boolean proposition, with the action reversed and the field cleared if validation fails, which is what prevents the model from choosing coordinates or elements that do not exist in the local state. macOS automation permissions must be granted explicitly or synthetic clicks are quietly dismissed, and the terminal-cleaning mandate is real: the agent OCRs your whole screen, so leftover automation logs become clickable targets and trap it in a recursive loop. Run the clicker diagnostic first (it prints the intended action sequence without sending hardware events), then add --act to move from read-only planning to execution — and relinquish the mouse, because fighting the agent corrupts its coordinate mapping. Two emergency stops exist: Ctrl-C if the terminal keeps focus, or shoving the physical mouse to the upper-left corner, which breaches a monitored coordinate boundary and terminates the thread instantly. Every run writes a flight-data-recorder folder — JSON states, screen bitmaps, per-phase timings — so you debug from annotated bounding-box screenshots without spending API credits. Known edges: two identical on-screen labels split Jev’s probability (0.4/0.4 against a 0.5 safety threshold) and the loop auto-stops, the agent is stateless between runs, and instructions like “click the submit button in the upper-right corner” beat “click submit” every time.

Video source

Alex Hitt

6:20fXAJRNfOUb4

Step-by-step walkthrough

  1. 1

    The old economics: 5-second screenshots at 32 cents a click

    Autonomous desktop control has been the most expensive way to use a computer: the agent screenshots the whole display (several megabytes), uploads it to a large multimodal model, and waits for generated reasoning — more than 5 seconds and up to 32 cents for a single click. The video’s balance-scale frame makes the contrast concrete before any code appears: raw screenshots on the heavy side of the scale, a sub-second local state on the other.

    Whiteboard balance scale comparing raw screenshots at 5.0 seconds or more against local state under 1 second for desktop agent perception.
    Shipping pixels to a reasoning model is the slow, expensive side of the scale.Watch at 0:32
  2. 2

    The new economics: $0.0002 per decision, under a second

    typesafe-computer-use abandons deliberative thought chains for a System One decision model. The desktop state is analyzed on your machine, a candidate list is compiled locally, and Jev answers a pure classification question about the next action. The headline numbers: execution cost drops to fractions of a cent — $0.0002 per step — and decision latency falls under one second, moving desktop automation out of research-demo territory into per-step pricing you can actually run a loop on.

    Hand-drawn headline reading $0.0002 and under one second, the per-decision cost and latency of Jev classification in the typesafe-computer-use walkthrough.
    One classification per step — the price of a click falls four orders of magnitude.Watch at 0:48
  3. 3

    The perception stack: Native OCR, Jev classification, AX tree parsing

    Instead of sending raw pixel data anywhere, the framework reads the screen locally: Apple Vision OCR (the reason for the macOS 14+ requirement), the accessibility tree parser, and CoreGraphics — wired together with the TypeSafe SDK and Python 3.12+ via uv. Jev’s role in the stack is exactly one box: classify the next action from the parsed local state. No pixels leave the machine for the decision itself.

    Three labeled stages of the local perception pipeline: Native OCR, Jev Classification, and AX Tree Parsing connected along a pipeline line.
    Perception is local; only the classification question travels to Jev.Watch at 0:42
  4. 4

    The candidate list: coordinates and targets, compiled up front

    The local analysis produces a hard list of candidate actions — an explicit table of X/Y coordinates mapped to named targets like Form Field 1, Button A, Dropdown C — alongside the parsed UI. That structural constraint is the safety rail: the model cannot choose coordinates or select interface elements that do not exist in the local state, because its choices are constrained to the compiled list rather than invented from a screenshot.

    Layout coordinates and targets list mapping X/Y positions to form fields, buttons, and dropdowns next to a web form sketch with one field highlighted as the chosen target.
    The agent picks from a compiled menu of real targets — hallucinated coordinates are structurally impossible.Watch at 2:02
  5. 5

    System hygiene: clear your terminal before every run

    The video’s non-obvious operational rule. The agent captures the visual state of your entire screen — so if old automation logs are still sitting in the terminal, the OCR engine detects the agent’s own previous output as clickable UI, and the agent gets stuck trying to interact with its own history. The mandate: physically clear the active terminal window before starting any execution. The SYSTEM HYGIENE frame is the reminder worth pinning next to your runbook.

    System hygiene headline hand-lettered in red marker above a whiteboard grid, the rule to clear terminal logs before starting typesafe-computer-use runs.
    Your agent reads the whole screen — yesterday’s logs are today’s phantom buttons.Watch at 2:45
  6. 6

    Dry-run first, then --act for real execution

    Verification comes before autonomy. Running the clicker command with a natural-language target prints the intended sequence of actions without sending hardware events — a read-only diagnostic that shows exactly what the agent plans to do. Only when the plan looks right do you add the --act flag, which moves the script from planning to execution. From that moment you relinquish physical control: moving the mouse or typing mid-loop creates focus conflicts that corrupt the coordinate mapping and derail the run.

    Terminal prompt sketch with the --act flag highlighted in orange, the switch from read-only diagnostic planning to live execution in typesafe-computer-use.
    Print the plan first; add --act only when the sequence reads correctly.Watch at 3:38
  7. 7

    Two emergency stops: Ctrl-C and the upper-left corner

    Delegating an operating system to an agent is only viable if you can take control back instantly. If your terminal retains focus, a standard Ctrl-C interrupt safely stops the Python loop. If focus is elsewhere, push the physical mouse firmly toward the upper-left corner of the screen — the framework continuously monitors coordinate boundaries and terminates the execution thread the moment the cursor breaches them. Anchoring the kill switch to a physical gesture means software errors cannot block your override.

    Dark keyboard illustration with the Ctrl and C keys highlighted in red, the keyboard interrupt that safely stops the typesafe-computer-use execution loop.
    Ctrl-C for the terminal, the corner shove for everything else — both stop the loop instantly.Watch at 4:28
  8. 8

    Debug from the flight recorder, not the API

    Analysis failures in borderline cases are inevitable, so every run writes a timestamped local folder: JSON states, screen bitmaps, and per-phase time metrics. Developers review screenshots annotated with numeric bounding boxes to confirm which elements Vision OCR and the accessibility tree actually detected — the AX Node highlight in the video — replaying failures offline without consuming a single API credit.

    Laptop sketch covered in numbered red and green bounding boxes with a paper dart, representing Vision OCR and accessibility-tree detections recorded for offline debugging.
    Every run records what the perception layer saw — failures become replayable, not mysterious.Watch at 5:02
  9. 9

    Probabilistic collisions stop the loop for you

    When an interface presents two identical labels, Jev evaluates both valid options and splits the probability evenly — 0.4 and 0.4 in the video’s example, under the 0.5 safety threshold. The vote ties, confidence collapses, and the loop automatically stops rather than guessing. The fix is on you: write highly explicit natural-language targets (“click the submit button in the upper-right corner”, never just “click submit”), remember the agent is stateless between runs, and give every run deterministic rules, strict probability limits, and extreme environmental clarity.

    Two identical Target Button labels each scoring probability 0.4 under a dashed safety threshold line at 0.5, illustrating a Jev probabilistic collision that stops the automation loop.
    A 0.4/0.4 tie against a 0.5 threshold halts the run — ambiguity is a stop sign, not a coin flip.Watch at 5:30

Frequently asked questions

What is typesafe-computer-use?

An open repository for local desktop automation built on TypeSafe’s Jev. Instead of sending multi-megabyte screenshots to a multimodal model, it parses the screen locally with Apple Vision OCR and the accessibility tree, compiles a hard list of candidate actions with coordinates, and asks Jev to classify the next action — about $0.0002 and under one second per decision.

What are the system requirements?

macOS 14 or later (for Apple Vision OCR), Python 3.12 or later, and the uv package manager to install dependencies. You also need to grant macOS automation/accessibility permissions explicitly — without them, the system’s synthetic clicks are quietly dismissed and nothing appears to happen.

Why does it need both a TypeSafe API key and an Anthropic key?

The dual API configuration separates structural logic from language: the TypeSafe key powers Jev’s action selection and verification, while the Anthropic key is used solely for text generation such as filling in search fields. Every free-text entry is verified by a Jev boolean proposition afterward — if validation fails, the action is reversed and the field cleared.

Why does my computer-use agent click on its own logs?

The agent OCRs your entire screen, so leftover automation output in a terminal window is detected as clickable interface text — the recursive-loop failure the video demonstrates. The fix is the terminal-cleaning mandate: clear the active terminal before starting any execution run.

How do I stop the agent in an emergency?

Two ways. If the terminal window keeps focus, Ctrl-C interrupts the Python loop safely. Otherwise, push the physical mouse firmly toward the upper-left corner of the screen — the framework watches coordinate boundaries and terminates the execution thread immediately on the breach. The hardware-anchored stop works even when the software is misbehaving.

How is this different from a browser agent?

Scope and layer. A browser agent automates one application over the network; computer use parses the operating system’s full visual state and synthesizes real clicks and keystrokes. The trade-offs match: OS-level reach comes with permissions, hygiene, focus, and safety-threshold rules that browser agents never face.

Related guides

More video walkthroughs