Guides / illustrated walkthrough

Build Your Own Jev: Train a Free Open-Source Zero-Shot Classifier (That Plays Doom)

A step-by-step breakdown of Micah’s 4:54 build: distill a 45×-larger teacher into a Qwen3 0.6B student on 1,000+ Doom situations in a two-minute training run, watch agreement climb from ~50 to 93 out of 100 on unseen rounds, and gate automation on the confidence scores — down to 12 decisions per second on one GPU.

Quick takeaway

Typesafe’s Jev plays Doom by receiving a situation plus three options and answering instantly — no text generated. Jev itself is closed and waitlisted, but the same shape is reproducible at home: Micah’s open-source openjev pipeline distills a teacher model 45× larger into a Qwen3 0.6B student over 1,000+ recorded Doom situations, training in about two minutes on one GPU. The student that once claimed 95% confidence while agreeing with the teacher only half the time now picks the teacher’s answer 93 times out of 100 on rounds it never saw — and confidence scores become the automation gate: act on confident answers, hand doubtful ones to a bigger model or a human. The teacher needs about 9 seconds of thinking per question; the student answers in roughly a hundredth of a second (~900× faster) and makes 12 moves per second in-game. The honest limits: it only knows the one Doom room it trained on, and open copies score a few points below Jev on TypeSafe’s own published test cases.

Video source

Micah

4:54JsVM4vgspEU

Step-by-step walkthrough

  1. 1

    See the shape first: typed answers all at once, not a chat stream

    The video opens on a split terminal. On the TYPE side, every answer and probability arrives in one response — there is no generation to wait for. On the LLM side, the client sits "waiting for first token" before a single word of output can be parsed. That difference is the whole premise: a decision model receives a situation and a list of answers and returns a percentage for each, which is why a Doom player can act twelve times a second while a chat loop cannot.

    Split terminal comparing a TYPE panel that returns every answer at once with an LLM panel still waiting for its first token
    Decisions come back complete in one response — no tokens, no stream.Watch at 0:10
  2. 2

    Feed the model state plus a strategy written in plain English

    What the model reads is not a prompt — it is a state blob. The green panel shows the Doom situation as data: enemy types, weapon effective range, the player’s body measured in map units, movement context, and a strategy line written in plain English ("point blind ambushing, slow reload" — aim before the target is in view, hold fire until it is close enough to hit). Alongside the state go three questions at once: which enemy to aim at, whether to hold the trigger, whether to dodge. The model scores every option of every question in a single pass.

    openjev state blob listing enemy types, effective range, movement context, and a plain-English strategy line reading point blind ambushing, slow reload
    State is data and strategy is a sentence — both go in together with the questions.Watch at 1:38
  3. 3

    Teach the student with labeled percentages, not prompts

    A small model answers fast but badly — the untrained version never dodged a fireball and died in half a minute. The fix is distillation: a teacher 45× larger answers every question across 1,000+ recorded Doom situations with a confidence percentage on each option, and the student is trained to reproduce those percentages. The training-record terminal shows what one labeled example looks like: a question, the chosen label, and the teacher’s score. The whole run takes about two minutes on one GPU, and everything down to the Doom copy is freely downloadable.

    openjev training record terminal showing a labeled Doom question with the teacher’s 91 percent hold-fire score used to train the student model
    The teacher’s probabilities are the labels — two minutes of training, all free to download.Watch at 1:43
  4. 4

    Start from the open-source copy on GitHub

    Jev itself is closed — its design is deliberately kept secret, and TypeSafe’s CEO only confirmed online that calling it a "zero-shot classifier" is "absolutely true." But the shape is reproducible: within two days of Jev’s launch, open-source copies appeared on GitHub. The one in the video, TheoLeeCJ/openjev, opens with the question the whole video answers — "Can we run something like Jev on a 3060 at home?" — and ships the pipeline, benchmarks, and demo configs.

    GitHub repository page for TheoLeeCJ openjev titled can we run something like Jev on a 3060 at home
    openjev is a third-party open-source replica, not an official TypeSafe product.Watch at 1:45
  5. 5

    Calibrate before you trust the confidence numbers

    The video’s most useful minute is about honesty of scores. Before training, the student claimed to be 95% sure while agreeing with the teacher only half the time — the slide crosses that claim out next to a dartboard. After distillation, when the student says it is confident, it agrees with the teacher 99 times out of 100, and on rounds it never saw during training it picks the teacher’s answer 93 times out of 100 (up from about half). That calibration is what makes confidence scores usable as an automation gate: act on confident answers, and pass doubtful ones to a larger model or a human.

    Whiteboard slide crossing out the claim that a model 95 percent sure is right half the time beside a dartboard calibration drawing
    Confidence means nothing until agreement is measured — then it becomes the gate.Watch at 3:20
  6. 6

    Price the teacher honestly: 9 seconds per question

    The speed claim has a fine print worth reading. When the teacher — the model 45× larger — is allowed to actually think about each question, it takes about 9 seconds to answer. The distilled student answers in roughly a hundredth of a second on the same map: almost 900× faster, at 12 moves per second in-game. That ratio is the economic argument for the whole pattern — pay the teacher once, offline, to label situations; let the student answer forever, online, in milliseconds.

    Slide timing the teacher model at 9 seconds of thinking per question before the distilled student takes over the answers
    The teacher thinks for 9 seconds so the student never has to.Watch at 4:22
  7. 7

    Watch it survive: 316ms per decision, 30 kills

    The closing run puts numbers on the glass: a Qwen3 0.6B trained on a bigger model’s answers plays the room at 316 milliseconds per decision across 453 decisions with 30 kills on the board. Given a new strategy — "don’t shoot, just dodge" — it stopped shooting and started sidestepping with no retraining at all, because the strategy is input, not weights. The honest edges: this student only knows the one room it trained on, and the best open copy still scores a few points below Jev on TypeSafe’s own published test cases.

    Doom run overlay reading Qwen3 0.6B trained on a bigger model answers at 316 milliseconds per decision across 453 decisions with 30 kills
    12 moves per second on one GPU — inside the room it was trained for.Watch at 4:28

Frequently asked questions

Is openjev the same thing as Jev?

No — and the video is careful about this. Jev is Typesafe AI’s closed, hosted decision model; its architecture is deliberately secret. openjev is a third-party open-source project that reproduces the usage shape: state plus typed questions in, calibrated probabilities per option out, in a single forward pass. TypeSafe’s CEO confirmed publicly that "zero-shot classifier" accurately describes Jev, which is exactly the category openjev operates in. The best open copy scores a few points below Jev on TypeSafe’s own published test cases.

How much does it cost to build your own Jev-style model?

In the video: nothing but electricity and one GPU you already own. The teacher is a free larger model, the training set is 1,000+ recorded Doom situations, the training run takes about two minutes, and everything — including the Doom copy — is freely downloadable. The video’s title claim is that you can build it in an evening. Your real cost is labeling situations for your own domain, which is the same data work any distilled classifier needs.

Why distill from a teacher instead of prompting a small model?

Because the untrained small model is fast but wrong — it never dodged the fireball and died in half a minute. Prompting cannot install judgment; distillation can. The teacher answers each of 1,000+ situations once, offline, with calibrated percentages; the student learns to reproduce those percentages so well that on unseen rounds it matches the teacher 93 times out of 100. You pay the 9-seconds-per-question teacher cost once, then serve hundredth-of-a-second answers forever.

What do the confidence scores actually gate?

Automation. After calibration, a confident student answer is almost always the teacher’s answer, so the program can act on it immediately; a doubtful one gets escalated to the larger model or a human. That is the same confidence-gated pattern used with hosted Jev — thresholds only work because the scores were checked against ground truth, which is why the video spends a full minute on the "95% sure, right half the time" trap.

Can the student handle situations it was never trained on?

Barely — and the video says so plainly. The student only knows the one Doom room it trained on, and zero-shot means it scores answers you write at ask-time, not that it understands new environments. Changing the strategy ("don’t shoot, just dodge") works with no retraining because strategy is input. Changing the room, the game, or the domain means recording new situations and re-distilling. Jev is sold as a solution for situations no one prepared it for; your home-built copy is honestly a specialist.

What hardware do I need to follow along?

One consumer GPU — the repo’s title question is "can we run something like Jev on a 3060 at home?", and the video’s answer is yes: an open copy on a six-year-old card answers 21 questions per second, and the trained student makes 12 moves per second in-game. Training the student on 1,000+ teacher-labeled situations takes about two minutes on the same card.

Related guides

More video walkthroughs