Guides / illustrated walkthrough
Train Your Own Jev: Fine-Tune a Jev-Style Decision Model for $5–$17 (What You Can and Cannot Train)
A frame-by-frame guide to training your own Jev-style decision model: why the hosted Jev itself has zero fine-tuning, the three levers you actually control, where 38,000 training questions come from, the $4.17 and $17 real-world training bills, Laya fine-tuned from 0.362 to 0.766 accuracy, calibration that survives an audit, and when the whole pipeline pays off.
Quick takeaway
The hosted Jev cannot be fine-tuned at all — TypeSafe serves identical weights to every account and offers no fine-tuning endpoint, so “training your own Jev” really means one of two things: fine-tuning a small open LLM into a Jev-style classifier, or fine-tuning an open decision engine such as Laya. TerraNet’s explainer frames every option as one of three levers: question design (guides both Jev and Laya), recalibration (adjusts confidence so an 80% score means 80% accuracy on your data), and fine-tuning (changes weights — Laya only). Two published replicas prove the economics: Hassan (Nutlope) fine-tuned Qwen3.5 4B on Together AI with about 38,000 questions from eight public datasets in roughly 25 minutes for a $17 training job, and CoderOne ported the same tev1 repo to Modal, training Qwen3.5 4B with Unsloth + LoRA on an L40S for $4.17 in about two hours — roughly 38,000 training examples (~17M tokens), 5,000 dev, and 4,000 untouched staging examples drawn from MultiNLI, BoolQ, and Banking77 plus synthetic policy, routing, and research-taxonomy sets, landing about 88% accuracy on the held-out staging data with warm inference around 600–900 ms end-to-end (about 160 ms of actual model time) served by vLLM. On the open-engine side, TerraNet’s benchmark card shows Laya’s 421M English checkpoint going from 0.362 to 0.766 accuracy across 2,000 separate test decisions — against a teacher model that agreed with itself only 73.5% of the time — and warns that calibration temperatures must be fitted on held-out data only (Laya’s old notebook approach, fitting on training data, left temperatures within 6% of doing nothing). The hard rules: split data three ways before anything else, keep every prompt under 20 options, count test samples per question (300 examples at 80% accuracy carry ±4.5 points of uncertainty), and if your daily volume is modest, stick with an LLM and start logging structured decisions today.
Video source
TerraNet Technologies LLC
Step-by-step walkthrough
- 1
Be clear about what you can and cannot train
The hosted Jev is off the table: TypeSafe serves identical weights to every account and provides no fine-tuning endpoint — TerraNet’s opening slide calls this “Jev: Zero fine-tuning,” while Laya sits on the other side of the frame as “retrainable on a free GPU.” That boundary is why “training your own Jev” always means one of two honest projects. Path one: fine-tune a small open LLM to imitate Jev’s behavior as a classifier — exactly what Hassan’s $17 Together AI tutorial and CoderOne’s $5 Qwen3.5 4B replica did. Path two: fine-tune an open decision engine such as Laya directly, which TerraNet notes trains in minutes on Kaggle’s free pair of T4 GPUs. The slide’s third bullet is the real headline: the bottleneck is not compute — it is gathering valid labels and fitting probabilities you can actually trust. Everything else on this page is about those two problems.

The hosted Jev cannot be fine-tuned — replicas and open engines are the two honest paths.Watch at 0:20 - 2
Pick your lever: question design, recalibration, or weights
TerraNet frames everything you can change as three levers. Question design guides both Jev and Laya: because TypeSafe gives every account the same weights, you shape answers through the request itself — breaking one broad judgment into several narrow questions and writing the policy into your code. Recalibration adjusts confidence so that an 80% score actually matches 80% accuracy on your specific distribution; both vendors train with RLCD, but calibration only holds on data resembling what it was calibrated on. Fine-tuning changes internal weights and is available for Laya only. If your goal is a Jev-style classifier, the fine-tuning lever points at a small open base model: both replicas chose Qwen3.5 4B for its open weights and workable license, and CoderOne’s port kept the training format Jev-shaped from day one — a state, a question, and answer choices in, one letter out.

Two of the three levers work on the hosted Jev; only fine-tuning requires your own model.Watch at 0:50 - 3
Label first, split three ways, then touch nothing
TerraNet’s pipeline figure is the whole discipline in one strip: collect and label, split three ways, fine-tune, fit the confidence, set the threshold — with escalations routed back in as fresh training data. Labels come from past operational decisions, a teacher LLM, or human review. The split must happen before any training: a training set, a calibration set, and a test set you never peek at early. The two replicas show what a real split looks like: Hassan’s tev1 setup trained on about 38,000 questions sampled from eight public datasets, and CoderOne’s port carried roughly 38,000 training examples (about 17M tokens), nearly 5,000 dev examples (about 2M tokens), and a staging set of around 4,000 examples left untouched until after the model shipped. The public sets behind those numbers, as listed in CoderOne’s video: MultiNLI for entailment, BoolQ for yes/no questions, Banking77’s 77 banking support intents, plus synthetic decision datasets — pop for policies, policy V2, routing V2, and a research taxonomy that classifies articles by primary contribution rather than topic.

Label, split three ways, fine-tune, calibrate, set the threshold — then escalations become tomorrow’s labels.Watch at 2:25 - 4
Run the fine-tune — and read the benchmark honestly
The training itself is now the easy part. CoderOne ran Unsloth with LoRA on official Qwen3.5 4B inside a CUDA container, did a one-minute smoke test on a data subset, then let the full run cook for about two hours on a Modal L40S — the training line item came to $4.17 (an H100 would cost more but maybe finish in 30 minutes, he estimates). Hassan’s Together AI version finished in roughly 25 minutes for a $17 training job. On the open-engine side, TerraNet’s benchmark card reports Laya’s 421M English checkpoint going from 0.362 base accuracy to 0.766 after fine-tuning across 2,000 separate test decisions, trained on 1,200 cases spanning four workflows. The catch sits on the same slide: the benchmark’s ground truth came from a teacher model that agreed with itself only 73.5% of the time, and Laya matched the noisy teacher more often than the teacher matched itself — evidence it learned the task pattern rather than the noise. CoderOne’s replica scored about 88% on its untouched staging set. One more boundary from the video: the replica is strongest at picking one option from a list (its demo also pushed a boolean-style true/false check through in about 600 ms), while Jev’s richer typed primitives would need additional labeled datasets before they can be trained properly.

0.362 to 0.766 looks great until you notice the teacher agreed with itself only 73.5% of the time.Watch at 1:50 - 5
Calibrate on held-out data only, then audit the evaluation
A raw fine-tuned score is not a probability you can gate code on. TerraNet’s recalibration lever means fitting temperature scaling or Platt scaling on the held-out calibration set — never on training data. The cautionary tale from the video: Laya’s own notebook used to fit calibration temperatures directly on a slice of training data, leaving temperatures barely moved — within 6% of doing nothing — and creating a false sense of calibrated certainty. Then audit the evaluation itself. The most frequent failure is measuring on the benchmark you trained on, or on a test set too small to yield statistical power: at 80% accuracy across 300 test examples, uncertainty spans about ±4.5 percentage points — too coarse to tell two competing fine-tunes apart — so count test examples per question rather than per case. Finally, respect the option budget: accuracy degrades when a single prompt presents more than 20 options, so split wide taxonomies into hierarchical decisions in your code instead of fine-tuning on dozens of simultaneous branches.

Three evaluation traps: training-set evals, underpowered test sets, and oversized option lists.Watch at 3:10 - 6
Deploy where latency pays you back — and start logging either way
The work is front-loaded, so it only pays at volume: TerraNet puts the break-even at thousands of daily queries, where LLM calls create noticeable latency and compute costs. CoderOne’s serving numbers show what a deployed replica looks like: vLLM behind Modal on both an L40S and an H100, where the small 4B model actually ran better on the cheaper L40S — about 160 ms of model time per decision (14 ms to build the input plus 145 ms of execution), 600–900 ms end-to-end when warm, and cold starts of roughly 90 seconds whenever the standby GPU woke up. His whole experiment — training plus smoke tests plus heavy inference — consumed about $10 of Modal’s $30 signup credit. His closing comparison table is the cleanest rule of thumb: a pretrained LLM like Qwen3.5 stays flexible, explains its answers, and carries a roughly 256K-token context, but costs more to serve and is not accurate enough for structured decisions; a decision model (Jev or Laya) returns typed decisions and probabilities in one fast pass; a fixed BERT-style classifier wins outright when you need the same labels every time. And if your volume is modest today, TerraNet’s parting advice costs nothing: record decisions alongside inputs in a structured format now, so the training data is ready the day you decide to train.

Break-even sits at thousands of daily queries — below that, log decisions and wait.Watch at 3:50
Frequently asked questions
Can you fine-tune the Jev model itself?
No. TypeSafe serves identical weights to every account and provides no fine-tuning endpoint for the hosted Jev — TerraNet’s explainer summarizes it as “Jev: zero fine-tuning.” Your only levers on the hosted product are question design and recalibration. “Training your own Jev” therefore means training a Jev-style classifier of your own (a fine-tuned small LLM such as Qwen3.5 4B) or fine-tuning an open decision engine such as Laya. Nothing on this page turns a self-trained replica into the official Jev, and none of the source projects claim parity with it.
How much does it cost to train your own Jev-style model?
Two published data points: Hassan trained Qwen3.5 4B on Together AI for a $17 training job (roughly 25 minutes, about 38,000 questions), and CoderOne trained the same class of model on a Modal L40S for $4.17 in about two hours — about $10 total once he included smoke tests, validation, and a lot of inference queries, all inside Modal’s $30 signup credit. TerraNet adds the floor: Laya fine-tunes in minutes on Kaggle’s free pair of T4 GPUs. Serving is the ongoing cost — either an always-on GPU or a standby mode that trades cheaper idle time for slow cold starts.
Where does the training data come from?
Three sources, per TerraNet’s pipeline: past operational decisions, a teacher LLM, or human review. Both replicas started from the open tev1 repo, whose roughly 38,000 training questions draw on eight public datasets — MultiNLI (entailment), BoolQ (yes/no), Banking77 (77 banking support intents), plus synthetic decision sets covering policies, routing, and a research taxonomy. Every example is formatted Jev-style: a state, a question, and answer choices, with the model learning to emit a single letter. Critically, the calibration and test sets are held back before any training happens.
Is a self-trained model as good as the official Jev API?
Treat them as different tiers. CoderOne’s Qwen3.5 4B replica scored about 88% on a staging set it had never seen and answered in under a second when warm — useful, but it covers the multiple-choice core rather than Jev’s full typed vocabulary, and no one has shown it matching the hosted reference quality. The benchmark numbers carry their own noise: TerraNet’s Laya fine-tune hit 0.766 accuracy against a teacher that only agreed with itself 73.5% of the time. If you need guaranteed quality, use the hosted Jev API and keep a self-trained model as a cheap fallback or a data-stays-in-house deployment.
What is calibration, and why do you need a separate calibration set?
Calibration adjusts the model’s confidence so that an 80% score really is right about 80% of the time on your traffic. The practical technique is temperature scaling or Platt scaling fitted on a held-out calibration set — fitting those parameters on training data barely moves them: TerraNet points out that Laya’s old notebook approach left temperatures within 6% of doing nothing while creating a false sense of calibrated certainty. Both vendors train with RLCD for calibrated decisions, but calibration only holds on data resembling the calibration distribution — so re-fit when your traffic drifts, then pick an operating threshold and route low-confidence escalations to an LLM or a human.
Should you fine-tune, or just use an LLM with prompts?
TerraNet’s break-even rule: the front-loaded engineering only pays off at thousands of daily queries, where LLM calls create noticeable latency and compute costs. CoderOne’s closing table narrows it further — pick the pretrained LLM (Qwen3.5, roughly 256K context, explains its answers) when you need flexibility and explanations; pick a decision model (Jev or Laya) for real-time typed decisions with low latency; pick a fixed BERT-style classifier when you need the same fixed labels every time. If your volume is modest today, stay on the LLM but start logging decisions alongside inputs in structured formats so the dataset is ready when you are.
Related guides
When to Use Jev: Decision Model vs Plain LLM
The decision framework behind the final step — when typed decisions beat prompting a general LLM.
ReadRun Jev Locally: Kev, SemIf & Von on Your Own GPU
Self-hosting playbooks for open decision routers — the deployment half of this pipeline.
ReadJev API Examples
The hosted reference behavior your Jev-style replica imitates, request by request.
ReadOpen Jev Alternatives Catalog
Where Laya and the other fine-tunable engines sit in the open decision-engine landscape.
ReadJev Benchmarks
How decision models are scored — the context behind numbers like 0.362 to 0.766.
ReadMore video walkthroughs
- Jev Classification Quickstart: OpenRouter API, Primitives & Real Probabilities
- Jev Architecture Explained: Why 70ms Decision Models Beat LLMs for Workflow Automation
- Ultra-Fast Browser Agents with Jev: 178ms DOM Loops & Dual Model Orchestration
- Open Jev Models Are Here: Semif, Nimble, Decider, DiffusionGemma & Laya Hands-On
- Jev vs LLM: Will Jev Replace LLMs? Krish Naik's Whiteboard Explainer
- Jev Trader Tutorial: Build a Subsecond AI Trading Bot on Monad
- Jev Model Router: Build a Privacy-Gated LLM Router with Jev & OpenJev
- Jev Tutorial for Beginners: State, Questions & the TypeScript SDK
- Run Jev Locally: Kev, SemIf & Von on Your Own GPU (OpenJev Guide)
- Jev RAG Reranker: Policy-Steered Reranking for Retrieval-Augmented Generation
- When to Use Jev: An Engineer's Audit of Claims, Gates, and Failure Modes
- LangChain + Jev Integration Tutorial: Routing, Guardrails & Evals
- Jev MCP Server: Connect Jev Decisions to Claude Code & Cursor
- Jev vs Luna: Independent Benchmarks Put "Better, Faster, Cheaper" to the Test
- Jev Agent Harness: Where the Decision Gate Sits in Your LLM Loop
- Jev Playground Walkthrough: The Hotdog Lesson, Criteria, and a Four-Console Token Test
- Jev Text Classification API: Zero-Shot CLI & REST with classifier.dev
- Jev API Examples: First Request, curl & All Three Question Types
- Jev Log Triage with Expanso Edge
- Jev Lead Enrichment with Treg: ICP and Signup Scoring
- Use Jev Decision Nodes in Heym for Model Routing
- Laya Tutorial: Open-Source AI Routing With Calibrated Probabilities (Laya vs Jev Setup)
- Jev Tips: 8 Best Practices for Better Decisions (State, Questions, Criteria & Thresholds)