Guides / illustrated walkthrough

Train Your Own Jev: Fine-Tune a Jev-Style Decision Model for $5–$17 (What You Can and Cannot Train)

A frame-by-frame guide to training your own Jev-style decision model: why the hosted Jev itself has zero fine-tuning, the three levers you actually control, where 38,000 training questions come from, the $4.17 and $17 real-world training bills, Laya fine-tuned from 0.362 to 0.766 accuracy, calibration that survives an audit, and when the whole pipeline pays off.

Quick takeaway

The hosted Jev cannot be fine-tuned at all — TypeSafe serves identical weights to every account and offers no fine-tuning endpoint, so “training your own Jev” really means one of two things: fine-tuning a small open LLM into a Jev-style classifier, or fine-tuning an open decision engine such as Laya. TerraNet’s explainer frames every option as one of three levers: question design (guides both Jev and Laya), recalibration (adjusts confidence so an 80% score means 80% accuracy on your data), and fine-tuning (changes weights — Laya only). Two published replicas prove the economics: Hassan (Nutlope) fine-tuned Qwen3.5 4B on Together AI with about 38,000 questions from eight public datasets in roughly 25 minutes for a $17 training job, and CoderOne ported the same tev1 repo to Modal, training Qwen3.5 4B with Unsloth + LoRA on an L40S for $4.17 in about two hours — roughly 38,000 training examples (~17M tokens), 5,000 dev, and 4,000 untouched staging examples drawn from MultiNLI, BoolQ, and Banking77 plus synthetic policy, routing, and research-taxonomy sets, landing about 88% accuracy on the held-out staging data with warm inference around 600–900 ms end-to-end (about 160 ms of actual model time) served by vLLM. On the open-engine side, TerraNet’s benchmark card shows Laya’s 421M English checkpoint going from 0.362 to 0.766 accuracy across 2,000 separate test decisions — against a teacher model that agreed with itself only 73.5% of the time — and warns that calibration temperatures must be fitted on held-out data only (Laya’s old notebook approach, fitting on training data, left temperatures within 6% of doing nothing). The hard rules: split data three ways before anything else, keep every prompt under 20 options, count test samples per question (300 examples at 80% accuracy carry ±4.5 points of uncertainty), and if your daily volume is modest, stick with an LLM and start logging structured decisions today.

Video source

TerraNet Technologies LLC

4:06Qx9CaCi9Zgo

Step-by-step walkthrough

  1. 1

    Be clear about what you can and cannot train

    The hosted Jev is off the table: TypeSafe serves identical weights to every account and provides no fine-tuning endpoint — TerraNet’s opening slide calls this “Jev: Zero fine-tuning,” while Laya sits on the other side of the frame as “retrainable on a free GPU.” That boundary is why “training your own Jev” always means one of two honest projects. Path one: fine-tune a small open LLM to imitate Jev’s behavior as a classifier — exactly what Hassan’s $17 Together AI tutorial and CoderOne’s $5 Qwen3.5 4B replica did. Path two: fine-tune an open decision engine such as Laya directly, which TerraNet notes trains in minutes on Kaggle’s free pair of T4 GPUs. The slide’s third bullet is the real headline: the bottleneck is not compute — it is gathering valid labels and fitting probabilities you can actually trust. Everything else on this page is about those two problems.

    TerraNet slide titled The Decision Model Dilemma listing Jev with zero fine-tuning, Laya retrainable on a free GPU, and labels plus calibration as the real bottleneck.
    The hosted Jev cannot be fine-tuned — replicas and open engines are the two honest paths.Watch at 0:20
  2. 2

    Pick your lever: question design, recalibration, or weights

    TerraNet frames everything you can change as three levers. Question design guides both Jev and Laya: because TypeSafe gives every account the same weights, you shape answers through the request itself — breaking one broad judgment into several narrow questions and writing the policy into your code. Recalibration adjusts confidence so that an 80% score actually matches 80% accuracy on your specific distribution; both vendors train with RLCD, but calibration only holds on data resembling what it was calibrated on. Fine-tuning changes internal weights and is available for Laya only. If your goal is a Jev-style classifier, the fine-tuning lever points at a small open base model: both replicas chose Qwen3.5 4B for its open weights and workable license, and CoderOne’s port kept the training format Jev-shaped from day one — a state, a question, and answer choices in, one letter out.

    TerraNet card titled Three Levers: What You Can Change contrasting question design and recalibration available for Jev with weight-changing fine-tuning offered only by Laya.
    Two of the three levers work on the hosted Jev; only fine-tuning requires your own model.Watch at 0:50
  3. 3

    Label first, split three ways, then touch nothing

    TerraNet’s pipeline figure is the whole discipline in one strip: collect and label, split three ways, fine-tune, fit the confidence, set the threshold — with escalations routed back in as fresh training data. Labels come from past operational decisions, a teacher LLM, or human review. The split must happen before any training: a training set, a calibration set, and a test set you never peek at early. The two replicas show what a real split looks like: Hassan’s tev1 setup trained on about 38,000 questions sampled from eight public datasets, and CoderOne’s port carried roughly 38,000 training examples (about 17M tokens), nearly 5,000 dev examples (about 2M tokens), and a staging set of around 4,000 examples left untouched until after the model shipped. The public sets behind those numbers, as listed in CoderOne’s video: MultiNLI for entailment, BoolQ for yes/no questions, Banking77’s 77 banking support intents, plus synthetic decision datasets — pop for policies, policy V2, routing V2, and a research taxonomy that classifies articles by primary contribution rather than topic.

    Five-box TerraNet loop diagram reading label, split three ways, fine-tune, calibrate, then set the threshold, with escalation cases feeding back in as new training data.
    Label, split three ways, fine-tune, calibrate, set the threshold — then escalations become tomorrow’s labels.Watch at 2:25
  4. 4

    Run the fine-tune — and read the benchmark honestly

    The training itself is now the easy part. CoderOne ran Unsloth with LoRA on official Qwen3.5 4B inside a CUDA container, did a one-minute smoke test on a data subset, then let the full run cook for about two hours on a Modal L40S — the training line item came to $4.17 (an H100 would cost more but maybe finish in 30 minutes, he estimates). Hassan’s Together AI version finished in roughly 25 minutes for a $17 training job. On the open-engine side, TerraNet’s benchmark card reports Laya’s 421M English checkpoint going from 0.362 base accuracy to 0.766 after fine-tuning across 2,000 separate test decisions, trained on 1,200 cases spanning four workflows. The catch sits on the same slide: the benchmark’s ground truth came from a teacher model that agreed with itself only 73.5% of the time, and Laya matched the noisy teacher more often than the teacher matched itself — evidence it learned the task pattern rather than the noise. CoderOne’s replica scored about 88% on its untouched staging set. One more boundary from the video: the replica is strongest at picking one option from a list (its demo also pushed a boolean-style true/false check through in about 600 ms), while Jev’s richer typed primitives would need additional labeled datasets before they can be trained properly.

    Laya benchmark slide showing base accuracy of 0.362 lifted to 0.766 after fine-tuning while the teacher model self-agreement sits at just 73.5 percent.
    0.362 to 0.766 looks great until you notice the teacher agreed with itself only 73.5% of the time.Watch at 1:50
  5. 5

    Calibrate on held-out data only, then audit the evaluation

    A raw fine-tuned score is not a probability you can gate code on. TerraNet’s recalibration lever means fitting temperature scaling or Platt scaling on the held-out calibration set — never on training data. The cautionary tale from the video: Laya’s own notebook used to fit calibration temperatures directly on a slice of training data, leaving temperatures barely moved — within 6% of doing nothing — and creating a false sense of calibrated certainty. Then audit the evaluation itself. The most frequent failure is measuring on the benchmark you trained on, or on a test set too small to yield statistical power: at 80% accuracy across 300 test examples, uncertainty spans about ±4.5 percentage points — too coarse to tell two competing fine-tunes apart — so count test examples per question rather than per case. Finally, respect the option budget: accuracy degrades when a single prompt presents more than 20 options, so split wide taxonomies into hierarchical decisions in your code instead of fine-tuning on dozens of simultaneous branches.

    Avoiding Flawed Evaluations checklist slide telling teams to split data into train, calibration, and test sets, count test samples per question, and split questions exceeding 20 options.
    Three evaluation traps: training-set evals, underpowered test sets, and oversized option lists.Watch at 3:10
  6. 6

    Deploy where latency pays you back — and start logging either way

    The work is front-loaded, so it only pays at volume: TerraNet puts the break-even at thousands of daily queries, where LLM calls create noticeable latency and compute costs. CoderOne’s serving numbers show what a deployed replica looks like: vLLM behind Modal on both an L40S and an H100, where the small 4B model actually ran better on the cheaper L40S — about 160 ms of model time per decision (14 ms to build the input plus 145 ms of execution), 600–900 ms end-to-end when warm, and cold starts of roughly 90 seconds whenever the standby GPU woke up. His whole experiment — training plus smoke tests plus heavy inference — consumed about $10 of Modal’s $30 signup credit. His closing comparison table is the cleanest rule of thumb: a pretrained LLM like Qwen3.5 stays flexible, explains its answers, and carries a roughly 256K-token context, but costs more to serve and is not accurate enough for structured decisions; a decision model (Jev or Laya) returns typed decisions and probabilities in one fast pass; a fixed BERT-style classifier wins outright when you need the same labels every time. And if your volume is modest today, TerraNet’s parting advice costs nothing: record decisions alongside inputs in a structured format now, so the training data is ready the day you decide to train.

    Is It Worth the Pipeline slide concluding that front-loaded engineering effort pays off at high daily volume when you begin by logging structured decisions.
    Break-even sits at thousands of daily queries — below that, log decisions and wait.Watch at 3:50

Frequently asked questions

Can you fine-tune the Jev model itself?

No. TypeSafe serves identical weights to every account and provides no fine-tuning endpoint for the hosted Jev — TerraNet’s explainer summarizes it as “Jev: zero fine-tuning.” Your only levers on the hosted product are question design and recalibration. “Training your own Jev” therefore means training a Jev-style classifier of your own (a fine-tuned small LLM such as Qwen3.5 4B) or fine-tuning an open decision engine such as Laya. Nothing on this page turns a self-trained replica into the official Jev, and none of the source projects claim parity with it.

How much does it cost to train your own Jev-style model?

Two published data points: Hassan trained Qwen3.5 4B on Together AI for a $17 training job (roughly 25 minutes, about 38,000 questions), and CoderOne trained the same class of model on a Modal L40S for $4.17 in about two hours — about $10 total once he included smoke tests, validation, and a lot of inference queries, all inside Modal’s $30 signup credit. TerraNet adds the floor: Laya fine-tunes in minutes on Kaggle’s free pair of T4 GPUs. Serving is the ongoing cost — either an always-on GPU or a standby mode that trades cheaper idle time for slow cold starts.

Where does the training data come from?

Three sources, per TerraNet’s pipeline: past operational decisions, a teacher LLM, or human review. Both replicas started from the open tev1 repo, whose roughly 38,000 training questions draw on eight public datasets — MultiNLI (entailment), BoolQ (yes/no), Banking77 (77 banking support intents), plus synthetic decision sets covering policies, routing, and a research taxonomy. Every example is formatted Jev-style: a state, a question, and answer choices, with the model learning to emit a single letter. Critically, the calibration and test sets are held back before any training happens.

Is a self-trained model as good as the official Jev API?

Treat them as different tiers. CoderOne’s Qwen3.5 4B replica scored about 88% on a staging set it had never seen and answered in under a second when warm — useful, but it covers the multiple-choice core rather than Jev’s full typed vocabulary, and no one has shown it matching the hosted reference quality. The benchmark numbers carry their own noise: TerraNet’s Laya fine-tune hit 0.766 accuracy against a teacher that only agreed with itself 73.5% of the time. If you need guaranteed quality, use the hosted Jev API and keep a self-trained model as a cheap fallback or a data-stays-in-house deployment.

What is calibration, and why do you need a separate calibration set?

Calibration adjusts the model’s confidence so that an 80% score really is right about 80% of the time on your traffic. The practical technique is temperature scaling or Platt scaling fitted on a held-out calibration set — fitting those parameters on training data barely moves them: TerraNet points out that Laya’s old notebook approach left temperatures within 6% of doing nothing while creating a false sense of calibrated certainty. Both vendors train with RLCD for calibrated decisions, but calibration only holds on data resembling the calibration distribution — so re-fit when your traffic drifts, then pick an operating threshold and route low-confidence escalations to an LLM or a human.

Should you fine-tune, or just use an LLM with prompts?

TerraNet’s break-even rule: the front-loaded engineering only pays off at thousands of daily queries, where LLM calls create noticeable latency and compute costs. CoderOne’s closing table narrows it further — pick the pretrained LLM (Qwen3.5, roughly 256K context, explains its answers) when you need flexibility and explanations; pick a decision model (Jev or Laya) for real-time typed decisions with low latency; pick a fixed BERT-style classifier when you need the same fixed labels every time. If your volume is modest today, stay on the LLM but start logging decisions alongside inputs in structured formats so the dataset is ready when you are.

Related guides

More video walkthroughs