Guides / illustrated walkthrough
Jev Tips: 8 Best Practices for Better Decisions (State, Questions, Criteria & Thresholds)
A screenshot-by-screenshot Jev tips guide from Codevolution’s 43-minute best-practices deep dive: shape the state as an object, compute facts in code before you call, trim irrelevant context, pick Noul vs Choice vs Score correctly, write criteria that define every answer, batch independent questions into one request, and match each confidence threshold to the cost of the action.
Quick takeaway
One 43-minute screen-recorded deep dive, every demo re-run and screenshot-verified. The video walks a TypeScript SDK project (three folders — 1-state, 2-questions, 3-answers) through 13 scenarios on one running example: a customer named Maya whose Aria wireless headphones stopped charging. The working rules it teaches: represent the state as an object so fields stay nameable and referenceable in questions (accuracy is a wash — 0.92/0.94 string vs 0.9/0.9 object — but the object scales); let code compute facts before the call (raw purchase dates score 0.93/0.74/0.38 on coverage while code-computed “13 months ago” scores 0.08); trim the state to what the question needs (full ticket history answers “is the customer frustrated?” at 0.77 on 811 input tokens, latest ticket alone at 0.19 on 440 — and irrelevant details are a documented weakness, with input tokens at 4.2 cents per million and output free); use one Choice when code needs a single pick (plus an “other” option — without it a store inquiry forces “orders” at 0.46, with it 0.99) and separate Noul questions when many labels can be true at once; never read a Noul probability as intensity (0.19/0.93/0.99 on “is she frustrated” means likely-or-not, not how angry — that is a Score’s job, with each level described); write criteria that define true, false, every option, and every level (sarcasm like “Great, the second pair broke too. Love that for me.” lands at 1.30 between plain levels at confidence 0.54, then snaps to 1.00 at confidence 1.00 once levels carry summary-plus-signals); batch independent questions into one request (7 calls, 2008 ms, 2829 tokens becomes 1 call, 368 ms, 639 tokens); and give every action its own confidence bar (suggest a reply at 0.5, issue the refund automatically at 0.95 — the same 0.93 answer clears one and not the other).
Video source
Codevolution
Step-by-step walkthrough
- 1
Write the state as an object, not a sentence
The first scenario takes the same support ticket — Maya’s Aria wireless headphones stopped charging three weeks after purchase — and encodes it twice: 001-plain-string.ts packs everything into one sentence, 002-ticket-object.ts splits it into ticket, customer, order, and message fields. The accuracy difference is a wash (the string scored defective-product 0.92 and covered-by-plan 0.94; the object 0.9 for both), so the recommendation is not about precision. TypeSafe recommends the object for most queries because it is easier to maintain — you add or drop a field without rewriting prose; the field names carry the semantics (plan benefits belong to the customer, purchase details belong to the order); it maps straight onto the records your database and API already return; and your questions can reference fields directly, like `order.item` and `customer.plan_benefits`. Keep the plain string only when you are checking a single piece of text.

Same facts as the string version — but now every field has a name your questions can point at.Watch at 2:50 - 2
Let your code do the math before you call
Maya’s Plus plan covers faulty items for one year from purchase. The first attempt, 001-raw-dates.ts, sends the raw purchase dates and asks whether each order is still covered — and Jev hedges: 2 months ago scores 0.93, 10 months 0.74, and 13 months just 0.38, even though that last purchase is past the one-year line. Calendar arithmetic is not a judgment about meaning, so the fix in 002-computed-in-code.ts is a small monthsBetween() function that turns each date into “N months ago” before the request. With the computation done in code, the same question returns 0.96, 0.94, and 0.08 — the out-of-warranty case finally collapses toward no. The general rule: anything your code can calculate (date differences, totals, plan lookups) or resolve (a user ID becomes “customer” or “agent” before it ships) should never be left to the model. Jev judges meaning; your code checks facts.

Warranty math moved into a for-loop — the model only answers whether the words still mean covered.Watch at 10:10 - 3
Send only the state the question needs
The check is “is the customer frustrated?”. With Maya’s entire ticket history attached — an delayed replacement, a double charge, real anger — the answer comes back 0.77 on 811 input tokens, because the old complaints drag the score up even though her latest message says the replacement arrived and works perfectly. Resending only the newest ticket drops the same question to 0.19 on 440 input tokens — the right answer at roughly half the input cost. Two reasons to trim, per the video: TypeSafe lists irrelevant details as a known weakness, so accuracy can drop when the state is full of context unrelated to the question; and you pay per input token — 4.2 cents per million, with output tokens free — so the savings compound across every ticket. Making the question sharper (“frustrated right now?”) also helps, but the durable habit is: send what helps answer this question, cut everything else, and only include history when the question is about the customer over time.

The whole history leans “yes, frustrated” at 0.77 — the latest ticket alone says 0.19 on half the tokens.Watch at 12:50 - 4
Choice picks one team; separate Nouls label every topic
Routing demo: a customer needs an invoice for an expense claim but cannot log in, and the password-reset email never arrives. Four Noul questions (“relevant to billing? orders? account? product?”) both fire — billing 0.97 and account 0.99 — which describes the message but decides nothing. One Choice question (“which team should handle it?”) returns account at 0.91, correctly putting the login blocker first. Three corollaries follow in the same folder. Choice returns exactly one option, so when a message contains several topics you need one Noul per topic and keep every tag above a keep-threshold (0.5 in the demo) — that recovered both the invoice and the tracking complaint that a single Choice missed. And when nothing fits, Jev still must choose: asked where to send a “where can I test headphones in Amsterdam?” message, it forced orders at 0.46; adding an explicit other option pushed the same message to 0.99. So: one answer → Choice plus an other escape hatch; many possibly-true labels → one Noul each.

Four Nouls describe the message; one Choice has to pick a team — account wins at 0.91.Watch at 16:40 - 5
Noul answers whether; Score answers how much
Three messages, one Noul question — “is the customer annoyed?” — come back 0.19, 0.93, and 0.99. It is tempting to read 0.99 as “twice as angry” as 0.5, but a Noul probability is the chance the answer is yes: 0.5 means equally likely yes or no, not medium annoyance. When the goal is prioritizing the angriest customers, the right primitive is a Score with described levels — not annoyed, disappointed or upset, angry and complaining — which returns 0, 1, and 2: small comparable integers your code can sort by. The video adds a refinement: swapping the described levels for generic low/medium/high produces similar values but with visibly lower confidence on the middle message, because the model has to guess what “medium” means. Decide first whether you are asking whether or how much — then describe every rung of the ladder.

The video’s own one-line rule for choosing between the two primitives.Watch at 21:50 - 6
Write criteria that define every answer
Every helper takes an instruction plus criteria, and criteria is where definitions live. A bare “does the customer want a refund?” Noul scores a partial-refund request 0.8 and “can I return it?” 0.66; after defining true as a refund to the card including partial refunds, and false as exchanges, replacements, store credit, or just asking about options, the returns question falls to 0.29. Choice options can be objects saying what each covers, what it does not, and an example — which settled a returns-belong-to-orders versus refunds-belong-to-billing boundary that split two refund-status messages. Score levels need the same treatment: with plain low/medium/high, sarcasm like “Great, the second pair broke too. Love that for me.” lands between levels at confidence 0.54; once each level is an object with a summary plus signals — sarcasm about a related problem without threats stays at level 1 — both sarcastic messages snap to 1.00 at confidence 1.00 and the threat “If this breaks again I’m done with you guys” holds level 2 at 0.99. Only true and false are reserved keys on a Noul; summary and signals on Choice and Score levels are yours to design.

Plain levels leave sarcasm stranded between rungs — summary-plus-signals pins it.Watch at 30:30 - 7
Batch independent questions into one request
Seven questions about the same ticket — team, annoyance, urgency, product fault, what she tried, refund intent, wrong charge — asked one per call take 7 requests, 2008 ms, and 2829 input tokens. The same seven combined into a single systemOne call: 1 request, 368 ms, 639 input tokens. The state ships once instead of seven times, and each question is still scored independently — a question can use the state but never another question’s answer — which is also the only good reason to split requests: when answer A prepares question B, like identifying the order first and asking about it afterwards. The demo sneaks in one more pattern: you may include questions that only matter in some cases and let code ignore the unused answers — the run prints “Not about a charge, so chargedIncorrectly is ignored” right under team: product.

Same seven answers, one call, under a quarter of the tokens — and code ignores what does not apply.Watch at 32:15 - 8
Match the confidence threshold to the action
A Choice or Score answer carries a confidence; a Noul answer is its own probability — either way, your code needs a bar, and the bar should scale with what a wrong action costs. The video’s reply-or-refund script says it in a comment: “The riskier the action, the higher the bar.” Suggesting a reply to the agent needs confidence 0.5; issuing the refund automatically needs 0.95. “I’d like my money back, please” returns confidence 1.00 and clears both. “I’d like my money back, unless a replacement can ship quickly” returns 0.93 — the suggestion fires, the automatic refund does not, and a human confirms the customer’s preference first. An earlier scenario bands the same idea: act on confidence 0.9 and above, ask the customer to confirm between 0.6 and 0.9, route below 0.6 to a person; for Nouls, act above 0.8, do nothing below 0.2, and hand the middle to a human. One global threshold for every action is the mistake — cheap actions can tolerate doubt, irreversible ones cannot.

Confidence 1.00 clears both bars; 0.93 only earns the suggestion — a human settles the rest.Watch at 40:05
Frequently asked questions
What are the best practices for getting better results from the Jev API?
The video groups them into three decisions. Shape the state: use an object with named fields, compute facts like date math in your own code, and send only the context the question needs. Ask the right questions: Choice when code needs one pick (with an other option as the escape hatch), separate Noul questions when several labels can be true, Score with described levels when you need intensity, and criteria that define every option and level. Use the answers deliberately: batch independent questions into one request, and give every action its own confidence threshold based on what a wrong action costs.
When should I use Noul vs Choice vs Score in Jev?
Noul answers whether: it returns the probability that a yes/no statement is true, so 0.5 means equally likely either way — not a medium amount. Choice answers which one: it picks exactly one option from your list and returns a confidence, which makes it right for routing to a single team — but wrong for tagging multiple topics in one message. Score answers how much: it places the input on an ordered scale you describe, returning a comparable level you can sort by. The video’s litmus test: frustration as three Noul probabilities (0.19/0.93/0.99) tells you who is annoyed; a Score (0/1/2) tells you who to call first.
How do I reduce the cost of calling Jev?
Two levers dominate, because billing is per input token — the listed price is 4.2 cents per million input tokens and output tokens are free. First, trim the state: dropping Maya’s old ticket history cut one frustration check from 811 to 440 input tokens while making the answer more accurate (0.77 to 0.19). Second, batch: seven independent questions about one ticket cost 2829 input tokens and 2008 ms as seven calls but 639 tokens and 368 ms as one systemOne request. Split into separate requests only when an earlier answer is needed to prepare a later question.
What confidence threshold should I use on Jev decisions?
One per action, not one global number — the video’s comment is “the riskier the action, the higher the bar.” Its worked example suggests a reply to a support agent at confidence 0.5 or higher but only issues a refund automatically at 0.95, so an answer at 0.93 triggers the suggestion and not the money movement. For routing decisions it bands act at 0.9+, confirm-with-the-customer at 0.6–0.9, and human review below 0.6; for Noul probability it acts above 0.8, skips below 0.2, and escalates the band in between. Calibrate the exact cutoffs against your own traffic — high confidence still does not guarantee a correct answer.
Do criteria and level descriptions really change Jev’s answers?
Measured in the video, yes, by double-digit points. Defining true and false for a refund Noul pulled “can I return the product?” down from 0.66 to 0.29. Describing what each team’s Choice option covers (and does not) moved two refund-status messages that had split across orders and billing into billing together. And on a frustration Score, sarcastic messages like “Oh wonderful, another week without headphones” straddled plain low/medium/high levels at confidence 0.54, then landed cleanly on level 1 at confidence 1.00 once each level carried a summary plus signals such as “sarcasm about a related problem, without threats”. The pattern: never let the model guess what your answers mean.
Related guides
Jev Getting Started Guide
The zero-to-first-call walkthrough from the same channel — what System One decisions are and how the state-plus-questions payload works, assumed by this tips page.
ReadJev Playground Walkthrough
A guided tour of the console where you can re-run the Noul, Choice, and Score experiments from this page interactively.
ReadWhen to Use Jev (and When Not To)
The engineering assessment view: which workloads fit a decision model and which still belong to an LLM.
ReadSupport Ticket Routing Recipe
The customer-support routing pattern from this video’s scenarios, shipped as a copyable recipe with state and question templates.
ReadTry the Primitives in the Playground
Fire Noul, Choice, and Score questions at a live endpoint and watch the probabilities and confidence move as you edit criteria.
ReadMore video walkthroughs
- Jev Classification Quickstart: OpenRouter API, Primitives & Real Probabilities
- Jev Architecture Explained: Why 70ms Decision Models Beat LLMs for Workflow Automation
- Ultra-Fast Browser Agents with Jev: 178ms DOM Loops & Dual Model Orchestration
- Open Jev Models Are Here: Semif, Nimble, Decider, DiffusionGemma & Laya Hands-On
- Jev vs LLM: Will Jev Replace LLMs? Krish Naik's Whiteboard Explainer
- Jev Trader Tutorial: Build a Subsecond AI Trading Bot on Monad
- Jev Model Router: Build a Privacy-Gated LLM Router with Jev & OpenJev
- Jev Tutorial for Beginners: State, Questions & the TypeScript SDK
- Run Jev Locally: Kev, SemIf & Von on Your Own GPU (OpenJev Guide)
- Jev RAG Reranker: Policy-Steered Reranking for Retrieval-Augmented Generation
- When to Use Jev: An Engineer's Audit of Claims, Gates, and Failure Modes
- LangChain + Jev Integration Tutorial: Routing, Guardrails & Evals
- Jev MCP Server: Connect Jev Decisions to Claude Code & Cursor
- Jev vs Luna: Independent Benchmarks Put "Better, Faster, Cheaper" to the Test
- Jev Agent Harness: Where the Decision Gate Sits in Your LLM Loop
- Jev Playground Walkthrough: The Hotdog Lesson, Criteria, and a Four-Console Token Test
- Jev Text Classification API: Zero-Shot CLI & REST with classifier.dev
- Jev API Examples: First Request, curl & All Three Question Types
- Jev Log Triage with Expanso Edge
- Jev Lead Enrichment with Treg: ICP and Signup Scoring
- Use Jev Decision Nodes in Heym for Model Routing
- Laya Tutorial: Open-Source AI Routing With Calibrated Probabilities (Laya vs Jev Setup)
- Train Your Own Jev: Fine-Tune a Jev-Style Decision Model for $5–$17 (What You Can and Cannot Train)