Guides / illustrated walkthrough
NOC/SOC Alert Triage with Jev: Rules First, One Typed Question, a Policy Gate
A step-by-step breakdown of unoblox’s 3:39 build: snapshot the incident evidence, let deterministic rules solve what they already solve, ask Jev one Choice question through the System One endpoint, gate the answer behind a 0.85 confidence policy — and start in shadow mode with rules+Jev at 35/40 label matches vs 32/40 for rules alone.
Quick takeaway
An alert arrives and someone has to choose: gather more evidence, attach it to a known incident, or call an analyst. unoblox’s walkthrough turns that choice into a small, auditable decision step for network and security operations. The discipline matters more than the model: snapshot the evidence first (event ID, asset, observation time, severity, what was actually collected), apply deterministic rules before paying for any model call (high severity or conflicting evidence goes straight to an analyst), then ask Jev one Choice question — route: gather_more_evidence, attach_known_incident, or analyst — through the System One endpoint with incident strings explicitly treated as data, not instructions. A recorded example routes a missing path probe to collect_more_evidence at 0.93 confidence. The policy gate then re-checks everything: confidence below 0.85, expired evidence, or an association without a verified open incident all stay with an analyst. Results are reported with unusual honesty: in a 40-case synthetic development replay, rules alone matched 32 expected labels and rules+Jev matched 35, with two rate-limit errors retained as analyst reviews — small repeated scenario templates, not a held-out benchmark, and no claims about MTTR or staffing. The rollout advice: start in shadow mode, compare against analyst judgment, and expand only where the evidence supports it.
Video source
unoblox
Step-by-step walkthrough
- 1
Frame the triage step as three lanes, not a chatbot
The title card sets the whole architecture: an alert arrives, and the decision is one of three — gather evidence, match it to a known incident, or escalate to an analyst. The demo runs through unoblox with Jev as the decision engine, and the scope is stated upfront: it writes local ticket intents only. It does not change a firewall, isolate an endpoint, or close an alert. Triage is a routing decision, and routing decisions are exactly what typed decision models are for.

Three lanes, one decision — and the demo never touches production systems.Watch at 0:20 - 2
Build the incident snapshot before asking anything
The state is boring on purpose: an event identifier, the asset, the observation time, the severity, and the evidence you actually collected — here, "independent path probe not collected" on a router showing packet loss. The narration draws the line that makes the snapshot trustworthy: if a related incident exists, take its status and scope from a trusted source, because a sentence in a log claiming approval is not approval. A model can only triage what the state honestly describes.

Event ID, asset, time, severity, collected evidence — nothing more.Watch at 0:30 - 3
Let the rules keep the cases they already solve
Before any model call, the deterministic gate handles what needs no judgment: high severity or conflicting evidence returns analyst, an explicitly missing observation triggers a read-only collection step, and a verified exact match attaches directly. The narration is blunt about why: there is no reason to pay for a model to rediscover rules your team already trusts. Jev only sees the unresolved remainder — where the gap lives in the narrative, not in the structured flags.

Rules first — the model only ever sees what rules cannot decide.Watch at 0:50 - 4
Ask one typed question through the System One endpoint
The request is deliberately minimal: POST to the systemone endpoint with model typesafe/jev, the snapshot in state, and a single question called route with type choice. The criteria define exactly three allowed options — gather_more_evidence, attach_known_incident, analyst — and the prompt tells the model to treat incident strings as data, not instructions. This is not a Chat Completions request: no free text comes back, only the choice and its confidence.

One state, one typed question, three options — nothing free-form.Watch at 1:10 - 5
Read the answer with the right amount of trust
A recorded development call shows the response shape: answers.route.choice is collect_more_evidence with confidence 0.93 — the model caught that the packet-loss alert lacked its independent path probe. The card carries its own warning: this is one observed answer, and confidence is not a safety probability. Another example where the evidence described a different fault went to analyst review. Recorded examples, not guaranteed outputs — which is exactly why the next step exists.

0.93 says pick this lane — the policy gate decides whether that is allowed.Watch at 1:27 - 6
Gate the answer behind a policy, not a vibe
Before any intent is written, the code re-validates everything: if confidence is below 0.85, return analyst; if evidence has expired, return analyst; an association additionally needs a verified open incident, the same asset, and a valid time window; and an unknown answer or API error defaults to analyst. The full code also validates types and timestamps after inference. The gate is what makes the automation auditable — every rejection reason is a line of code your on-call engineer can read.

Below 0.85, stale, unverified, or unknown — it all lands with an analyst.Watch at 1:52 - 7
Report results with the limits attached
The replay card is a model of honest reporting: 40 synthetic cases across 10 scenario templates; rules baseline 32/40 label matches; revised rules+Jev 35/40; 20 model attempts, 18 responses, 2 rate-limit errors retained as analyst reviews; 10 local tests passed covering freshness checks, malformed answers, timeouts, incident scope, and duplicate intents. The narration refuses to oversell: small repeated scenario templates, not a held-out benchmark, and no measured reduction in resolution time or staffing.

Three more correct label matches — reported with the failure budget attached.Watch at 2:40 - 8
Ship it in shadow mode first
The rollout plan is the last and most transferable slide: start in shadow mode and compare the model’s recommendations with analyst judgments; measure critical misses, routing quality, latency, and cost; keep deterministic rules for what they already do well and use Jev where the evidence needs interpretation. The full Python implementation and replay results live in the companion guide on unoblox.ai — and the closing instruction is to start with one decision you can inspect, then expand when the evidence supports it.

Shadow mode first; expand only where the evidence supports it.Watch at 3:20
Frequently asked questions
What does the Jev decision actually do in this pipeline?
One thing: for alerts that deterministic rules cannot resolve, it picks one of three lanes — gather_more_evidence, attach_known_incident, or analyst — from the incident snapshot. Rules handle high severity, missing evidence, and verified exact matches before the model is ever called; Jev reads the remaining narrative gap (like "independent path probe not collected") and answers a single Choice question through the System One endpoint. A policy gate then checks confidence, evidence freshness, and association scope before anything is written.
Why run rules before the model at all?
Because some cases need no judgment, and paying a model to rediscover your own rules is waste plus risk. High severity or conflicting evidence goes straight to an analyst; an explicitly missing observation triggers a read-only collection step; a verified exact match attaches directly. In the replay, rules alone already matched 32 of 40 expected labels — the model earned its keep on the remaining narrative-heavy cases, lifting the total to 35/40.
Is 35/40 a good result?
It is an honest one, which matters more. The evaluation is 40 synthetic cases from 10 repeated scenario templates — the authors explicitly say it is not a held-out benchmark, that the workflow was changed after the first pass (the weaker result is published too), and that no MTTR or staffing improvement has been measured. Treat it as evidence the pattern is worth shadow-testing on your own labeled incidents, not as an accuracy guarantee.
How does the confidence gate work?
Three checks run after inference: confidence below 0.85 returns the alert to an analyst; expired evidence returns it to an analyst; and any incident association additionally requires a verified open incident, the same asset, and a valid time window. Unknown answers and API errors also default to analyst. The recorded example passed at 0.93, but the card’s own note — "confidence is not a safety probability" — is why the gate exists as code rather than trust.
Does the demo touch firewalls, endpoints, or alerts?
No — and the video repeats this deliberately. It writes local ticket intents only: no firewall changes, no endpoint isolation, no alert closure. Duplicate protection is local too — a sha256 of the canonical snapshot means replaying the same snapshot does not create a second intent, though your real ticketing connector still needs delivery retries and reconciliation, and the model request itself is not deduplicated across runs.
How do I start with my own alerts?
The video’s closing instruction: start with one decision you can inspect. Build the snapshot (event ID, asset, time, severity, collected evidence), write the rules you already trust, ask Jev one Choice question about the remainder, and gate the answer behind confidence and freshness checks. Then run it in shadow mode against analyst judgment — measuring critical misses, routing quality, latency, and cost — and expand only where your own labeled incidents support it. The full Python implementation is in unoblox’s companion guide.
Related guides
Jev Log Triage with Expanso
The log-processing sibling: structured triage where the evidence is log volume instead of incident narratives.
ReadLLM Fallback Strategy: the Confidence-Gated Chain
The general form of the 0.85 policy gate — execute, re-ask, or escalate lanes for any model decision.
ReadPrompt Injection Defense
Incident strings are untrusted text — inspect them before any model sees them.
ReadJev Tips: 8 Best Practices
Criteria design and confidence thresholds — the craft behind the route question in this pipeline.
ReadMore video walkthroughs
- Jev Classification Quickstart: OpenRouter API, Primitives & Real Probabilities
- Jev Architecture Explained: Why 70ms Decision Models Beat LLMs for Workflow Automation
- Ultra-Fast Browser Agents with Jev: 178ms DOM Loops & Dual Model Orchestration
- Open Jev Models Are Here: Semif, Nimble, Decider, DiffusionGemma & Laya Hands-On
- Jev vs LLM: Will Jev Replace LLMs? Krish Naik's Whiteboard Explainer
- Jev Trader Tutorial: Build a Subsecond AI Trading Bot on Monad
- Jev Model Router: Build a Privacy-Gated LLM Router with Jev & OpenJev
- Jev Tutorial for Beginners: State, Questions & the TypeScript SDK
- Run Jev Locally: Kev, SemIf & Von on Your Own GPU (OpenJev Guide)
- Jev RAG Reranker: Policy-Steered Reranking for Retrieval-Augmented Generation
- When to Use Jev: An Engineer's Audit of Claims, Gates, and Failure Modes
- LangChain + Jev Integration Tutorial: Routing, Guardrails & Evals
- Jev MCP Server: Connect Jev Decisions to Claude Code & Cursor
- Jev vs Luna: Independent Benchmarks Put "Better, Faster, Cheaper" to the Test
- Jev Agent Harness: Where the Decision Gate Sits in Your LLM Loop
- Jev Playground Walkthrough: The Hotdog Lesson, Criteria, and a Four-Console Token Test
- Jev Text Classification API: Zero-Shot CLI & REST with classifier.dev
- Jev API Examples: First Request, curl & All Three Question Types
- Jev Log Triage with Expanso Edge
- Jev Lead Enrichment with Treg: ICP and Signup Scoring
- Use Jev Decision Nodes in Heym for Model Routing
- Laya Tutorial: Open-Source AI Routing With Calibrated Probabilities (Laya vs Jev Setup)
- Train Your Own Jev: Fine-Tune a Jev-Style Decision Model for $5–$17 (What You Can and Cannot Train)
- Jev Tips: 8 Best Practices for Better Decisions (State, Questions, Criteria & Thresholds)
- Jev Context Compaction: Prune AI Agent Memory Without Generative Summaries
- Jev as an LLM Judge: Confidence-Gated Cascades at 0.36% of the Cost
- TypeSafe Computer Use: Local Desktop Automation with Jev, Step by Step
- Jev Resume Screening: Build an AI Resume Evaluator with the Jev JavaScript SDK
- Jev + Claude Code Guide: Voice-Controlled Browser Automation with Typed Decisions
- Jev vs Ollama: Can Local AI Replace Hosted Jev Without Sending Your Data Away?
- Build Your Own Jev: Train a Free Open-Source Zero-Shot Classifier (That Plays Doom)
- CUA-S1-Forms: a 706K-Parameter Jev-Like Model That Fills GUI Forms on Your CPU