Guides / illustrated walkthrough
A Session Drift Guard for Pi Agent: Let the Jev Model Propose, Let Code Decide
A hands-on build of pi-jev-router 0.3.0: hook Jev into Pi Agent’s input event, compress the conversation into an ~800-token state, route each message through a four-way Choice with probability gates, and learn the Choice-question trap — semantic overlap splits probabilities 0.45/0.47 — that a 52-scenario labeled test caught and the sum-in-code fix solved (14/14 caught, 0/38 interrupted, 277ms average).
Quick takeaway
Sessions drift: you are twelve messages into debugging a checkout service OOM when you casually ask about a robot vacuum — and that question now rides in the context of every turn that follows. This build (from a Chinese-language creator, a first for this site) puts a typed Jev decision at Pi Agent’s input event, after Enter and before the message reaches the agent. A local filter handles slash commands, shell commands, file paths, sub-15-character inputs, images, queued messages, and first messages for free; everything ambiguous goes to Jev as an ~800-token compressed state (title, last three user messages, reply tail, new input) against a four-way Choice — continue, side_chat, fork, new session. The policy is three lines: irrelevant under 0.6 passes, on-topic over 0.6 passes, the rest triggers a reminder — “the model provides probabilities, the code decides.” The honest engineering story is the evaluation: 52 hand-labeled scenarios run against Jev 1.13 on OpenRouter found the worst version missing 2–3 of 14 drifts and interrupting 1 of 38 normal inputs, because a casual-question case split its probability across two “irrelevant” options (0.45 + 0.47, neither over 0.6). Summing related options in code (0.92 ≥ 0.6 → remind) fixed it to 14/14 caught, 0/38 interrupted, at a 277 ms average and P95 under 0.35 s. Failure handling is fail-open: no key disables the plugin, API errors and over-2.5s responses pass through, and a failed move returns your message to the editor — one missed reminder is the worst case, never a lost message.
Video source
01Coder
Step-by-step walkthrough
- 1
The pain point: sessions drift, and the drift pollutes every turn after it
You are mid-investigation — the checkout pod keeps getting OOMKilled, and you have traced the memory spike to its image-processing worker — when you remember you need to replace your robot vacuum, and ask about it right there. The agent happily answers. But the question is now permanent chat history: every subsequent turn carries it, and the context gets less and less focused on the problem you were actually solving. Long sessions, casual asides, polluted context — the creator’s three-tag summary of why this needs fixing before the message ever enters the agent.

聊着聊着,就跑偏了 — the drift happens mid-session, not at the start.Watch at 1:07 - 2
Why a decision model fits: one paragraph vs probabilities
The comparison the whole design hangs on. Let a large model judge: output is a paragraph of text, latency is seconds, cost per judgment is high. Let Jev judge: output is a fixed answer space with probabilities, latency is about 0.3 seconds, and 300k judgments cost about a dollar. The judgment needs to run before every message, be cheap, and return something the program can use directly — a fixed answer space evaluated in one forward pass is exactly that shape.

Fixed answer space, one forward pass — the shape that fits an input hook.Watch at 2:40 - 3
Hook the input event — then filter locally before spending a call
Pi Agent (also called PIagent or the Pi coding agent) exposes an input event that fires after Enter but before the message reaches the agent. Hooking there means the drift check runs before the context is polluted. But Jev is not free, so the plugin passes locally everything that cannot drift: slash commands, shell commands, file paths, short instructions under 15 characters, images, messages queued while the agent is running, and the first input of a chat. If a session has been idle over 12 hours or 85% of the context is used, code decides directly. Only genuinely ambiguous messages go to the model.

Enter → Input event → local filter → only the ambiguous cases reach Jev.Watch at 3:32 - 4
Give Jev a compressed view, not the whole conversation
The RouterState type is deliberately small: session title (300 chars), the last three user messages (300 chars each), the tail of the previous assistant reply (last 1500 chars), and the new input (up to 2000 chars). That is roughly 800 tokens per call — because judging drift only needs to know what the session is currently about, not every detail of it. Smaller state means faster, cheaper, and less noise diluting the decision. In production, the terminal breakdown shows the real budget: questions ~14%, summary ~55%, the live tail the rest.

Don’t send the session — send the compressed view (~800 tokens).Watch at 4:17 - 5
Ask one four-way Choice plus the Noul questions
The route question is a single Choice with four options: continue (keep going in this session), side_chat (an unrelated follow-up question), fork (spin a new version off the current content), and new_session (start an unrelated task). Alongside it sit Noul questions like “is this still on the current topic?” — every answer comes back with a probability. One request resolves the whole routing judgment; the answers are values, not prose to parse.

Four fixed options, each with a probability — the whole question.Watch at 5:02 - 6
The policy is three rows: model proposes, code decides
Jev only outputs probabilities; whether to remind the user is written in code. Row one: irrelevant probability — P(side_chat) + P(new_session) — under 0.6, let it pass. Row two: on_topic at or above 0.6, let it pass. Row three: everything else, remind. That division — the model supplies probabilities, regular code makes the decision — keeps the policy testable, versionable, and independent of any model update. When the reminder fires, the user chooses: stay, fork, start new, or mute for this session.

让模型出概率,让代码做决定 — the whole gate in three rows.Watch at 5:32 - 7
Read the raw answer: the fork case that correctly passes
Mid-way through writing an 8-minute video script, the creator asks how a 60-second vertical version would be structured. /route:status shows the raw answer: on_topic 0.85 — still on theme — with continue at 0.47 and most of the remainder on fork; irrelevant sits near zero, under the 0.6 threshold. Fork is not in the “irrelevant” set, so the code passes it and the topic model just answers. The judgment about whether to branch stays with the user, exactly as designed.

on_topic 0.85, fork ≈ 0.5 — the code passes it; branching is the user’s call.Watch at 8:46 - 8
The trap the test caught: semantic overlap splits probabilities
The labeled evaluation — 52 scenarios, human-labeled, run two rounds each against Jev 1.13 on OpenRouter — found the worst version (0.2.0) missing 2–3 of 14 real drifts while interrupting 1 of 38 normal inputs. The missed case is the lesson: “another project suddenly errors: ECONNREFUSED 127.0.0.1:16379, help me investigate.” Jev’s judgment was correct — unrelated to the current session — but “unrelated” was split across two options: side_chat 0.45 and new_session 0.47. Each sat under the 0.6 gate, so neither fired and the message passed.

0.45 + 0.47 and neither crosses 0.6 — the classic Choice-question trap.Watch at 9:47 - 9
Two fixes: exclusive options, or sum in code
The slide names both cures. Method one: design options to be mutually exclusive so probability never splits. Method two — the one he shipped: sum the probabilities of same-class options in code. Whether the message is a casual question or a core task from another project, both mean the same thing for the reminder decision, so 0.45 + 0.47 = 0.92 ≥ 0.6 fires the reminder. The choice between stay/fork/new still belongs to the user; summing only fixes the reminder gate, which is why it is safe.

0.45 + 0.47 = 0.92 ≥ 0.6 → remind. Sum in code, decide in code.Watch at 10:32 - 10
After the fix: 14/14 caught, 0/38 interrupted, 277 ms average
The corrected confusion matrix (shipped as 0.3.0) is clean: all 14 drift scenarios reminded, zero of the 38 normal inputs interrupted. Latency: 277 ms average, P95 under 0.35 s — cheap enough to sit on every message. The creator’s caveats are part of the lesson: 52 self-written, self-labeled scenarios are not a public benchmark, and the compressed conversation snippets go to a third-party API (OpenRouter), so think twice before pointing this at sensitive code.

14/14 reminded, 0/38 interrupted — at 277 ms average, P95 < 0.35 s.Watch at 10:57 - 11
Failure handling: fail-open, always
A check on the critical path of every message must answer “what happens when it breaks?” The table is the answer: no API key — plugin disabled, no impact; API error or rate limit — pass through, cost is one missed reminder; response over 2.5 seconds — pass; no UI mode — no interruption; move failure — the message returns to the editor, never lost. Jev judges, but the final decision always rests with the user, and no failure mode can block an input or destroy a message. That fault-tolerance stance is what makes putting a model on the hot path viable.

Worst case in every failure row: one missed reminder — never a lost message.Watch at 11:22
Frequently asked questions
What is session drift and why does it matter for agents?
Session drift is asking something unrelated mid-task — a robot-vacuum question twelve messages into debugging a checkout service. The agent answers, but the question becomes permanent context: every later turn carries it, and the conversation gets progressively less focused on the original problem. Catching drift at the input event — before the message enters the agent — keeps the context clean without policing what users may ask.
What is pi-jev-router and how do I install it?
A Pi Agent extension (Pi’s official name for plugins) that hooks the input event — after Enter, before the message reaches the agent — and uses a typed Jev decision to remind you when a message looks unrelated to the current session. Version 0.3.0 is the one tested in the video; the plugin source, test scenarios, and data are linked in the video description. You need a Jev API key (the video uses Jev 1.13 through OpenRouter); without one the plugin simply stays disabled.
Why a four-way Choice instead of a yes/no question?
Because “unrelated” is not one thing. continue, side_chat, fork, and new_session describe four genuinely different user intents, and the probabilities over all four carry more signal than a single yes/no: the fork case (on-topic, but branching) must pass while an unrelated question must remind. The trap to avoid — which the video demonstrates with real data — is semantic overlap between options: side_chat 0.45 and new_session 0.47 both mean “irrelevant” but neither crosses a 0.6 gate alone. Either make options mutually exclusive or sum same-class options in code.
What does “the model proposes, code decides” mean here?
Jev returns probabilities, never actions. The policy is ordinary code: irrelevant (P(side_chat) + P(new_session)) under 0.6 passes, on_topic at or above 0.6 passes, everything else triggers a reminder where the user chooses to stay, fork, start a new session, or mute for this session. Keeping the policy in code makes it testable and versionable — and it is where the probability-summing fix lives.
How accurate is it, and what did it miss before the fix?
On 52 hand-labeled scenarios run two rounds each: the worst version (0.2.0) caught 11–12 of 14 drifts, missed 2–3, and interrupted 1 of 38 normal inputs. The miss came from probability splitting across overlapping options. After summing same-class options in code, version 0.3.0 reminds on all 14 drifts and interrupts none of the 38, at 277 ms average and P95 under 0.35 s. The honest caveats: the scenarios are self-written and self-labeled — not a public benchmark — and results do not guarantee real-world accuracy.
What happens when the Jev API is down or slow?
Everything fails open. Missing API key: the plugin disables itself. API error or rate limiting: the message passes, costing one missed reminder. Response over 2.5 seconds: pass. No UI mode: no interruption. Failed session moves: the message returns to the editor untouched. The design principle: a drift guard that can block or lose messages is worse than one that occasionally stays silent.
Related guides
Jev Model Router: Building a Model Router with Jev
The other routing axis — using Jev to pick which LLM answers, not whether the session should continue.
ReadJev + Claude Code: Typed Decisions in Agent Orchestration
The same typed-decision-in-the-agent-loop pattern, wired into Claude Code instead of Pi Agent.
ReadJev Context Compaction: the Context Economy
The state side of this design — what the compressed 800-token view is protecting.
ReadJev Tips: 8 Best Practices
Criteria and option design practices — including keeping answer options semantically clean.
ReadMore video walkthroughs
- Jev Classification Quickstart: OpenRouter API, Primitives & Real Probabilities
- Jev Architecture Explained: Why 70ms Decision Models Beat LLMs for Workflow Automation
- Ultra-Fast Browser Agents with Jev: 178ms DOM Loops & Dual Model Orchestration
- Open Jev Models Are Here: Semif, Nimble, Decider, DiffusionGemma & Laya Hands-On
- Jev vs LLM: Will Jev Replace LLMs? Krish Naik's Whiteboard Explainer
- Jev Trader Tutorial: Build a Subsecond AI Trading Bot on Monad
- Jev Model Router: Build a Privacy-Gated LLM Router with Jev & OpenJev
- Jev Tutorial for Beginners: State, Questions & the TypeScript SDK
- Run Jev Locally: Kev, SemIf & Von on Your Own GPU (OpenJev Guide)
- Jev RAG Reranker: Policy-Steered Reranking for Retrieval-Augmented Generation
- When to Use Jev: An Engineer's Audit of Claims, Gates, and Failure Modes
- LangChain + Jev Integration Tutorial: Routing, Guardrails & Evals
- Jev MCP Server: Connect Jev Decisions to Claude Code & Cursor
- Jev vs Luna: Independent Benchmarks Put "Better, Faster, Cheaper" to the Test
- Jev Agent Harness: Where the Decision Gate Sits in Your LLM Loop
- Jev Playground Walkthrough: The Hotdog Lesson, Criteria, and a Four-Console Token Test
- Jev Text Classification API: Zero-Shot CLI & REST with classifier.dev
- Jev API Examples: First Request, curl & All Three Question Types
- Jev Log Triage with Expanso Edge
- Jev Lead Enrichment with Treg: ICP and Signup Scoring
- Use Jev Decision Nodes in Heym for Model Routing
- Laya Tutorial: Open-Source AI Routing With Calibrated Probabilities (Laya vs Jev Setup)
- Train Your Own Jev: Fine-Tune a Jev-Style Decision Model for $5–$17 (What You Can and Cannot Train)
- Jev Tips: 8 Best Practices for Better Decisions (State, Questions, Criteria & Thresholds)
- Jev Context Compaction: Prune AI Agent Memory Without Generative Summaries
- Jev as an LLM Judge: Confidence-Gated Cascades at 0.36% of the Cost
- TypeSafe Computer Use: Local Desktop Automation with Jev, Step by Step
- Jev Resume Screening: Build an AI Resume Evaluator with the Jev JavaScript SDK
- Jev + Claude Code Guide: Voice-Controlled Browser Automation with Typed Decisions
- Jev vs Ollama: Can Local AI Replace Hosted Jev Without Sending Your Data Away?
- Build Your Own Jev: Train a Free Open-Source Zero-Shot Classifier (That Plays Doom)
- CUA-S1-Forms: a 706K-Parameter Jev-Like Model That Fills GUI Forms on Your CPU
- NOC/SOC Alert Triage with Jev: Rules First, One Typed Question, a Policy Gate
- Ollama Decision Models: Run tev1 and Nimble Locally (Tested on an 8 GB Card)
- Jev Guardrails in Production: A Five-Step Playbook for Decision Automation