Guides / illustrated walkthrough

A Session Drift Guard for Pi Agent: Let the Jev Model Propose, Let Code Decide

A hands-on build of pi-jev-router 0.3.0: hook Jev into Pi Agent’s input event, compress the conversation into an ~800-token state, route each message through a four-way Choice with probability gates, and learn the Choice-question trap — semantic overlap splits probabilities 0.45/0.47 — that a 52-scenario labeled test caught and the sum-in-code fix solved (14/14 caught, 0/38 interrupted, 277ms average).

Quick takeaway

Sessions drift: you are twelve messages into debugging a checkout service OOM when you casually ask about a robot vacuum — and that question now rides in the context of every turn that follows. This build (from a Chinese-language creator, a first for this site) puts a typed Jev decision at Pi Agent’s input event, after Enter and before the message reaches the agent. A local filter handles slash commands, shell commands, file paths, sub-15-character inputs, images, queued messages, and first messages for free; everything ambiguous goes to Jev as an ~800-token compressed state (title, last three user messages, reply tail, new input) against a four-way Choice — continue, side_chat, fork, new session. The policy is three lines: irrelevant under 0.6 passes, on-topic over 0.6 passes, the rest triggers a reminder — “the model provides probabilities, the code decides.” The honest engineering story is the evaluation: 52 hand-labeled scenarios run against Jev 1.13 on OpenRouter found the worst version missing 2–3 of 14 drifts and interrupting 1 of 38 normal inputs, because a casual-question case split its probability across two “irrelevant” options (0.45 + 0.47, neither over 0.6). Summing related options in code (0.92 ≥ 0.6 → remind) fixed it to 14/14 caught, 0/38 interrupted, at a 277 ms average and P95 under 0.35 s. Failure handling is fail-open: no key disables the plugin, API errors and over-2.5s responses pass through, and a failed move returns your message to the editor — one missed reminder is the worst case, never a lost message.

Video source

01Coder

14:19WrJAkrfWuS8

Step-by-step walkthrough

  1. 1

    The pain point: sessions drift, and the drift pollutes every turn after it

    You are mid-investigation — the checkout pod keeps getting OOMKilled, and you have traced the memory spike to its image-processing worker — when you remember you need to replace your robot vacuum, and ask about it right there. The agent happily answers. But the question is now permanent chat history: every subsequent turn carries it, and the context gets less and less focused on the problem you were actually solving. Long sessions, casual asides, polluted context — the creator’s three-tag summary of why this needs fixing before the message ever enters the agent.

    Yellow card naming the daily agent pain point: chatting along until the session drifts, tagged with long sessions, casual questions, and polluted context
    聊着聊着,就跑偏了 — the drift happens mid-session, not at the start.Watch at 1:07
  2. 2

    Why a decision model fits: one paragraph vs probabilities

    The comparison the whole design hangs on. Let a large model judge: output is a paragraph of text, latency is seconds, cost per judgment is high. Let Jev judge: output is a fixed answer space with probabilities, latency is about 0.3 seconds, and 300k judgments cost about a dollar. The judgment needs to run before every message, be cheap, and return something the program can use directly — a fixed answer space evaluated in one forward pass is exactly that shape.

    Comparison table of letting a large model judge versus letting Jev judge: paragraph versus fixed answers with probabilities, seconds versus about 0.3 seconds, high cost versus about one dollar per 300 thousand
    Fixed answer space, one forward pass — the shape that fits an input hook.Watch at 2:40
  3. 3

    Hook the input event — then filter locally before spending a call

    Pi Agent (also called PIagent or the Pi coding agent) exposes an input event that fires after Enter but before the message reaches the agent. Hooking there means the drift check runs before the context is polluted. But Jev is not free, so the plugin passes locally everything that cannot drift: slash commands, shell commands, file paths, short instructions under 15 characters, images, messages queued while the agent is running, and the first input of a chat. If a session has been idle over 12 hours or 85% of the context is used, code decides directly. Only genuinely ambiguous messages go to the model.

    Sequence diagram of user, Pi TUI, pi-jev-router, Jev API and current session with the input event hook and a local pass-through card for slash commands shell commands and short inputs
    Enter → Input event → local filter → only the ambiguous cases reach Jev.Watch at 3:32
  4. 4

    Give Jev a compressed view, not the whole conversation

    The RouterState type is deliberately small: session title (300 chars), the last three user messages (300 chars each), the tail of the previous assistant reply (last 1500 chars), and the new input (up to 2000 chars). That is roughly 800 tokens per call — because judging drift only needs to know what the session is currently about, not every detail of it. Smaller state means faster, cheaper, and less noise diluting the decision. In production, the terminal breakdown shows the real budget: questions ~14%, summary ~55%, the live tail the rest.

    RouterState type slide listing session title 300 characters, recent user messages last three by 300, last assistant reply tail 1500, new input 2000 with the compressed view title
    Don’t send the session — send the compressed view (~800 tokens).Watch at 4:17
  5. 5

    Ask one four-way Choice plus the Noul questions

    The route question is a single Choice with four options: continue (keep going in this session), side_chat (an unrelated follow-up question), fork (spin a new version off the current content), and new_session (start an unrelated task). Alongside it sit Noul questions like “is this still on the current topic?” — every answer comes back with a probability. One request resolves the whole routing judgment; the answers are values, not prose to parse.

    Route choice card listing continue to stay, side chat for unrelated follow-up, fork to branch current content, and new session for unrelated task, each with descriptions
    Four fixed options, each with a probability — the whole question.Watch at 5:02
  6. 6

    The policy is three rows: model proposes, code decides

    Jev only outputs probabilities; whether to remind the user is written in code. Row one: irrelevant probability — P(side_chat) + P(new_session) — under 0.6, let it pass. Row two: on_topic at or above 0.6, let it pass. Row three: everything else, remind. That division — the model supplies probabilities, regular code makes the decision — keeps the policy testable, versionable, and independent of any model update. When the reminder fires, the user chooses: stay, fork, start new, or mute for this session.

    Policy table with three rows: irrelevant probability under 0.6 passes, on topic at least 0.6 passes, and everything else reminds, marked pass pass and remind
    让模型出概率,让代码做决定 — the whole gate in three rows.Watch at 5:32
  7. 7

    Read the raw answer: the fork case that correctly passes

    Mid-way through writing an 8-minute video script, the creator asks how a 60-second vertical version would be structured. /route:status shows the raw answer: on_topic 0.85 — still on theme — with continue at 0.47 and most of the remainder on fork; irrelevant sits near zero, under the 0.6 threshold. Fork is not in the “irrelevant” set, so the code passes it and the topic model just answers. The judgment about whether to branch stays with the user, exactly as designed.

    Raw route status answer showing on_topic 0.85 still on theme, the under 0.6 pass rule noting fork is not in the irrelevant set, and probability bars for continue 0.47, fork about 0.5, irrelevant near zero
    on_topic 0.85, fork ≈ 0.5 — the code passes it; branching is the user’s call.Watch at 8:46
  8. 8

    The trap the test caught: semantic overlap splits probabilities

    The labeled evaluation — 52 scenarios, human-labeled, run two rounds each against Jev 1.13 on OpenRouter — found the worst version (0.2.0) missing 2–3 of 14 real drifts while interrupting 1 of 38 normal inputs. The missed case is the lesson: “another project suddenly errors: ECONNREFUSED 127.0.0.1:16379, help me investigate.” Jev’s judgment was correct — unrelated to the current session — but “unrelated” was split across two options: side_chat 0.45 and new_session 0.47. Each sat under the 0.6 gate, so neither fired and the message passed.

    A missed case slide showing an ECONNREFUSED debugging question judged unrelated to the session but split into two options with side chat at 0.45 and new session at 0.47 below the 0.6 threshold
    0.45 + 0.47 and neither crosses 0.6 — the classic Choice-question trap.Watch at 9:47
  9. 9

    Two fixes: exclusive options, or sum in code

    The slide names both cures. Method one: design options to be mutually exclusive so probability never splits. Method two — the one he shipped: sum the probabilities of same-class options in code. Whether the message is a casual question or a core task from another project, both mean the same thing for the reminder decision, so 0.45 + 0.47 = 0.92 ≥ 0.6 fires the reminder. The choice between stay/fork/new still belongs to the user; summing only fixes the reminder gate, which is why it is safe.

    Fix slide for overlapping option semantics with method one making options mutually exclusive and method two summing same-class probabilities in code where 0.45 plus 0.47 equals 0.92 crossing the 0.6 reminder threshold
    0.45 + 0.47 = 0.92 ≥ 0.6 → remind. Sum in code, decide in code.Watch at 10:32
  10. 10

    After the fix: 14/14 caught, 0/38 interrupted, 277 ms average

    The corrected confusion matrix (shipped as 0.3.0) is clean: all 14 drift scenarios reminded, zero of the 38 normal inputs interrupted. Latency: 277 ms average, P95 under 0.35 s — cheap enough to sit on every message. The creator’s caveats are part of the lesson: 52 self-written, self-labeled scenarios are not a public benchmark, and the compressed conversation snippets go to a third-party API (OpenRouter), so think twice before pointing this at sensitive code.

    Confusion matrix after the fix showing all 14 drift cases reminded and zero of 38 normal inputs interrupted, with average latency 277 milliseconds and p95 within 0.35 seconds
    14/14 reminded, 0/38 interrupted — at 277 ms average, P95 < 0.35 s.Watch at 10:57
  11. 11

    Failure handling: fail-open, always

    A check on the critical path of every message must answer “what happens when it breaks?” The table is the answer: no API key — plugin disabled, no impact; API error or rate limit — pass through, cost is one missed reminder; response over 2.5 seconds — pass; no UI mode — no interruption; move failure — the message returns to the editor, never lost. Jev judges, but the final decision always rests with the user, and no failure mode can block an input or destroy a message. That fault-tolerance stance is what makes putting a model on the hot path viable.

    Failure handling table listing missing API key disables, API errors and rate limits pass with one missed reminder, responses over 2.5 seconds pass, no UI mode stays silent, and move failures return the message to the editor
    Worst case in every failure row: one missed reminder — never a lost message.Watch at 11:22

Frequently asked questions

What is session drift and why does it matter for agents?

Session drift is asking something unrelated mid-task — a robot-vacuum question twelve messages into debugging a checkout service. The agent answers, but the question becomes permanent context: every later turn carries it, and the conversation gets progressively less focused on the original problem. Catching drift at the input event — before the message enters the agent — keeps the context clean without policing what users may ask.

What is pi-jev-router and how do I install it?

A Pi Agent extension (Pi’s official name for plugins) that hooks the input event — after Enter, before the message reaches the agent — and uses a typed Jev decision to remind you when a message looks unrelated to the current session. Version 0.3.0 is the one tested in the video; the plugin source, test scenarios, and data are linked in the video description. You need a Jev API key (the video uses Jev 1.13 through OpenRouter); without one the plugin simply stays disabled.

Why a four-way Choice instead of a yes/no question?

Because “unrelated” is not one thing. continue, side_chat, fork, and new_session describe four genuinely different user intents, and the probabilities over all four carry more signal than a single yes/no: the fork case (on-topic, but branching) must pass while an unrelated question must remind. The trap to avoid — which the video demonstrates with real data — is semantic overlap between options: side_chat 0.45 and new_session 0.47 both mean “irrelevant” but neither crosses a 0.6 gate alone. Either make options mutually exclusive or sum same-class options in code.

What does “the model proposes, code decides” mean here?

Jev returns probabilities, never actions. The policy is ordinary code: irrelevant (P(side_chat) + P(new_session)) under 0.6 passes, on_topic at or above 0.6 passes, everything else triggers a reminder where the user chooses to stay, fork, start a new session, or mute for this session. Keeping the policy in code makes it testable and versionable — and it is where the probability-summing fix lives.

How accurate is it, and what did it miss before the fix?

On 52 hand-labeled scenarios run two rounds each: the worst version (0.2.0) caught 11–12 of 14 drifts, missed 2–3, and interrupted 1 of 38 normal inputs. The miss came from probability splitting across overlapping options. After summing same-class options in code, version 0.3.0 reminds on all 14 drifts and interrupts none of the 38, at 277 ms average and P95 under 0.35 s. The honest caveats: the scenarios are self-written and self-labeled — not a public benchmark — and results do not guarantee real-world accuracy.

What happens when the Jev API is down or slow?

Everything fails open. Missing API key: the plugin disables itself. API error or rate limiting: the message passes, costing one missed reminder. Response over 2.5 seconds: pass. No UI mode: no interruption. Failed session moves: the message returns to the editor untouched. The design principle: a drift guard that can block or lose messages is worse than one that occasionally stays silent.

Related guides

More video walkthroughs