Guides / illustrated walkthrough
Jev Context Compaction: Prune AI Agent Memory Without Generative Summaries
A frame-by-frame guide to fast-jev-compaction: how Alex Hitt’s 7-minute walkthrough replaces lossy generative autocompaction with Jev’s typed yes/no pruning — anchored messages, 8-stage compression, a 0.5 keepThreshold, and an 86% token cut.
Quick takeaway
Alex Hitt’s 7-minute motion-graphics walkthrough (fast-jev-compaction on GitHub) shows why generative autocompaction is the weak link in long coding-agent sessions: when the context window fills, the agent pauses for 30–60 seconds while an LLM rewrites its own history, and the rewrite routinely drops exact file paths, error stack traces, and standing user constraints like “never edit this generated directory.” The replacement is non-destructive pruning powered by Jev’s System One parallel sampling. The session state is reshaped in three moves: the first instructional messages and the most recent turns are permanently anchored and never evaluated; unanchored plain text is mathematically reduced by an 8-stage progressive compression algorithm until it fits below a static maxStateTokens cap (25,000 in the repo’s default), with token counts estimated by character-weight heuristics (letters ≈ 0.16, digits 0.50, symbols 1.00) instead of a tokenizer library; and if the state still will not fit, the run safely reverts to native summarization. Every unanchored tool call then faces two typed Noul questions evaluated in parallel — Question A: is merely knowing that this call happened, with its arguments, still relevant? Question B: is the exact textual output strictly necessary, or would re-running the tool suffice? A configurable keepThreshold (0.5 in the demo) turns the probability pairs into a three-way verdict: both below → tool call and result are deleted; only the call relevant → keep the log line, truncate the heavy output; the result needed → preserve the text verbatim. Measured on dense TypeScript sessions, the model identified repetitive bash output as safe to delete and cut 86% of tokens, pruning session arrays beyond 187,000 tokens in under 1.5 seconds — and because output tokens are free with parallel request routing, evaluating 400 simultaneous tool calls resolved at 1.2 cents total.
Video source
Alex Hitt
Step-by-step walkthrough
- 1
See the saturation problem: 84,455 tokens and climbing
A continuous engineering session accumulates terminal commands, grep results, and file reads, and every one of them consumes a slice of the base model’s context window. The walkthrough opens with the counter that matters — a context window meter pushing past 84,455 of 128,000 tokens while whole documents of tool logs pile up underneath it. This is the moment every agent user knows: the session is about to hit its limit, and something is about to rewrite your history.

Every bash command and grep result earns a permanent seat in the window — until something intervenes.Watch at 0:26 - 2
Compare the two architectures: lossy summary vs verbatim pruning
The industry-standard response to saturation is generative self-compaction: pause, have an LLM read the entire conversation, and write a condensed narrative. The flowchart contrasts that with fast-jev-compaction’s System One pruning. On the left, an autoregressive LLM destructively combines raw context — tool calls plus user constraints — into a lossy summary; digests routinely extract exact file paths, stack traces, and explicit restrictions like “never edit a dynamically generated directory.” On the right, a System One parallel sampler evaluates the same raw context and emits per-item Noul probabilities (0.78, 1.00 in the diagram) that route each piece to keep or discard — the constraints bypass deletion entirely, word for word.

Generation rewrites memory; evaluation grades it. The right-hand path never paraphrases a constraint.Watch at 2:36 - 3
Anchor the edges, compress the middle in 8 stages
Before Jev can evaluate anything, the repo protects the two ends of the conversation: the first instructional messages (the system prompt and rules you negotiated at start) and the most recent turns are permanently anchored and exempt from pruning. Everything in between is reduced by an 8-stage progressive compression algorithm — large blocks of plain text are mathematically truncated, tool inputs cut, and messages grouped into single blocks — until the whole state fits below a static maxStateTokens ceiling.

Compression is staged, not one-shot — each pass buys another chunk of headroom.Watch at 3:25 - 4
Estimate tokens with character weights, not a tokenizer
Loading a full tokenization library just to decide whether to prune is its own tax, so the repo estimates with character-count heuristics: roughly six letters per token, digits at half a token, symbols at one. The video’s weight table prices alphabetical characters at about 0.16 tokens, numerics at 0.50, and specials at 1.00. If the gross estimate lands above the 25,000 maxStateTokens limit, an aggressive reduction pass fires before anything is sent for evaluation — Jev only spends compute on states that are structurally sound.

Cheap arithmetic stands in for a tokenizer — the estimate only has to be good enough to gate the request.Watch at 3:31 - 5
Confirm the state fits before spending a Jev call
The checkpoint frame shows the happy path: after compression the state is certified “structurally sound” at 18,500 tokens against the 25,000 cap. If the eight stages had failed to get under the limit, an exception protocol would interrupt compaction and revert to the native summarization mechanism instead — a deliberate escape hatch so the deterministic pipeline never sends an oversized state to the evaluator.

Fit check passed at 18,500 tokens — the evaluator only ever sees states that comply.Watch at 3:42 - 6
Ask two typed questions for every unanchored tool call
Each tool call that is not anchored faces a dual yes/no evaluation, and the matrix is the heart of the system. Question A grades the call’s continued relevance: is merely knowing that this call happened, with these arguments, still worth keeping? Question B grades the result: is the exact textual output strictly necessary to preserve, or would running the tool again later be enough? In the worked example a log-read scores 0.67 on relevance but only 0.34 on necessity, while a passed validation keeps 0.92 on the result side — two independent probabilities instead of one opaque summary verdict.

Relevance and necessity are different questions — and the pipeline asks them separately, in parallel.Watch at 5:00 - 7
Gate deletion on a keepThreshold
The probability pairs become a three-way verdict through one configurable number. In the demo a document pins at 0.87 against a 0.5 keepThreshold, which lands it firmly on the preserve side. The full policy: if the result’s probability clears the threshold, the literal text is preserved verbatim; if only the call itself is relevant, the execution log line stays but the heavy output is truncated; if both probabilities fall below, the call and its result are obsolete data and are permanently removed — with every surviving operation keeping its chronological identity, unmodified.

One threshold, three outcomes: keep verbatim, keep the log and truncate, or delete outright.Watch at 5:08 - 8
Read the results: 86% of tokens gone, 1.5 seconds to prune
The benchmarks panel reports what the policy earns. In dense TypeScript sessions, identifying repetitive bash output as safe to delete removed up to 86% of tokens — Terminal 1 kept roughly 15% of its session, Desktop about 20%. Because the model is not autoregressive, the whole evaluation is fast: session arrays beyond 187,000 tokens were pruned in under 1.5 seconds, against the 30–60 second freeze a generative summary imposes three to five times in a two-hour session.

Retained versus pruned per session — the repetitive bash output never survives the gate.Watch at 5:58 - 9
Check the economics: 400 tool-call evaluations for 1.2 cents
The API cost scaling table prices the continuous evaluation cycle: 40 tool calls in 1 request cost $0.001, 150 calls $0.005, and 400 extreme simultaneous tool calls resolve for exactly $0.012 — 1.2 cents — because output tokens are free and requests route in parallel. The video’s closing argument: as autonomous sessions grow longer, delegating memory management to a deterministic, non-generative evaluator keeps the agent’s history accurate without anyone paying the summarization tax again.

The full pruning cycle for a 400-call session costs about one-tenth of a single frontier-model summary.Watch at 6:18
Frequently asked questions
What is generative autocompaction, and why replace it?
It is the standard mechanism where an agent near its context limit pauses and asks an LLM to rewrite the conversation history as a condensed narrative. The video’s critique: the rewrite is lossy (exact file paths, stack traces, and standing user constraints get paraphrased away), slow (a 30–60 second interruption, three to five times per two-hour session), and fragile — you are trusting a generative model to faithfully summarize its own memory.
How does Jev decide which tool calls to delete?
Every unanchored tool call is evaluated against two typed Noul questions: A — is knowing this call happened (with its arguments) still relevant? B — is the exact output text strictly necessary, or could the tool simply be re-run? A keepThreshold (0.5 in the demo) converts the probability pair into keep verbatim, keep the log but truncate the output, or delete both.
Does compaction lose my user constraints, like “never edit this folder”?
No — that is the core design win. The first instructional messages and the most recent conversation turns are permanently anchored and never evaluated for deletion, so standing restrictions survive word for word. Only unanchored middle-of-session tool traffic is subject to the dual-question pruning.
What is maxStateTokens in fast-jev-compaction?
It is the static cap (25,000 tokens by default in the repo) that the session state must fit under before Jev evaluates it. Unanchored text passes through an 8-stage progressive compression algorithm, with token counts estimated by character-weight heuristics (letters ≈ 0.16, digits 0.50, symbols 1.00) instead of loading a tokenizer library. If the state still will not fit, the run safely reverts to native summarization.
How much does continuous Jev compaction cost?
In the video’s measured table, 400 simultaneous tool-call evaluations cost $0.012 total — output tokens are free and requests are routed in parallel. The pruning itself ran at up to 86% token reduction and processed a 187,000-token session array in under 1.5 seconds.
Does this only work with Claude Code?
The empirical data in the video comes from Claude Code terminal sessions, but the mechanism is agent-agnostic: any harness that accumulates tool calls and can pose two yes/no questions per record — “was this call relevant, is the result necessary?” — can delegate pruning to Jev the same way.
Related guides
Building an Agent Harness with Jev
The architecture companion: typed gates around a coding agent, from allow/block verdicts to confidence-gated tools.
ReadJev Tips: 8 Best Practices for Better Decisions
Includes the trim-the-state rule this pipeline industrializes — send only the context the question needs.
ReadWhat Is Jev? System One Models Explained
The concept page behind the parallel-sampling architecture: typed questions, label probabilities, no generation.
ReadMore video walkthroughs
- Jev Classification Quickstart: OpenRouter API, Primitives & Real Probabilities
- Jev Architecture Explained: Why 70ms Decision Models Beat LLMs for Workflow Automation
- Ultra-Fast Browser Agents with Jev: 178ms DOM Loops & Dual Model Orchestration
- Open Jev Models Are Here: Semif, Nimble, Decider, DiffusionGemma & Laya Hands-On
- Jev vs LLM: Will Jev Replace LLMs? Krish Naik's Whiteboard Explainer
- Jev Trader Tutorial: Build a Subsecond AI Trading Bot on Monad
- Jev Model Router: Build a Privacy-Gated LLM Router with Jev & OpenJev
- Jev Tutorial for Beginners: State, Questions & the TypeScript SDK
- Run Jev Locally: Kev, SemIf & Von on Your Own GPU (OpenJev Guide)
- Jev RAG Reranker: Policy-Steered Reranking for Retrieval-Augmented Generation
- When to Use Jev: An Engineer's Audit of Claims, Gates, and Failure Modes
- LangChain + Jev Integration Tutorial: Routing, Guardrails & Evals
- Jev MCP Server: Connect Jev Decisions to Claude Code & Cursor
- Jev vs Luna: Independent Benchmarks Put "Better, Faster, Cheaper" to the Test
- Jev Agent Harness: Where the Decision Gate Sits in Your LLM Loop
- Jev Playground Walkthrough: The Hotdog Lesson, Criteria, and a Four-Console Token Test
- Jev Text Classification API: Zero-Shot CLI & REST with classifier.dev
- Jev API Examples: First Request, curl & All Three Question Types
- Jev Log Triage with Expanso Edge
- Jev Lead Enrichment with Treg: ICP and Signup Scoring
- Use Jev Decision Nodes in Heym for Model Routing
- Laya Tutorial: Open-Source AI Routing With Calibrated Probabilities (Laya vs Jev Setup)
- Train Your Own Jev: Fine-Tune a Jev-Style Decision Model for $5–$17 (What You Can and Cannot Train)
- Jev Tips: 8 Best Practices for Better Decisions (State, Questions, Criteria & Thresholds)
- Jev as an LLM Judge: Confidence-Gated Cascades at 0.36% of the Cost
- TypeSafe Computer Use: Local Desktop Automation with Jev, Step by Step