Skip to content
Made with Jev

How does Jev compaction work?

Jev compaction replaces the summary with a score. Instead of asking a model to rewrite old turns, a plugin asks Jev whether each tool call and result still matters, deletes what does not, and keeps the rest word for word. One reported run took a Claude session from nearly a million tokens to 86,000 in about one second. This page covers how it works, how it differs from a summary, the 6 builds, and the case against.

Updated 2 Oct 2026 · by Made with Jev

In short

  • Nothing is rewritten. The plugin only deletes or truncates, so what stays is exact.
  • Each tool call gets two yes-or-no questions: should the call stay, and should its result stay in full.
  • The 86K figure is one reported session. No one has published how often an agent still finishes after being compacted this way.
  • The risk is the opposite of a summary. You lose whole results, not words, and a deleted result can be context the agent needed.

How does Jev compaction work?

The best-documented plugin is fast-jev-compaction. It builds one state from the conversation, oldest first, with every tool result replaced by a short note such as “ok, 4213 chars (omitted)”. That state is fitted into a 25,000-token budget in stages, each applied only if the last was not enough: tool inputs cut to 1,000, then 200, then 60 characters, long text shortened to head and tail, old messages collapsed to a note. If it still does not fit, the plugin stops rather than guess.

For every call that is not pinned, Jev answers two questions. Should the call itself stay? Should its result stay verbatim? Calls in the first message and in the newest messages are pinned and never touched. The questions are split across as many requests as needed to stay under Jev’s 32k request limit. The price of each call is on Jev pricing.

Jev compaction versus summary compaction

Summary compactionJev compaction
Who writes new textA model rewrites the old turnsNo one. It only deletes or truncates
What is keptA paraphraseThe original words, in order
How it failsA path, error or constraint is reworded or lostA whole result is dropped that the agent needed later
Reported timeNot publishedAbout one second for a session of nearly 1M tokens

N01ennn’s walk-through of agent loops gives the reasoning for why scoring should beat summarising. Summarising is general compression, which is hard. Compressing for a known task is easier, because if you know what you are looking for, deciding what to drop is simple. It adds that standard compaction rests on one assumption, that every later turn wants the same shared state, and that query-aware filtering does better once you drop it. Treat that as the author’s argument. The plugins below test a small part of it.

Jev compaction plugins and tools

Plugins that drop stale context instead of summarising it

Claude Code first, then Pi and the DeepSeek harness. What is kept stays word for word, and the original can be fetched again.

Alex Volkov

@altryne

This is actually insane. This uses @typesafeai Jev model, as a plugin in Claude to review all the un-nesseasary tool calls, and it takes 1s to run! Like, literally, 1 second to take my Claude session from nearly 1M to ... 86K tokens! 😮 Ask your claude to install it and be amazed Use this prompt ``` Install, and configure : https://github.com/tamaratran/fast-jev-compaction ```

GitHubContext and memory

Pick

fast-jev-compaction

Stars
7.1k
Reported
~1M → 86K tokens

tamara

@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

XContext and memory

Pick

Instant compaction for Claude

Ishwar

@WesternCube

Instant Claude Code compaction is my favorite use of Jev so far For people going to complain about me letting my context get so long, this was a server redeploy task because Oracle terminated my free VPS with no explanation so it had a lot of moving pieces

XContext and memory

Instant Claude Code compaction

GitHubContext and memory

pi-jev-compaction

Automatic context clearing for Pi, without losing the conversation.

Nour

4

GitHubContext and memory

dsh-jev-prune

Context compaction for DeepSeek Harness, judged by Jev.

yangyu666

3

Deciding when to compact

The question before the question: is this session at a point where cutting is safe. One build, tuned against sessions someone labelled by hand.

Kun Chen

@kunchenguid

almost every day i hear people ask "when should i /compact my session" there's no easy answer because it depends on how likely your future action will need detailed context in the existing window but we have Jev now! introducing compact-adviser - an agent plugin you can use in claude and pi today to help determine whether you're likely at a task boundary that's safe to compact https://github.com/kunchenguid/compact-adviser i built a private eval set from 40 real sessions and manually labeled all the safe vs unsafe checkpoints to evaluate this, and hillclimbed the Jev prompt till it performed quite well i also made it so that the classifier will - optimize for precision (not triggering a compaction prematurely) when context window is small - and gradually shift to optimize for recall (not missing an opportunity to compact) when context window fills up, because the cost of not compacting becomes higher, and at the end the agent will be forced to compact anyway it supports a "hint" mode (just give you a hint and it's up to you to run /compact) vs "auto" mode which runs compaction whenever Jev says it's safe to do so if you have Jev and want to put your compaction on autopilot, try this out and let me know how it goes! support for more harness is coming soon as well

GitHubContext and memory

compact-adviser

Eval set
40 sessions

When to compact with Jev: the timing question

Dropping context is only safe at a boundary. Kun Chen’s compact-adviser asks Jev whether a session is at a task boundary where cutting is safe. He built an eval set from 40 real sessions, hand-labelled safe and unsafe checkpoints, and tuned the question against it. The classifier favours precision while the window is small and shifts toward recall as it fills. The set is private, so the result cannot be re-run.

The case against Jev compaction

Theo’s argument against it is the most useful counterweight on this site. His three points: per-tool-call filtering is not compaction, it loses reasoning, it raises cache-write costs, and it leaves agents that loop. All three are reported, not measured here.

There is also a design answer. The guide above proposes replacing the keep-or-drop choice with a policy per chunk: show nothing, show a short summary, show a long summary, or show the full text. The plugins in this directory drop or truncate. None of them implements that yet.

What Jev compaction has not proved yet

Task success after compaction. Tokens removed is an easy number. Whether the agent still finishes the job is the one that matters, and no build here has published it as a repeatable run.

The size of the saving. The one reported session was tried and shared by Alex Volkov. It is a result, not a typical case.

The cost of rebuilding the cache. Theo reports higher cache-write costs, and no one has published the net effect on a bill.

How to try Jev compaction safely

  1. Try it on a session you can replay. Keep the original so you can compare the finished work.
  2. Pin what must never go. The first message and the newest ones are pinned by default in the plugin above.
  3. Prefer a plugin that keeps originals retrievable. pi-jev-compaction lets the agent fetch a dropped result without running the command again.
  4. Watch for loops. If the agent repeats a step it already did, the result it needed was dropped.

Compaction is one half of the saving. The other half is choosing the model for each turn, on the Jev router page. The whole stack in a coding agent is on Jev with Claude Code, and the single-agent layer behind both is on the agentic harness page. A paste-into-agent version of these steps is in the Jev compaction prompt.

Where the Jev compaction figures come from

The one reported run, the case against it, and the guide behind the design argument.

Alex Volkov

@altryne

This is actually insane. This uses @typesafeai Jev model, as a plugin in Claude to review all the un-nesseasary tool calls, and it takes 1s to run! Like, literally, 1 second to take my Claude session from nearly 1M to ... 86K tokens! 😮 Ask your claude to install it and be amazed Use this prompt ``` Install, and configure : https://github.com/tamaratran/fast-jev-compaction ```

Guidex.com

The compaction plugin, tried

Alex Volkov: a Claude session went from nearly 1M tokens to 86K in about one second.

Theo - t3.gg

@theo

This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the agent focused, not just deleting noise. It should be used sparingly when context gets too long, not constantly to keep context small. 2. Jev doesn't even know what it's deciding on! Models use the context of the thread to decide what to keep or not keep in a summary. This implementation goes through on a "line-by-line" (per tool call) basis to decide what should be left or deleted. Not only does this 32k token context model know very little of what happened before, but (in this implementation) it doesn't even know what the result of the tool call is! Deleting these things randomly will keep the model from knowing what it's tried and dooms you to end up in "stupid loops" where the model keeps trying the same thing over and over. 3. You're giving up the reasoning entirely Frontier models from OpenAI, Anthropic, XAI, and Google do not share reasoning traces over the API. They share encrypted payloads, which Jev cannot see (and often will drop). Anthropic is even stricter with this, requiring you to preserve the entire history in order to get any of the reasoning data. As a result, using this in Claude Code guarantees the model will act way dumber. 4. Models are tuned on their compaction flows For the last year, Frontier Labs have been including compaction and long runs as part of the training process. These models have learned ways to compact that are more effective than any rudimentary solution. Fun fact: If you switch models in Codex and compaction is necessary, compaction will run on the model that was previously used in the thread. 5. Cache writes are more expensive than cache reads. Cache writes are the biggest cost by far for agents. I often see cache write costs go over 60% of my total LLM spend in my personal use of Claude Code and Codex. Cache writes are insanely expensive when data earlier in the history is changed (because the old cache is invalidated when things change at the top). Every history edit requires a cache rewrite for ANY data past the history edit. If your history is "1,2,3,4,5,6" and you delete "2", you have to rewrite "3,4,5,6". This is more expensive than leaving "2" in the history. Good news. Since we're already killing all of the reasoning tokens by doing this stupid compaction strategy, the rewrite cost won't actually be that high because the model is missing so much data! 🙃🙃 6. The implementation is hot garbage. > "Whatever is not kept is deleted permanently, but the assistant can always re-run a tool or re-read a file." Good luck with that one. To be clear: this is a cool experiment and I find it genuinely interesting. That said, if you think this style of bs filtering on a probability threshold is actually a compaction strategy, I highly recommend you just use the defaults in tools like Claude Code and Codex. You're much less likely to hurt yourself that way.

Guidex.com

The case against Jev compaction

Theo on why per-tool-call filtering is not compaction: lost reasoning, higher cache-write costs and agents stuck in loops.

Common questions

What is Jev compaction?
A way to shrink an agent's context without writing a summary. Jev answers a yes-or-no question for each tool call and each tool result, such as whether it is still relevant, and the plugin deletes or truncates what fails. Everything kept stays word for word.
How much does Jev compaction save?
One reported session went from nearly 1M tokens to 86K in about one second, as Alex Volkov described it. That is one run, not a benchmark, and tokens removed say nothing about whether the agent still finished its task.
Does Jev compaction lose information?
It can. It never paraphrases, so a path or an exact error cannot be reworded, but a deleted tool result is gone unless the plugin lets you fetch the original. Theo's case against it reports lost reasoning, higher cache-write costs and agents stuck in loops.
When should I compact a session?
At a task boundary, not mid-task. One plugin asks Jev whether the session is at such a point, and its author tuned the question against 40 real sessions he labelled by hand.
Does it work outside Claude Code?
Yes. There are versions for the Pi agent and for the DeepSeek harness, built by other people on the same idea.

Made with Jev is independent and not affiliated with TypeSafe AI. Every figure on this page is the one its author published, linked to where it can be checked.