Prune tool output by rule, leave the reasoning chain alone
My agent hit the context ceiling last week at turn 61. The framework did what every framework does now: it called a model to summarize the conversation, waited eleven seconds, and handed back a paragraph that had quietly
My agent hit the context ceiling last week at turn 61. The framework did what every framework does now: it called a model to summarize the conversation, waited eleven seconds, and handed back a paragraph that had quietly dropped the exact error string I needed. The task failed. Not because the model was dumb ā because the compaction step decided what mattered and decided wrong.
I've since ripped the summarizer out of my own loop and replaced it with about eighty lines of bookkeeping. Here's the reasoning, and the rule.
Where the tokens actually are
Pull the token breakdown from any long agent trace and the shape is always the same. The assistant's reasoning ā the text where it thinks, plans, decides ā is a thin ribbon. The tool results are the ocean. File reads, grep dumps, test output, JSON blobs, directory listings, stack traces. In my own traces, tool results run 70ā85% of the window on a real coding task. The reasoning chain is usually under 10%.
So the default compaction strategy spends its most expensive, most lossy, least deterministic operation on the part of the context that is cheapest to keep and most valuable to preserve. That's backwards.
Why summarization is the wrong first move
Three problems, in order of how much they've cost me.
It's slow and it's blocking. You hit the ceiling precisely when you're mid-task, and now you insert a full model round-trip before the next step. On a local model that's seconds. On a hosted one it's seconds plus a bill, and you pay it again every time you cross the threshold.
It's lossy in a way you can't audit. The summarizer doesn't know what the next turn needs. It compresses by salience, and salience is a guess. The error string, the exact file path, the off-by-one in the test name ā those look like noise to a summarizer and like everything to the next tool call.
It's non-deterministic. Same conversation, two runs, two different summaries. That kills prompt caching, kills reproducibility, and makes a bug report impossible to replay. I've spent an afternoon on "it did something weird yesterday" and the summary was different the second time.
The rule: liveness, not importance
Here's the reframe. You don't need to decide which tool results are important. You need to decide which ones are still referenced. That's a much easier question, and it has a mechanical answer.
A tool result is live if the live reasoning chain still points at it. Concretely, walk backwards from the newest message and mark a tool result live when:
- its tool-call ID appears in an assistant message that's still in the window, or
- a later assistant message quotes or paraphrases its content, or
- it's one of the last K results (I use K=3), or
- it's explicitly pinned.
Everything else is stale. Not unimportant ā stale. The distinction matters, because "unimportant" requires judgment and "unreferenced" doesn't.
The reasoning chain ā every assistant message, every tool call (the request, not the result), every user turn ā stays untouched. It's small. It's the actual state of the task. Deleting it is how agents forget what they were doing.
What a pruner looks like
Roughly:
- Take the message list.
- Walk it backwards, building a set of live tool-call IDs.
- For every tool result not in that set and not in the last K, replace its content with a tombstone.
- Recompute the token count. If you're still over budget, prune harder ā shrink K, then drop the oldest complete tool-call/result pairs.
- Only if you're still over budget after that, reach for a summarizer.
Step 3 is the whole trick and it's cheap. No model call, no latency, no variance. Same input, same output, every time.
Tombstones, not deletions
Don't delete the tool result message. Replace its body with something like [pruned: read_file src/parser.py, 4.2k tokens].
Two reasons. First, if you delete the pair entirely, the model has no record that it already ran that call and will happily run it again ā you've turned a context problem into a loop. Second, the tombstone is a cheap index. When the agent needs that file again, it knows it read it, and re-reading is one tool call instead of a rediscovery.
Keep the tool call message intact either way. It's tiny and it's reasoning.
What to pin
A few things should never be pruned by rule alone:
- The current plan or todo list, if the agent maintains one. That's the spine.
- The most recent error or test failure. It's usually the reason for the next three turns.
- Anything the agent explicitly wrote down as a note. If it took the trouble to record it, respect it.
- The original task statement. Obviously.
I mark these with a flag at write time rather than trying to detect them later. Detection is another judgment call, and judgment calls are what I'm trying to remove.
When summarization is still right
I'm not saying delete your summarizer. I'm saying it shouldn't be the first thing you reach for, and it shouldn't be the only thing.
Summarization earns its keep when the reasoning chain itself is the bulk ā long research or planning sessions where the agent has produced pages of deliberation and you genuinely need the gist. It's also fine as a second tier, applied to the oldest third of the conversation where nothing is live anyway.
Pruning first makes summarization better, too. Summarize a context that's already been pruned and the summarizer spends its budget on reasoning instead of on a directory listing from forty turns ago.
The payoff
The obvious win is speed and cost. The less obvious win is that my agent's context is now a deterministic function of its history. Same trace in, same context out. I can log it, diff it between runs, and replay a failure exactly. Prompt caching works again because the prefix is stable.
And the failures got boring, which is what I wanted. The agent doesn't forget the file it's editing halfway through editing it. It forgets the ls output from turn 12, which it was never going to look at again.
Tool results are most of the tokens and least of the meaning. Prune them by rule. Leave the thinking alone.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes ā full credit and traffic to the original publisher.