Your Agent Rereads Every Tool Result. Build a Tiny Context Compactor in TypeScript.
An agent loop has a quiet habit. Every time it calls the model, it sends the whole conversation again. The system prompt. The task. Every tool call. Every tool result. So a 20,000-token test log is not paid for once.
An agent loop has a quiet habit.
Every time it calls the model, it sends the whole conversation again.
The system prompt. The task. Every tool call. Every tool result.
So a 20,000-token test log is not paid for once.
It is paid for on every turn after it lands.
That is the part I keep thinking about, because the last three weeks of agent news kept pointing at it.
- On September 17, 2026, researchers published An Empirical Study of Harness Design for Coding Agents. Across 176 matched settings, they found context management matters more as the window gets tighter, and most of its benefit comes from preventing context-overflow failures. Rule-based elision before LLM summarization gave the best overall efficiency.
- On September 21, 2026, the Strands Agents team released Strands harness and reported 28% lower token cost across six benchmarks on the same Claude or GPT models. In their words, "Our default context management largely drove the token-efficiency and accuracy": tool results over about 1,500 tokens get truncated, and compaction triggers above 85% of the window. That is their benchmark, not an independent one.
- On September 22, 2026, CliffCompaction reported up to 50% lower cost under a bounded context. Its rule is strict: only truncate or drop content, never rephrase it, and never compact a compaction.
- On October 3, 2026, DeepSeek published Harness v0.2.1-alpha.1, the latest build of its open-source, everything-is-a-plugin harness. In the Strands benchmark, DeepSeek Harness was the most token-efficient overall, and it also typically had the lowest accuracy.
Different teams. Same lever.
That last point matters too.
Spending fewer tokens is easy. Spending fewer tokens without losing the one line that mattered is the actual job.
So let's build a tiny version.
By the end, you'll run one command:
npx tsx compact.ts
And watch one scripted agent run replayed under four context policies, with the tokens each one actually sends to the model.
No API key.
No model.
Just TypeScript.
One honesty note: this is my small model of the idea, not the Strands or DeepSeek implementation. The tools, the outputs and the token counter are mocked, and every number in the output comes from example inputs.
Table of Contents
- What We Are Building
- Project Setup
- Step 1: A Scripted Agent Run
- Step 2: Cap Big Tool Results
- Step 3: Elide Old Results
- Step 4: Build the View
- Step 5: Replay and Count
- Where It Breaks Down
- The Bigger Idea
What We Are Building
The agent keeps a full log. That never changes.
What changes is the view: the slice of that log the model reads on each call.
Full log (never edited)
β
Cap: big tool results become a preview + a ref
β
Keep: failure lines always survive the cut
β
Elide: old tool results become one-line stubs
β
Drop: over 85% of the window, drop the oldest whole messages
β
View (what the model reads this turn)
Four rules. No summarizer.
Project Setup
mkdir tiny-context-compactor && cd tiny-context-compactor
npm init -y
npm install -D tsx typescript @types/node
Save the following TypeScript blocks in order as compact.ts.
Step 1: A Scripted Agent Run
// compact.ts: a tiny context compactor for an agent loop.
// Everything is mocked: the tools, their output, and the token counter (about 4 characters per token).
// The run, the 32,000-token window and the thresholds are example inputs. No model, no API key.
// Step 1: a scripted agent run and a rough token counter
type Msg = { turn: number; role: "system" | "user" | "assistant" | "tool"; tool?: string; text: string };
const tokens = (s: string) => Math.ceil(s.length / 4); // a heuristic, not a real tokenizer
const lines = (n: number, f: (i: number) => string) => Array.from({ length: n }, (_, i) => f(i)).join("\n");
const FAIL = "FAIL src/checkout.test.ts > applies 10% coupon: expected 90, received 100";
const testLog = (failAt: number | null) =>
lines(1800, (i) => (i === failAt ? FAIL : `PASS src/suite-${i % 97}.test.ts > case ${i} (${(i * 37) % 90 + 3} ms)`));
const source = (name: string, n: number) =>
lines(n, (i) => ` const ${name}${i} = applyRule(cart.items[${i % 12}], rules.${name}); // line ${i + 1}`);
const RUN: [tool: string, call: string, output: string][] = [
["list_files", "list_files src/", lines(40, (i) => `src/module-${i}.ts`)],
["run_tests", "run_tests", testLog(1137)],
["read_file", "read_file src/checkout.ts", source("checkout", 220)],
["grep", "grep -n coupon src/", lines(30, (i) => `src/module-${i}.ts:${i * 7 + 3}: coupon`)],
["read_file", "read_file src/coupon.ts", source("coupon", 160)],
["read_file", "read_file CHANGELOG.md", lines(900, (i) => `- v2.${900 - i}.0: internal release notes, item ${i}`)],
["edit_file", "edit_file src/coupon.ts", "ok: 1 line changed"],
["run_tests", "run_tests", testLog(null)],
["git_diff", "git diff", "- return price;\n+ return price * (1 - coupon.percent / 100);"],
];
const LOG: Msg[] = [
{ turn: 0, role: "system", text: "You are a coding agent. Use tools. Verify before you finish." },
{ turn: 0, role: "user", text: "The checkout test is failing. Find the bug and fix it." },
...RUN.flatMap(([tool, call, output], i): Msg[] => [
{ turn: i + 1, role: "assistant", text: `call ${call}` },
{ turn: i + 1, role: "tool", tool, text: output },
]),
];
const CALLS = RUN.length + 1; // the model is called once per turn, plus once to write the answer
The run is a coding agent fixing a failing checkout test. Nine tool calls.
Three of them are big: two run_tests logs of 1,800 lines each and a 900-line CHANGELOG. Only one line in the first test log matters.
The token counter is a rough heuristic, about four characters per token. Real tokenizers differ, so treat every count here as relative.
Step 2: Cap Big Tool Results
// Step 2: cap big tool results and keep the full text out of the context
const store = new Map<string, string>();
const CAP = 1500; // offload a tool result above this many tokens
const PREVIEW = 750; // keep this many tokens of it in context
function headTail(text: string, keep: string[] = []): string {
const all = text.split("\n");
const pick = (from: string[]) => {
const out: string[] = [];
for (const l of from) {
if (tokens(out.join("\n")) + tokens(l) > PREVIEW / 2) break;
out.push(l);
}
return out;
};
const head = pick(all);
const tail = pick([...all].reverse()).reverse();
const hidden = all.length - head.length - tail.length;
return [...head, `... ${hidden} lines offloaded ...`, ...keep, ...tail].join("\n");
}
// Same budget, but lines that look like failures always survive the cut.
const isFailure = (line: string) => /FAIL|ERROR/.test(line);
const errorsFirst = (text: string) =>
headTail(text, text.split("\n").filter(isFailure).slice(0, 5));
function capped(m: Msg, preview: (t: string) => string): string {
if (m.role !== "tool" || tokens(m.text) <= CAP) return m.text;
const ref = `offload://turn-${m.turn}`;
store.set(ref, m.text);
return `${preview(m.text)}\n[full result: ${tokens(m.text)} tokens at ${ref}]`;
}
// A tool the model can call to get back what was cut.
const retrieve = (ref: string, pattern: RegExp) =>
(store.get(ref) ?? "").split("\n").filter((l) => pattern.test(l)).slice(0, 5);
The most important function is errorsFirst().
headTail() keeps the start and the end of a big result, which is what most truncation does.
But a failing test does not politely sit at the top or the bottom of the log.
errorsFirst() spends the same budget and pins anything that looks like FAIL or ERROR into the preview.
The full text goes into a store, with a ref the model can pass to retrieve().
Step 3: Elide Old Results
// Step 3: elide old tool results with a rule, before any summarizing
function elided(m: Msg, now: number, keepRecent: number): string | null {
if (m.role !== "tool" || now - m.turn <= keepRecent) return null;
store.set(`offload://turn-${m.turn}`, m.text);
return `[elided: ${m.tool} from turn ${m.turn}, ${tokens(m.text)} tokens, ref offload://turn-${m.turn}]`;
}
Once a tool result is more than two turns old, it becomes one line: which tool, which turn, how big, and where it lives.
The model still knows the result existed.
It just stops paying for it.
Step 4: Build the View
// Step 4: build the view for each model call. Truncate or drop. Never rewrite.
type Policy = { name: string; preview?: (t: string) => string; keepRecent?: number; window?: number };
const WINDOW = 32_000;
function view(call: number, p: Policy): Msg[] {
// Always rebuilt from the original log,
// so a compaction never compacts a compaction.
let v = LOG.filter((m) => m.turn < call).map((m) => {
const stub = p.keepRecent === undefined ? null : elided(m, call, p.keepRecent);
return { ...m, text: stub ?? (p.preview ? capped(m, p.preview) : m.text) };
});
const size = () => v.reduce((n, m) => n + tokens(m.text), 0);
const pinned = (m: Msg) => m.turn === 0 || call - m.turn <= (p.keepRecent ?? 0);
const over = () => !!p.window && size() > p.window * 0.85;
while (over() && v.some((m) => !pinned(m))) {
v.splice(v.findIndex((m) => !pinned(m)), 1); // drop oldest, whole
}
return v;
}
The important part is the first line of view().
The view is rebuilt from the original log on every call. Nothing is edited in place, so a stub is never made from another stub.
Then, if the view is still above 85% of the window, it drops the oldest unpinned message, whole. The system prompt, the task and the last two turns are pinned.
Truncate or drop. Never rewrite.
Step 5: Replay and Count
// Step 5: replay the same run under each policy and count what the model actually reads
const POLICIES: Policy[] = [
{ name: "naive" },
{ name: "cap", preview: headTail },
{ name: "cap+errors", preview: errorsFirst },
{ name: "full", preview: errorsFirst, keepRecent: 2, window: WINDOW },
];
console.log(`${CALLS} model calls, ${RUN.length} tool results, window ${WINDOW.toLocaleString("en-US")} tokens (example run)\n`);
console.log(`${"policy".padEnd(12)} ${"billed".padStart(9)} ${"peak".padStart(9)} ${"fits".padEnd(14)} FAIL seen at call 3`);
for (const p of POLICIES) {
let billed = 0, peak = 0, overflow = 0;
let sawFail = false;
for (let call = 1; call <= CALLS; call++) {
const v = view(call, p);
const size = v.reduce((n, m) => n + tokens(m.text), 0);
billed += size;
peak = Math.max(peak, size);
if (size > WINDOW && !overflow) overflow = call;
if (call === 3) sawFail = v.some((m) => m.text.includes(FAIL));
}
const fits = overflow ? `no (call ${overflow})` : "yes";
const n = (x: number) => x.toLocaleString("en-US").padStart(9);
console.log(`${p.name.padEnd(12)} ${n(billed)} ${n(peak)} ${fits.padEnd(14)} ${sawFail ? "yes" : "no"}`);
}
console.log(`\nretrieve("offload://turn-2", /FAIL/):`);
for (const l of retrieve("offload://turn-2", /FAIL/)) console.log(` ${l}`);
Run it:
npx tsx compact.ts
Real output from my run:
10 model calls, 9 tool results, window 32,000 tokens (example run)
policy billed peak fits FAIL seen at call 3
naive 290,208 58,205 no (call 7) yes
cap 22,905 4,242 yes no
cap+errors 23,057 4,261 yes yes
full 9,367 1,637 yes yes
retrieve("offload://turn-2", /FAIL/):
FAIL src/checkout.test.ts > applies 10% coupon: expected 90, received 100
Here is how I read that table.
naive would read 290,208 input tokens across ten calls, and it never gets the chance. It passes the 32,000-token window on call 7, right after the CHANGELOG read lands on top of the first test log.
cap fits easily. But the head-and-tail preview cut the one FAIL line out of the first test log. On the very next call, the model cannot see why the test failed.
cap+errors costs 152 more tokens than cap across the run, and the failure line is back.
full adds elision and the 85% drop rule. It reads 9,367 tokens across the run, with a peak of 1,637.
And here is a detail I like. In this run, the 85% rule never fires. Capping and elision already kept every view small.
That lines up with the harness study. Cheap rules first. The expensive machinery is a safety net.
Where It Breaks Down
This demo is small on purpose. Here is what sits right outside it.
Recoverable is not the same as recovered. My
retrieve()tool can get theFAILline back. The harness study found that making elided content recoverable "adds machinery that models rarely use and yields no accuracy gain." A ref helps only if the model thinks to follow it. Put the signal in the preview.Regex is a guess about what matters.
FAIL|ERRORworks for this test runner. A stack trace, a warning that turns into an outage, or a JSON field namedstatuswill not match. Each tool deserves its own preview rule.Compaction can break your prompt cache. Providers cache a request's reused prefix. When a result turns into a stub, everything after it changes, and the cache has to warm up again. Strands ships prompt caching and context management as defaults side by side, and the two pull against each other. Elide at stable boundaries, not on every turn.
Dropping whole messages has rules. Real chat APIs expect each tool result to follow its tool call. Drop one without the other and the request can fail. My demo drops plain text.
The counter is fake. Four characters per token is a heuristic. Use your provider's token counting before you trust a threshold.
Summaries drift. I left LLM summarization out on purpose. CliffCompaction's rule of never rephrasing exists because a summary of a summary slowly stops being true.
The Bigger Idea
Tool result
β
Full log (the truth, never edited)
β
Policy (cap, keep, elide, drop)
β
View (what the model reads)
β
Model call
The model provides reasoning.
The tools provide evidence.
The log provides the truth.
The view decides what the model is allowed to pay attention to.
That is why the harness is becoming the product. Two agents on the same model can read very different conversations.
A cheap agent that forgot the failing line is not cheap. It is wrong at a discount.
Your context window is a budget. Spend it on evidence.
Try Roster
I'm building Roster around this idea: AI employees with real responsibilities, tools, memory and schedules, and a harness that decides what they need to see on every step.
If the same follow-ups, handoffs, and waiting loops keep eating your week, give them to an AI employee.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.
