Why did one AI agent replay about 386M input tokens, and is a compression proxy the fix or is it stopping agents from living too long?
Subagent Context Guard
Lifetime, not payload. One runner lived 837 turns and replayed about 386M input tokens while its tool output was only 0.7 MB. A compression proxy would have saved just 10 to 25 percent, so we skipped it. A hook now warns a subagent at 140k tokens, blocks it at 200k, and forces a handoff so a fresh agent takes the next task.
- Claude Code hooks
- Python
- Bash
Relevant services: Thrivbe AI
Hypothesis
Agent cost is dominated by replay: every turn re-reads the whole context, so the longer one agent lives, the more each further step costs, and total cost grows faster than the number of turns. If that is the real driver, a compression proxy that trims tool output is a partial fix at best, and the better lever is to stop agents from living too long. We wanted to test that against our own transcripts before installing anything.
What we built
- A measurement of the incident. One github-flow runner, the kind of subagent that works a single pull request, ran 837 turns for four PRs. Its context peaked at 927k tokens and the whole run replayed about 386M input tokens. The median context across its life was about 435k. The tool output it produced totalled only 0.7 MB, so the bulk was the agent carrying its own history from task to task, not large command output.
- A PreToolUse guard hook. It fires only inside subagents, reads the
agent's own transcript, and takes the context size from the last assistant
usage record (input, cache read, cache write and output tokens). Behaviour:
- Below the soft line it stays silent.
- From the soft line (default 140k tokens) it warns once per additional 20k tokens: finish the step in flight, write the handoff, do not start new exploration.
- At the hard line (default 200k tokens) it blocks every tool call except reading and writing that agent's own handoff file, and tells the agent exactly what to write.
- Smaller-window models get lower lines (Haiku: 90k soft, 120k hard), because its 200k window auto-compacts first, which hid it from the guard in the first live test.
- Limits come from a config file or environment overrides, and the guard fails open: if it errors, it never wedges an agent.
- A forced handoff. The agent writes a short note (under 60 lines) with its
brief, goal, what is done, what remains, the exact next command, state on
disk and gotchas, then ends with a single
HANDOFF:line pointing at it. A second hook, running in the owning chat, tells the parent once that the agent was stopped and that a fresh agent should be spawned from the handoff. - A working rule. One runner per pull request or task, never send a finished runner a new task, and every runner brief carries output rules (tail long output, ask for compact JSON, read files by range, roughly a 100-turn budget).
- Tests. A 20-check test script for the guard, with the checks verified to fail when the guard is deliberately broken.
- A Headroom evaluation. Headroom is a local proxy that compresses tool output before the model sees it. Rather than install it, we measured our own transcripts.
Learnings
- Lifetime, not payload, was the cost. 0.7 MB of tool output against 386M replayed input tokens makes the point: the same history was read again on every turn of an agent that never ended.
- A compression proxy attacks the smaller share. On our transcripts, tool output was 39 to 64 percent of visible session content, while agent-written commands, edits and prompts were another 30 to 52 percent, which a proxy cannot shrink. Its own published benchmarks save 21 to 57 percent on tool output only, so we estimated about 10 to 25 percent of what remains after the real fix. Everything also passes through a local proxy, including code and any secrets in tool output, for a small gain. We decided to skip it for now and to revisit only if guard logs show tool output is still the dominant share.
- Prose limits get ignored. The runner's brief had already asked it to stay small. A hard block that forces a handoff is what changed behaviour.
- Guards can hide behind other mechanisms. The first live test on a small-window model showed auto-compaction firing before the guard's line, so limits had to be set per model.
- Every fresh agent has a fixed startup cost. We measured about 37k tokens before any work: roughly 9k of rule files, 8k to 10k of skill listing, about 5k of agent listing, about 1.7k of MCP notes and about 10k for the tool's own prompt. Of 444 skills, only 77 were used in six months. Marking unused skills name-only did not change what subagents see, so we reverted it. Rotating agents is not free, which is why the lines sit at 140k and 200k and not far lower.
- Transcript file size overstates context. About 30 percent of the file is hook metadata that never reaches the model, so the guard uses token counts from usage records, not bytes.
- The guard is global. On its first day it also stopped a subagent in an unrelated chat at 205k tokens. That is the intended behaviour, but it is a surprise worth knowing about.
- Open item: the main chat. The orchestrating chat is now the biggest replay cost once runners rotate. One session was at 485k tokens of context after 636 turns. The guard does not cover it; a separate rotation watcher does.
Log
- 2026-09-21 — Guard hook, handoff relay and tests shipped, with the rule added to the github-flow skill. What is still pending is judging daily use from several days of guard logs, so we mark this In progress. Headroom measured and skipped. Next step: read a few days of guard logs and decide whether tool output is still the dominant share.
