Most engineering teams working on long-context agents hit the same billing wall around turn twenty. A coding agent runs twenty shell commands, reads twelve files, and runs pytest three times. By turn twenty-five, the prompt is 80,000 tokens long. Over eighty percent of those tokens are terminal dumps, compiler warnings, grep outputs, and directory trees.
The default reaction across research and devtools has been to compress the transcript. Teams summarize older turns, drop middle messages, or project prompt tokens into learned latent soft embeddings.
A paper from Peking University titled "Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents" (arXiv:2609.31430, by Zhensheng Zou, Guoqing Wang, and Dan Hao) puts hard numbers on why soft compression usually ruins coding agents.
When you compress an entire agent transcript into latent vectors, you compress two fundamentally different categories of text: what the environment printed, and what the agent decided.
The exact-match penalty
A software agent does not read historical context the way a human reads an essay. It reads context to copy exact file paths, variable names, line offsets, git commit hashes, and regex patterns.
If an agent needs to edit src/core/connection_manager.py at line 412, a compressed semantic summary that says "the user reviewed the connection pooling setup earlier" is useless. The agent needs the exact string src/core/connection_manager.py. If that string gets blurred into soft tokens, the model either invents a nearby path or spends another tool call running find or ls to recover what it already saw.
Compressing the agent's own actions creates a second failure mode: behavioral drift. Agents trained with standard instruction tuning rely on the exact surface forms of their tool definitions and scratchpads. Once you feed them soft tokens representing prior thoughts, their syntax degrades. They drop closing brackets, mangle JSON arguments, or repeat earlier failed actions.
The split: LOHA and ACD
The authors break the problem into two specific techniques:
First, a context layout called Latent Observations, Hard Actions (LOHA). Instead of compressing everything, LOHA preserves:
- Every turn written by the agent in raw text.
- The system prompt and instructions in raw text.
- The most recent K tool observations in raw text.
- Older tool observations compressed into soft latent tokens.
The separation addresses the actual workload. The agent retains full, uncorrupted access to its own previous reasoning and tool syntax. It retains byte-exact access to whatever the last few tools just printed. Historical observations (such as a 2,000-line grep run from ten turns ago) remain accessible in latent space so the model knows which files were touched, without occupying thousands of raw token slots.
Second, a training objective called Anchored Context Distillation (ACD). Fine-tuning an agent to read soft tokens typically degrades its out-of-distribution performance on standard text. ACD trains the agent on latent-observation histories while simultaneously anchoring its output distributions against the original base model running on plain text.
The numbers on SWE-bench Verified
The team evaluated the approach on SWE-bench Verified across two open models: Qwen3-4B and SWE-Master-4B-RL.
Setting the uncompressed observation window to K=3 produced these results:
- Context reduction: 43% fewer tokens per call on Qwen3-4B, and 57% fewer on SWE-Master-4B-RL.
- Task completion: Qwen3-4B resolved 12.1% of issues with LOHA versus 14.5% uncompressed. SWE-Master-4B-RL resolved 21.8% versus 27.5%.
- Recency scaling: Expanding the exact observation window to K=8 pushed resolve rates back up to 14.4% for Qwen3 and 23.0% for SWE-Master.
The critical test comes under memory limits. Under a strict 32K token budget, where a standard uncompressed agent truncates or crashes on long tasks, Qwen3 with K=3 resolved 21.1% on a 199-instance long-horizon subset, compared to only 11.1% for the same adapted agent trying to run on truncated plain text.
On concurrent single-GPU serving, the smaller context footprint boosted instance throughput by 1.9x.
What to take away for agent harnesses
If you are running long-horizon agent workflows on local hardware or self-hosted models, soft context distillation offers a real path to doubling concurrency. But the architectural rule matters more than the specific weights:
Never compress the agent's own action history. If your compression scheme touches the model's scratchpad, tool calls, or immediate working memory, you will lose task resolution. Let the environment outputs absorb the lossy compression, and keep the agent's decisions exact.
Top comments (17)
To me the architectural split between "environment output" and "agent intent" makes perfect sense from an infrastructure perspective. If we treat a coding agent like a system administrator, it needs an exact log of its own executed commands (the hard actions) to understand what broke.
However, the 10,000-line output of a crashed grep or a full core dump can be compressed.
It’s essentially separating the audit log from the payload.
That means keeping the decision layer byte-exact while distilling the "environment noise" is a much more robust design pattern than blind semantic truncation.
The audit log versus payload distinction is clean, but the edge case that always bites is where stdout turns into error diagnostics. If a harness collapses a 5,000-line build failure into a one-sentence summary, the agent loses the exact linker symbol or missing header on line 4,812 that tells it why the build broke.
What has worked well in my scripts is keeping the exact exit code plus a bounded head and tail of the stream in context, while dumping the full unedited output to a scratch file on disk. The agent gets the failure trace immediately without burning 20k tokens on compiler noise, and it can grep the scratch path directly if it actually needs the full traceback.
You map classic Linux log management directly to the LLM context window.
Effective troubleshooting never involves reading massive log files sequentially.
You isolate the exit code.
You read the tail.
You grep for specific linker symbols.
Dumping unedited stdout streams to disk treats the agent like a standard Unix process.
This shifts operational burden from expensive context memory to cheap local storage.
The agent uses standard diagnostic tools instead of brute-force reading comprehension.
Solid infrastructure architecture solves model limitations.
One thing I'd add from building in this space: past turn twenty, most of what's worth keeping isn't the raw history at all, it's a handful of facts. What was decided, what was tried and failed, what's still open. Pulling those into a small state block and starting fresh with it has worked better for me than compressing the transcript, both on cost and on the agent repeating old mistakes. Did you measure answer quality alongside token savings?
The ablation compared keeping K recent raw observations alongside compressed LOHA state blocks against naive head-truncation of plain text. At K=1, performance dropped because the model lost immediate execution feedback, while moving from K=3 to K=8 yielded diminishing returns on resolve rates while inflating token usage. The jump to 21.% came from retaining those three recent verbatim tool outputs so the agent had precise file paths and compiler errors for immediate next steps, while the compressed state block preserved long-range task intent.
The benchmark numbers in the post come from the Zou et al. paper on SWE-bench Verified, and they tracked resolve rate directly against token reduction. On unbounded windows with K=3 observations kept raw, resolve rate dipped from 14.5% to 12.1% on Qwen3-4B. But on long-horizon runs capped at 32K tokens, preserving latent observations resolved 21.1% compared to 11.1% for hard truncation.
That structured state block (decisions, failed attempts, open tasks) works well for keeping high-level intent intact. Where I see it hit friction in a coding harness is patch application. If an extraction pass summarizes an earlier traceback instead of keeping the exact line numbers and symbol names, the agent has to re-run the inspection tools just to get the raw strings back into scope.
That matches what I hit building Deiko. The state block is great for intent and bad for exact strings, so I keep two layers: a short note (decided, tried, open) goes into the next chat, and each line links back to the original brief, with the exact screen text, error output and file paths, which the agent fetches on demand through an MCP tool. In my tests, agents did pull the full originals when a question needed exact details, so the note could stay small. The 21.1% vs 11.1% at 32K is a big gap. Did the paper split out how much came from exact observations versus recency?
Keeping the small "decided / tried / open" state and fetching the exact details only when needed makes a lot of sense. A summary is great for intent, but it's a pretty bad place to store exact error text or file-level facts.
Exact error messages and stack traces degrade quickly once paraphrased. A model summarizing an error tends to strip out column offsets, exception classes, and exit codes, turning a concrete compiler failure into a generic complaint. Leaving the raw execution artifacts in a local SQLite table or structured append log while keeping only the high-level intent in the prompt window preserves both token budget and diagnostic accuracy.
I ran into this the opposite way. For a while I was compressing everything past a certain turn count because the bill was hurting, and I noticed that accuracy on any step requiring an exact identifier tanked while more open-ended reasoning seemed fine. It took a while to isolate because the failures looked like the model getting confused, not like it was missing a specific string it had seen earlier. What I actually needed was to keep the raw text for anything the model would need to quote back verbatim, and let everything else compress.
That exact identifier loss is where the failure chain usually starts. The model keeps the high-level intent, but file paths, UUIDs, or compiler flags turn into plausible hallucinations. I had the same issue with patch application, where a summarized diff dropped the exact leading whitespace and git rejected the hunk. Keeping raw stdout for tool calls and only compressing the reasoning turns saves the budget without breaking the downstream edits.
We ran into this with a document retrieval pipeline last year: I'd summarized older turns to cut costs and the agent started generating file paths that were close but not quite right, then wasting two or three extra tool calls to re-confirm what it had already seen. Didn't connect it to compression until I looked at which turns I'd trimmed. It's the agent's own reasoning steps that need to stay verbatim, not the grep outputs or test logs. Expanding the exact window back a bit fixed it almost immediately.
That path hallucination loop happens because lossy summaries replace exact identifier strings with generalized descriptions. The model remembers it inspected a router file, but loses the exact directory depth or filename extension, triggering redundant search tool calls to recover state it already discovered. Retaining structured reasoning traces while trimming raw tool output preserves the exact nouns without context bloat.
The 21.1% vs 11.1% result under the 32K limit is the part that really caught me.
It suggests context compression isn’t just about saving tokens so it’s about deciding what must remain exact. A small state summary plus on demand access to the original logs might be the better architecture.
The failure mode with pure summarization is that exact line numbers, compiler flags, and git hashes vanish first. When an agent needs to apply a patch three turns later, lossy prose in the context window forces it to guess the indentation or re-run the whole test suite just to recover stdout. Dumping raw command output to a local scratch directory and leaving only a two-line index in the active window gives you the token savings of a tight budget without turning file paths into approximations.
The recency scaling row is the underrated one: moving K from 3 to 8 takes Qwen3 from 12.1% back to 14.4% — essentially full recovery against 14.5% uncompressed — which means nearly all the recovered accuracy comes from simply widening the raw window, not from the latent compression itself. The honest headline of this paper is the 32K row: 21.1% vs 11.1% against truncation. LOHA doesn't preserve accuracy through compression; it preserves viability when the alternative is a truncated prompt.
The K scaling delta makes that mechanic very clear. Dropping K below three starves the model of the immediate tool outputs it needs to anchor file edits, which is why widening K to eight does the heavy lifting. The 32K row works as an insurance policy against hard context cliffs. An agent pays an overhead penalty across every intermediate turn to keep the task trajectory intact when truncation would otherwise drop the initial prompt.