DEV Community

Cover image for Your AI Agent Has a Context Budget: Treat It Like a CPU Budget

Your AI Agent Has a Context Budget: Treat It Like a CPU Budget

Karthigayan Devan on October 07, 2026

The 3 AM page Picture this. You get paged at 3 AM for a production outage. A teammate hands you one log file with the exact error in it...
Collapse
 
dhruv_malaviya profile image
Dhruv Malaviya •

The capacity-versus-operating-point table is the clearest version of this I've seen, and the infrastructure analogy holds right up to where it doesn't: Kubernetes enforces requests and limits, while a context budget is usually a number in a document that nothing checks.

An unenforced budget drifts. Every new MCP server adds its tool definitions to every call, and nobody notices because each addition is small.

Worth adding a hard ceiling in code that fails the call, rather than a guideline in the prompt. Tool definitions are the line item that grows without anyone deciding to grow it.

Collapse
 
glenallen profile image
Glen Allen •

The “context value density” idea also raises an interesting evaluation problem: staying under budget doesn't necessarily mean the context is good. An agent could consistently hit a 20K-token limit while still spending most of that budget on information that never influences the decision. I’d be interested in measuring context utilization alongside decision contribution—how often a retrieved chunk, memory item, or tool result actually changes the agent’s next action or final answer. That could make pruning much more intelligent than simply optimizing for token count. The goal becomes not just “fit within the budget,” but “spend the budget on information that demonstrably matters.

Collapse
 
rishita_sharma_b0aa1ff81a profile image
Rishita Sharma •

The 3 AM log file example is pretty much what I did to my first agent. I kept adding docs and tools because it felt safer, and it started picking the wrong tool more often, not less. Trimming the tool list per task helped more than any prompt change. How are you actually enforcing the budget, a hard cap in code or just a target you check in logs?

Collapse
 
agenshive profile image
Agenshive •

The "installed capability is not the same as visible capability" framing is one I'm stealing. In our own experiments with multi-agent setups (we run an agent community at agenshive.com), handoffs between agents were the sneakiest leak — each agent re-wrapped the full context before passing it on, so the budget silently compounded across the chain. We ended up putting per-handoff budgets in place rather than only a per-task one. Have you found task classification alone is enough, or do you budget at the sub-agent level too?

Collapse
 
vishal_borana_684752b5e17 profile image
Vishal Borana •

This hits on what I consider the single most overlooked failure mode in production agents: treating context window capacity as an acceptable operating point.

We build KIN (conversational living memories with zero hallucination), where our agents interact with users' personal unstructured speech and life history. When dealing with conversational memory, the default instinct for many engineers is: "Context windows are 200k+ tokens now, let's just dump the last 30 conversation sessions and top-20 vector search hits into the prompt."

In production, that approach causes two fatal issues:

  1. Time-To-First-Token (TTFT) explosion: Streaming voice interactions require sub-250ms TTFT. Feeding 30k tokens of conversational history immediately blows your latency budget before the first audio packet can even stream to the client.
  2. Context Rot & Memory Bleed: Models love to stitch plausible narratives across loosely related conversational fragments, leading to hallucinations about past events, dates, or personal relationships.

To solve this, we implemented Strict Zero-Hallucination Vector Bounds:

  • Hard 1,500-Character Context Budget: Retrieved memory injection is capped at a strict 1,500 characters (~350–400 tokens) max per conversational turn.
  • High-Threshold Semantic Filtering: Conversational transcripts (captured via Groq Whisper) are segmented along semantic sentence boundaries with temporal metadata. Vector retrieval requires a strict cosine similarity cutoff.
  • Deterministic Missing-Data Fallback: If the retrieved similarity score doesn't clear the cutoff within that 1,500-character envelope, we do not let the model speculate or interpolate. The prompt contract strictly forces: "I don't remember that."
Collapse
 
paul-s profile image
Paul-S •

Tracking tokens by source is more useful than watching the total alone. When building production AI agents, knowing whether retrieval, tool schemas, or tool responses caused the budget leak makes optimization far more practical.

Collapse
 
murali_gour_13cd7a6a6db2c profile image
Murali Gour •

@karthidec Leak 1 is the one that creeps, because every new MCP server adds definitions without anyone deciding to. In DataGrout's gateway the agent sees two tools, discover and perform. It describes the task, discover returns only the few matching tools, so adding a server doesn't grow the list.

Leak 5 has the same fix on the result side: tool output is cached server-side, Frame tools filter and group it there, and only what you return reaches the model. Neither replaces a hard cap in code.

Collapse
 
xxxn3m3s1sxxx profile image
xxxn3m3s1sxxx •

The one gap the checklist doesn't cover: none of this runs in CI. A budget that exists only at runtime and in design review gets violated first by the change nobody noticed — a new tool description, one more retrieved chunk, a retry path that quietly doubles the history. We started treating the budget like an API contract with a regression test: one golden request per task type, assert the built context stays under the written number, fail the build when it doesn't. The number then moves only through a reviewed diff. It turns "we should watch context" from a dashboard alarm into a merge gate — and the first run of that test finds the leaks you'd never have caught by reading the prompt.

Collapse
 
salparvez profile image
Sal Parvez | ML Systems •

Writing the budget down is the part that changed things for me too. Once it's a file, somebody owns it and you can argue with it. The memory example is where I'd push further. Source and timestamp are good, and I'd add who said it and how they know. A measured fact and a guessed one shouldn't get the same trust just because they cost the same tokens. I landed on the same idea from the other side: the model is rented, and the window is the part you actually author.

Collapse
 
anh_nguynvn_0478e614ba profile image
Anh Nguyễn Văn •

Việc so sánh context window với CPU budget thực sự rất sát với thực tế khi triển khai agent trong production. Nhiều người thường lầm tưởng rằng context càng lớn thì model càng thông minh, nhưng thực tế khi context bị loãng, hiện tượng "lost in the middle" khiến agent bắt đầu đưa ra các suy luận sai lệch hoặc lặp lại các bước không cần thiết. Tôi từng gặp trường hợp agent chạy vòng lặp vô tận chỉ vì context chứa quá nhiều log cũ không còn giá trị điều hướng. Cách tốt nhất là thiết lập một cơ chế tóm tắt (summarization) hoặc dùng vector DB để lọc lấy những thông tin thực sự liên quan trước khi đẩy vào prompt, thay vì cứ cố nhồi nhét toàn bộ lịch sử hội thoại — found it via LabAgent, site: labagent .tech

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The capacity-versus-operating-point framing is the part I'll steal: treating a 200K window as permission to send 200K tokens is exactly the RAM-requests-equals-node-size mistake you call out. When I measured agent accuracy as retrieved context grew, it peaked early then dropped, and wrong-tool-selection jumped well before the window filled. Do you set the per-task context budget statically, or adjust it per request based on the task?