My Claude Code week started ending on Wednesday. I blamed the September 14 cut, like everyone on r/ClaudeAI did. Then I opened the usage screen and the biggest line was not me typing. It was the helpers I had told it to launch.
So I read a month of my own logs before touching a single setting. 455 sessions, 63,000 requests, and a number next to every saving tip I had been told to apply.
Which of the five fixes everyone lists gives a measurable share of the week back, and which one costs quality?
What I actually counted
Claude Code writes one JSONL file per session under ~/.claude/projects/, and one file per subagent run next to it. Every assistant line carries the usage of its request: input, output, cache reads, cache writes.
The first trap is double counting. The log writes the same answer about twice (1.96 lines per request in my month), so summing every line gave 18.6 billion tokens. De-duplicated on message id and request id, it was 9.3 billion. Every number below is the de-duplicated one.
The second limit is bigger. Anthropic publishes the weekly limit as percentages and plan multipliers, never as a token count. So everything here is tokens from one workload, not a share of your week. The order may change for you. The usage screen will tell you.
The cut itself is simple arithmetic. Call the old weekly limit 100. The summer promotion made it 150. The permanent level since September 14 is 125. From 150 to 125 is the 17% people feel, and Anthropic's "+25%" is true at the same time.
Subagents took 48% of everything
2,631 subagent runs in the month. They took 48.1% of all tokens and wrote 64% of all output, which is the heaviest line on the meter.
Each one pays an entry price. Before a subagent does anything, its opening request already carries a median of 47,117 tokens: the instructions, the tool list, the skill list, all sent again. Someone else measured 16 to 21k on a different machine with tiny agent prompts, and their line stuck with me: your agent file is a rounding error inside its own launch cost.
The model is the other half. A subagent inherits the main session's model unless the agent file says otherwise. Switch to the biggest model for yourself and every helper runs on it. In my logs the smallest model handled under 1% of subagent requests.
---
name: test-runner
model: haiku
---
Two habits, then. Skip the subagent for a job you could do in place. Pin a small model on the ones you keep. Nobody has measured what pinning saves as a share of the week, and a small model that needs more turns can cost more, so this stays a habit with a direction rather than a percentage.
The five-minute cache nobody lists
The main session's prompt cache lives one hour on a subscription. A subagent's lives five minutes. My logs agree with the docs on this: every cache write from a subagent landed on the 5-minute tier, every write from a main session on the 1-hour tier.
One setting moves it:
{ "subagentPromptCacheTtl": "1h" }
A Reddit user whose subagents wait on long builds saw cache writes go from 12 million tokens to 3 million after the change. In my own logs it barely matters: 2 requests in 1,000 arrived after a wait of five minutes, though each one rewrote about 75,000 tokens. If your subagents idle, turn it on. If they run in short bursts, leave it, because an hour of cache costs more to write.
One lunch break rewrote 130,000 tokens
This is the one I did not expect. A message sent within five minutes of the previous one wrote about 1,200 tokens of cache. After a pause of more than an hour, the median was 130,332. The session at that moment held about 175,000 tokens, so most of it was written again.
gap before the message n cache written (median)
< 5 min 18,029 1,176
5 to 60 min 414 1,327
> 60 min 79 130,332
Two cheap habits. Clear the session when a task is done, while the cache is still warm, so the next task starts small. And when Claude Code offers to resume a big session from a summary after a break, take it. A model switch mid-session empties the cache too, since each model keeps its own.
The caveat: 79 cold returns is a small sample, and some of them follow a compaction rather than pure idle.
Effort is the fix that can cost you quality
One developer ran the same 29 real tasks at all five effort levels. Mean cost per task went from $2.50 at low to $8.84 at max. Medium passed 28 of 29, more than any level above it. His words: the curve appears to peak at medium.
Then the catch. On the hard problems he picked, low passed 0 of 5 and high passed 5 of 5. A low attempt took two minutes, a high one thirty-three. So medium to build, high when a mistake is expensive, max almost never. Those costs are in dollars on an older model, not a share of your week.
The popular three weigh less than advertised
Removing MCP servers is the tip everybody repeats. Tool definitions are deferred by default now, and one article counted 1,350 tokens for 51 tools across three servers. A single server with one tool cost 18.
Turning off prompt suggestions made the rounds with a claim of up to 10% of the week. That was one account with enormous contexts. Another user on the same thread measured 3 to 4%.
Filtering shell output: the developer who measured it on his own usage found about a tenth of one percent.
All three grow with the size of your context. Switch them off if you like. Don't expect the week back.
What this doesn't prove
It is one month of one person's work, with heavy subagent fan-out. A solo session workflow will see a different split. The effort numbers are someone else's, on an older model, in dollars. And nothing here converts to a percent of the weekly cap, because the cap is not published in tokens.
The full run
The video shows the ranked table and the method on screen, including the counting trap.
What does your usage screen blame? A) subagents B) long sessions C) something I did not count. I would like to know whether 48% is my workload or the norm.
I used an AI assistant to tidy the prose. The logs, the counting and the opinions are mine.
Top comments (2)
The lunch break table is the part nobody talks about. 1,200 tokens warm versus 130k cold is wild, and it explains why long meeting days feel so expensive. Pinning haiku on the test runner agents is such an easy win too. Did you notice whether the 47k launch cost drops much when you trim the skill list, or is most of it the tool definitions?
The 47k entry overhead per subagent matched what I ran into until I stopped passing the full tool and skill catalog into child runs.A subagent spawned just to run tests or parse a diff doesn't need search APIs, browser tools, or twenty procedural instructions in its opening frame. Trimming the child agent's tool whitelist down to the three commands it actually executes cuts that initial payload by more than half before the first turn starts.Pairing that with pinned smaller models on sidecars kept the weekly budget alive. Letting every subagent inherit the flagship planner model turns background tasks into a token leak faster than long chat sessions ever could.