Build and Learn

Context Hygiene Worth Automating

Anil Sen

Senior Fullstack Engineer

Context Hygiene Worth Automating

Introduction

You probably already know that long conversations with a coding agent get expensive. Every message, and every tool call behind it, makes the model process the entire conversation again. But knowing isn't the same as acting on it. Sometimes we just forget, because we're in the middle of a task. And sometimes we don't really know what it costs, so it never feels urgent.

So I stopped guessing and looked at my own data. The numbers were clear enough that I added two tiny automatic checks to my setup. The ideas work with any agent. The implementation below is for Claude Code, since that's what I use.

The Numbers

I went through my recent Claude Code sessions to see where the usage actually went. The numbers are rough estimates, not a bill, but the pattern was obvious.

Here's what stood out:

  • ~82% of the spend happened on calls where the context was already over 300K
  • Just 2 sessions, each left open for weeks, made up half of the total usage
  • ~11% came from the first message after a break of over an hour, when a giant session had to be cached again from scratch
  • Work under 100K context was only ~3%

One more trap: reviewing screenshots inside a 400K+ session was, in my logs, almost always followed by a full cache rewrite. Visual reviews are better off in a short, separate session.

When I replayed my history with a reset at 200K, weighted usage dropped by more than half. That's an optimistic simulation of one person's habits, so expect less in practice. But the direction is hard to argue with.

Two Small Habits

1. Reset Big Sessions at a Natural Stopping Point

When a session gets big, reset it when a piece of work is finished and verified, never in the middle of a task. Then pick the cheapest way forward:

  • Same task continues - Compact the conversation and keep going
  • Next task is unrelated - Start a fresh session. No handoff note needed
  • Already compacted 2–3 times - Fresh session with a short handoff note, because summaries of summaries get blurry

You could also just lower the auto-compaction threshold. That works, but it fires on a number, not at the right moment. It can compact in the middle of a task, and it always compacts, even when the next task is unrelated and a fresh session would be cleaner.

2. Don't Come Back to a Big Session Cold

The prompt cache is what makes long sessions affordable, and it doesn't last forever. In the current Claude Code version, subscriptions get a 1-hour cache, while API keys (and usage past your plan's limit) get 5 minutes. Once it expires, your next message re-reads the whole conversation at full price. On a 400K session, a simple "ok, where were we?" is one of the most expensive messages you'll send.

So before you continue a big session after a break, decide first. A fresh session or /clear costs nothing. /compact still has to read everything once, but everything after it is cheaper. Just continuing pays the full price now and keeps paying for the big context afterwards. If you're on an API key and tend to pause between messages, setting ENABLE_PROMPT_CACHING_1H=1 is also worth a look.

None of this is hard. The hard part is remembering it. Nobody watches the token counter mid-task, "one more quick fix" always feels cheaper than starting over, and after lunch you just type into the session that's already open. That's what the hook below is for.

Automating It

Claude Code supports a UserPromptSubmit hook. It's a small local script that runs every time you send a message, before anything reaches the model. It can pass extra context to the model, or block the message entirely. I use it for both habits:

  • Size nudge - Once the context crosses 200K, Claude gets a note to suggest a reset at the next natural stopping point. It fires once per crossing (and again every +100K), resets after a compaction, and never blocks anything
  • Break check - If the session is over 200K and the last reply is older than the cache lifetime, the first message is held back once with a short explanation. Send it again to continue anyway. Slash commands like /clear and /compact are never held back

The break check costs nothing, because the script runs locally. I checked the transcript after a blocked message: there was no model call at all, and the 194K context was never touched. The desktop app also keeps your message with an "Edit prompt" button, so you don't have to type it again.

You don't need to write the hook yourself. Paste this prompt into Claude Code, and it will write the script, register it and test it:

Set up a UserPromptSubmit hook that helps me keep long sessions cheap.

- Put the script in ~/.claude/hooks/context-pressure.sh and register it in ~/.claude/settings.json with a 5s timeout. Merge it with my existing hooks, don't overwrite them.
- Read transcript_path, session_id and prompt from the hook's stdin JSON.
- Current context = input_tokens + cache_read_input_tokens + cache_creation_input_tokens of the last main-thread (non-sidechain) assistant message in the transcript. If a compact_boundary entry comes after it, treat the context as 0. Only read the tail of the file, since transcripts get huge.
- Threshold: 200K (100K if I'm on a 200K context window model).

Size nudge:
- Fire once when the threshold is crossed and again every +100K, not on every message. Keep that state per session in a temp file and clear it when the context drops below the threshold.
- When it fires, return additionalContext telling Claude not to interrupt the current work. At the next natural stopping point, it should end its reply with one short line suggesting the cheapest option: /compact if the same work continues, a fresh session if the next task is unrelated, and a fresh session with a short handoff note only if this session was already compacted 2+ times.

Break check:
- If the context is over the threshold and the last assistant message is older than my prompt cache lifetime (1 hour on a subscription, 5 minutes on an API key), block the prompt with decision "block" and a short reason: the cache has likely expired, /clear or a fresh session is cheapest, /compact reads everything once, and sending the message again continues anyway.
- Block only once per break: remember per session which reply you already blocked for, so the resend goes through. Never block prompts that start with a slash.

General:
- Fail silently. On any error, let the message through.
- Test it with fake transcripts: the nudge fires once and stays silent on a second run, stays silent right after a compaction, the break check blocks once and lets the resend through, slash commands always pass, and bad input exits cleanly.

That's it. From then on, Claude will suggest a reset when a session gets too big, and you'll get a heads-up before you wake up a giant session after a break.

Using another agent? If it supports hooks, paste the same prompt and ask it to adapt the details to its own hook system and log format. If it doesn't, keep its context meter visible and follow the two habits above by hand.

Bonus: Check Your Own Numbers

Before you install anything, check whether this is even your problem. This script buckets your Claude Code usage by context size:

import glob, json, os
from collections import Counter

# Relative API prices: cache reads are cheap, cache writes and output are not
WEIGHTS = {"input_tokens": 1, "cache_creation_input_tokens": 1.25,
           "cache_read_input_tokens": 0.1, "output_tokens": 5}
BUCKETS = [100, 200, 300, 500, float("inf")]  # context size in K tokens

spend, total = Counter(), 0
for path in glob.glob(os.path.expanduser("~/.claude/projects/*/*.jsonl")):
    seen = set()
    for line in open(path):
        try:
            entry = json.loads(line)
        except ValueError:
            continue
        msg = entry.get("message") or {}
        usage = msg.get("usage")
        if entry.get("type") != "assistant" or not usage or msg.get("id") in seen:
            continue
        seen.add(msg.get("id"))
        context = (usage.get("input_tokens", 0) + usage.get("cache_read_input_tokens", 0)
                   + usage.get("cache_creation_input_tokens", 0))
        cost = sum((usage.get(k) or 0) * w for k, w in WEIGHTS.items())
        bucket = next(b for b in BUCKETS if context < b * 1000)
        spend[bucket] += cost
        total += cost

for b in BUCKETS:
    label = "500K+" if b == float("inf") else f"< {b}K"
    print(f"{label:>7}  {spend[b] / total:6.1%}")

Here's my output:

 < 100K    2.9%
 < 200K    8.8%
 < 300K    6.2%
 < 500K   18.7%
  500K+   63.5%

If most of your spend sits in the bottom rows, the habits are worth it. If it's mostly under 200K, you're already fine, and the hook would rarely fire anyway.

Compatibility Note

Tested with: Claude Code 2.1.228 on macOS. The generated script will usually use jq to read the transcript, so install it if you don't have it.

The 200K default assumes a 1M context window. On 200K-window models, auto-compaction kicks in earlier, which is why the prompt drops the threshold to 100K there. The break check can't see which cache lifetime you actually have, so it goes by what you tell it in the prompt. One known limitation: the hook only runs when you send a message. A single long autonomous run won't be interrupted, and you'll get the nudge after it finishes.


How to learn more or get in touch

  • Visit our Resources page to get the latest NetFire product news, company events, research papers, branding guidelines, and much more.
  • Explore our Support Center for overviews and guides on how to use NetFire products and services.
  • For partnerships, co-marketing, or general media inquiries, email marketing@netfire.com.
  • For all sales inquiries, email sales@netfire.com to get setup with an account manager.

Find help fast with guides and
resources, on our

Join the NetFire newsletter

Get our latest announcements, industry insights, product news,
and much more. It’s free to join.

Top reasons to subscribe

  • Expert tips on tech and security best practices
  • Early access to cutting edge AI and data science research
  • Discover real-world use cases and customer success stories
  • Special offers and insider perks