Skip to content
Go back

Two Things I Do to Keep Claude Code's Context Cheap

I wrote before about teaching Claude Code my workflow: CLAUDE.md, skills, memory, modes. What I didn’t say is that the same setup also keeps my token usage down. Turns out the two problems share one answer: don’t load everything, load what’s relevant.

Here are the two things that actually matter.

1. Keep an index, not the payload

My memory isn’t one growing file that Claude re-reads in full every session. It’s a short index file, one line per memory:

- [No em dashes](feedback_no_em_dashes.md): never use em dashes in any output
- [Daily Schedule](project_schedule_change.md): 05:00 wake, gym first, goals block until 09:30

Claude reads the index every time. It only opens the full file behind a line if that memory turns out to be relevant to the current task. A hundred memories cost a hundred lines in the index, not a hundred files in context.

Same idea in the sprint system. I don’t keep one running task file for the year. Every week gets its own file. Claude reads this week’s file, not twelve months of history.

2. Load on demand, not upfront

I have several “modes”: coach (accountability, goal check-ins), health (medical context, training log), learn (new topics, step-by-step explanations). Plus a handful of other skills. None of that lives in my main instruction file. A mode is just a name written in a separate file. Claude checks if it’s set, and only then pulls in the rules for that mode.

Skills work the same way. Claude sees a one-line description of each one. The full instructions only load when a skill actually fires.

The alternative would be pasting every mode and every skill into the system prompt so it’s “always available.” That’s simple to build and expensive to run. Every session pays for content it never uses.

The pattern

Both tricks are the same trick: separate the pointer from the content. Keep pointers cheap and always loaded. Keep content expensive and loaded only when it’s actually needed.

It’s not a prompting hack. It’s how you’d design any system that has to stay fast as it grows.


Share this post on:

Previous Post
Deterministic Scripts Before Prompts: Cheaper, More Reliable AI Workflows
Next Post
How to Teach Claude Code to Know Your Workflow