Context Engineering for Claude Code: CLAUDE.md, Skills, Docs

Context engineering is deciding what an agent sees before it acts: which instructions load into every session, which load only when they are relevant, and what never enters the context window at all. In Claude Code that decision has four surfaces, CLAUDE.md, docs, agent files and Claude Code skills, plus hooks that run outside the model entirely. This post is the context engineering half of my agentic coding setup: a short always-loaded CLAUDE.md that points at depth-on-demand docs, a brief template that carries the spec, skills as slash-only pipelines, and a list of what I keep out of context on purpose.

Everything below is from my own agent config, running today. Where it falls short of the practice it is copying, I say so.

Context engineering vs prompt engineering vs harness engineering

Prompt engineering is wording one message well. Context engineering curates every token the LLM holds while it works: the system prompt, the instruction files, the tool definitions, the retrieval results, the file an agent read three steps ago and still carries. Harness engineering is the layer above both: the scaffolding deciding which agent runs, with what tools, at what level of autonomy, and what must be true before a turn ends.

The three are nested, not rival. A prompt engineering fix is rewording one line of a brief. A context engineering fix is moving a section out of CLAUDE.md into a doc that loads only when a rule points at it. A harness fix is a Stop hook that blocks the turn while the project’s checks are red.

That last one is the interesting case, because a hook consumes no context window: the model never reads the check, only the failure if there is one. Work moved into the harness is work removed from context.

Most of what people call prompt engineering in agentic coding is really context engineering with a worse name. The model is not underperforming because a sentence is clumsy. It is underperforming because its working memory is packed with information that is not relevant to the next tool call, and rewording fixes none of that. The repair is to move that information onto a surface that loads only when it becomes relevant.

What lives where: CLAUDE.md, docs, agents, skills, hooks

Each surface has a different trigger, and choosing the trigger is the act of context engineering.

Surface When it loads What it holds In my setup
CLAUDE.md every session, always resident cross-cutting rules that apply to most turns delegation ladder, review contract, code style, the brief template
docs/*.md only when a rule points at it procedures too long to keep resident least-feasible.md, github.md, local.md
agents/*.md when that subagent is spawned one role’s job, tool list and model researcher, implementer, reviewer, builder, scout, mechanic
skills/*/SKILL.md when it is invoked multi-step pipelines with control flow /feature, /bugscan, /deep-review
hooks on a harness event, outside the model deterministic gates format on write, project checks on stop

Context window pressure drops at every row. CLAUDE.md is the only file paid for on every turn, so it is the one to keep thin. A SKILL.md or an agent file is paid for only by the session that uses it, and only for the time it is loaded. Hooks cost nothing in context terms, since no part of them enters the model’s memory, which is why anything expressible as a check belongs there rather than as an instruction the model must remember.

A short CLAUDE.md that points outward

Anthropic’s Claude Code best practices are blunt about the failure mode: “Bloated CLAUDE.md files cause Claude to ignore your actual instructions!” The same page recommends progressive disclosure, keeping the resident file short and loading depth on demand. That is the CLAUDE.md best practice I try to follow, and the one I break most.

The skeleton looks like this:

# Global Instructions

## Before writing code
Stop at the first rung that holds: does it need to exist · reuse it ·
stdlib · platform feature · installed dep · new library · write it.
Workers follow `~/.claude/docs/least-feasible.md`.

## Delegation
Ladder: haiku fan-out · sonnet scoped · opus default worker · frontier
for judgment. Briefs use the template below.

## Code review
Round 1 opus, 2+ sonnet deltas, escalate after 3.
Transport specifics: `~/.claude/docs/github.md` or `local.md`.

Every section is a rule plus a pointer. The rule must be in context on every turn, because it changes what the model does next. The pointer loads only when the work reaches that step: a worker about to edit files reads the least-feasible doc, a reviewer about to post findings reads the transport doc, nobody else pays for either.

Where it falls short: my real CLAUDE.md is 223 lines, not 15. It carries a code style section, a documentation section and a change-description template that most turns never touch. Each earns its place individually and the set is still too long. The resident file should hold only rules that change behavior on a typical turn, and mine is not cut down to that yet.

Depth on demand: the docs a rule points at

least-feasible.md is a 61-line contract for anyone about to write code: checkpoint the plan before editing, prefer depth over count, run a shrink pass once the tests are green, end the report with Not built: … so the compromises stay visible. It opens with the rule that the whole doc exists to enforce: “Cost is files touched and concerns braided together, not lines — a line count is a prompt for a reason, never a gate.” An implementer loads it. The orchestrator, which writes no code, never does.

github.md and local.md are the same review procedure under two transports, one with a forge and one without. The context engineering trick is that they carry identical section names: Preflight, Diff, Findings, Fix replies, Resolving, Verdict, Change description, Merge signal, Merge. A rule in CLAUDE.md can then say “reply per § Fix replies” and be correct under either transport, with neither transport’s commands resident: the agent detects the transport once, loads one file, and the reference resolves. Identical headings also keep the docs honest against each other, since a section in one and not the other is visibly a gap. And the deferred weight is real: the GitHub doc is full of gh api invocations, JSON field selectors and a GraphQL query for resolving review threads, none of which belongs in a session refactoring a parser.

The brief is the spec, and the spec is the context

The best practices page describes what a good spec contains: “The most useful specs are self-contained: they name the files and interfaces involved, state what is out of scope, and end with an end-to-end verification step that proves the feature works.” Self-contained is a context engineering property. A worker that has to re-derive the task from a conversation it cannot see has no spec, it has a hint.

Every implementation brief in my setup uses one template:

Outcome:      one sentence, observable behavior
Acceptance:   one runnable check
Mechanism:    one sentence, or "implementer's call"
Touch:        files expected; anything else needs a checkpoint
Not in scope: explicit list (hardening, docs sections, refactors)
Tests:        behaviors to prove, by name; nothing else
Checkpoint:   yes (default) | no

Acceptance is the end-to-end verification step. Not in scope does real work: hardening ideas get recorded there rather than built, which is how a scoped task stays scoped. Touch bounds the blast radius, and anything outside it forces a checkpoint back to the user instead of a quiet expansion. One logical change is one brief, and an “and” in the Outcome line means the task is really two.

Where it falls short: no acceptance criterion is linked to a named test in a machine-checkable way, so nothing outside the orchestrator’s judgment verifies that the delivered change matches the Outcome line. The reviewer checks each hunk against a Scope block copied from the brief, which is close, but still one model reading another’s claim.

Claude Code skills as slash-only pipelines

Skills are the unit of context engineering I underuse the most and rate the highest. A skill is a folder holding a SKILL.md with a name, a description and an optional allowed-tools list naming the specific tools it may use. Its body loads when the skill runs, not before, so a 142-line pipeline costs nothing until someone invokes it.

Two configuration choices matter more than the contents.

First, the two pipelines that write to the forge set disable-model-invocation: true, which makes them slash-only: /feature, which creates branches and pull requests, and /bugscan, which files issues. Either one should start because a user typed the command, never because the model inferred from context that now would be a good moment to open a PR. /deep-review carries no such flag and is model-invocable on purpose: it only reads a diff and applies quality fixes inside work that is already underway. Autonomy is fine inside a pipeline and a bad idea as the trigger for one that creates artifacts other people see.

Second, pipelines are skills rather than agents on purpose. A skill runs inside the orchestrator session, so it can spawn subagents, message a running one, poll the forge and react to a human signal mid-flight. An agent cannot orchestrate other agents. The orchestration logic therefore has to live in a skill, and the skill’s own body is the only part of it that occupies the main session’s context window.

/deep-review is the clearest example: it fans out eight quality specialists in parallel, each given exactly one dimension to look for, then validates and applies what survives. Eight narrow contexts find more than one wide one: each specialist holds only the data relevant to its own dimension, and none carries the other seven’s findings. A wide context is not a better context, it is a diluted one.

Per-agent files are scoped context, not personalities

Each subagent is a file with a role, a tools list and a model. The reviewer and the researcher get Read, Grep, Glob, Bash and nothing else, a permission boundary and a context boundary at once: an agent whose tools cannot write files carries no write-shaped instructions, and the tool list is itself context the model reads on every turn. The cheapest workers are the shortest files. My mechanic brief is 12 lines, the whole of it “execute exactly what the brief specifies, report facts, return questions instead of guessing”.

Suraj Khaitan, after building and cutting a hundred subagents, calls a subagent a context firewall rather than a personality, quoted in full in my post on Claude Code agents and the model ladder. A worker gets the brief and the codebase, does its job, returns a report; the files it read and the dead ends it walked down stay on its side of the wall.

The corollary: review fixes go back to the same worker through a message, never a fresh spawn. It already holds the brief, the diff and the reasoning, and respawning pays for that context twice. Cognition makes the same argument about sharing full traces rather than individual messages, quoted in the agents post.

Isolation, plans and the compaction problem

Two smaller pieces of context engineering round it out, and one gap.

Worktrees are isolation you can see. Every implementer works in <repo>/.worktrees/<branch>, ignored through .git/info/exclude, so parallel workers never collide in one codebase checkout. Cognition’s write-up of managing parallel agents states the underlying problem: “when one agent tries to handle too many things in a single session, context accumulates, focus degrades, and the quality of each subtask suffers.” A worktree is the filesystem half of that split, a subagent the context half.

The plans directory plus an output style handles the other end. Plans go to ./plans, and a custom output style makes each one terse: bullets over paragraphs, no “Background” section, file paths referenced directly, every section that does not change what gets implemented dropped. A plan is read back into context later, sometimes by another session, so padding in one is a tax paid every time.

The gap is the handoff. Anthropic’s harness design post names it: “While compaction preserves continuity, it doesn’t give the agent a clean slate, which means context anxiety can still persist.” When a long session compacts, my setup relies on the brief plus whatever survived. There is no handoff file, no written state of the work a fresh session could load instead of inheriting a summarized transcript, which is the long-term memory such a file would buy. It is the clearest hole in my context engineering today.

What is deliberately not in CLAUDE.md

Saying what to leave out is as much a part of context engineering as saying what to include. Out of my always-loaded file, on purpose:

  • Setup and onboarding instructions. Generating per-repo hook templates is a skill’s job, examples included. Those steps need not live in context every session.
  • Transport-specific commands. Every gh invocation sits in one of the two transport docs.
  • The verification commands. The Stop hook runs the project’s check script and blocks the turn on failure, so neither the model nor the user has to remember the command for a specific project.
  • Per-project stack details. Each project owns its own CLAUDE.md and check script. The global file stays generic.
  • Anything a permission list can enforce. The deny list covers secrets, curl, sudo, rm -rf, force pushes and hard resets. A denied tool call fails without any instruction being read and without asking the user.
  • A cross-vendor AGENTS.md. The agents.md convention answers the one-file-per-vendor problem, but my setup is Claude-only, so adopting it would add a file to the codebase without removing one.

The pattern across those six: a rule the harness can enforce, a permission can express, or a doc can supply on demand should not be resident. Resident context is the scarcest resource in the system, and instructions compete with the work for it.

What I still cannot measure

Every claim above is a design argument, not a result. I have no evals for my context engineering and no data on it: when I move a section of CLAUDE.md into a doc, or tighten an agent brief, nothing tells me whether it helped. Anthropic’s post on evals for AI agents is direct about the cost: “Without them, it’s easy to get stuck in reactive loops—catching issues only in production.” The same post sets a reachable bar, “In reality, 20-50 simple tasks drawn from real failures is a great start”, and I have not built even that.

So take the structure and not the certainty. The trigger-per-surface table is the part I would defend hardest, because it makes every other decision mechanical: decide when a thing must load, and where it lives follows. Effective context engineering, at this maturity, is mostly subtraction you cannot score yet.

FAQ

What is the difference between prompt engineering and context engineering?

Prompt engineering optimizes the wording of a single message; context engineering curates everything in the model’s context window while it works. The practical test is what you would change to fix a bad answer: rewording a sentence is prompt engineering, changing what loads and when is context engineering. The reason the second one dominates in agentic setups is arithmetic, since a session’s instruction files, tool definitions, retrieved data and prior tool outputs outweigh any single message by an order of magnitude.

What is the difference between context engineering and harness engineering?

Context engineering decides what the model sees; harness engineering decides what runs around it. The connection is a budget: every rule the harness can enforce is a rule the resident file does not have to carry in memory, so harness work is context work by subtraction. The practical marker is whether a rule can fail loudly without the model’s cooperation, which is the test I apply to decide where a new rule goes.

Do skills work in Claude Code?

Yes. A skill is a directory containing a SKILL.md with a name, a description and optional frontmatter such as an allowed-tools list, and Claude Code loads its body only when the skill runs. Setting disable-model-invocation to true makes it slash-only, which is what I use for any pipeline that creates branches or pull requests. Claude Code skills are the cheapest way to keep a long procedure available without keeping it resident at the session level.

What should go in CLAUDE.md?

Only rules that change behavior on a typical turn, plus pointers to the depth. CLAUDE.md best practices reduce to one test: would this line change what the model does on a random turn, and if not, which surface should own it instead? The cheapest way to apply that test to an existing file is to delete a section for a week and see whether anything regresses, which is how mine lost its setup instructions and its gh command reference.

How can I learn context engineering?

Start with one real session: notice what the agent is holding when it makes a bad call, and move that class of information to another surface. Read Anthropic’s Claude Code best practices and the subagents docs for the mechanics, then change one surface at a time so you can attribute the difference. Effective context engineering is mostly the habit of asking, for every instruction you are about to write, whether it must be present on every turn or only on one.

Sources

claude codecontext engineeringagentic codingcase study
← Back to the blog