Claude Code Agents and Agent Teams: How I Run Them

Claude Code agents, in my setup, are six files. Each file names a model, a tool list and a one-line contract, and not one of them can start another agent. That single constraint shapes everything else: one orchestrating session delegates, the workers run, a separate reviewer reads what they wrote, and a human merges. This post is the agent half of agentic coding in practice, the parts that decide who does the work. The short version: I run Claude Code subagents plus skills, not an agent team, and the reason is that a skill can spawn agents while an agent cannot.

The six Claude Code subagents I actually keep

A subagent in Claude Code is a Markdown file with front matter. The front matter carries a name, a description that tells the main agent when to pick it, an optional tools: list and a model; the body is that subagent’s system prompt. Claude Code subagents run in their own context window, so nothing a subagent reads on its way to an answer lands in the main conversation, and the main session gets back only the final report. Anthropic documents the mechanics in create custom subagents; what follows is the roster I converged on after deleting a lot of others.

Agent Model Tools Contract
researcher opus Read, Grep, Glob, Bash Trace the real flow and return a brief an implementer can act on without re-investigating
implementer opus all One subtask, one worktree branch, one minimal verified diff opened for review
reviewer opus Read, Grep, Glob, Bash Adversarially review one diff for scope, correctness, verification and quality, then return a verdict
builder opus all Default worker for a scoped change that does not need the full pipeline
scout sonnet all Routine scoped work: one search, one small specified edit, one summary of a file or diff
mechanic haiku all Mechanical fan-out: grep and report, file inventories, rote transforms, git plumbing

Two of the six are read-only on code by tool access, not by instruction. The researcher and the reviewer get Read, Grep, Glob, Bash and nothing else, so neither can write a file even if its prompt drifts. The other four inherit the full tool set. Tool access is a security boundary I set once per file, not a request repeated in every prompt.

The descriptions are load-bearing in a way that is easy to miss. Claude Code picks a subagent by reading the description field and little else, so “use for implementation work that has a clear spec” and “use for batched small steps that need no judgment” are routing instructions, not documentation. A vague description means the wrong subagent runs, or that the main agent quietly does the work itself.

What is missing from the list matters as much. There is no test-writer agent, no documentation agent, no separate security reviewer stacked on the code reviewer, no architect. Every one of those I tried collapsed into either the code reviewer’s rubric or the implementer’s brief, and the best single line I have read on the subject is Suraj Khaitan’s, after he built a hundred of them: “A subagent is not a personality. It’s a context firewall.” An agent earns a file when it needs a different context, different tool access or a different model. Not when it needs a different tone.

Orchestrator, workers, reviewer, human

The main Claude Code session is the orchestrator. It plans, decomposes, delegates, sequences and decides, and it writes no code at all. Everything mechanical goes to a worker: reading a file to report a fact, git plumbing, running a command, applying a fully specified edit. The reason is arithmetic before it is architecture. The session runs on the most expensive model available, so every tool call it makes directly bills at frontier rates.

Orchestrator session delegates to a researcher and an implementer; the pull request goes to a reviewer, fixes return to the implementer, the human merges.

The reviewer is a separate agent for one reason, and Anthropic’s harness design post states it better than I can: “When asked to evaluate work they’ve produced, agents tend to respond by confidently praising the work—even when, to a human observer, the quality is obviously mediocre.” The generator is not the evaluator. The implementer that just wrote four hundred lines will call them excellent lines. A reviewer with a fresh context window, a rubric and no write tools will not.

That the reviewer is also a Claude model has a mechanism behind it. Panickssery et al. (2024) found a causal link between how well a model recognizes its own text and how much it over-scores it, and Wataoka et al. (2024) traced the effect to perplexity: judges over-reward text that is easy for them to predict.

Three things push back. The reviewer runs only after the Stop hook executed the tests, and execution is the strongest documented mitigation, “increased its agreement with ground truth from below 42% to approximately 72%” in the SE literature review of Findeis et al. (2025). It runs in a clean context with no memory of writing the code and no mention of who did, the blind condition under which Chae et al. (2026) found self-preference disappears. And its brief is adversarial rather than approving. Why it stays on Claude, with the cross-model numbers, is in the umbrella post.

Verification is not left to either of them. A Stop hook runs the repository’s own check script, lint and typecheck and the fast test suite, before an agent may end a turn with unverified changes, and a red check blocks the stop. The skip conditions that decide when there is something to check are in my post on the Stop hook that enforces verification. The code reviewer then audits whether the author verified at all: a missing test for non-trivial logic is a finding, and so is documentation in the diff that contradicts the code next to it. Two different mechanisms, deliberately, because an instruction in a prompt is advisory and an exit code is not.

The model ladder, and the rule that makes it cheap

Model choice is per subagent, not per session. It lives in one line of front matter per file, and the ladder is short.

Tier What runs there
haiku Mechanical fan-out: grep and report, file inventories, rote transforms, format checks
sonnet Routine scoped work: single-purpose searches, small specified edits, summarizing a diff
opus Default worker: implementation, review round 1, multi-file analysis, verification
frontier session model Orchestration, and any subtask that itself needs frontier reasoning

The selection rule is written to be lazy on purpose: “When unsure between two tiers, take the cheaper one — a failed worker is cheap to rerun one tier up.” That is a cost statement and a diagnostic one. When a sonnet worker fails, the usual cause is an ambiguous brief, and rerunning it at opus hides that.

Haiku is excluded from both of my pipelines. Not because it cannot do the work in principle, but because I judged its output not worth the savings for those specific tasks. It keeps its place in the general ladder, as the mechanic agent, where a wrong answer costs one rerun rather than a bad pull request.

How agents are dispatched, resumed and stopped

Agents run asynchronously by default. The orchestrator spawns a worker through the agent tool, naming the subagent type it wants and passing the brief, then keeps planning while that worker runs; it goes synchronous only when the next step genuinely depends on the result. Background subagents are the normal case here, not an optimization: the orchestrating session always has plan-shaped work of its own. Each implementer works in its own git worktree under .worktrees/<branch>, excluded from version control locally, so two agents editing one project at the same time never collide in a single checkout.

Then comes the part I got wrong first. When the reviewer requests changes, the fix goes back to the same implementer through SendMessage, addressed by that agent’s name or agent id, never to a freshly spawned one. The original agent already holds the brief, the diff and the reasoning behind both; respawning pays for that context twice and loses the nuance. The same applies when a worker returns a genuine unknown as a question instead of guessing: I answer and resume that agent, rather than restarting the subtask from zero. Cognition’s Don’t Build Multi-Agents is blunt about why context-poor handoffs fail: “Share context, and share full agent traces, not just individual messages”.

Three limits keep the loop from running forever:

  • Three review rounds. Round 1 is opus; later rounds are sonnet deltas against the known findings only. After three, the orchestrator escalates to me, because three failed rounds means the brief is wrong, not the code.
  • A signature on every agent comment. Findings start with 🤖 reviewer · round N and fix replies with 🔧 fix · <sha>, so my own comments are the unsigned ones and the unfixed findings are literally my to-do list.
  • A human merge signal. Agents open and fix pull requests. A label from me merges them. A comment from me counts as request-changes and does not burn an agent round.

Skills versus agents, and why there is no conductor

The pipelines that drive all of this are skills, not agents, and that is the most consequential design decision in the whole setup. A skill runs inside the orchestrator session, which means it can spawn subagents, message a running one, poll the forge for a label and react to a human signal mid-flight. An agent can do none of that. My design doc states the constraint in one line: “An agent cannot orchestrate other agents.”

So the shape is fixed by the tool, not chosen: orchestration lives in the session, work lives in the agents. I also tried the obvious extension, a conductor agent between the session and the workers. It duplicated the orchestrator with no added capability and cost an extra context window. “It was tried and reverted.” Cognition’s note on managed agents describes the real problem a conductor was supposed to solve, which is different: a single session handling too many things accumulates context until each subtask degrades, quoted in full in my post on context engineering. The fix for that is narrower briefs, not another layer of management.

Claude Code agent teams versus subagents plus skills

Anthropic has since shipped a first-class agent teams feature, documented in Orchestrate teams of Claude Code sessions. What that page describes is several Claude Code sessions coordinated as a team: a lead that delegates, a shared task list the team works against, and teammates that can message each other. The team lead is the role my orchestrator already plays, and it is a friendlier surface than a skill file that spawns workers by hand.

I have not moved my pipelines onto it, for three reasons that are about my workload rather than the feature’s quality.

First, my coordination is not generic. A /feature run is a fixed sequence with gates in it: research before a branch exists, subtasks ordered so shared code goes first, a review round cap, a label poll, a merge. That is a program, and I would rather keep it in a skill file I can read and diff than in a team’s emergent negotiation.

Second, teammates talking to each other is the property I am most suspicious of. Every message between two agents is a chance for a plan to drift with nobody holding the original. In my setup the orchestrator is the only thing that talks to everyone, which makes the trace of any decision one conversation.

Third, cost. Anthropic’s multi-agent write-up reports that “Agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats.” Parallel teammates multiply that, and most of my subtasks are sequenced by shared code anyway, so they would be waiting rather than working.

None of that is an argument against agent teams for someone else. If your work decomposes into genuinely independent parallel tasks, a team of teammates with a shared task list is less machinery than a skill that spawns Claude Code subagents by hand, and it survives across sessions as a unit in a way a one-shot subagent does not. Mine does not decompose that way, so I pay the explicitness tax instead.

When not to use multiple agents at all

The honest boundary comes from the same company contradicting itself in public. Cognition published Don’t Build Multi-Agents in June 2025, arguing that parallel agents lose each other’s context and produce incoherent work. In April 2026 the same author published Multi-Agents: What’s Actually Working, which narrows rather than retracts: “multi-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions.”

That line is the rule I run. One implementer writes per branch. The researcher and the reviewer contribute intelligence and touch no files. Anthropic’s multi-agent research system post reaches the same place from the research side: “Most coding tasks involve fewer truly parallelizable tasks than research, and LLM agents are not yet great at coordinating and delegating to other agents in real time.”

Concretely, I do not reach for a second agent when the task is a single edit whose result I must immediately reason about, when writing the brief would cost more than doing the step, or when the work is a chain of dependent steps. A chain goes to one worker, start to finish. It does not become a relay of five.

Where this setup falls short is measurement. I have no ledger of review rounds, rework rate or time to merge, so a change to an agent’s description or contract ships unverified. The agent files are tuned by impression. That is the gap I would close first.

FAQ

What is an agent team in Claude Code?

An agent team, per Anthropic’s agent teams documentation, is several Claude Code sessions coordinated as a team: a lead delegates against a shared task list and teammates can message each other. It differs from plain subagents in that the teammates coordinate with one another, rather than each being spawned for one task and returning one report to whoever spawned it.

When to use Claude agent teams?

Use a team when the work splits into tasks that are genuinely independent, so teammates are working rather than waiting on each other. The test I apply before reaching for one is whether two subtasks can touch disjoint files: if the second one edits what the first one is still writing, the team’s shared task list turns into a queue with extra token cost. My own pipelines fail that test, which is why they stayed as skills.

Do Claude Code agents talk to each other?

In my setup, no: every subagent reports to the orchestrator session and only the orchestrator talks to everyone. That is deliberate, since one conversation holding the whole trace is easier to debug than several agents negotiating. Claude Code’s agent teams feature does allow teammates to message one another directly, and the orchestrator can also resume a specific agent by name or id with SendMessage rather than spawning a new one.

Is Claude Code an agent itself?

Yes. The main Claude Code session is an agent: it holds a context window, calls tools, runs commands and loops until the work is done. Subagents are additional agents it can start, each with its own context window, its own tool access and its own model. The distinction that matters is the context window, not the name on the file: a subagent whose only difference from the session is its wording is a prompt, not an agent.

Which model should each subagent use?

Match the tier to the judgment the task still needs after the brief is written: haiku for mechanical fan-out, sonnet for routine scoped work, opus for implementation and first-round review, and the strongest model for the orchestrating session that only decides. When you are unsure between two tiers, take the cheaper one, because rerunning a failed worker one tier up is cheap and usually reveals that the brief, not the model, was the problem.

Sources

claude codeagentic codingsubagentscase study
← Back to the blog