Agentic Coding in Practice: Tools, Workflow and Agents

Agentic coding is delegating a whole task to an AI coding agent that plans, edits files, runs the tests, reads the failure and iterates until a check it cannot fake turns green. In my setup that looks like this: a session that plans and writes no code, workers on cheaper models that write all of it, a hook that blocks an agent from ending a turn with unverified changes while the check is red, a reviewer agent that reads code and writes only review notes, and a human whose only merge signal is a label. The whole arrangement serves one goal, stated in my own design doc: the human never reviews one 10k-line PR.

This is the overview of that system: the tools, the workflows, and what is still missing.

What agentic coding is (and how it differs from AI-assisted and vibe coding)

AI-assisted coding is autocomplete with a longer memory. I hold the plan, an AI coding tool supplies the next chunk, and I accept or reject each piece of generated code. Vibe coding is the same loop with the reviewing removed: you keep what the model wrote because it runs, not because you read it. Agentic coding is neither. The unit of work is a task with an acceptance check, handed to a coding agent that plans, edits files across the codebase, runs the test suite, reads the failure and tries again, and comes back only when something external is required: a decision, a merge, a genuine unknown.

The dividing line between all three is where verification lives. In AI-assisted coding it lives in your eyes, in vibe coding nowhere, and in agentic coding in a command the agent has to pass before it may stop. That is the only one of the three that scales past what one person can read in a day.

Put another way: AI code generation changed how fast you write code, and agentic AI changes who runs the loop. You state the outcome in natural language, an AI agent turns it into a code change across the codebase, and your part of software development becomes specifying and gating.

That makes me the owner of a harness rather than a faster typist. “Harness engineering” is the vocabulary circulating for this right now, and Anthropic’s own piece on the subject is titled Harness design for long-running application development: the design surface is not the prompt, it is the scaffolding around the model. Which agent gets which model, what a brief may contain, which commands are denied outright, what must be true before a turn may end. That is also why “which of the agentic coding tools is best” is the wrong question at this high level: the tool is the runtime, the harness is the product, and the harness is the part you own.

The agentic coding tools I actually use

The field is wide now. Cursor, GitHub Copilot, Gemini CLI, Codex and a long tail of AI tools all run the same agentic loop over the same three or four frontier models. I have not used Cursor or Copilot’s agent mode long enough at this level of autonomy to compare them honestly, so this is one developer’s stack rather than a ranking of AI coding tools. The differences between AI coding agents matter less than what you wrap around them. Each tool below links to the post that covers it in depth.

Claude Code as the runtime. Everything below is a feature of one CLI, which is why I stopped trying to build an agentic workflow out of separate tools: the terminal runtime gets the same file, shell and git access I have. One session is the orchestrator, and it plans, decomposes, delegates, decides, and writes no code at all.

Subagents on a model ladder. Workers are Claude Code subagents, each a file with a brief, a tool list and a model. The orchestrator runs on the most expensive model available, so every direct tool call it makes bills at frontier rates: reading a file, git plumbing, running a command and applying a specified edit all go to a cheaper worker. Mechanical fan-out goes to the smallest model, routine scoped work one tier up, implementation and first-round review to the default worker tier. The rule is deliberately lazy: “When unsure between two tiers, take the cheaper one — a failed worker is cheap to rerun one tier up.” The ladder, the read-only researcher and reviewer roles, and why there is no conductor agent above the orchestrator are in Claude Code agents and agent teams.

Hooks that enforce verification. The highest-leverage piece of the setup is about 200 lines of bash across three files. A Stop hook (and its SubagentStop twin) runs the repository’s own check.sh before an agent may end a turn with unverified changes; if it fails, the hook exits 2 and the failure goes back into the agent’s context, so the turn cannot end. When there is nothing to check, a clean pushed tree or a repository with no check.sh, it skips and says so in a trace log. Anthropic’s best-practices page puts it in one sentence: “Unlike CLAUDE.md instructions which are advisory, hooks are deterministic and guarantee the action happens.” A second hook formats every file an agent writes and always exits 0, because “Formatting can fall back safely; running the wrong test suite cannot.” The fingerprint cache, the five skip conditions and the rest of the failure policy are in Claude Code hooks that enforce verification.

CLAUDE.md and skills. The always-loaded context file holds only cross-cutting rules and points at depth-on-demand documents, because a long one gets ignored. The pipelines are skills, not agents: a skill runs inside the orchestrator session and can spawn agents, message a running one and react to a human signal mid-flight, while “An agent cannot orchestrate other agents.” The two that write to the forge, /feature and /bugscan, are slash-only by a flag in their front matter: they create branches, pull requests and issues, which should start because a person typed something. How the context file, the skills and the on-demand docs fit together is in context engineering for Claude Code.

MCP, honestly. MCP is the weakest part of my stack. The shared config declares no MCP server at all: agents reach the outside world through the shell and gh, not through tools. I add MCP servers per project, for development work a CLI genuinely cannot do, such as driving a browser session, because an MCP server is permanent context in every session of that project. In an enterprise setup, where the data sits behind internal APIs instead of in the repository, that balance would come out very different.

Worktrees. Each implementer works in a git worktree under <repo>/.worktrees/<branch>, excluded locally rather than through the shared ignore file. Two agents editing one checkout is the cheapest way to corrupt a run, and a worktree per branch removes that without a container, though it is still weaker than a network-isolated VM.

The gh CLI. The forge is an API, not a website, so agents work through gh for issues, labels, pull requests and CI status. It is also the only surface where autonomy was expanded: gh api, gh issue, gh label and gh pr are allowlisted, while the deny list stays closed on destructive commands (rm -rf, sudo, force pushes) and on exfiltration (curl, wget, any read of .env or *.pem). An AI agent that cannot make an outbound request and cannot read a key has no path from a secret it found to a secret leaving the machine.

The workflow end to end

One task moves through six states of the agentic workflow, and only two of them involve me.

Six-stage agentic workflow: brief, checkpoint, implement, hook-gated stop, review and merge-me, with review findings looping back to implement.

The brief names one high level outcome, a runnable acceptance check, the files expected to change and an explicit not-in-scope list, never a list of deliverables. The checkpoint comes before any edit: the worker returns the units it would add, each with one clause saying why, and the orchestrator deletes whatever the acceptance line does not require. Then it implements, and the hook decides when the turn may end. A reviewer agent that never wrote the code reviews the result, because the author is the worst possible judge, as Anthropic’s harness design post puts it: “When asked to evaluate work they’ve produced, agents tend to respond by confidently praising the work—even when, to a human observer, the quality is obviously mediocre.” Findings go back to the same implementer for at most three rounds; three failed rounds means the brief was wrong, not the code, and that escalates to me. Working this way, review notes are the only agent output I read closely.

The merge signal is a label. The agents act through my own credentials, and “GitHub forbids approving your own PR”, so the approve button is greyed out on every pull request the system opens. One carrying merge-me has my consent. If CI is red the orchestrator does not merge and does not ask again either: “Consent is to the change, not to a broken build; re-labelling after every red run would be exactly the human interaction the system is trying to remove.” A human comment counts as request-changes, never against the three-round limit. Without a forge, review state is files under .review/<branch>/ and the signal is me typing “merge”.

The failure mode I did not anticipate is that agent pull requests arrive too big. Not wrong. Too big. A brief saying “add a caching layer” comes back as a module, a config class, an abstract base for two implementations I might want later, and a docs section explaining all of it. Reviewers check that what exists is correct, never whether it should exist.

So the contract has three teeth. The checkpoint above, a shrink pass once the check is green (the worker drops what the acceptance check does not need and ends its report with a Not built: … line), and a reviewer scope pass that runs before correctness, where “Every n is a finding whose default direction is drop it”.

The tendency is not just mine: correct-and-concise patching is an open research problem in LLM-based program repair, not a prompt I got wrong.

Pull request (2026-09-15) First cut Change actually needed Ratio
PR A (small API change) +455 / −47 ~80 lines ~5.7x
PR B (CLI tweak) +322 / −43 ~120 lines ~2.7x

Both projects are private, so these rows are illustrative rather than independently checkable. Source: git diff --stat on those branches against my estimate for each brief when it was written. Target is a first cut within 1.5x, re-measured on the next three pull requests. Two data points and a subjective denominator is a thin baseline, and I would rather publish it thin than use adjectives.

What the contract does not do is gate on line count. “Cost is files touched and concerns braided together, not lines — a line count is a prompt for a reason, never a gate.” The only number in it is a ceiling, “~400 changed lines is the hard ceiling for reviewability, never the target.”

What surprised me in real world use is that code quality comes mostly from the check, not the model: the same agent with the same brief produces better code in a project with a fast check.sh. My own time now goes into writing briefs and reading review threads, and the hardest skill is saying what “done” means in plain language before anything starts.

Agentic coding best practices: a scorecard against the 2026 literature

I scored the setup against published practice in September 2026. It comes out strong on process and weak on evidence, which I suspect is the common shape for developers assembling agentic workflows alone rather than inside an enterprise platform team. Each row is a practice the 2026 literature on agentic software development converged on.

Practice Status Where it lives
Orchestrator / worker split Implemented Agents and agent teams
Model tiering per subagent Implemented Agents and agent teams
Generator is not the evaluator Implemented Read-only reviewer agent
Review by a separate agent Implemented Adversarial brief and rubric
Enforced verification via hooks Implemented Claude Code hooks
Context engineering Implemented Context engineering
Small scoped changes Implemented Least-feasible contract, ~400-line ceiling
Avoid over-engineering Implemented Checkpoint, shrink pass, Not built: line
Worktree isolation Implemented One worktree per implementer
Human-in-the-loop consent Implemented merge-me label, three-round escalation
Permission and security boundaries Implemented Deny list on destructive and exfil surfaces
Spec or brief before code Partial Brief mandatory; no acceptance criteria
Measurement and telemetry Not built Two data points, one trace log
Evals for agent changes Not built Nothing tests a prompt change
Cross-model review Not built Deliberately, see below
AGENTS.md cross-vendor file Not built Claude-only today

What I haven’t built

A measurement ledger. There is no record of lines, review rounds, human comments, time to merge or rework rate. The plan is one row appended by a worker after each merge. Until it exists, a regression from editing an agent brief is invisible, and Anthropic’s evals post names the cost: “Without them, it’s easy to get stuck in reactive loops—catching issues only in production.”

Evals. Every change to a context file or an agent brief ships unverified, and the bar the evals post sets is not high: “In reality, 20-50 simple tasks drawn from real failures is a great start.” Months of transcripts have already handed me the task list. I have not frozen it.

A spec artifact. The brief describes the work but never states, before code starts, what “done” means as acceptance criteria mapped to tests. Without that contract the review-and-fix loop has nothing to converge to, which is how a loop converges confidently on the wrong fix.

Cross-model review, a deliberate no rather than a not-yet. A Claude reviewer over Claude-written code invites the obvious objection of self-preference. Panickssery et al. (2024) found a causal link between recognizing one’s own text and over-scoring it, and Wataoka et al. (2024) traced it to perplexity: judges over-reward text that feels familiar.

The strongest mitigation is verification, not a second vendor. Grounding a judge in code execution (Findeis et al. 2025) “increased its agreement with ground truth from below 42% to approximately 72%”, per the SE literature review, and Anthropic’s harness post agrees the problem is worst where there is “no binary check equivalent to a verifiable software test.” My reviewer runs after the test and lint hooks, in a separate agent with no memory of writing the code, the blind condition under which Chae et al. (2026) found self-preference disappears. It is never told the author’s model: Saraf et al. (2025) found the “‘Claude’ label consistently elevated scores regardless of actual authorship.”

The one controlled cross-model study, Xiang et al. (2026), does not support the swap: Claude reviewing Codex raised pass rates from 71.6% to 89.7%, Codex reviewing Claude dropped them from 91.4% to 82.8%, and Claude self-review left the baseline unchanged. A vendor swap bets the other model reviews better; here it was weaker. The live hole: blind review removes the authorship label but not familiarity, and nobody has tested a blind same-family reviewer. I would revisit it for work no test can check, where Pombal et al. (2026) show the bias survives objective rubrics.

Sandboxing. Worktrees on my own machine are a weaker boundary than the network-isolated VMs the vendors run. Cognition’s description of how each managed Devin is isolated is the standard I am not meeting: “Each managed Devin gets a clean slate, a narrow focus, its own shell, and its own test runner.” The deny list and the human merge gate are the only compensating controls, adequate as long as agents have no deploy rights.

Explicitly rejected

  • A conductor agent above the orchestrator. A frontier agent under a frontier session duplicates it with no added capability. “It was tried and reverted.”
  • The smallest model inside the pipelines. The savings did not cover the output quality on these tasks. It keeps its place in the general ladder for grep-and-report fan-out.
  • Pull request approval as the merge signal. Structurally impossible when agents act through my own credentials.
  • Parallel writers on one branch. Cognition published Don’t Build Multi-Agents in June 2025 and then Multi-Agents: What’s Actually Working in April 2026, narrowing rather than retracting: “multi-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions.” That is the multi agent boundary my setup landed on: one implementer writes per branch, researcher and reviewer touch nothing.
  • A line-count merge gate. Reviewability is about braided concerns, and a numeric gate just teaches workers to split badly.

FAQ

What is the best setup for agentic coding?

The best setup is the simplest agentic workflow with a real verification gate: one agent, one runnable check per repository, and a hook that blocks the turn when the check fails. Most developers build it backwards, starting from the agents. Add a role only when a failure demands one, which in my experience meant a separate reviewer agent. The choice of AI tools matters far less than whether “the agent said it was done” is your only signal.

How do I get started with agentic coding?

Start with the check, not the agents. Write a check.sh that runs lint, typecheck and unit tests in seconds, wire it to a Stop hook, and let one agent work against it for a week on low-stakes tasks. Then add a short context file with the commands and conventions of the codebase, and only after that a second agent to review.

Is agentic AI good for coding?

It is good at bounded tasks with an executable definition of done, and unreliable without one. The 2025 DORA report’s line is the one I keep coming back to: “AI doesn’t fix a team; it amplifies what’s already there.” Whether it makes anyone faster is a question I cannot answer from this setup, because I have no ledger: the only controlled study I know of, METR’s 2025 trial with experienced open-source developers, found them “19% longer to complete issues” while believing they had sped up. What a fast test suite and clear boundaries buy is not speed but the ability to tell; a codebase without them gets more AI-generated code faster and a slower review queue. The long term question for software engineering is not whether AI models can write code, but whether your repository can prove it is correct.

What is the difference between agentic coding and AI-assisted coding?

The unit of work differs. An AI-assisted coding tool hands you one change at a time and you accept or reject it; a coding agent takes a whole task and returns when a check passes or when it needs a decision from you. The tell in a codebase is what the tools force you to write: assisted coding needs no acceptance check, and an agent is unusable without one, which is why adopting it starts with a check.sh rather than with a model.

Which tools are used for agentic coding?

The common agentic coding tools split into three groups: terminal AI coding agents such as Claude Code, Gemini CLI and Codex; editor-based coding assistants such as Cursor and the VS Code Copilot agent; and cloud agents that open pull requests on their own. I run Claude Code, and what I depend on are its harness features rather than the chat: subagents, hooks, skills, a context file, git worktrees and the gh CLI. The group matters less than the question to ask of any of them, which is what the tool will not let an agent do.

Next in the queue: the ledger, because everything else I claim here is unmeasured; then a frozen set of replayable tasks drawn from failures I have transcripts for; then acceptance criteria in the brief, mapped onto tests. A harness that enforces verification amplifies a person who verifies. One without measurement amplifies whatever I am wrong about, faster.

Sources

agentic codingagentic coding toolsclaude codeagentic aicase study
← Back to the blog