Spec, Plan, Build, Review: A Workflow for Agent Coding
Agents write code quickly. The slow part is finding out what they got wrong. This workflow keeps the agent on a short leash: a written spec, a plan where every task ends in a test, and one fresh reviewer at the end. We used it to build this site's playbook section. All 1,277 tests passed, and the fresh reviewer still found two critical bugs.
The workflow at a glance
flowchart LR
S[Write the spec] --> P[Plan tasks with tests]
P --> B[Build one task at a time, test first]
B --> R[Fresh agent reviews the whole branch]
R --> F{Critical or important findings?}
F -- yes --> X[Fix each one with a failing test first]
X --> T[Run the full test suite]
F -- no --> T
T --> D[Ship]
When to use it
- Changes that touch several files or more than one system.
- Work where a silent bug costs more than an extra review: billing, publishing, user data.
- Any request with gaps. An agent rarely stops to ask; it picks an interpretation and builds on it. The spec and plan are where those guesses surface while they are cheap to change.
Skip it for a one-file fix or a copy change. A quick test and a read of the diff is enough.
Step 1: Write the spec
- Do: Describe the outcome, who it is for, and every decision you have already made. Ask the agent to list its assumptions before it writes anything, then approve the spec before any code exists.
- Output: A short document in
docs/specs/. It is the authority when anything later disagrees.
Step 2: Turn it into a plan
- Do: Break the work into tasks that each end with a test you can run. Name what each task hands to the next one, such as function names and fields.
- Add a review focus: List the five inputs or failures your tests will not cover, like "GitHub is down" or "the model invents a link".
- Output: A plan file the agent works through in order.
Step 3: Build task by task, test first
- Do: For each task, write the failing test, watch it fail, write the code, and watch it pass. Then commit.
- Keep a ledger: Whenever the agent decides something the plan did not cover, it writes one line: the decision, the reason, and the cost if it is wrong.
Step 4: Review with fresh eyes
- Do: Start a new agent on your strongest model. Give it the spec, the plan, the full diff, the review focus list and the ledger.
- Ask for: Findings graded Critical, Important or Minor, each with a concrete failure scenario.
- Why fresh: The agent that wrote the code shares its own blind spots. A second agent with the same context would too.
Step 5: Fix in one pass
- Do: Re-grade each finding by what a user would actually see. Fix the critical and important ones, each with a test that fails first. Log the minor ones for later.
- Then: Run the full suite once more before you ship.
What the review caught
Two of its findings from our build:
- When a preview build failed, the automatic fix re-drafted a playbook post as an ordinary article and published it without review.
- After its research step, a post read as "pick an angle", so a stalled run could not be reopened or retried.
Both slipped past every test, because no test walked those paths. The review focus list pointed the reviewer straight at them.
Variations
- Medium change: Skip the written spec. Describe the design in chat and wait for a yes.
- Tight budget: Build everything in one session and pay only for the final reviewer.
- Ready-made: The open-source Superpowers plugin for Claude Code packages this workflow as skills.