Files
claude-foundations/best-practices/spec-driven-development.md
Paul O'Reilly e94417b896 Add best practices, hooks, memory files, and scripts from recent sessions
Includes: spec-driven and test-driven development best practices,
reproduce-before-fixing debugging workflow, require-plan-file hook,
find-project-root script, session logs, memory files for decisions/
gotchas/process-lessons, and updates to existing best practice topics.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-17 09:47:47 +13:00

10 KiB

Spec-Driven Development with AI Agents

Best practices for using structured specifications to coordinate AI agent implementation work. Extracted from real project experience (agent-runtimes) and industry research (OpenSpec, Codified Context paper, Addy Osmani's workflow guides).

Why Specs Matter for AI Agents

AI agents trust documentation absolutely. A well-written spec gives an agent everything it needs to implement a subsystem without reading the entire codebase. A stale or vague spec causes silent failures where agents generate code that is structurally valid but architecturally wrong.

Specs serve three functions that CLAUDE.md alone cannot:

  1. Compressed context — an agent reads one spec, not 300 lines of mixed concerns
  2. Testable contracts — numbered requirements and scenarios translate directly to pytest
  3. Handoff boundaries — an agent working on the dispatcher doesn't need to understand the entrypoint internals, just the interface between them

Spec Structure

Each spec follows a consistent template. Sections are ordered so an agent can read top-down and build understanding progressively.

Required Sections

  1. Overview — What this subsystem does. 2-3 sentences. An agent should know if this spec is relevant after reading this.
  2. Responsibilities — What this subsystem owns and what it delegates. Prevents scope creep during implementation.
  3. Dependencies — Which other specs to read first. Keeps the reading list minimal.
  4. Data Model — Types, schemas, state machines, interfaces. The concrete contract.
  5. Requirements — Numbered functional requirements (e.g., E-1, E-2). Each must be independently testable.
  6. Scenarios — Concrete given/when/then examples. These become test functions.

Optional Sections

  1. Interface — API surface, function signatures, HTTP endpoints. Include when the subsystem has an external-facing API.
  2. Extension Points — How to add new capabilities without modifying existing code. Step-by-step instructions.
  3. Error Handling — Failure modes and expected behaviour. Prevents agents from inventing their own error strategies.

Writing Guidelines

  • Be specific, not comprehensive. A spec that says "handle errors appropriately" is useless. A spec that says "return exit code 124 on timeout" is testable.
  • Include the why. Design intent and constraints prevent agents from making structurally valid but architecturally wrong changes. Requirements without rationale are followed mechanically — agents can't judge edge cases or make trade-offs. Every constraint needs a "Why:" line. Example: "Secrets never in payload" needs "because payloads may be logged and stored in task history".
  • Use concrete examples. Every data model section should include a realistic JSON/code example, not just a schema.
  • Cross-reference, don't duplicate. If two specs share a concept (e.g., the payload schema), one spec owns it and the other links to it.
  • Keep each spec self-contained. An agent should be able to implement a subsystem by reading the target spec plus its listed dependencies. If it needs to read CLAUDE.md, the spec is incomplete.

Requirement Numbering

Each spec uses a short prefix derived from its name, followed by a sequential number:

Spec Prefix Example
payload.md P P-1, P-2
entrypoint.md E E-1, E-2
actions.md A A-1, A-2
runners.md R R-1, R-2
dispatcher.md D D-1, D-2
control-plane.md CP CP-1, CP-2

Requirements must be:

  • Independently testable — each maps to one or more test functions
  • Unambiguous — an agent can determine pass/fail without human judgement
  • Stable — changing a requirement number invalidates tests, so avoid renumbering

Scenarios as Test Blueprints

Every scenario in a spec should be directly translatable to a test function. Use this format:

### Scenario: Pre-action failure
**Given:** Payload with clone pre-action (invalid repo URL)
**When:** Clone fails (git returns non-zero)
**Then:** on_error actions run, container exits 1. Runner never executes.

This becomes:

def test_scenario_preaction_failure_runs_on_error_and_exits_1(...):
    """Given clone fails, on_error runs and exits 1. Runner never executes."""

Guidelines:

  • Each scenario tests one behaviour, not a combination
  • Include both happy paths and error paths
  • Name the scenario descriptively — it becomes the test function's docstring
  • Include enough setup detail that an agent can write the test without guessing

The Spec → Test → Code Workflow

This is the core development loop. Tests are written from the spec before code exists.

1. Write or Update the Spec

Define requirements and scenarios. Get them reviewed. The spec is the source of truth for what the system should do.

2. Write Tests from the Spec

Translate requirements and scenarios into pytest functions. Tests should:

  • Map to requirement IDs in their names: test_e4_preaction_failure_skips_remaining
  • Use the scenario's given/when/then as the test body structure
  • Mock external dependencies (subprocess, HTTP, filesystem)
  • Run fast (no Docker, no network, no real APIs)

3. Run the Tests — They Should All Fail

This confirms the tests are actually testing something. If a test passes before implementation, it's either testing the wrong thing or the feature already exists.

4. Implement Until Tests Pass

Write the minimum code to make tests pass. The spec defines what, the tests verify it, the code implements it.

5. Update Spec if Implementation Reveals Issues

Sometimes implementation reveals that a requirement is unworkable or incomplete. Update the spec, update the test, then update the code. The spec stays authoritative.

Spec Maintenance

Preventing Drift

Specs drift from code when they're treated as planning documents that are "done" after implementation. They must be treated as living contracts.

Rules:

  • Spec changes require test changes. If a requirement changes, its test must change in the same commit.
  • Code changes that affect interfaces require spec changes. If a function signature, API endpoint, or data schema changes, the relevant spec must be updated in the same commit.
  • New features require spec-first. Add the requirement and scenario to the spec, write the test, then implement.

CI Enforcement

Enforce spec hygiene with automated checks:

  1. Pre-commit hook — run pytest, block commit on failure (already implemented)
  2. Spec coverage check — a script that verifies every numbered requirement has at least one test function referencing it
  3. Orphan test detection — tests referencing requirement IDs that no longer exist in specs

Review Checklist

When reviewing a PR that touches a spec subsystem:

  • Spec updated if interface or behaviour changed
  • Test added/updated for new/changed requirements
  • Cross-references still valid
  • No requirements removed without deprecation note

Context Architecture for Agents

Based on the Codified Context paper (108k-line system, 283 sessions), structure project knowledge in three tiers:

Tier 1: Hot Context (Always Loaded)

CLAUDE.md — conventions, env vars, repo structure, scripts. Loaded every session. Keep under ~300 lines by linking to details elsewhere.

Tier 2: Spec Context (Per-Task)

spec/ files — loaded based on what the agent is working on. An agent implementing a new action reads spec/actions.md + spec/payload.md. An agent working on the dispatcher reads spec/dispatcher.md + spec/container-backends.md.

The spec index (SPEC.md) has a "read this when..." column to guide selection.

Tier 3: Cold Context (On-Demand)

memory/ files — gotchas, reflections, decisions. Loaded only when relevant. An agent hitting a weird Cilium issue checks memory/gotchas-cilium.md.

Routing Context to Agents

When launching an agent to work on a subsystem:

  1. Point it at the relevant spec(s) via its prompt
  2. Include CLAUDE.md for conventions
  3. Let it pull from memory/ on-demand if it hits issues

Don't load everything — agents perform better with focused context than with a 50-page dump.

Testing Depth

The spec→test→code workflow defines when to write tests. For how to write comprehensive tests — edge case discovery, property-based testing, mutation testing, AI agent testing patterns — see Test-Driven Development.

Post-Write Spec Audit

After writing specs, audit them against best practices before implementation. Common gap categories:

  1. Missing rationale — Constraints without "Why:" lines. Agents follow them mechanically but can't judge edge cases.
  2. Missing error/failure scenarios — Happy paths are covered but failure modes aren't specified.
  3. Cross-spec interface misalignment — Two specs describe the same interface differently.
  4. Vague requirements — "Handle errors appropriately" instead of specific error codes and behaviours.
  5. Missing specs for discovered subsystems — Implementation reveals components that weren't planned for.

Write-then-audit is more productive than trying to get specs perfect on the first pass. The audit step catches systematic gaps across all specs at once.

Anti-Patterns

Specs as documentation, not contracts

Symptom: Specs describe what was built, updated after the fact. Tests don't reference spec IDs. Fix: Write specs before code. Tests reference requirement IDs. Specs are the input, not the output.

Mega-spec

Symptom: One large spec covering the entire system. Agents must read thousands of lines to find what they need. Fix: Split by subsystem. Each spec should be readable in under 5 minutes.

Spec without scenarios

Symptom: Requirements are abstract ("handle errors gracefully"). No concrete examples. Fix: Every requirement needs at least one given/when/then scenario with specific inputs and outputs.

Implementation details in specs

Symptom: Spec dictates variable names, algorithm choices, internal data structures. Fix: Specs define what and why, not how. The interface is specified; the implementation is free.

Untested requirements

Symptom: Requirements exist in the spec but no test references them. They drift without anyone noticing. Fix: Spec coverage check in CI. Every requirement ID must appear in at least one test function name.