Distill best practices from agent-runtimes M1-M3 memory files
12 additions/updates across 5 best-practice files: - docker-uid-matching: userdel simplification, SSH agent socket UID match - debugging: GIT_SSH_COMMAND scope limitation - test-driven-development: subprocess mock gotcha, routing callables, Pydantic v2 field_validator defaults, sys.exit at module level - spec-driven-development: multi-agent orchestration practices (commit WIP, self-verify, import conventions, assembly budget) - validation: test pre-commit hooks after adding dependencies Source: agent-runtimes/memory/ (decisions, gotchas-docker, gotchas-python, process-lessons, m1/m2/m3 reflections) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -178,6 +178,32 @@ After writing specs, audit them against best practices before implementation. Co
|
||||
|
||||
Write-then-audit is more productive than trying to get specs perfect on the first pass. The audit step catches systematic gaps across all specs at once.
|
||||
|
||||
## Planning Session Limits
|
||||
|
||||
Architecture decisions, infrastructure research, and spec refinement each get one planning session. After three sessions of planning, start implementation. Specs are hypotheses that need code to validate them — extended planning without implementation produces diminishing returns and theoretical designs that don't survive contact with reality.
|
||||
|
||||
## Categorize Findings Before Acting
|
||||
|
||||
When a spec review or audit produces many findings, categorize them by priority (high/medium/low) before making changes. Present the categorized list for alignment before editing. Starting edits without prioritization leads to scope creep — low-priority cosmetic fixes consume time that should go to high-priority structural gaps.
|
||||
|
||||
## Multi-Agent Orchestration Practices
|
||||
|
||||
### Commit WIP Before Decomposing Tasks
|
||||
|
||||
Untracked and uncommitted files are NOT available in git worktrees. If agents work in worktrees (or container-mounted worktrees), they won't see specs, plans, or dependency outputs that haven't been committed. Commit to a staging branch before decomposition — this eliminates the dominant overhead of manually copying files into each worktree.
|
||||
|
||||
### Agents Must Self-Verify with Tests
|
||||
|
||||
Add "Run tests and fix any failures" to every implementation agent prompt. Agents that write code without running tests produce bugs that only surface during assembly. Self-verification catches issues while the agent still has full context of what it wrote.
|
||||
|
||||
### State Import and Style Conventions Explicitly
|
||||
|
||||
Agents default to standard language conventions (e.g., relative Python imports, standard packaging). If the project uses non-standard patterns (bare imports, specific naming conventions, module-level structure), state them explicitly in the prompt. A single line like "Use `from harness import X`, not `from .harness import X`" prevents import mismatches during assembly.
|
||||
|
||||
### Budget for Assembly Fixups
|
||||
|
||||
Parallel agent work produces ~3 fixups per orchestration run, each under 5 minutes. Common fixup categories: import conventions, module-level side effects, SDK exception constructor signatures, validator patterns. This is the expected cost of parallel work, not a failure. Budget 15-20 minutes for assembly and fixup after each orchestration run.
|
||||
|
||||
## Anti-Patterns
|
||||
|
||||
### Specs as documentation, not contracts
|
||||
|
||||
Reference in New Issue
Block a user