8 Commits

Author SHA1 Message Date
Paul O'Reilly
c8691368e4 distill: flush prior-session best-practice additions
Additive/refining entries left uncommitted in the working tree from an earlier
distill session (found during the 2026-07 sweep commit): agent-repos 5-step
plan pattern; ai-parallel-agents cheap-model scope; spec-driven spec-inversion
and caller/callee cross-reference rules.
2026-07-02 15:58:52 +12:00
Paul O'Reilly
bd3a99e2ad docs(airouter): correct failure-path picture + F70 vs local divergence
The previous "Dogfood Failure Path: Branch + Logs Lost" section was wrong
on two counts (verified by M16 Wave B3, 2026-05-08):

1. The agent's branch IS pushed on F70-failure (dispatcher logs "branch
   will still push (partial work preserved)"). The earlier "no branch"
   claim was a fetch-refspec mistake — both B2 and B3 have task-<id>
   branches on the agents fork.

2. The CP task record's logs ARE populated (~50 KB on B3). What's actually
   missing is the pytest stdout/stderr from the F70 invocation — only the
   high-level "Tests failed" line is logged.

Section retitled "What's Visible, What Isn't" with the corrected picture.
Operator pattern updated: fetch the task branch, apply the diff locally,
run the test — if local passes, cherry-pick to main.

New section "F70 Pytest Can Disagree with Local Pytest" captures the B3
finding: airouter's 15-line MN-14 validator passed 5/5 locally but F70
reported failure twice. Possible causes listed; workaround is
--max-test-iterations 1 + local apply-and-run after failure.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-08 20:27:41 +12:00
Paul O'Reilly
52cd0a3cb6 docs(airouter): 2026-05-08 dogfood batch + test design rules
Three additions to agent-repos.md based on the post-A1-incident dogfood
batch (8 successes + 1 operator-induced "failure"):

1. Airouter Qwen3.6 section: pattern reconfirmed across M16 Wave A1/A2/B1
   and M25 Waves A1-A5 + B1-B2. Time-to-success bands recorded for cost
   calibration (1m30s for git rm; ~5 min for Pydantic regex; ~12 min for
   class addition). Default --max-test-iterations 1 for cheap probes.

2. New section: Test Design for AI Agent Dogfood Pipelines. Triggered by
   the M16 Wave B2 (MN-4 prompt cap) failure — a 14-minute airouter run
   blamed on the agent that was actually an over-strict test asserting on
   sanitised 422 body content. CP's RequestValidationError handler strips
   Pydantic detail for security; tests asserting body content for that
   path are structurally impossible. Rules: verify test passes against a
   reference impl before pushing; status-code-only ceiling for validator-
   driven 422s; model on previous successes not stricter variants; F70
   retries don't recover structurally impossible tests.

3. New section: Dogfood Failure Path: Branch + Logs Lost. When all F70
   retries exhaust, the agent's last attempt is not pushed to the agents
   fork, the CP task record's logs field is empty, and the pod is gone.
   Operator must reproduce locally — until F70 finalize-on-failure pushes
   the failed branch.

BESTPRACTICES.md index updated to reflect the new sub-topics.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-08 19:50:45 +12:00
Paul O'Reilly
7bfabceaac docs(airouter): confirm single-file-unit pattern with D4a+D4b results
D4a (harness.yaml, 1 file) = success
D4b (CLAUDE.md, 1 file) = success
Both succeeded where original D4 (2 files) = silent no-output.

Pattern confirmed: single-file scope is the reliable unit for airouter/Qwen3.6.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-04 15:52:55 +12:00
Paul O'Reilly
e618ca4d22 docs(airouter): add scope-decomposition guidance for Qwen3.6 reliability
Airouter Qwen3.6 reliable for single-file tasks; unreliable for multi-file/multi-rule.
D7 (single harness + one rule) = success. D4 (two files, multiple rules) = silent no-output.

Pattern: decompose airouter tasks to one file per dispatch.
BUG-21 reference included.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-04 15:06:43 +12:00
Paul O'Reilly
22d49b2c9a distill: best practices from 2026-04-19 cross-project run
Adds 3 new topic files (ai-parallel-agents, api-integration,
python-patterns) and extends 21 existing topic files with new gotchas
and patterns surfaced from memory across tracked projects. Index
updated accordingly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-25 13:41:47 +12:00
Paul O'Reilly
8aa400a5d4 distill: 49 best practices from 5 projects (2026-03-27..2026-04-05)
Add 37 new entries and update 7 existing entries across 13 topic files.
Major contributions from agent-runtimes (K8s secrets, CI, Docker gotchas),
cluster-bootstrap (ArgoCD SSA, etcd tuning, DB migrations, Compose networking),
and cluster-apps/octopus-deploy (Helm vs raw manifests, ArgoCD source types).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-06 01:10:07 +12:00
Paul O'Reilly
bff46b9182 Add agent repos & container agent operations best practice
Comprehensive guide covering task submission to the agent-runtimes
control plane, available harnesses, monitoring, multi-model workflows,
agent repo forks with workspace layout, and artifact extraction patterns.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-05 11:32:48 +12:00