Files
claude-foundations/best-practices/debugging.md
Paul O'Reilly 1b5e73dc54 Distill best practices from agent-runtimes M1-M3 memory files
12 additions/updates across 5 best-practice files:
- docker-uid-matching: userdel simplification, SSH agent socket UID match
- debugging: GIT_SSH_COMMAND scope limitation
- test-driven-development: subprocess mock gotcha, routing callables,
  Pydantic v2 field_validator defaults, sys.exit at module level
- spec-driven-development: multi-agent orchestration practices (commit WIP,
  self-verify, import conventions, assembly budget)
- validation: test pre-commit hooks after adding dependencies

Source: agent-runtimes/memory/ (decisions, gotchas-docker, gotchas-python,
process-lessons, m1/m2/m3 reflections)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-25 11:11:54 +13:00

4.8 KiB

Debugging Methodology

Check Before You Act

  • Before writing firewall/network rules, check actual routing (ip route get <dest>)
  • Before running config management with variables, ensure values are real, not placeholders
  • Before assuming a container has a shell, docker inspect it
  • Before creating API tokens, research all required scopes upfront — iterating one scope at a time costs a push-debug cycle each

Routing and Networking

  • Always run ip route get <dest> on the forwarding host first
  • macvlan, Docker bridge, and other virtual interfaces mean the "obvious" physical interface is often wrong
  • Test from both in-cluster and external perspectives

Full-Chain Testing

After wiring up any new service:

  1. Test direct to backend (bypass all proxies)
  2. Test through reverse proxy (bypass DNS)
  3. Test end-to-end as a user would

Use curl --resolve to test specific paths without depending on DNS propagation.

When Something Doesn't Sync/Apply

  • Check resource exclusions in the GitOps controller immediately
  • Check if the resource type requires special permissions or labels
  • Check if ServerSideApply conflicts are preventing field changes
  • Don't try workarounds before understanding the root cause

OIDC Integration Checklist

Before starting any OIDC integration, research:

  1. What format is the sub claim (UUID? username?)
  2. Which claims are in the ID token vs userinfo endpoint
  3. How the consumer matches RBAC identities (groups? email? username?)

Log-First Diagnosis

  • CrashLoopBackOff: check logs first. Error messages in pod logs usually point directly to the fix. Don't tweak configuration or security contexts blindly — kubectl logs <pod> first.
  • Discriminate transient from persistent errors. CSI lock contention, etcd timeouts during first install, and brief connectivity blips are self-healing. Don't spend time debugging errors that resolve on retry. If you see retry/backoff patterns in logs, wait before intervening.
  • Trust controller retry logic. CSI controllers, operators, and reconciliation loops have built-in retry. Transient failures during rapid provisioning are expected, not bugs.

Reproduce Before Fixing

When a bug is discovered or reported, do not start by trying to fix it. The first step is always to write a test that reproduces the failure:

  1. Write a failing test. Capture the bug as a test case that demonstrates the broken behaviour. This forces you to understand the bug precisely — what input triggers it, what the wrong output is, and what the correct output should be.
  2. Fix the bug in isolation. Use a subagent or a separate session to write the fix. The fixing agent gets the failing test as its success criterion — it's done when the test passes. This separation prevents the fixer from unconsciously weakening the test to match a broken implementation.
  3. The test stays forever. The reproduction test becomes a permanent regression test. It proves the fix works and prevents the bug from returning.

This workflow has several advantages:

  • Forces precise understanding. Writing a test means you know exactly what's broken, not just "it doesn't work."
  • Prevents partial fixes. The test defines "done" objectively — the fix either passes or it doesn't.
  • Parallelises work. While one agent fixes the bug, you can continue other work.
  • Catches regressions. The test remains in the suite, guarding against the same class of failure.
# Step 1: Write the failing test FIRST
def test_regression_issue_427_empty_payload_crashes():
    """Bug #427: Empty payload causes unhandled TypeError in dispatcher.
    Should return a 400 validation error, not crash."""
    response = client.post("/dispatch", json={})
    assert response.status_code == 400  # Currently crashes with 500

# Step 2: Hand to a subagent/session: "Make this test pass without breaking others"

GIT_SSH_COMMAND Only Affects Git-Invoked SSH

GIT_SSH_COMMAND (e.g., ssh -o StrictHostKeyChecking=no) only applies when git invokes SSH internally (clone, push, fetch). Direct ssh calls — such as ssh -T git@host for connectivity testing — ignore it entirely. When working in containers or CI environments where host keys aren't pre-trusted, direct SSH commands need explicit flags: ssh -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null.

Pattern Mining Before Authoring

Before building a new service, component, or script, read existing patterns in the codebase first. This matches conventions on the first attempt and avoids rework on naming, structure, and integration points. Applies to K8s manifests, CI pipelines, skill authoring, and script structure.

Grep Your Own Docs

Known issues documented in CLAUDE.md or MEMORY.md but not applied to new scripts/configs waste debugging time. Search your own documentation before writing automation that touches areas with known gotchas.