Promotions from reflecting 21 projects' session logs (incl. agent-runtimes 122-log drain). Adds coverage across networking (eBPF VIP/VPN SNAT/VLAN bridge/forward-auth preflight/ingress TLS), kubernetes (CSI hotplug/PodSecurity debug/self-managed GitOps/runtime annotations), CI (dispatch tokens/runner death/base image), git (CI-rebase/shallow reset/PR governance), python (async session pool/httpx redirects/logging), TDD (AsyncMock/xfail lifecycle), api-integration (SDK parse/token-scope 404/schema probing), plus docker, scripting, debugging, security-architecture, secrets, react, octopus. State: .distill-state.json refreshed with current HEADs + 5 newly-tracked projects.
19 KiB
Debugging Methodology
Check Before You Act
- Before writing firewall/network rules, check actual routing (
ip route get <dest>) - Before running config management with variables, ensure values are real, not placeholders
- Before assuming a container has a shell,
docker inspectit - Before creating API tokens, research all required scopes upfront — iterating one scope at a time costs a push-debug cycle each
Routing and Networking
- Always run
ip route get <dest>on the forwarding host first - macvlan, Docker bridge, and other virtual interfaces mean the "obvious" physical interface is often wrong
- Test from both in-cluster and external perspectives
Full-Chain Testing
After wiring up any new service:
- Test direct to backend (bypass all proxies)
- Test through reverse proxy (bypass DNS)
- Test end-to-end as a user would
Use curl --resolve to test specific paths without depending on DNS propagation.
Split-Horizon DNS Can Hide Bugs from Local Testing
When /etc/hosts or internal DNS points a public hostname at an internal IP, local curl bypasses the external path (VPS, CDN, cloud LB) and masks bugs that are only visible to external users. Always verify production behaviour through the actual public path:
curl --resolve domain:443:<public-ip> https://domain/...to force the real external IP- Or test from an external machine (phone on cellular, a cloud VM, etc.)
Applies to reverse-proxy routing bugs, HTTP/2 SAN mismatches, and TLS configuration that differs between internal and external ingress.
Trace multi-hop DNS chains end-to-end, not just the endpoints. For split-horizon / VPN DNS that flows through several resolvers (e.g. VPN MagicDNS → local forwarder → authoritative server), a record existing in the authoritative server does NOT mean a client resolves it. Each hop can drop the query: a missing conditional-forward rule, a resolver pushed to clients that doesn't know a downstream zone, or a forwarder pointed at a target with no authoritative zone. Walk the full resolution path hop-by-hop (client → each forwarder → authoritative), querying each resolver directly (dig @<resolver> <name>), rather than only confirming the record exists at the source. Every individual server can look healthy while the chain is broken at one forwarding link.
Corollary — conditional forwarding needs an authoritative target. server=/zone/<ip> (dnsmasq) or any conditional-forward rule returns empty answers if the forward target has no real zone for that domain — e.g. a DHCP-only resolver that answers bare hostnames but has no SOA. The forward is syntactically valid but there is nothing authoritative to forward to; work around it by ingesting the records another way (poll the source API, write a hosts file).
When Something Doesn't Sync/Apply
- Check resource exclusions in the GitOps controller immediately
- Check if the resource type requires special permissions or labels
- Check if ServerSideApply conflicts are preventing field changes
- Don't try workarounds before understanding the root cause
OIDC Integration Checklist
Before starting any OIDC integration, research:
- What format is the
subclaim (UUID? username?) - Which claims are in the ID token vs userinfo endpoint
- How the consumer matches RBAC identities (groups? email? username?)
Log-First Diagnosis
- CrashLoopBackOff: check logs first. Error messages in pod logs usually point directly to the fix. Don't tweak configuration or security contexts blindly —
kubectl logs <pod>first. - Discriminate transient from persistent errors. CSI lock contention, etcd timeouts during first install, and brief connectivity blips are self-healing. Don't spend time debugging errors that resolve on retry. If you see retry/backoff patterns in logs, wait before intervening.
- Trust controller retry logic. CSI controllers, operators, and reconciliation loops have built-in retry. Transient failures during rapid provisioning are expected, not bugs.
- Framework-sanitised error bodies are unreliable to assert on. Security-conscious frameworks strip informative detail from response bodies (FastAPI's
RequestValidationErrorhandler returns{"detail": "Request validation failed"}regardless of the actual cause; GraphQLformatErrorcan sanitise messages). Don't write tests or downstream parsers that depend on the stripped body — only status code is reliable. The informative version is usually written to server-side logs; check there, not the wire response.
Reproduce Before Fixing
When a bug is discovered or reported, do not start by trying to fix it. The first step is always to write a test that reproduces the failure:
- Write a failing test. Capture the bug as a test case that demonstrates the broken behaviour. This forces you to understand the bug precisely — what input triggers it, what the wrong output is, and what the correct output should be.
- Fix the bug in isolation. Use a subagent or a separate session to write the fix. The fixing agent gets the failing test as its success criterion — it's done when the test passes. This separation prevents the fixer from unconsciously weakening the test to match a broken implementation.
- The test stays forever. The reproduction test becomes a permanent regression test. It proves the fix works and prevents the bug from returning.
This workflow has several advantages:
- Forces precise understanding. Writing a test means you know exactly what's broken, not just "it doesn't work."
- Prevents partial fixes. The test defines "done" objectively — the fix either passes or it doesn't.
- Parallelises work. While one agent fixes the bug, you can continue other work.
- Catches regressions. The test remains in the suite, guarding against the same class of failure.
# Step 1: Write the failing test FIRST
def test_regression_issue_427_empty_payload_crashes():
"""Bug #427: Empty payload causes unhandled TypeError in dispatcher.
Should return a 400 validation error, not crash."""
response = client.post("/dispatch", json={})
assert response.status_code == 400 # Currently crashes with 500
# Step 2: Hand to a subagent/session: "Make this test pass without breaking others"
GIT_SSH_COMMAND Only Affects Git-Invoked SSH
GIT_SSH_COMMAND (e.g., ssh -o StrictHostKeyChecking=no) only applies when git invokes SSH internally (clone, push, fetch). Direct ssh calls — such as ssh -T git@host for connectivity testing — ignore it entirely. When working in containers or CI environments where host keys aren't pre-trusted, direct SSH commands need explicit flags: ssh -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null.
Pattern Mining Before Authoring
Before building a new service, component, or script, read existing patterns in the codebase first. This matches conventions on the first attempt and avoids rework on naming, structure, and integration points. Applies to K8s manifests, CI pipelines, skill authoring, and script structure.
Grep Your Own Docs
Known issues documented in CLAUDE.md or MEMORY.md but not applied to new scripts/configs waste debugging time. Search your own documentation before writing automation that touches areas with known gotchas.
Read the Spec Before Proposing a Workaround
When a mid-implementation design question arises in a subsystem that already has a written spec, read the spec before proposing a bridge hack. The correct design is often already documented. One session burned hours considering Phase 1 bearer-token bypasses before realising the spec already defined the Phase 2 design (bootstrap tokens + mTLS).
Rule: specs exist to prevent this — grep spec/ or re-read the relevant spec file before inventing a workaround.
Budget Infrastructure-Recovery Time After Disruptions
After any disruption (power cut, network outage, cluster reboot, registry migration), start the next session with an infrastructure health check before planning feature work. Snap confinement quirks, stuck Terminating pods, read-only filesystems, and unreachable Git remotes each consume meaningful time to diagnose. Budget recovery as an explicit first phase rather than discovering it mid-task.
Parallel Research Agents for Broad Topic Coverage
When researching a topic with multiple independent facets, dispatch parallel research agents (e.g., one per sub-topic or source type) rather than sequential queries. Scope each agent narrowly (e.g., jurisdiction, domain filter, doc set) to reduce noise and improve signal. Cheap when facets are independent; poor fit when later queries depend on earlier results.
API Token Scope Errors
When an API endpoint returns a permission/scope error, read the error response body before guessing. Many APIs (Gitea, GitHub, GitLab) explicitly state the required scope in the error message (e.g., required=[write:admin]). This is faster and more reliable than consulting documentation or iterating one scope at a time.
Minimal Container Images Have No Debug Tools
Distroless and single-binary containers (Garage, distroless Go images, etc.) have no shell, curl, wget, or other debug tools. kubectl exec commands will fail.
For HTTP checks: Use kubectl port-forward svc/<name> <local-port>:<svc-port> and run curl locally.
For verification scripts: Don't assume exec-based checks will work. Design health checks around port-forward + local tools, or use Kubernetes-native probes.
Use Container Logs to Narrow the Failure Boundary
When a service appears down, check its logs for successful requests from other clients before assuming a total outage. If some clients are connecting successfully (e.g., HTTP/3 but not TCP, or internal but not external), the issue is narrower than "service is down." This distinction dramatically reduces debugging scope.
Structured JSON Logging from Application Entry Points
Web frameworks like uvicorn don't configure application-level loggers — only access logs appear by default. Named loggers have no handler unless logging.basicConfig() is called explicitly. This makes application logs invisible in production (K8s, Docker) with no error — just silence. Always call logging.basicConfig() with a structured format (JSON) early in application startup, before any getLogger() calls.
Verify DB Schema Matches Application Models After Every Deployment
After deploying a new version of an application that uses an ORM or schema migration tool, verify that the live database schema matches what the application expects. Common failure mode: a migration ran in dev/staging but not in production, or a new field was added to a model without a corresponding migration.
Quick check: run the application's schema validation command, or compare alembic current vs alembic head, or run SELECT column_name FROM information_schema.columns WHERE table_name='<table>' and diff against the model definition.
When to check: after every deployment that touches models or migrations — not just on explicit migration commits. An ORM auto-create (e.g., SQLAlchemy create_all) can silently succeed while leaving optional columns missing, causing subtle bugs rather than hard crashes.
Lost-Webhook Zombie Pattern
Distributed task/job systems that depend on a webhook or callback to advance state can leave records stuck in running indefinitely when the callback is lost. The state record looks active; nothing is actually happening.
Pattern signature:
state=runningfor >20 min with no progress- Claim/lease set, but
current_load=0on the worker - Zero log progress since the claim event
- The work artifact (output branch, file, queue entry) exists or is partially populated
Diagnosis rules:
- Don't trust the orchestrator's state record — verify side effects directly. Check whether the work-artifact branch was pushed, the output file written, the queue entry produced.
- Don't expect the system's DELETE endpoint to recover — most refuse to terminate non-terminal records. You will need to manually transition the state or wait for a timeout that may never come.
- Cross-reference dispatcher/worker logs (if still alive) for the claim event and any subsequent webhook attempt. A missing "callback succeeded" log entry confirms the lost-webhook hypothesis.
Prevention: make the orchestrator poll for terminal artifacts in addition to listening for webhooks. The webhook is an optimisation; the artifact poll is the source of truth. Generalises to CI pipelines, async job queues, agent runtime systems — anything where worker completion depends on an out-of-band callback.
Detection Before Auto-Remediation
For any recurring failure mode (storage hotplug failures, network policy drops, stale leases, queue zombies, rate-limit hits), ship a detection signal — Prometheus alert, log-pattern check, scheduled audit — BEFORE building any auto-remediation.
Rationale: auto-remediation has its own failure modes (drain timeouts, PodDisruptionBudget conflicts, recursive failures, race conditions with the underlying bug). Shipping it without observability hides those failures; shipping observability first lets you measure how often the bug occurs and how often a manual fix succeeds before you trust automation with the same fix.
Order of operations:
- Detect — alert with sufficient context to recover manually (resource name, pod, host, last-known state)
- Document the manual fix — capture the working recovery sequence as a runbook or script the alert links to
- Build automation behind a feature flag — auto-remediation in shadow mode (log what it would do; don't act)
- Compare shadow decisions against operator actions — if they agree at high rate, flip the flag
Detection alone turns multi-hour incidents into ~15-minute ones at near-zero risk. Automation built without the prior alert layer is unauditable.
Distinguish Mitigations from Cures in Writing
When a recurring bug is reduced but not eliminated by a fix, explicitly label the fix as a "mitigation" in the gotcha file and track the recurrence vector. The temptation to mark "fixed" leads to surprise when the bug returns and to wasted effort re-debugging from scratch.
Conventions:
- Use the phrase "mitigations, NOT a full fix" in the gotcha file's title or first paragraph when the underlying root cause persists.
- Link to the long-term replacement plan (e.g., a FUTURE.md item or upstream issue).
- Record the recurrence vector: under what conditions does the mitigated bug come back? "Recurs when load > X", "recurs on host rebuild", "recurs after Y days".
- When the bug recurs, append the new incident to the same gotcha entry — don't open a new one. The history of recurrence is the evidence that the fix is a mitigation, not a cure.
Generalises across all projects that track incidents in memory/gotchas-*.md. Prevents future sessions (and future Claude instances) from misreading a mitigation as a cure.
Hypervisor- or Platform-Level Diagnostics Before Application-Level Blame
When a symptom appears at the application layer (pod stuck ContainerCreating, device missing, mount failing, port unreachable) but the platform underneath has its own lifecycle model, check the platform's task/event API first.
Concrete examples:
- Pod stuck attaching a volume →
qm pending <vmid>andpvesh get /nodes/<host>/tasks --vmid <id>reveal Proxmox hotplug failures invisible fromqm config - AWS PV not attaching → EC2 attachment-state events reveal failures invisible from
kubectl - Systemd unit failed →
journalctl -ureveals failures invisible frompsor service-level health checks - VM unresponsive → hypervisor console output / serial-line buffer reveals kernel panics invisible from inside the guest
Rule: five seconds on the platform API beats fifteen minutes debugging the wrong layer. Build a habit of "check one layer down" before forming a hypothesis at the application layer. Applies to any layered stack: K8s-on-hypervisor, container-on-host, application-on-systemd, agent-on-orchestrator.
Test From Outside the Broken Thing
When something is unreachable, test from a known-good external vantage point — different host, cellular network, cloud VM, curl --resolve with the public IP — before assuming the local machine is at fault.
Symptoms this catches:
- "No route to host" from your laptop while remote users can reach the service fine (your VPN dropped)
- "DNS gives public IPs but I expected internal" (your
/etc/hostsor split-horizon DNS is bypassed) - "Latency spike" that's actually a single ISP-side route flap (test from a different ISP)
The inverse trap: when split-horizon DNS or local /etc/hosts overrides bypass the production path, your local "it works" tells you nothing about what real users see. Always verify production behaviour through the actual public path (curl --resolve domain:443:<public-ip> https://domain/..., or test from a phone on cellular / a cloud VM in a different region).
Cross-cutting diagnostic discipline applicable to any distributed system, web-facing service, or VPN/proxy stack.
Verify User-Reported Identifiers Before Recovery
When a user reports a problem by name ("the X pod is stuck", "the Y volume is broken", "the Z VM won't start"), confirm the actual identifiers with kubectl get, qm config, docker inspect, pvesh ls, or the equivalent on whichever platform owns the resource — BEFORE running any recovery command.
Why: humans under pressure confuse names (VM IDs, node hostnames, PVC names, container names). A 30-second verification prevents executing the wrong recovery on the wrong resource. Recovery commands are often destructive (qm reset, kubectl delete pod --force, volume detach); running one on a healthy resource turns a small incident into a bigger one.
Verification pattern:
- Restate the user's claim: "You said pod
foois stuck on nodebar." - Verify each identifier exists and is in the claimed state:
kubectl get pod foo -o wide(does it exist? is it onbar? is it actually stuck?) - Confirm with the user before destructive action: "Pod
fooonbaris inContainerCreating. Runningkubectl delete pod foo --force --grace-period=0. OK?"
General operational-safety rule, applicable in any high-pressure recovery situation.
Discover API Endpoints via Swagger/OpenAPI Before Re-Reading Prose Docs
When an API call returns 404 or "endpoint not found", fetch the actual paths from the live spec (/swagger.v1.json, /openapi.json, /api-docs) with curl + jq before consulting prose documentation. The spec is generated from the running code; the prose drifts. Five seconds with curl beats fifteen minutes of doc spelunking. See API Integration section 6 for the full pattern.
Grep Minified Third-Party Source to Learn Its Runtime Protocol
When you must understand how a minified third-party library behaves at runtime (popup/auth detection, postMessage protocol, event names, expected window.name), grep the minified bundle for concrete string/pattern signatures (grep -oP 'window\.opener[^;]*', event-name literals, magic constants) instead of guessing or reading upstream docs that may not match the shipped build. Minified code still contains the literal strings and property accesses that reveal the protocol reliably — enough to integrate against it without modifying its source.