distill: 48 cross-project best-practices from 2026-07 reflection sweep

Promotions from reflecting 21 projects' session logs (incl. agent-runtimes
122-log drain). Adds coverage across networking (eBPF VIP/VPN SNAT/VLAN
bridge/forward-auth preflight/ingress TLS), kubernetes (CSI hotplug/PodSecurity
debug/self-managed GitOps/runtime annotations), CI (dispatch tokens/runner
death/base image), git (CI-rebase/shallow reset/PR governance), python (async
session pool/httpx redirects/logging), TDD (AsyncMock/xfail lifecycle),
api-integration (SDK parse/token-scope 404/schema probing), plus docker,
scripting, debugging, security-architecture, secrets, react, octopus.

State: .distill-state.json refreshed with current HEADs + 5 newly-tracked projects.
This commit is contained in:
Paul O'Reilly
2026-07-02 15:57:42 +12:00
parent 5e67cbcfbb
commit 7e348f5ee3
16 changed files with 577 additions and 57 deletions

View File

@@ -1,101 +1,126 @@
{
"version": 1,
"last_run": "2026-04-19T11:13:09.870821Z",
"last_run": "2026-07-02T03:24:50Z",
"projects": {
"agent-runtimes": {
"path": "/home/paul/dev/claude/projects/agent-runtimes",
"last_sha": "f0a5c22477fe3ed9d97b95f9fecb02ac6671d170",
"last_run": "2026-04-19T11:13:09.870821Z"
"last_sha": "47d15bbbe17cbd28d4cbeb47083184bc58df40bd",
"last_run": "2026-07-02T03:24:50Z"
},
"claude-foundations": {
"path": "/home/paul/dev/claude/projects/claude-foundations",
"last_sha": "6dfa20c47ccd293d6c5de99524ed474f8f3ceef8",
"last_run": "2026-04-19T11:13:09.870821Z"
},
"cluster-apps/octopus-deploy": {
"path": "/home/paul/dev/claude/projects/cluster-apps/octopus-deploy",
"last_sha": "3f23b8a2c3d4a120599f5a38a654e4af2419672f",
"last_run": "2026-04-19T11:13:09.870821Z"
},
"cluster-bootstrap": {
"path": "/home/paul/dev/claude/projects/cluster-bootstrap",
"last_sha": "9536d0880c308310b71d5e6b41c00bd0a3ef73c3",
"last_run": "2026-04-19T11:13:09.870821Z"
},
"custom-claude-skills": {
"path": "/home/paul/dev/claude/projects/custom-claude-skills",
"last_sha": "7d7856a07971958c3e0b98720edeb685bac68678",
"last_run": "2026-04-19T11:13:09.870821Z"
},
"hugo-accelerator": {
"path": "/home/paul/dev/claude/projects/hugo-accelerator",
"last_sha": "206ab9bc11f60305beede2a2522dbcf794b6306b",
"last_run": "2026-04-19T11:13:09.870821Z"
},
"small-scripts": {
"path": "/home/paul/dev/claude/projects/small-scripts",
"last_sha": "5a1b4b14bc13fbc0a626f3450b913b3f788e4de4",
"last_run": "2026-04-19T11:13:09.870821Z"
"agent-runtimes-deploy": {
"path": "/home/paul/dev/claude/projects/agent-runtimes-deploy",
"last_sha": "d11ec8196d1a57c68df2f28d9cf3a5a4579018a7",
"last_run": "2026-07-02T03:24:50Z"
},
"ai-image-gen": {
"path": "/home/paul/dev/claude/projects/ai-image-gen",
"last_sha": "89c8605470e0188a5ad3765e6ce3b055adefba67",
"last_run": "2026-04-19T11:13:09.870821Z"
"last_run": "2026-07-02T03:24:50Z"
},
"brainiac-app": {
"path": "/home/paul/dev/claude/projects/brainiac-app",
"last_sha": "7ff7a4ce605275fc7bf8b9fc664604af3e55e381",
"last_run": "2026-04-19T11:13:09.870821Z"
"last_run": "2026-07-02T03:24:50Z"
},
"claude-foundations": {
"path": "/home/paul/dev/claude/projects/claude-foundations",
"last_sha": "d13beca5d7320b431111a3781012c14623a0a1da",
"last_run": "2026-07-02T03:24:50Z"
},
"cluster-apps/agent-runtimes": {
"path": "/home/paul/dev/claude/projects/cluster-apps/agent-runtimes",
"last_sha": "bf2413512576fb771ec9670ba6834faabacd8ce4",
"last_run": "2026-07-02T03:24:50Z"
},
"cluster-apps/octopus-deploy": {
"path": "/home/paul/dev/claude/projects/cluster-apps/octopus-deploy",
"last_sha": "3f23b8a2c3d4a120599f5a38a654e4af2419672f",
"last_run": "2026-07-02T03:24:50Z"
},
"cluster-bootstrap": {
"path": "/home/paul/dev/claude/projects/cluster-bootstrap",
"last_sha": "e54ddb121dd7a07da97fd1239c4f529ccbe028c2",
"last_run": "2026-07-02T03:24:50Z"
},
"contracts": {
"path": "/home/paul/dev/claude/projects/contracts",
"last_sha": "d56c64f6ae95224c5ae68284ff7d3c21ede7a0c2",
"last_run": "2026-04-19T11:13:09.870821Z"
"last_run": "2026-07-02T03:24:50Z"
},
"crud-accelerator": {
"path": "/home/paul/dev/claude/projects/crud-accelerator",
"last_sha": "b61b64b723ace6f9c7e6f4be94766955625cf90c",
"last_run": "2026-04-19T11:13:09.870821Z"
"last_sha": "98468f61ab214017db26654d7611952494a86c0d",
"last_run": "2026-07-02T03:24:50Z"
},
"custom-claude-skills": {
"path": "/home/paul/dev/claude/projects/custom-claude-skills",
"last_sha": "ca77d3873aecaade8fbd6343c112d77d4746b04f",
"last_run": "2026-07-02T03:24:50Z"
},
"dns-manager": {
"path": "/home/paul/dev/claude/projects/dns-manager",
"last_sha": "f4ce7850a338e9ea35aebf64a1fd6461f955f011",
"last_run": "2026-04-19T11:13:09.870821Z"
"last_sha": "605fba676dd471737f58590f8b2d4d6cf88b9dd5",
"last_run": "2026-07-02T03:24:50Z"
},
"hugo-accelerator": {
"path": "/home/paul/dev/claude/projects/hugo-accelerator",
"last_sha": "ae059fdd175d778f9dc7760843a5136f1d48a94d",
"last_run": "2026-07-02T03:24:50Z"
},
"hugo-gabby-oreilly-counselling-content": {
"path": "/home/paul/dev/claude/projects/hugo-gabby-oreilly-counselling-content",
"last_sha": "baf97e53d2b7756ae974288f012580c9e0cbfb2e",
"last_run": "2026-07-02T03:24:50Z"
},
"hugo-oreillyconsulting-integration": {
"path": "/home/paul/dev/claude/projects/hugo-oreillyconsulting-integration",
"last_sha": "877032aa839661d359ef382f011ac405c28aa55e",
"last_run": "2026-07-02T03:24:50Z"
},
"small-scripts": {
"path": "/home/paul/dev/claude/projects/small-scripts",
"last_sha": "b76e56f7502ae0dabec670516ca856ff19ebda33",
"last_run": "2026-07-02T03:24:50Z"
},
"octopus/ai-assisted-migration": {
"path": "/home/paul/dev/claude/octopus/ai-assisted-migration",
"last_sha": "761b7cb744a6eba575437a08d19844a7ceb0e5ab",
"last_run": "2026-04-19T11:13:09.870821Z"
"last_run": "2026-07-02T03:24:50Z"
},
"octopus/customer-issue-sync": {
"path": "/home/paul/dev/claude/octopus/customer-issue-sync",
"last_sha": "ffeb6ea2409b35bbfcdf71db2710181ea44012a8",
"last_run": "2026-04-19T11:13:09.870821Z"
"last_sha": "0a1c4b6405a9b02a3d26add665ccb039a3f6f694",
"last_run": "2026-07-02T03:24:50Z"
},
"octopus/customers/duel-image": {
"path": "/home/paul/dev/claude/octopus/customers/duel-image",
"last_sha": "176c129af07ffaae4392ca9b82fa6c28c5be0909",
"last_run": "2026-04-19T11:13:09.870821Z"
},
"octopus/customers/slb": {
"path": "/home/paul/dev/claude/octopus/customers/slb",
"last_sha": "bf4bf160a492ce6ea8880f9564a3b4d71e2fe2f5",
"last_run": "2026-04-19T11:13:09.870821Z"
"octopus/duel-image": {
"path": "/home/paul/dev/claude/octopus/duel-image",
"last_sha": "7b2261a2e63ca74167ce7373fe0ae266205638c1",
"last_run": "2026-07-02T03:24:50Z"
},
"octopus/goes": {
"path": "/home/paul/dev/claude/octopus/goes",
"last_sha": "b39e9c77f9e214bdb71f47543ad5eb821ae31c32",
"last_run": "2026-04-19T11:13:09.870821Z"
"last_sha": "fb49ce95b8fd404ca792eb3c1fdf3079cb223bff",
"last_run": "2026-07-02T03:24:50Z"
},
"octopus/PlatformHub-Demo": {
"path": "/home/paul/dev/claude/octopus/PlatformHub-Demo",
"last_sha": "d8d039180b424ed0b0e56bc400228d2a51899804",
"last_run": "2026-04-19T11:13:09.870821Z"
"last_sha": "30ba9aea3b6245f5963ab44661893da0c903eb67",
"last_run": "2026-07-02T03:24:50Z"
},
"octopus/policy-demo": {
"path": "/home/paul/dev/claude/octopus/policy-demo",
"last_sha": "8f652e13ee42a06a3f99c173b029f4a124ee3002",
"last_run": "2026-07-02T03:24:50Z"
},
"octopus/slb": {
"path": "/home/paul/dev/claude/octopus/slb",
"last_sha": "f9e1adbd7410616cfb653fac798d209a34dd9e34",
"last_run": "2026-07-02T03:24:50Z"
},
"octopus/the-case-for-agentic": {
"path": "/home/paul/dev/claude/octopus/the-case-for-agentic",
"last_sha": "69f3ca10de9994ed5386bd76f44fad7b1c23e8c4",
"last_run": "2026-04-19T11:13:09.870821Z"
"last_run": "2026-07-02T03:24:50Z"
}
}
}
}

View File

@@ -129,6 +129,52 @@ Cross-references: [API Design](api-design.md) covers server-side API design (the
---
## 6. Discover API Endpoints via Swagger/OpenAPI Before Re-Reading Prose Docs
**Principle:** When an API call returns 404 or "endpoint not found", fetch the actual paths from the live spec (`/swagger.v1.json`, `/openapi.json`, `/api-docs`, `/.well-known/openapi`) with `curl + jq` before consulting documentation.
**Why it matters:** Documentation prose drifts from the live API faster than the OpenAPI spec does. The spec is generated from the running code; the prose is hand-maintained. Five seconds with curl beats fifteen minutes of doc spelunking and beats minutes of guessing at path variants.
**How to implement:**
- Grep the spec for the resource/verb: `curl -s https://api.example.com/swagger.v1.json | jq '.paths | keys[]' | grep -i <resource>`
- For verb-specific lookups: `jq '.paths | to_entries[] | select(.value.post) | .key'`
- For required body fields: `jq '.components.schemas.<TypeName>.required'`
- If the API requires auth even for the spec endpoint, fetch via your existing credential — the spec is not sensitive.
**Anti-patterns:**
- Reading the docs page-by-page when a 30-character `jq` query finds the answer.
- Trying URL variants (`/foo`, `/foos`, `/foo/v1`, `/v1/foo`) without checking the spec first.
- Trusting a path from training data that returns 404 instead of consulting the live spec.
Generalises across any REST API consumer — Gitea, GitHub, Octopus, Kubernetes apiserver, cloud provider APIs. The 4xx response is a strong signal that you have the wrong path; treat it as a prompt to fetch the spec, not as a prompt to guess again.
## 7. Generated OpenAPI SDKs: Bypass `_parse_response` for Untyped/Error Bodies
Clients generated by `openapi-python-client` (and similar generators) only handle the status codes and response shapes that were in the spec at generation time. Two silent failure modes result:
- A valid RFC-9457 (422) problem body raises `ValueError: Unexpected status code` inside the generated `_parse_response` — the caller sees it as a 502, not a structured 422.
- An endpoint whose 200 body wasn't captured in the schema (untyped/streaming/`Any`) returns `.parsed is None` on success — no error, the value is just absent.
Fix: for any call where you need to handle 4xx/5xx bodies as data, or where the 200 body isn't in the generated model, bypass parsing. Use the generated `_get_kwargs` to build the request, then call the raw httpx client and inspect `.status_code`/`.content` yourself: `sdk_fn._get_kwargs(**kwargs)``sdk_client.get_httpx_client().request(**req_kwargs)`. Do NOT use `*_detailed()` / `sync_detailed()` for those endpoints — they route through `_parse_response` and will raise or silently return `None`.
## 8. A 404 on a Known-Good Endpoint May Be a Token-Scope Miss, Not a Missing Resource
Some APIs (Gitea, and others that avoid confirming resource existence to unauthorized callers) return **404 rather than 403** when the token lacks the required scope for an operation. A `workflow_dispatch` (or any write) call to an endpoint you *know* exists, returning 404, is a strong signal that the token is missing a scope (e.g. `write:actions`), not that the path is wrong.
- Before re-checking the path or the API version, check the token's scopes against what the operation needs.
- Keep a scoped-token map: know which token has which scopes, and switch to the correctly-scoped credential rather than debugging the URL.
This is a distinct 404 root cause from section 6 (wrong path) — when the path is known-good, suspect scope before re-fetching the spec.
## 9. Discover Undocumented Schemas by Probing With Minimal Writes
When a resource's required fields and accepted enum values aren't documented, send deliberately minimal `POST`/`PUT` requests and read the validation errors — they enumerate required fields and reveal which enum values are accepted vs rejected. Create-then-delete a throwaway resource to confirm the full response shape, then delete it so no residue remains. Always confirm the running version first (e.g. `GET /api` for version) rather than trusting a schema from memory or training data, since schemas drift between versions. This complements OpenAPI discovery (section 6): use the spec when one exists; fall back to error-driven probing when the resource is under-specified or the spec omits body requirements.
## 10. Check the Resource/Link Map + Feature-Toggles Before Assuming a Feature Exists
Before concluding an API offers a given primitive (a "policies" endpoint, a "required steps" feature, a webhook family), inspect the service's root/link map or feature-toggle list rather than guessing at paths. Absence of the link means the primitive isn't exposed. When a documented feature seems missing from the live API, check the feature-toggle endpoint (e.g. `/api/configuration/feature-toggles`) before concluding it's absent — toggles frequently gate visibility of an entire endpoint family.
---
## Summary
Integrating with a third-party API is an exercise in compensating for its limitations. Verify capabilities before designing, build application-layer compensation for missing features, use git + SOPS as a state store when you have no persistent compute, and always declare a system of record for bidirectional syncs. Idempotency and observability are non-negotiable.

View File

@@ -169,3 +169,78 @@ When separating image tiers by registry (e.g., `org-nonprod/app` on push-to-main
- Give CI a **service account that is a member of both registry orgs**, and store one credential per registry (do not share a single token across tiers).
- Keep tag formats **distinct per tier** (e.g., `sha-<8>` for non-prod, `prod-sha-<8>` for prod) so downstream systems — Octopus channels, Kustomize overlays, audit tooling — can reason about provenance from the tag alone.
- Gate the prod publish workflow on `v*` tags or an explicit release event, never on main-branch pushes.
## Codegen-Freshness Scripts Must Be Observation-Only and Sandboxed
CI scripts that regenerate generated code from manifests and compare it to committed output (e.g., "the OpenAPI client matches the spec", "the generated K8s manifests match the Helm chart") must obey two rules:
1. **Regenerate into a per-job temp directory** — use `${RUNNER_TEMP}` or `mktemp -d` inside the job's workspace. **Never `/tmp`** on shared multi-tenant runners — other jobs can plant `conftest.py`, `package.json` postinstall hooks, or shell rc files that exfiltrate from this job.
2. **Compare with `diff -ru` only.** No `pip install`, `pytest`, `npm ci`, `make`, or any other command that executes code from the regenerated tree. A malicious manifest can plant code that runs during install/test/build and exfiltrates secrets or pivots through the runner.
```yaml
# Pattern: regenerate into sandbox, diff against committed copy, never execute
- name: Check codegen freshness
run: |
SANDBOX=$(mktemp -d)
scripts/codegen.sh --output "$SANDBOX/generated"
diff -ru generated/ "$SANDBOX/generated/" || {
echo "Generated output is stale. Run scripts/codegen.sh and commit."
exit 1
}
```
Applies to any "verify the generated artifact matches the source" CI step: OpenAPI codegen, Helm chart generators, Terraform-from-manifest, protobuf compilation, GraphQL schema diff. The threat model is: a PR author plants a malicious manifest that, when regenerated, contains code that runs at test or install time.
## Substring-Matching Security Tests Trip on User-Facing Comments
When a CI-policy validator scans workflow or source files for forbidden binary names (`pip install`, `npm ci`, `pytest`, `curl http://`), it must distinguish prohibited **invocations** from echo'd error messages, code comments, and synonyms in user-facing strings.
**Common failure mode:** the validator forbids `pip install` in workflows, but the workflow's failure-message handler contains the exact phrase as part of an error string:
```yaml
- name: Codegen check
run: |
diff -ru ... || {
echo "Stale output. Run pip install . then rerun codegen." # ← matches the forbidden pattern
exit 1
}
```
The validator fires on its own remediation hint. **Fixes:**
- Phrase user-facing text without naming the binaries literally: "package-install", "test runner", "build entry point".
- Or use a multi-line form that breaks substring boundaries: `echo "Run: $(echo pip)" install`.
- Or whitelist explicitly: `if [[ "$line" =~ pip[[:space:]]+install ]] && [[ "$line" != *"#"* ]] && [[ "$line" != *"echo"* ]]`.
Same trap applies to any policy enforcement that matches code patterns in source files: pre-commit hooks scanning for secrets, lint rules forbidding API calls, banned-import walkers. Substring matchers must distinguish active invocations from quoted strings, comments, and documentation. A real AST/lexer-based check is harder to write but immune to the trap.
## `if: always()` Steps Don't Survive Early Runner-Process Death
A dispatch/notify step gated with `if: always()` is meant to run even when an earlier job step fails. But `always()` only runs the step while the **job process is still alive**. If the runner itself dies early (act-runner/agent crashes during checkout, pip, or setup before reaching the step), the whole job is torn down and `always()` steps never execute — so a downstream build dispatch silently never fires.
- Symptom: intermittent — the same workflow fires the downstream build on most pushes but not all, with no error in the failed run's log.
- Root cause is a runner/env flake, not a workflow bug; not reliably reproducible.
- Do not treat `always()` as a guaranteed "run no matter what" — it is scoped to a living job.
- Workaround / unblock: dispatch the downstream build manually via `workflow_dispatch`.
- For a hard guarantee, trigger the downstream from an event the runner can't swallow (a separate scheduled/webhook-driven job, or the platform's native workflow-completed event), not from an in-job step.
## Don't Trust the Job-Listing API's `run_id` Filter to Map Jobs to Runs
On some CI backends (observed on Gitea Actions 1.25.x), the job/task listing API filtered by `?run_id=N` returns job entries belonging to *other* runs — you cannot reliably map a job back to its run through that filter. When correlating a job to a specific run programmatically, identify jobs by descending job ID or by timestamp, not by trusting the `run_id` filter.
## A Bare Language Base Image Breaks Node-Based CI Actions
Setting `container: image: python:3-slim` (or any minimal single-language base) as the job container removes tooling that CI actions assume is present. Node-based actions like `actions/checkout@v4` fail with "executable file not found" because there's no Node.js runtime, and steps using `curl` (e.g. API dispatch calls) break because curl isn't installed either.
Fix: don't override the job container just to get a language runtime — the standard `ubuntu-latest` runner already ships Node.js, curl, and python3. Install extra language deps as a step (`pip install ...`) rather than swapping the whole container image.
## Create CI Secrets Before Pushing Code That Uses Them
A CI secret referenced by a workflow (`DISPATCH_TOKEN`, `ARGOCD_TOKEN`, a deploy PAT) must exist on the repo/org **before** you push the code that depends on it. The first run after the push otherwise executes against a missing secret — and many CI systems treat a missing secret as an empty string rather than an error, so the step **fails silently** (empty auth header, `workflow_dispatch` that no-ops). Create the secret via the platform API first, then push. Verify the secret exists rather than assuming the push "must have" picked it up.
## Verify Which Branch the Deploy Controller Actually Tracks
When porting a CI workflow from a template, confirm which branch the GitOps controller (ArgoCD/Flux) tracks before wiring the image-tag write-back. Templates assume `main`, but real repos often track `staging` or a release branch. If the build writes the new image tag to `deployment.yaml` on `main` while the app tracks `staging`, the deploy succeeds, CI goes green, and the pod stays on the old image — a silent no-op. Read/write the tag on the tracked branch (`?ref=staging`, `"branch":"staging"`).
## Semver-Style Release Tools Skip Non-`main` Release Branches
Tools that gate release-version emission on "the main branch" (semver-ci and similar) output an **empty version** when run on another release branch (e.g., `staging`), producing an invalid empty image tag. When you deliberately cut releases from a non-`main` branch, tell the tool that branch is a release branch (e.g., `--main-branch staging`), or it silently skips and downstream tagging breaks.

View File

@@ -31,6 +31,10 @@ When `/etc/hosts` or internal DNS points a public hostname at an internal IP, lo
Applies to reverse-proxy routing bugs, HTTP/2 SAN mismatches, and TLS configuration that differs between internal and external ingress.
**Trace multi-hop DNS chains end-to-end, not just the endpoints.** For split-horizon / VPN DNS that flows through several resolvers (e.g. VPN MagicDNS → local forwarder → authoritative server), a record existing in the authoritative server does NOT mean a client resolves it. Each hop can drop the query: a missing conditional-forward rule, a resolver pushed to clients that doesn't know a downstream zone, or a forwarder pointed at a target with no authoritative zone. Walk the full resolution path hop-by-hop (client → each forwarder → authoritative), querying each resolver directly (`dig @<resolver> <name>`), rather than only confirming the record exists at the source. Every individual server can look healthy while the chain is broken at one forwarding link.
**Corollary — conditional forwarding needs an authoritative target.** `server=/zone/<ip>` (dnsmasq) or any conditional-forward rule returns empty answers if the forward target has no real zone for that domain — e.g. a DHCP-only resolver that answers bare hostnames but has no SOA. The forward is syntactically valid but there is nothing authoritative to forward to; work around it by ingesting the records another way (poll the source API, write a hosts file).
## When Something Doesn't Sync/Apply
- Check resource exclusions in the GitOps controller immediately
@@ -50,6 +54,7 @@ Before starting any OIDC integration, research:
- **CrashLoopBackOff: check logs first.** Error messages in pod logs usually point directly to the fix. Don't tweak configuration or security contexts blindly — `kubectl logs <pod>` first.
- **Discriminate transient from persistent errors.** CSI lock contention, etcd timeouts during first install, and brief connectivity blips are self-healing. Don't spend time debugging errors that resolve on retry. If you see retry/backoff patterns in logs, wait before intervening.
- **Trust controller retry logic.** CSI controllers, operators, and reconciliation loops have built-in retry. Transient failures during rapid provisioning are expected, not bugs.
- **Framework-sanitised error bodies are unreliable to assert on.** Security-conscious frameworks strip informative detail from response bodies (FastAPI's `RequestValidationError` handler returns `{"detail": "Request validation failed"}` regardless of the actual cause; GraphQL `formatError` can sanitise messages). Don't write tests or downstream parsers that depend on the stripped body — only status code is reliable. The informative version is usually written to server-side logs; check there, not the wire response.
## Reproduce Before Fixing
@@ -128,3 +133,92 @@ After deploying a new version of an application that uses an ORM or schema migra
**Quick check:** run the application's schema validation command, or compare `alembic current` vs `alembic head`, or run `SELECT column_name FROM information_schema.columns WHERE table_name='<table>'` and diff against the model definition.
**When to check:** after every deployment that touches models or migrations — not just on explicit migration commits. An ORM auto-create (e.g., SQLAlchemy `create_all`) can silently succeed while leaving optional columns missing, causing subtle bugs rather than hard crashes.
## Lost-Webhook Zombie Pattern
Distributed task/job systems that depend on a webhook or callback to advance state can leave records stuck in `running` indefinitely when the callback is lost. The state record looks active; nothing is actually happening.
**Pattern signature:**
- `state=running` for >20 min with no progress
- Claim/lease set, but `current_load=0` on the worker
- Zero log progress since the claim event
- The work artifact (output branch, file, queue entry) exists or is partially populated
**Diagnosis rules:**
- Don't trust the orchestrator's state record — verify side effects directly. Check whether the work-artifact branch was pushed, the output file written, the queue entry produced.
- Don't expect the system's DELETE endpoint to recover — most refuse to terminate non-terminal records. You will need to manually transition the state or wait for a timeout that may never come.
- Cross-reference dispatcher/worker logs (if still alive) for the claim event and any subsequent webhook attempt. A missing "callback succeeded" log entry confirms the lost-webhook hypothesis.
**Prevention:** make the orchestrator poll for terminal artifacts in addition to listening for webhooks. The webhook is an optimisation; the artifact poll is the source of truth. Generalises to CI pipelines, async job queues, agent runtime systems — anything where worker completion depends on an out-of-band callback.
## Detection Before Auto-Remediation
For any recurring failure mode (storage hotplug failures, network policy drops, stale leases, queue zombies, rate-limit hits), ship a detection signal — Prometheus alert, log-pattern check, scheduled audit — BEFORE building any auto-remediation.
**Rationale:** auto-remediation has its own failure modes (drain timeouts, PodDisruptionBudget conflicts, recursive failures, race conditions with the underlying bug). Shipping it without observability hides those failures; shipping observability first lets you measure how often the bug occurs and how often a manual fix succeeds before you trust automation with the same fix.
**Order of operations:**
1. **Detect** — alert with sufficient context to recover manually (resource name, pod, host, last-known state)
2. **Document the manual fix** — capture the working recovery sequence as a runbook or script the alert links to
3. **Build automation behind a feature flag** — auto-remediation in shadow mode (log what it would do; don't act)
4. **Compare shadow decisions against operator actions** — if they agree at high rate, flip the flag
Detection alone turns multi-hour incidents into ~15-minute ones at near-zero risk. Automation built without the prior alert layer is unauditable.
## Distinguish Mitigations from Cures in Writing
When a recurring bug is reduced but not eliminated by a fix, explicitly label the fix as a "mitigation" in the gotcha file and track the recurrence vector. The temptation to mark "fixed" leads to surprise when the bug returns and to wasted effort re-debugging from scratch.
**Conventions:**
- Use the phrase "mitigations, NOT a full fix" in the gotcha file's title or first paragraph when the underlying root cause persists.
- Link to the long-term replacement plan (e.g., a FUTURE.md item or upstream issue).
- Record the **recurrence vector**: under what conditions does the mitigated bug come back? "Recurs when load > X", "recurs on host rebuild", "recurs after Y days".
- When the bug recurs, append the new incident to the same gotcha entry — don't open a new one. The history of recurrence is the evidence that the fix is a mitigation, not a cure.
Generalises across all projects that track incidents in `memory/gotchas-*.md`. Prevents future sessions (and future Claude instances) from misreading a mitigation as a cure.
## Hypervisor- or Platform-Level Diagnostics Before Application-Level Blame
When a symptom appears at the application layer (pod stuck `ContainerCreating`, device missing, mount failing, port unreachable) but the platform underneath has its own lifecycle model, check the platform's task/event API first.
**Concrete examples:**
- Pod stuck attaching a volume → `qm pending <vmid>` and `pvesh get /nodes/<host>/tasks --vmid <id>` reveal Proxmox hotplug failures invisible from `qm config`
- AWS PV not attaching → EC2 attachment-state events reveal failures invisible from `kubectl`
- Systemd unit failed → `journalctl -u` reveals failures invisible from `ps` or service-level health checks
- VM unresponsive → hypervisor console output / serial-line buffer reveals kernel panics invisible from inside the guest
**Rule:** five seconds on the platform API beats fifteen minutes debugging the wrong layer. Build a habit of "check one layer down" before forming a hypothesis at the application layer. Applies to any layered stack: K8s-on-hypervisor, container-on-host, application-on-systemd, agent-on-orchestrator.
## Test From Outside the Broken Thing
When something is unreachable, test from a known-good external vantage point — different host, cellular network, cloud VM, `curl --resolve` with the public IP — before assuming the local machine is at fault.
**Symptoms this catches:**
- "No route to host" from your laptop while remote users can reach the service fine (your VPN dropped)
- "DNS gives public IPs but I expected internal" (your `/etc/hosts` or split-horizon DNS is bypassed)
- "Latency spike" that's actually a single ISP-side route flap (test from a different ISP)
**The inverse trap:** when split-horizon DNS or local `/etc/hosts` overrides bypass the production path, your local "it works" tells you nothing about what real users see. Always verify production behaviour through the actual public path (`curl --resolve domain:443:<public-ip> https://domain/...`, or test from a phone on cellular / a cloud VM in a different region).
Cross-cutting diagnostic discipline applicable to any distributed system, web-facing service, or VPN/proxy stack.
## Verify User-Reported Identifiers Before Recovery
When a user reports a problem by name ("the X pod is stuck", "the Y volume is broken", "the Z VM won't start"), confirm the actual identifiers with `kubectl get`, `qm config`, `docker inspect`, `pvesh ls`, or the equivalent on whichever platform owns the resource — BEFORE running any recovery command.
**Why:** humans under pressure confuse names (VM IDs, node hostnames, PVC names, container names). A 30-second verification prevents executing the wrong recovery on the wrong resource. Recovery commands are often destructive (`qm reset`, `kubectl delete pod --force`, volume detach); running one on a healthy resource turns a small incident into a bigger one.
**Verification pattern:**
1. Restate the user's claim: "You said pod `foo` is stuck on node `bar`."
2. Verify each identifier exists and is in the claimed state: `kubectl get pod foo -o wide` (does it exist? is it on `bar`? is it actually stuck?)
3. Confirm with the user before destructive action: "Pod `foo` on `bar` is in `ContainerCreating`. Running `kubectl delete pod foo --force --grace-period=0`. OK?"
General operational-safety rule, applicable in any high-pressure recovery situation.
## Discover API Endpoints via Swagger/OpenAPI Before Re-Reading Prose Docs
When an API call returns 404 or "endpoint not found", fetch the actual paths from the live spec (`/swagger.v1.json`, `/openapi.json`, `/api-docs`) with `curl + jq` before consulting prose documentation. The spec is generated from the running code; the prose drifts. Five seconds with curl beats fifteen minutes of doc spelunking. See [API Integration](api-integration.md) section 6 for the full pattern.
## Grep Minified Third-Party Source to Learn Its Runtime Protocol
When you must understand how a minified third-party library behaves at runtime (popup/auth detection, `postMessage` protocol, event names, expected `window.name`), grep the minified bundle for concrete string/pattern signatures (`grep -oP 'window\.opener[^;]*'`, event-name literals, magic constants) instead of guessing or reading upstream docs that may not match the shipped build. Minified code still contains the literal strings and property accesses that reveal the protocol reliably — enough to integrate against it without modifying its source.

View File

@@ -132,6 +132,29 @@ When a module will run inside a container, create its `_cli.py` entry point and
Treat "module + CLI shim + ENTRYPOINT verification" as one atomic unit of work.
## Pin All Container Images by Digest, Never by Tag
Mutable tags (`:latest`, `:v1.2.3`, `:stable`) can silently resolve to a different image after a rebuild. An image pinned to `myimage:v1.2.3` at deploy time may pull a newly-built `myimage:v1.2.3` months later that contains different code — no signal to the deployer that anything changed.
**Rule:** always pin to a content-addressable digest. Obtain the digest at build time:
```bash
docker buildx build --push -t reg.io/img@sha256:<digest> .
```
In Dockerfiles, reference the digest directly:
```dockerfile
FROM reg.io/base@sha256:a1b2c3d4e5f6...
```
In Kubernetes manifests and Helm values, use:
```yaml
image: reg.io/img@sha256:a1b2c3d4e5f6...
```
If using a tag is unavoidable, add a pre-flight check that fails if the tag resolves to a different digest than the one baked into the deployment config.
The semver versioning system (where available) encodes expectations about change severity — a patch bump should not silently include unrelated changes. Digest pinning makes the version-content contract verifiable.
## PostgreSQL Alpine Image Runs as UID 70, Not 999
The `postgres:*-alpine` images run postgres as **UID 70**, not the 999 used by Debian-based `postgres:*` images. Setting data-directory ownership to 999 (or 1000) via host `chown` or Ansible `file` modules silently breaks access — `pg_isready` may still pass while internal operations fail with "Permission denied" on WAL writes, replication slots, or extension installs.
@@ -142,3 +165,13 @@ docker exec --user postgres <container> id
```
Match the automation to the actual UID. If a project mixes Alpine and Debian Postgres images across environments, treat the UID as per-environment config, not a hardcoded constant.
## Drop Placeholder `require` Lines Before `hugo mod get`/`go get @branch`
When a `go.mod` pins a module to the placeholder version `v0.0.0` (common for a not-yet-tagged private dependency), `hugo mod get <module>@<branch>` (and `go get`) still tries to resolve the existing `v0.0.0` first and fails with `unknown revision v0.0.0` — even though a valid `@branch` override is supplied.
Fix: run `go mod edit -droprequire=<module>` for each placeholder dep *before* the `mod get`. Dropping the require line lets the `@branch` fetch proceed cleanly. Prefer this over editing `go.mod` with sed/bash arithmetic (no `set -e` footguns). Works inside any image bundling Go (e.g. the Hugo build image).
## Make a Script Self-Contained When It Runs in a Different Image
A helper script that imports the app's package (`from myapp import ...`) breaks with `ModuleNotFoundError` when copied into a separate, minimal image that doesn't install that package. When a script is destined to run inside a different (e.g. slimmer, single-purpose) image than the one it was authored against, inline its dependencies — call the underlying library (`httpx`, etc.) directly instead of importing the application package. Don't assume the target image shares the source image's installed packages.

View File

@@ -53,6 +53,7 @@ Separation yields:
- **Org repos require explicit collaborator grants.** Don't assume organizational membership implies write access — verify permissions before setting up automation or CI/CD.
- **Shallow clones break push operations.** `git clone --depth 1` is fine for read-only CI jobs, but pipelines that push artifacts, tags, or mirror to other remotes need full clones.
- **Reset shallow clones against `HEAD`, not `origin/main`.** A single-branch shallow clone (`--depth N`) creates no remote-tracking refs for other branches, so `origin/main` does not exist unless `main` is the branch being cloned. Init/reset logic that does `git reset --hard origin/main` fails on any other branch. Reset against `HEAD` instead — it is branch-agnostic and works regardless of which branch was shallow-cloned.
## Cross-Remote Hygiene for Multi-Remote Projects
@@ -87,3 +88,22 @@ When running parallel agents or tasks that modify the same repo:
- Tasks with multiple dependencies get an octopus merge base branch
- Worktrees share the `.git` object store — fast creation, minimal disk usage
- Keep containers detached (`docker run -d`, not `--rm`) so logs survive for inspection after exit
## Rebase Before Manual Commits to a CI-Auto-Bumped Branch/Deploy Repo
When CI auto-commits back to the same branch you push to — `[skip ci]` dependency-sync commits, deploy manifests with an image `newTag` bumped by a build job, ArgoCD `chore: deploy` commits — your push races those bots. A push made just after your fetch is rejected as non-fast-forward (`! [rejected] ... (fetch first)`), and on a deploy repo the running image can silently lag `origin/main` by many builds with no obvious error.
- **Standard sequence:** `git pull --rebase origin <branch> && git push origin <branch>`. Rebase (not merge) keeps history linear against the bot commits.
- **Expect conflicts in bot-managed files** (dependency manifests, version pins, image tags). Resolve by taking the higher/newer value.
- This race also fires immediately after *you* trigger a dependency-sync that CI auto-commits — pull-rebase before pushing your own follow-up.
## Gitea: Use `workflow_dispatch`, Not `repository_dispatch`, for API-Triggered Builds
On Gitea (observed 1.25.x), `POST /api/v1/repos/<org>/<repo>/dispatches` (the `repository_dispatch` trigger) returns 404 and never fires the workflow. Trigger builds instead via `POST /api/v1/repos/<org>/<repo>/actions/workflows/<file>.yaml/dispatches` with body `{"ref":"main","inputs":{...}}` (the `workflow_dispatch` trigger).
Also don't point a Gitea webhook at the `/dispatches` endpoint to chain builds: Gitea webhooks send the full push-event body (not `{"event_type": ...}`, which the endpoint ignores) and carry no auth header (so the call 401s). Chain builds from an in-CI dispatch step instead.
## PR Governance on Protected Branches
- **Don't bypass a `required_approvals` gate to unblock automation.** On a branch protected with `required_approvals: N`, do not self-approve a PR through a second bot account/token to let automation merge. That approval gate is a deliberate production-publish guardrail set by the repo owner; bypassing it defeats its intent. Leave the PR open for a human to approve. (This is distinct from a merge-whitelist misconfiguration, which is a config bug to fix — an approval gate is intentional.)
- **Promote a targeted change with a feature branch, never by merging the whole staging branch.** When a protected `main` requires approvals and you need to publish one targeted change, do not push/merge the entire integration branch (e.g. `staging → main`) — that promotes *all* accumulated changes at once. Push a feature branch containing only the targeted change and open a PR against `main`. Keeps the diff reviewable and prevents accidental bulk promotion.

View File

@@ -52,6 +52,10 @@ Manual bootstrap secrets (encryption keys, OIDC client secrets) must be document
- **Never use imperative operations on GitOps-managed resources.** `kubectl rollout restart` adds annotations that conflict with the GitOps controller's desired state, causing permanent OutOfSync. Use declarative paths instead — update a configmap hash annotation in Git, or change a pod template label.
- **ArgoCD reconciliation has latency.** New Application manifests don't appear immediately due to polling intervals. Use manual refresh annotations when automation needs immediate reconciliation.
- **Self-managed GitOps controllers revert their own live config.** When the GitOps controller manages itself via its own Helm chart with `selfHeal: true` (e.g., ArgoCD reconciling `argocd-cm`/`argocd-rbac-cm`), direct `kubectl patch`/`apply` on its ConfigMaps is reverted within seconds. Change the controller's Helm `values.yaml` in the deploy repo (accounts, RBAC policy CSV, server settings) — never the live ConfigMaps. Applies to any account/RBAC/config change on a self-managed controller.
- **A newly-created API account/token 403s until the config sync completes.** If you generate an API token for a GitOps-managed service account before the account exists in the live config (still mid-Helm-sync), the token returns 403. Wait for self-sync to complete, confirm the account is present, then generate the token.
- **Grant the narrowest role for the job.** A service that only reads controller/app status (e.g., a preview-readiness check) needs a read-only role, not a sync/deploy role. Scope the GitOps API account to exactly what it does.
- **Runtime-only "poke" annotations cause persistent OutOfSync.** Annotations added imperatively to trigger controller behaviour — e.g., a `force-sync`/`reconcile` annotation on an ExternalSecret to make External Secrets Operator refresh — are not present in Git, so the GitOps controller reports the resource `OutOfSync` indefinitely. Remove the annotation once it has done its job: `kubectl annotate <kind> <name> -n <ns> <annotation-key>-`. Applies to any "poke the controller" annotation not stored in the source manifest.
## PodSecurity Alignment
@@ -75,6 +79,25 @@ The `namespaceSelector: kube-system` rule is also ineffective for kube-apiserver
Using the wrong entity (or the wrong policy kind) results in silent policy drops. Symptom for apiserver-bound traffic: HTTPS calls hang until the client times out (typically 30s for Go HTTP defaults), then the upstream returns a generic error like `permission denied`. Test with `cilium monitor --type drop` to confirm the drop is at L3, or — if you can `exec` into the pod — try `wget --timeout=5 https://kubernetes.default.svc/healthz` and see if it `Terminated`s.
**Worked example (recurring incident pattern):** A secrets backend pod calls `auth/kubernetes/login` against the apiserver and hangs for exactly 30 seconds before returning a generic `permission denied`. The first hypothesis is RBAC or token misconfiguration — neither is the cause. Standard `NetworkPolicy ipBlock` rules listing the control-plane subnet don't match because cluster nodes carry the `kube-apiserver` / `remote-node` Cilium identity, and `ipBlock` only matches IPs without an identity. Fix is `CiliumNetworkPolicy` with `toEntities: [kube-apiserver]` on TCP 6443. Same pattern affects every workload that hits `https://kubernetes.default.svc` — External Secrets Operator, custom controllers, anything doing TokenReview / SubjectAccessReview.
## Every New K8s API Surface Needs Explicit RBAC in the Deploy Repo
When a service starts touching a new K8s API resource (CRDs, custom controllers, ExternalSecrets, Jobs, Leases), add a namespace-scoped `Role` + `RoleBinding` to the deploy repo in the **same commit** as the code that touches the API.
**Why this matters:**
- Manual `kubectl apply` during development masks the gap because the local kubectl admin context bypasses RBAC. Production then 403s silently and stalls reconcile loops.
- The GitOps reconciliation tries to apply the controller config but fails to read the underlying API; the symptom is a "permission denied" log line lost among thousands of other logs.
- The RBAC commit lands AFTER the feature commit, leaving an interval where the deployed code is broken in any environment that doesn't have the dev's admin context.
**Pattern:** every new API call requires either:
1. A namespace-scoped `Role` with the specific verbs (`get`, `list`, `watch`, `create`, `update`, `patch`, `delete` — only the ones actually used) on the specific resource
2. A `RoleBinding` to the service's `ServiceAccount`
Bundle both into the same deploy-repo commit as the code change. A reviewer should be able to grep for any new K8s API call in the code diff and find the matching `Role` verbs in the manifest diff.
**Common miss:** CRD custom resources need explicit `apiGroups` entries (e.g., `apiGroups: ["external-secrets.io"]`), not just `resources`. Forgetting the API group looks like a working Role definition but matches nothing.
## Kustomize Overlay `images:` Blocks Silently Override Base Tags
Kustomize `images:` blocks in an overlay apply to the entire rendered manifest, including any images defined in `base/`. If the base defines `image: my-app:v1.0.0` and the overlay has an `images:` block targeting `my-app`, the overlay's `newTag` silently wins — even if you intended the base tag to remain. When deploying a new image version via Kustomize, always update the `images:` block in the overlay, not just the base manifest. If the overlay doesn't have an `images:` block, add one rather than editing the base tag directly.
@@ -102,6 +125,10 @@ ArgoCD with ServerSideApply sometimes fails to detect changes to ConfigMap `data
CSI hot-plug of storage devices can fail silently — the VolumeAttachment object says `attached: true` but the device never appeared on the node. Pods get stuck in `ContainerCreating` with "device not found." Fix: delete the stale VolumeAttachment (`kubectl delete volumeattachment <name>`). The CSI driver recreates it and retries the attach.
**QMP-timeout leading indicator for hypervisor CSI hotplug failures.** When a hypervisor-backed CSI driver (Proxmox, vSphere) returns `ControllerPublishVolume` "published" but the block device never appears in the guest (`/dev/disk/by-id/wwn-0x...` absent) and the pod stays in `ContainerCreating`, check the hypervisor logs for a QMP command timeout during the preceding unpublish (e.g. `qmp command 'query-pci' failed - got timeout`). This points to an unstable backing VM/host, not a Kubernetes bug. Short-term workaround: add `nodeAffinity` with `NotIn: <bad-node>` on the deployment to steer the workload off the flaky VM while the hypervisor-side QMP stability is investigated.
**Targeted SCSI scan instead of full-bus rescan when hotplugged disks don't appear.** After a CSI hotplug, forcing a device rescan with a full-bus wildcard (`echo "- - -" > /sys/class/scsi_host/host*/scan`) hangs on some controllers. Use a targeted single-LUN scan instead: `echo "0 0 <lun>" > /sys/class/scsi_host/host<X>/scan` to probe a specific LUN without blocking the whole bus.
## Delete and Recreate ArgoCD Apps on Source Type Changes
When changing an ArgoCD Application's source type (e.g., multi-source Helm to single-source Kustomize), the repo-server may serve cached manifests from the old configuration, and old Helm hook resources become ghost entries that block deletion via finalizers. Delete the Application entirely and let the root app recreate it rather than patching source types in-place.
@@ -181,3 +208,15 @@ Pair periodic reconciliation with an authenticated POST `/reconcile` endpoint so
## kubernetes-py CustomObjectsApi: SSA Requires a Dedicated ApiClient
`CustomObjectsApi.patch_namespaced_custom_object(force=True)` fails with HTTP 422 on kubernetes-py v35 (`PatchOptions.meta.k8s.io is invalid: force: Forbidden: may not be specified for non-apply patch`). The default Content-Type is `application/merge-patch+json`; `force` is only valid on real server-side applies (`application/apply-patch+yaml`). The `_content_type` kwarg that older docs reference is not exposed in v35. Workaround: build a dedicated `ApiClient` and `set_default_header("Content-Type", "application/apply-patch+yaml")` on it; pass that client to a separate `CustomObjectsApi` used only for SSA patches. Reads/deletes use the default client (no body, default Content-Type harmless). The bug is silent in tests because mocks accept any kwargs — only a real apiserver round-trip surfaces it.
## `kubectl debug node/<node>` Fails Under Enforced PodSecurity
In namespaces/clusters enforcing PodSecurity `baseline` or `restricted`, `kubectl debug node/...` is rejected because its debug pod uses `hostPID` and `hostPath` (a baseline violation). Workaround: manually create a debug pod in a `privileged`-labelled namespace (e.g. the CSI driver's namespace) with `nodeSelector: kubernetes.io/hostname: <node>` and `securityContext.privileged: true`, rather than relying on `kubectl debug node`.
## `kubectl apply` of a New Image Tag May Not Roll Pods
Re-applying a Deployment with a bumped image tag does not reliably trigger a new rollout — the in-cluster image reference can stay cached, and `kubectl apply --force` does not fix it. When bumping an image version imperatively, `kubectl delete deployment <name>` before re-applying to guarantee a fresh pull. (In GitOps flows, prefer a digest pin or a template-hash annotation change; this delete-then-apply pattern is for imperative/dev workflows only.)
## A Correct CiliumNetworkPolicy Egress Rule Won't Help If Ingress Is Gated by a Plain NetworkPolicy
When a client pod times out reaching a service despite a correct `CiliumNetworkPolicy` egress rule on the client side, check the target's **ingress** policy — it may be a standard Kubernetes `NetworkPolicy` (not Cilium) that whitelists only specific source namespaces. Cilium and plain NetworkPolicy coexist and are additive; both directions must permit the flow. The ingress policy is often owned by the deploy repo, separate from the application and cluster-bootstrap repos, so grep there first. When debugging cross-namespace connectivity, enumerate both the client's egress rules and every ingress policy selecting the target pod.

View File

@@ -45,3 +45,43 @@ This breaks on-demand and "scale-to-zero" backends: the proxy has no route to th
## Cilium DNAT Resolves LB VIP Before NetworkPolicy Evaluation
Cilium performs DNAT on LoadBalancer VIP traffic before evaluating NetworkPolicy. Traffic to a VIP is rewritten to a backend pod IP before the policy check. For egress to LoadBalancer services in CiliumNetworkPolicy, use `toEndpoints` targeting the backend pods (by namespace/label), not `toCIDR` targeting the VIP.
## Don't Use ICMP/ping to Test eBPF LoadBalancer VIPs
Cilium (and other eBPF LB datapaths) only program TCP/UDP for a LoadBalancer VIP; ICMP is not handled by the eBPF LB and falls through to the kernel, which can generate redirect loops. Symptom: `ping <lb-vip>` returns "TTL exceeded" or "connection refused" even from a host on the same network, tempting a false "the load balancer is broken" conclusion — while TCP/HTTPS to the same VIP works fine.
**Rule:** test service-VIP reachability with `curl`/`nc` (TCP), never `ping`. A failed ping to an eBPF LB VIP is meaningless.
## VPN Subnet/Exit Routing Needs Return-Path SNAT
For any mesh/overlay VPN (Tailscale/Headscale, WireGuard subnet routers) that advertises LAN subnets or acts as an exit node, outbound reachability proves nothing on its own. LAN hosts have no route back to the VPN's client address range (e.g. Tailscale's `100.64.0.0/10` CGNAT range), so return traffic is silently dropped — the symptom is partial connectivity where some destinations answer and others don't.
**Fix:** MASQUERADE/SNAT traffic sourced from the VPN client range out each physical LAN interface on the router node, and persist it (`netfilter-persistent` or equivalent). **When testing any VPN or subnet-routing change, verify both directions** — confirm the SNAT rule exists before concluding a route works. Outbound success alone is a false positive.
## VLAN-Aware Bridge vs Raw Sub-Interface Coexistence (Proxmox/Linux bridges)
On a host that runs BOTH a VLAN-aware bridge (`vmbr0` on a bond/uplink) AND raw `vlanN@bondX` sub-interfaces on the same uplink, the kernel delivers tagged frames for VLAN N to the raw sub-interface, NOT to the vlan-aware bridge's VLAN handling. A VM NIC attached as `vmbr0,tag=N` then receives nothing. Symptom: guest has the right IP/routes and the switch port is healthy, but there is zero L2 connectivity (no DHCP lease → APIPA `169.254.x`, gateway un-ARPable) — the guest config looks perfect while frames go nowhere.
**Rules:**
- A genuinely-tagged VLAN that also has a raw `vlanN` sub-interface must attach to its **dedicated named bridge** (whose port is `vlanN`), never `vmbr0,tag=N`.
- A VM on the trunk's **native/untagged** VLAN must use untagged `vmbr0` (no `tag=`); a `tag=` NIC emits tagged frames that don't match the native VLAN and are dropped both directions.
- For a multi-VLAN VM, one NIC per VLAN on per-VLAN bridges is more reliable than Proxmox `trunks=` (which silently fails to route on some `vmbr0` configs).
**Diagnostic shortcut:** a guest with correct IP/routes but unreachable from a same-subnet host is an L2/VLAN problem, not guest config. Check which bridge the tap landed on (`ls /sys/class/net/<Bridge>/brif/`, `bridge vlan show dev <tap>`) and whether the VLAN is native(untagged) vs tagged on the uplink.
## Forward-Auth Proxy Redirects Break CORS Preflight (302 on OPTIONS → ERR_INVALID_REDIRECT)
A forward-auth layer (Authelia, oauth2-proxy, etc.) sitting in front of a backend intercepts *unauthenticated* requests and 302-redirects them to its login portal — including CORS OPTIONS preflight requests. The browser reports `ERR_INVALID_REDIRECT` (not `ERR_FAILED`), pointing at a 302 whose target is the auth portal.
**This is an auth-bypass gap, not a CORS bug.** The backend's CORS middleware handles OPTIONS correctly once the auth layer stops intercepting. Do not touch the CORS config.
- **Signature:** `ERR_INVALID_REDIRECT` + a 302 whose `Location` is the auth portal ⇒ bypass gap. A true CORS bug shows `ERR_FAILED` / missing `Access-Control-*` headers with no redirect.
- **Fix:** add the endpoint path to the auth layer's bypass/allow list, then reload the auth layer.
- **Bypass rules are path-scoped.** Every new backend endpoint needs its own bypass entry — one that works does not cover a sibling. Enumerate all public paths and confirm each is bypassed, including root-level paths a client library may call outside the expected prefix (e.g. `/schemas`, `/health`).
## `tlsv1 alert internal error` = DNS Resolves to the Ingress, but No Route Matches the Host
A domain resolves to the ingress IP, but `https://<domain>` aborts during the TLS handshake with `tlsv1 alert internal error`. Root cause: the ingress controller (Traefik, and equivalents) has no route matching `Host(<domain>)` — only related hosts exist (a different subdomain, a sibling domain). With no matching route there is no certificate to present, so the handshake aborts before any HTTP layer. This is the expected "nothing is served here" behaviour, not a cert or reverse-proxy misconfiguration.
- Diagnostic: if `curl -v https://<domain>` fails at TLS but a known-good sibling host on the same ingress succeeds, suspect a missing route, not a broken cert.
- Fix (if the site should be served): add an ingress route/IngressRoute matching that exact Host. Otherwise it is working as intended.

View File

@@ -255,6 +255,11 @@ To demonstrate a full CI-to-deployed flow, set lifecycle phases to auto-deploy b
7. **Process templates cannot reference the project's own Git repo** for scripts — use inline scripts or external URLs.
8. **Test templates** by creating a test project that consumes them before sharing widely.
## Deployment Freeze Gotchas
- **Freeze recurrence is Daily/Weekly/Monthly only — no sub-daily granularity.** `RecurringSchedule.Type` accepts only `Daily`, `Weekly`, `Monthly`; `Cron`, `OnceDaily`, `Custom`, and `None` are rejected by validation. A short-cycle rolling freeze (e.g. a few minutes in every ten) cannot be expressed as a native recurring freeze. Fallback: a scheduled runbook that rewrites the freeze's `Start`/`End` every N minutes.
- **Deployment freezes are instance-level, not space-scoped.** Use `/api/deploymentfreezes` — the space-scoped path 404s silently.
## References
- OCL Syntax: https://octopus.com/docs/projects/version-control/ocl-file-format

View File

@@ -203,6 +203,8 @@ requires = ["setuptools>=68"]
build-backend = "setuptools.build_meta"
```
**(d) New `[project.scripts]` console entry point not found after adding it.** A `command not found` for a script declared in `[project.scripts]`, despite the declaration being correct. Editable installs only register console scripts that existed *at install time* — adding an entry point after `pip install -e .` doesn't retroactively create the shim. Re-run `pip install -e .` in the active venv to register any newly added entry point.
## Bridging `threading.Event` to `asyncio` — Use a Polling Bridge
When an HTTP server or other handler running in a **background thread** needs to signal an `asyncio` controller loop (e.g., webhook handler waking a reconciler), a plain `threading.Event.set()` cannot directly wake `asyncio.wait_for` or `asyncio.sleep` in the event-loop thread. The event loop only wakes for tasks it scheduled.
@@ -285,3 +287,19 @@ spec.loader.exec_module(module)
```
The `sys.modules` registration must happen before `exec_module` — otherwise any `import my_tool` inside the loaded module creates a second, distinct module object.
## `httpx` Credential-Carrying Clients: Disable Redirect-Following
`httpx.Client(follow_redirects=True)` (the default) forwards the `Authorization` header to the redirect target on a 302 — leaking a bearer/OIDC token to whatever host the redirect points at. Set `follow_redirects=False` unconditionally on any client that carries credentials. Related: `httpx.MockTransport` fails on relative URLs, so always pass an explicit `base_url=` in tests; and a `base_url` with a path prefix (`https://x/api`) is silently dropped when the request path is absolute (`/v1/items``https://x/v1/items`) — end `base_url` with `/` and `lstrip("/")` the path before joining.
## Construct a Fresh HTTP Client Per Call for Per-Request Auth Headers
When each request needs its own headers (HMAC signature + timestamp, one-time nonce, per-call bearer), do NOT mutate or evolve a shared client. httpx's `Client.with_headers()` evolves a copy that still shares the underlying transport, so per-request headers leak across calls and race under concurrency. Build a fresh `Client(headers=...)` per call. Related header trap: `dict(request.headers).get("Authorization")` returns `""` — casting httpx's case-insensitive multidict to a plain dict loses the header; use `request.headers.get(...)` directly.
## SQLAlchemy async: Always Use `async with session_factory()` — Never a Raw Session
A raw `AsyncSession` created as `session = session_factory()` (not `async with`) that goes out of scope without `commit()`/`rollback()`/`close()` leaves psycopg3's implicit transaction dangling. The GC cannot async-close it, so the pool's checked-out counter never decrements — the slot becomes a permanent ghost. Once all `pool_size + max_overflow` slots are ghost-occupied, every new connection request blocks forever: health checks time out, the liveness probe kills the pod, and it crash-loops. Precursor signal in logs: SAWarning "garbage collector is trying to clean up non-checked-in connection." Fix: route every session through `async with session_factory() as session:` (or a `get_session()` helper). Audit every call site that calls `session_factory()` directly — the crash is invisible in local/memory-store testing and only bites production.
## `logging` Reserved LogRecord Attributes Crash the Formatter
Passing a reserved key in `logger.info(msg, extra={...})` raises `KeyError("Attempt to overwrite 'name' in LogRecord")`. Reserved names include `name`, `msg`, `args`, `levelname`, `levelno`, `pathname`, `filename`, `module`, `funcName`, `lineno`, `created`, `process`, `thread`, `message`, `asctime`, and the rest of the LogRecord fields. Rename your keys (e.g. `resource_name`). Worse failure mode: if this log line sits in a success path wrapped by an outer `try/except Exception`, the KeyError makes a *successful* operation look like a failure and rolls back / skips subsequent work. Rule: logging is best-effort — never put load-bearing logic in a log call, and audit `extra=` keys for reserved-name collisions before shipping.

View File

@@ -487,3 +487,7 @@ Gate on `application/problem+json` content type. Create an `ApiProblemError` cla
### Token injection
Validate token expiry before making API calls (proactive), not in a 401 response interceptor (reactive). Use the OIDC client's `getAccessTokenSilently()`.
### Run the Real Build (`tsc -b`), Not Just `tsc --noEmit`, Before Pushing
`tsc --noEmit` (what many `typecheck` scripts run) does NOT check test files or apply project-reference settings. `tsc -b` / `vite build` resolves project references and applies `noUnusedLocals` across ALL files including tests — so CI's `npm run build` fails on `TS6133 'X' is declared but never read` in a test file that `typecheck` passed clean. Make the pre-push hook run the actual build command CI runs, not the lighter check. Related strict-mode friction with `noUncheckedIndexedAccess: true` (recommended in §14): every `arr[i]` is `T | undefined` — guard the access or use a justified non-null assertion.

View File

@@ -52,6 +52,23 @@ Every script that modifies state should support `--dryrun` / `-n`:
- **Order matters in sed/regex transformation pipelines.** Process more specific patterns before general ones. For example, if both `![[image.png]]` and `[[page]]` are valid patterns, process the image embed first — otherwise the general wikilink regex matches the inner `[[image.png]]` and the `!` prefix is left orphaned.
- **Bare `except: pass` swallows `SystemExit` in Python.** `sys.exit(0)` inside a bare `except: pass` block is captured as a `SystemExit` exception and silently swallowed — the script continues instead of exiting. Use specific exception types in except clauses, or use `break`/`return` for loop early-exit, or re-raise after checking `isinstance(e, SystemExit)`. Applies to any script with loops that short-circuit on a condition inside exception handling.
- **Never use `GROUPS` (or other reserved names) as a bash variable.** `GROUPS` is pre-set by bash completion and session initialisation with numeric group IDs — assignment appears to succeed but the pre-existing value often persists in sourcing contexts, producing bizarre "array iterates over 1000, 24, 27..." bugs. Other reserved/built-in names to avoid: `UID`, `EUID`, `PWD`, `OLDPWD`, `SHLVL`, `RANDOM`, `SECONDS`, `LINENO`, `PIPESTATUS`, `IFS`. Prefix project variables (`PROJECT_GROUPS`, `TEMPLATE_SLUGS`).
- **Inline `VAR=val cmd "$VAR"` expands the pre-existing value, not the new one.** Bash inline env-var assignment sets `VAR` for the child process, but `$VAR` in argument position is expanded by the **calling** shell using its existing (often empty) value. The trap:
```bash
CP_TOKEN=$(get_token) curl -H "Authorization: Bearer $CP_TOKEN" ... # sends empty header
```
The Authorization header is empty because `$CP_TOKEN` is expanded before `CP_TOKEN=$(get_token)` takes effect. **Fixes:**
```bash
export CP_TOKEN=$(get_token) # set on a prior line
curl -H "Authorization: Bearer $CP_TOKEN" ...
# — or —
curl -H "Authorization: Bearer $(get_token)" ... # inline substitution at use site
```
Doesn't affect Python subprocesses launched with `env=...` because `os.environ` reads at runtime, not at command-parse time. Costs ~30 minutes per occurrence; common in CLI-tool authentication wrappers. See [Secrets Management](secrets-management.md) "Never Source .env Files" for the related safe-parser pattern when handling `.env`-style files.
- **`git mv <src> <dest>/<sub>` won't create missing parent directories.** `git mv admin static/admin` fails with `renaming 'admin' failed: No such file or directory` when `static/` doesn't exist yet, even though `admin/` does. Create the parent first: `mkdir static && git mv admin static/admin`.
## Grep All Consumers Before Removing or Keeping a Field
Before deleting — or deciding to keep — a config field, env var, or interactive prompt, grep every consumer across the tree (`grep -r VAR_NAME .` / `~`). Two outcomes: (1) dead fields accumulate silently when nothing reads them — the grep proves they're unused and safe to remove; (2) for fields you keep or rename, the grep enumerates every downstream file needing a matching edit (templates, status lines, other scripts), so you don't ship a half-applied rename. Auditing consumers is cheaper than shipping a change that leaves orphaned references.
## JSON Construction in Scripts

View File

@@ -42,10 +42,24 @@ The `ps`-visibility problem is especially easy to hit in container entrypoints t
Prefer stdin, env vars, or `@file` references over argv in every entrypoint script.
### Secrets in `kubectl exec` One-Liners
The same argv-visibility rule applies to ad-hoc debugging, not just entrypoints. `kubectl exec <pod> -- sh -c "... TOKEN=${X} ..."` places the secret in the pod's `ps` output (visible to every process in that PID namespace) and trips argv-based secret classifiers/blockers. Instead, write a small helper script to a scratch path and pass the value via stdin, and use the target tool's file-reference flag (e.g. `bao kv patch ... KEY=@/path/to/file`) rather than inline values.
## Bootstrap Secrets
Some secrets are chicken-and-egg (e.g., the age decryption key for ArgoCD's KSOPS). These must be created manually as a bootstrap step and documented clearly.
### Read Runtime Bootstrap Secrets from Their Live Store, Never Copy Them
Bootstrap/root/unseal credentials for a secrets backend (Vault/OpenBao root token, unseal keys) deliberately do not live in the general secrets directory. They live only in their runtime store — typically a Kubernetes Secret created at init — and are read on demand at execution time:
```
kubectl -n <ns> get secret <unseal-secret> -o jsonpath='{.data.<field>}' | base64 -d | <parse-the-json-field>
```
Document only *where* the credential lives and how to retrieve it — never copy the value into a memory file, the secrets directory, or any other file. Retrieving it ephemerally at point of use is the pattern; persisting a second copy defeats the single-source-of-truth and widens blast radius.
## Credential Lifecycle Management
- **Track credential expiry dates.** OAuth client secrets, API tokens, and certificates have expiry dates that can cause silent failures. Document expiry dates when creating credentials.

View File

@@ -78,6 +78,17 @@ Some scenarios genuinely require client-side credentials (e.g., direct S3 upload
**Right:** CMS sits behind Authelia (MFA). A proxy service intercepts the OAuth flow, issues a proxy session token (containing only the user's identity), and forwards all API calls to Gitea using a per-user server-side Gitea token. The Gitea token never leaves the server.
### Browser-Side OAuth Popup: COOP Severs `window.opener`
When a client-side app (CMS admin UI, SPA) runs an OAuth popup flow and the callback redirects through a different origin (e.g., an auth proxy on another subdomain), the Cross-Origin-Opener-Policy (COOP) mismatch severs `window.opener`. The popup can no longer `postMessage` its result back to the opener, and the parent's popup-detection logic (commonly `window.opener?.origin === location.origin && window.name === 'auth'`) silently fails — producing zero token exchanges and, if the callback falls through to re-init the login flow, an infinite popup loop.
Mitigations:
- Set `COOP: same-origin-allow-popups` on the page that opens the popup; keep COOP off the cross-origin callback/proxy page.
- Handle the OAuth callback explicitly: when `?code=&state=` are present, always `return` after processing — never fall through to auto-init/auto-click, which is what amplifies a broken flow into an infinite loop.
- When `window.opener` is severed, deliver the auth result via a `BroadcastChannel` fallback. Re-dispatch it as a synthetic `MessageEvent` so the framework's existing handler fires without patching third-party (often minified) source.
This is the browser-mechanics complement to the server-boundary rule above: even with a correct backend proxy, the popup handshake itself breaks on COOP. Check security headers (COOP/COEP) early when a login popup silently never completes — before investigating sessions, cookies, or backend state.
## Multi-Tenant Isolation Guard on Outbound Writes
When a service writes into per-tenant destinations — customer Slack channels, per-customer boards, per-org webhooks, tenant-scoped buckets — route every outbound write through a single **IsolationGuard** module.
@@ -111,3 +122,37 @@ CallSite ──→ IsolationGuard.send(tenant_id, destination, payload)
```
No call site should import the underlying transport directly. Lint or grep for direct imports as a CI check.
## File-Only Secret Delivery for Long-Lived Processes
For any container or process running longer than a few minutes, deliver secrets via tmpfs-backed file mounts (mode 0400, owned by root or the secrets operator), never via environment variables. Env vars leak into:
- `ps -auxe` and `/proc/<pid>/environ` for any process in the PID namespace
- Log aggregators (anything that captures the process tree)
- Accidentally-echoed error output (`echo "config: $DATABASE_URL"` leaks the secret to stdout)
- Child processes that inherit the environment by default
Env-var delivery is acceptable only for short-lived (<60s) ephemeral tasks where the exposure window is bounded and no other workloads share the PID namespace.
**Defense-in-depth pattern (External Secrets Operator + init):**
1. ESO mounts the Secret at `/var/run/secrets/...` as mode 0400, owned by root (controller's service account)
2. The pod's init container stages a 0600 copy to a path owned by the application's UID
3. The application reads from disk at invocation time, not at startup — so rotation lands without restart
**Two-layer scoping:** the ESO mount is the source-of-truth; the init copy is the application-readable view. A compromised application can read its copy but cannot read the ESO source. A compromised init container does not survive past startup.
## Server Never Holds the Private Key — CSR-Based Flow for Server-Issued Identity
When a server issues identity material to a client (mTLS cert, agent signing key, per-request token), design the protocol so the server signs a CSR rather than generating the keypair.
**Rationale:** Most managed-runtime languages (CPython, Java, Go via reflection) cannot reliably zeroise private bytes. CPython strings are immutable; there is no `memset` equivalent; GC residue, copy-on-write, and string interning all leave secret material in addressable memory after "deletion." Once the server has held the private key in process memory, you cannot prove it is gone.
**Flow:**
1. Client generates the keypair locally via the standard `cryptography` primitives
2. Client builds a CSR including its identity claims (SAN URI, subject CN)
3. Server receives the CSR over an authenticated transport, validates the identity claims, signs the CSR, returns the certificate
4. The server never sees the private key
Applies to any mTLS issuance, ACL/agent-signing keys, per-request token issuance, and any "server hands the client an identity" protocol. The pattern survives memory-dump forensic analysis of the server.

View File

@@ -423,6 +423,14 @@ def test_regression_crlf_corruption():
These are the highest-value tests because they catch proven failure modes.
### `xfail(strict=True)` Is the Right Red Primitive for Specs Ahead of Implementation
When a spec requirement has no implementation yet, write a test that asserts the not-yet-existing import/attribute/behaviour and mark it `@pytest.mark.xfail(strict=True)`. It keeps CI green while red; when the impl lands and the test passes, strict mode flips it to XPASS (a visible CI failure) that signals "remove the marker." Better than `@pytest.mark.skip` (never runs, false-green) or no marker (breaks CI immediately).
Two traps: a module-level `pytest.skip(allow_module_level=True)` on a failed import swallows every xfail-strict marker in the file (reports `skipped`, not `xfailed`) — put spec-ahead tests in a file with no module-level skip and lazy-import inside each test body. And `pytest.importorskip("mod.foo")` on the very module being implemented produces a false-green SKIP that lets an agent claim success without writing code — require a `pytest --collect-only` ImportError gate instead.
**Retire markers once the feature ships.** After each milestone, grep for `xfail(reason=` / stale `xfail` markers and remove the now-obsolete ones — leftover xfail state contributes coverage noise and can hold total coverage below the gate even though the code paths execute. Set `xfail_strict = true` globally so xpassing tests fail loudly and force the cleanup rather than silently rotting.
## Test Quality Metrics
### What to Measure
@@ -502,6 +510,10 @@ Patching an entire module (e.g., `patch("mod.kubernetes.config")`) replaces exce
When migrating a codebase from sync to async, helper functions get converted but test functions are often left as sync `def`. Every test that calls an async function needs `async def` + `@pytest.mark.asyncio` + `await`. After any async migration, run tests and grep for `RuntimeWarning: coroutine '...' was never awaited` to find remaining sync-to-async gaps.
### `AsyncMock.side_effect` on a Sync Call Is Silently Dead
`mock = AsyncMock(side_effect=SomeError)`; a *synchronous* `mock()` call returns an un-awaited coroutine and never raises — the `side_effect` is discarded and the test passes regardless of implementation (a `pytest.fail` line placed after the call executes, masking the real assertion). Fix: make the test `async def` + `@pytest.mark.asyncio` and `await` the call, or use a plain `MagicMock(side_effect=...)` when the production code is sync. Sibling trap: `getattr(MagicMock(), "attr", None)` returns a *new MagicMock*, not `None` — so `mock.account_id == "x"` is always False. Explicitly set every attribute the code reads (`m.account_id = "x"`) on test MagicMocks.
## Anti-Patterns
### Tests that mirror implementation

View File

@@ -164,3 +164,36 @@ For services that maintain in-memory state that should survive restarts (task qu
3. **Clean-slate fallback** — if the state file is missing, corrupt, or incompatible, start with empty state and log a warning rather than crashing
This pattern handles pod restarts, rolling deploys, and upgrade scenarios without requiring an external database. The clean-slate fallback is critical — a startup crash due to a corrupt state file is worse than losing the state.
## Phase-Done Gate: Wired, Integrated, AND Exercised End-to-End
A subsystem is not "complete" because its unit tests pass. A phase ships when ALL THREE hold:
1. **Unit tests pass** — the new code is exercised in isolation
2. **Wired into the application lifespan and covered by integration tests** — the new code is reachable from production startup paths (imported by `main.py`, registered in lifespan handlers, included in the DI graph) AND integration tests confirm it
3. **Exercised in one full end-to-end production-equivalent run** — the side effect is observed: resource materialised, event delivered, persistence verified, callback fired
The dominant failure mode is two-of-three: code exists and passes unit tests but was never imported from `main.py`; or wired up but never exercised live. Both silently look "done." Walking through these three checkpoints before stamping a phase complete catches the gap before it becomes a production incident. Generalises to any service with lifespan startup, scheduled drainers, background workers, or out-of-band reconciliation loops.
## Verify Agent-Chosen Version Pins Against the Live Registry
When an AI agent or research tool produces version pins (`react@18.2.0`, `httpx==0.28.0`, `<image>:1.4.2`), validate every pin against the live registry before committing — `npm view <pkg> version`, `pip index versions <pkg>`, `docker manifest inspect <image>:<tag>`, `helm search repo <chart>`. Agents (especially those without web access, including all container-bound agents) fabricate plausible-looking version numbers from training data with measured ~30% inaccuracy.
A 5-second registry check prevents a multi-cycle push-debug-fix loop. Same rule applies to release dates, "actively maintained" claims, and any other version-shaped assertion an agent produces.
Generalises the existing "Verify Container Image Tags Before Writing References" rule to cover npm/PyPI/Helm-chart pins.
## Restart Services That Read Config Only at Startup After a Config Change
Many services (Authelia, and any app that loads config at boot) read their config file exactly once at startup. Editing the mounted config — ConfigMap, file, env source — has **no effect until the process restarts**. A config change that "did nothing" almost always means the process was never restarted.
- **In GitOps/K8s:** a ConfigMap edit does not restart the consuming pod. Force it with a `kubectl.kubernetes.io/restartedAt` pod-template annotation (syncs through GitOps) or `kubectl rollout restart deployment/<name>`. The annotation is preferred when the change must land purely through GitOps.
- Before diagnosing a "config change had no effect" symptom, confirm the process actually restarted. Check pod age / start time, not just that the ConfigMap contains the new value.
## Long-Lived Compose/Service Stacks Must Survive Reboots
A Docker Compose stack (or any long-lived service) started manually does not come back after a power cut or reboot. Any long-lived stack on a machine that can lose power must be wired into the boot path — a systemd unit or an Ansible/config-management-managed equivalent — as part of bringing it up, not deferred to "later." Verify it restarts unattended (e.g. reboot the host, or `systemctl restart`) rather than assuming the unit works.
## Pause and Update the Plan When Verification Changes a Mechanism
When an early verification step (schema check, API probe, capability survey) reveals the real system cannot express a planned mechanism — e.g. native recurrence can't do the cadence you designed for, or a linking primitive doesn't exist — stop before building the affected milestone and revise the plan's mechanism table first. Building against a disproven assumption and reworking it later is far more expensive than a mid-plan pause. This is the corrective action that pairs with front-loaded capability verification.