Add 37 new entries and update 7 existing entries across 13 topic files. Major contributions from agent-runtimes (K8s secrets, CI, Docker gotchas), cluster-bootstrap (ArgoCD SSA, etcd tuning, DB migrations, Compose networking), and cluster-apps/octopus-deploy (Helm vs raw manifests, ArgoCD source types). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
10 KiB
Kubernetes Patterns
Volume Mounts
- Avoid
subPathvolume mounts for Secrets and ConfigMaps. The kubelet does not auto-updatesubPathmounts when the source changes — the pod must be restarted. Use directory mounts instead and adjust the application's config path. - Secret volume propagation is async. After updating a Secret, the kubelet takes seconds to sync mounted volumes. A
rollout restartissued immediately after may start pods with stale data. Add a short delay (5s) before restarting.
Deployment Strategies
- RWO PVC + RollingUpdate = Deadlock. New pod can't attach the volume while the old pod holds it. Use
strategy: Recreatefor single-replica deployments with RWO PVCs. - SSA + strategy change conflict. Switching from RollingUpdate to Recreate via ServerSideApply fails because SSA won't remove the old
rollingUpdatefield. Must patch the live resource first.
Naming
metadata.namemust be DNS-1035 compliant — no dots allowed. Replace dots with dashes (e.g.,oreillyit-nznotoreillyit.nz). Label values CAN contain dots.
Bootstrap Ordering
Some components have chicken-and-egg dependencies:
- CNI (e.g., Cilium) must be installed before anything else — nodes are NotReady without it
- GitOps controller (e.g., ArgoCD) installed second
- Root app applied last — the GitOps controller then "adopts" CLI-installed releases
Manual bootstrap secrets (encryption keys, OIDC client secrets) must be documented as explicit steps.
Network Policies
- DNS egress for
toFQDNsrules must usetoEndpointstargeting kube-dns pods withrules.dns— this triggers the DNS proxy. UsingtoCIDRSetfor DNS bypasses the proxy and FQDN rules never populate. - Cross-namespace policies need explicit namespace matching (e.g.,
matchExpressionson namespace label). - Always test from the actual consumer namespace, not same-namespace test pods.
Probe Strategy
- Liveness vs readiness probes serve different purposes. TCP checks confirm the process is listening (liveness). Exec/command checks confirm the application is ready to serve (readiness). Don't conflate them.
- Probes must match application host validation. Applications that validate Host headers (e.g., Next.js
ALLOWED_HOSTS) will reject probes sent to the pod IP. SethttpGet.httpHeaderswith the expected Host value. - Don't load credentials into liveness probes. If readiness requires an authenticated check (e.g.,
sqlcmd), use a simple TCP check for liveness and reserve the authenticated check for readiness only.
Init Container Patterns
- Writable config via init container + emptyDir. When apps require writable directories but ConfigMaps are read-only, use an init container to copy config into an emptyDir volume that the main container mounts read-write.
- Privilege separation. Init containers can run as root to create directories or set ownership, while the main container runs as a non-root UID. Prefer this over running the entire workload as root.
- Non-root images have hidden filesystem requirements. Many modern images (e.g., MSSQL 2022, UID 10001) need writable directories beyond the obvious ones. Always check image documentation or
docker inspectbefore writing manifests.
StatefulSet Edge Cases
- CrashLoopBackOff pods won't auto-replace on spec update. The StatefulSet controller won't delete and recreate a crashing pod when you update the spec — manual
kubectl delete podis required to force recreation. - Immutable field diffs can deadlock auto-sync. StatefulSet fields like
volumeClaimTemplatesare immutable after creation. GitOps controllers (ArgoCD) will show permanent OutOfSync if the desired state differs from the live immutable fields. Force sync or recreate the StatefulSet. - SSA causes perpetual OutOfSync from defaulted fields. Kubernetes defaults fields on StatefulSets (
persistentVolumeClaimRetentionPolicy,revisionHistoryLimit,updateStrategy.rollingUpdate.partition) that aren't in the Helm template. WithServerSideApply=true, GitOps controllers see these as diffs and report OutOfSync even though the app is Healthy. The app functions correctly — this is cosmetic. Consider ArgoCDignoreDifferencesfor these fields.
GitOps: Imperative vs Declarative
- Never use imperative operations on GitOps-managed resources.
kubectl rollout restartadds annotations that conflict with the GitOps controller's desired state, causing permanent OutOfSync. Use declarative paths instead — update a configmap hash annotation in Git, or change a pod template label. - ArgoCD reconciliation has latency. New Application manifests don't appear immediately due to polling intervals. Use manual refresh annotations when automation needs immediate reconciliation.
PodSecurity Alignment
- Namespace PodSecurity labels must match container security contexts. DinD, CSI drivers, and other privileged workloads need
pod-security.kubernetes.io/enforce: privilegedon their namespace. Abaselineorrestrictednamespace silently blocks privileged pods. - Document privileged namespace requirements. When a workload needs elevated privileges, document the specific requirement (e.g., "Docker-in-Docker for CI builds") alongside the namespace label.
ArgoCD Source Type Detection
- ArgoCD auto-detects Kustomize. When a source directory contains
kustomization.yaml, ArgoCD runs Kustomize automatically. Adding an explicitdirectory:source type overrides this detection and causes ArgoCD to try applyingkustomization.yamlas a raw K8s resource, which fails with schema errors. Remove explicit directory source types from Kustomize sources. - Credential template URL-prefix must match exactly. ArgoCD repo-creds secrets use URL prefix matching. When migrating Git server URLs (hostname, protocol, or port changes), update the credential template to match the new prefix. Stale credentials cause "authentication required" errors on all apps using that prefix.
Miscellaneous
enableServiceLinks: falsemay be needed when K8s-injected service env vars conflict with app config (e.g., Authelia interpretsAUTHELIA_*service vars as configuration).- Proxmox VM names must match K8s node hostnames for cloud controller manager integration.
- Metrics-server on Talos needs
--kubelet-insecure-tls(self-signed kubelet certs).
ArgoCD SSA + StatefulSet volumeClaimTemplates = Perpetual OutOfSync
Kubernetes injects apiVersion and kind fields into StatefulSet volumeClaimTemplates on apply. These don't exist in source manifests, causing ArgoCD with ServerSideApply to report perpetual OutOfSync. Fix: add ignoreDifferences on the ArgoCD Application targeting .spec.volumeClaimTemplates[]?.apiVersion and .spec.volumeClaimTemplates[]?.kind (using jqPathExpressions), plus RespectIgnoreDifferences=true in syncOptions.
ArgoCD SSA May Not Detect ConfigMap Data Changes
ArgoCD with ServerSideApply sometimes fails to detect changes to ConfigMap data values, reporting "Synced + Healthy" while the live ConfigMap has stale content. Root cause: SSA field ownership conflicts between Helm's managed fields and a prior kubectl apply annotation. After syncing ConfigMaps managed by Helm+SSA, verify content with kubectl get cm <name> -o jsonpath='{.data.<key>}'.
CSI VolumeAttachment Stuck After Hot-Plug Failure
CSI hot-plug of storage devices can fail silently — the VolumeAttachment object says attached: true but the device never appeared on the node. Pods get stuck in ContainerCreating with "device not found." Fix: delete the stale VolumeAttachment (kubectl delete volumeattachment <name>). The CSI driver recreates it and retries the attach.
Delete and Recreate ArgoCD Apps on Source Type Changes
When changing an ArgoCD Application's source type (e.g., multi-source Helm to single-source Kustomize), the repo-server may serve cached manifests from the old configuration, and old Helm hook resources become ghost entries that block deletion via finalizers. Delete the Application entirely and let the root app recreate it rather than patching source types in-place.
etcd on Slow Storage Requires Timeout Tuning
etcd requires sub-10ms fsync for stable operation. On slow storage (HDD, network-attached, overloaded SSD), default timeouts (heartbeat 100ms, election 1000ms) cause leader election flapping, cascading scheduler/controller-manager restarts, and widespread probe failures. Symptoms: "leader failed to send out heartbeat on time", "apply request took too long." Mitigation: increase heartbeat-interval (e.g., 500ms) and election-timeout (e.g., 5000ms). Fix: move etcd to SSD storage. Periodic defrag also helps.
CrashLoopBackOff Delays New Image Pickup
After CI builds a fix for a crashing pod, the CrashLoopBackOff exponential backoff (up to 5 minutes) means the kubelet won't pull the new image until the next retry window. Run kubectl rollout restart deployment/<name> immediately after CI completes to create a fresh pod instead of waiting.
K8s Secret Volumes Are Read-Only with Root Ownership
K8s Secret volume mounts are read-only — you cannot write or modify files in them. Files are owned by root regardless of fsGroup settings. Non-root processes need defaultMode: 0444 (world-readable) to access the files. Additionally, mounting a Secret at a parent path shadows any other Secret mounts at child paths — mount each Secret at its own non-overlapping path.
K8s Env Var Size Limits for Payloads
Kubernetes has a hard limit on environment variable sizes (~228KB base64). Large payloads embedded as env vars cause containers to crash with exit 255 and zero logs. For inter-task artifact passing, use git branches or mounted volumes instead of env var payloads.
Non-Blocking Registration in FastAPI Lifespan Handlers
Blocking operations (external API calls, service registration) in application lifespan handlers prevent the HTTP server from starting. K8s liveness probes fail and the pod enters CrashLoopBackOff before the operation completes. Use background tasks (e.g., asyncio.create_task) for registration so health endpoints respond immediately while registration happens asynchronously. This applies to any K8s-deployed app framework with startup hooks (FastAPI, Flask, etc.).