Adds 3 new topic files (ai-parallel-agents, api-integration, python-patterns) and extends 21 existing topic files with new gotchas and patterns surfaced from memory across tracked projects. Index updated accordingly. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
144 lines
6.4 KiB
Markdown
144 lines
6.4 KiB
Markdown
# Database Selection
|
|
|
|
## The Rule: SQLite Is Not a Production Database
|
|
|
|
**Any service that meets ANY of the following criteria MUST use PostgreSQL (or equivalent server-grade database) from day one:**
|
|
|
|
- Attached to a FQDN (has a real domain name, even internal)
|
|
- Serves traffic from more than one process (API consumers, CI runners, webhooks, polling)
|
|
- Backs infrastructure that other systems depend on (Git hosting, container registries, auth providers)
|
|
- Will be accessed concurrently by automated systems (ArgoCD, CI runners, cron jobs)
|
|
|
|
**Do not use SQLite for these workloads. Not temporarily. Not "to start with." Not "we'll migrate later."**
|
|
|
|
SQLite uses file-level locking — only one writer at a time, and writes block reads. Under concurrent access, requests queue up waiting for the write lock, causing cascading timeouts. The failure mode is insidious: the service appears to work fine under light load but becomes intermittently unresponsive under real workloads. By the time you notice, everything that depends on it is also failing.
|
|
|
|
## The Cost of "We'll Migrate Later"
|
|
|
|
The Gitea SQLite→PostgreSQL migration (2026-03-28) cost nearly a full day of productivity:
|
|
|
|
- **Hours of accumulated unresponsiveness** across multiple projects before root cause was identified
|
|
- **Planning and implementation** of the migration itself
|
|
- **Migration complexity** that didn't need to exist: Gitea 1.23 has no `restore` command, `doctor convert` only handles charset conversion, `docker cp` corrupted PostgreSQL directory permissions, SSH authorized_keys weren't regenerated
|
|
- **Downstream impact** on ArgoCD (20 apps polling a locked database), CI runners (continuous 500 errors), container registry pulls (timeouts)
|
|
|
|
The PostgreSQL container takes 5 minutes to add to a Docker Compose stack at initial setup time. The migration took a day. Always pay the 5 minutes upfront.
|
|
|
|
## When SQLite Is Acceptable
|
|
|
|
SQLite is fine for:
|
|
- Local development databases (single developer, single process)
|
|
- Embedded application data stores (mobile apps, desktop apps, CLI tools)
|
|
- Read-heavy workloads with rare writes and a single writer process
|
|
- Test fixtures and throwaway data
|
|
- Configuration stores read at startup (not at request time)
|
|
|
|
## Implementation Pattern
|
|
|
|
For Docker Compose services that need a database:
|
|
|
|
```yaml
|
|
services:
|
|
postgres:
|
|
image: postgres:17-alpine
|
|
restart: unless-stopped
|
|
environment:
|
|
POSTGRES_DB: myapp
|
|
POSTGRES_USER: myapp
|
|
POSTGRES_PASSWORD: {{ db_password }}
|
|
volumes:
|
|
- /opt/postgres-myapp:/var/lib/postgresql/data
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U myapp -d myapp"]
|
|
interval: 10s
|
|
timeout: 5s
|
|
retries: 5
|
|
start_period: 30s
|
|
networks:
|
|
- app_internal
|
|
deploy:
|
|
resources:
|
|
limits:
|
|
memory: 1G
|
|
|
|
myapp:
|
|
depends_on:
|
|
postgres:
|
|
condition: service_healthy
|
|
networks:
|
|
- app_internal
|
|
- external_network
|
|
|
|
networks:
|
|
app_internal:
|
|
driver: bridge
|
|
internal: true
|
|
```
|
|
|
|
Key points:
|
|
- PostgreSQL on an **internal bridge network** (no external access needed)
|
|
- Application **depends on PostgreSQL health** before starting
|
|
- **Resource limits** to prevent runaway memory usage
|
|
- **Separate data directory** per application (`/opt/postgres-myapp`, not shared)
|
|
- PostgreSQL container UID is **999** (not 1000) — set directory ownership accordingly
|
|
|
|
## For Kubernetes Deployments
|
|
|
|
Use the application's Helm chart PostgreSQL subchart, or deploy a standalone PostgreSQL instance:
|
|
- Bitnami PostgreSQL Helm chart for simple deployments
|
|
- CloudNativePG operator for production-grade PostgreSQL with HA, backups, and failover
|
|
- Never use SQLite with `emptyDir` or even PVC-backed volumes in multi-replica deployments
|
|
|
|
## Checklist for New Service Deployment
|
|
|
|
Before deploying any new service, check:
|
|
|
|
1. What database does the default configuration use?
|
|
2. If SQLite: does the service support PostgreSQL? (Almost all do — Gitea, Authelia, Headscale, Zulip, etc.)
|
|
3. Switch to PostgreSQL **before the first deployment**, not after problems appear
|
|
4. Add the database password to SOPS-encrypted secrets
|
|
5. Verify the database connection works before adding consumers
|
|
|
|
## `ON CONFLICT DO NOTHING` Requires a Real Unique Constraint
|
|
|
|
SQLAlchemy `on_conflict_do_nothing()` and raw `ON CONFLICT DO NOTHING` only work when there is a matching unique constraint or unique index. Without one, the statement either silently does nothing or inserts a duplicate, depending on the exact phrasing.
|
|
|
|
**Trap:** "URL" looks like a natural unique key for a crawl/ingest table, so the LLM or developer adds `UNIQUE(url)` to make the upsert work. But URL content changes over time — two rows for the same URL at different timestamps are semantically distinct. The fake unique constraint then corrupts the model (or blocks legitimate re-ingestion).
|
|
|
|
**Rule:**
|
|
|
|
- If the natural key is truly unique (user ID, slug, message hash), add the unique constraint and use `on_conflict_do_nothing()`
|
|
- If the natural key is *not* unique over time (URL, title, filename), use **query-before-insert** in a transaction, not a fake unique constraint
|
|
- Don't invent unique constraints to make `ON CONFLICT` work — you're encoding a false invariant into the schema
|
|
|
|
## Alembic Multi-Schema Migrations
|
|
|
|
To run Alembic against multiple Postgres schemas in one database:
|
|
|
|
1. Set `include_schemas=True` in `env.py` so autogenerate sees non-default schemas
|
|
2. Add an `include_name` filter so autogenerate only tracks the schemas you manage (otherwise it tries to "fix" `information_schema`, `pg_catalog`, etc.)
|
|
3. **Issue `CREATE SCHEMA IF NOT EXISTS <name>`** before `run_migrations()` — otherwise the first migration fails on a missing schema
|
|
|
|
```python
|
|
def include_name(name, type_, parent_names):
|
|
if type_ == "schema":
|
|
return name in {"app", "audit", "reporting"}
|
|
return True
|
|
|
|
def run_migrations_online():
|
|
connectable = engine_from_config(...)
|
|
with connectable.connect() as connection:
|
|
for schema in ("app", "audit", "reporting"):
|
|
connection.execute(text(f'CREATE SCHEMA IF NOT EXISTS "{schema}"'))
|
|
context.configure(
|
|
connection=connection,
|
|
include_schemas=True,
|
|
include_name=include_name,
|
|
...
|
|
)
|
|
with context.begin_transaction():
|
|
context.run_migrations()
|
|
```
|
|
|
|
Also set `version_table_schema` on `context.configure` if the Alembic version table should live in a specific schema rather than `public`.
|