Add API design and LLM code security best practices
Two new topic files from research: - api-design.md: Transport security, OAuth2/JWT/mTLS auth, API patterns (versioning, pagination, idempotency, rate limiting), input validation, secrets handling, zero-trust service mesh patterns. Maps to OWASP API Security Top 10. - llm-code-security.md: Common vulnerabilities in LLM-generated code (injection, hardcoded secrets, hallucinated packages, over-permissive defaults, IaC risks, crypto mistakes). Includes per-technology review checklists and cites 18 research sources (2024-2026). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -23,4 +23,6 @@ Generalised best practices extracted from real project work via the `/distill-be
|
||||
- [Docker UID Matching](docker-uid-matching.md) — UID wrapper entrypoint for mounted volumes, gosu pattern, when to use vs K8s securityContext
|
||||
- [Database Selection](database-selection.md) — SQLite is not a production database; always use PostgreSQL for services with FQDNs, multiple consumers, or concurrent access
|
||||
- [Docker](docker.md) — gosu PID 1, GIT_SSH_COMMAND scope, slim image health checks, buildx local images, default users, TTY flags, UID resolution
|
||||
- [API Design](api-design.md) — Transport security, auth (OAuth2/JWT/mTLS), versioning, pagination, error handling, idempotency, rate limiting, input validation, zero-trust patterns
|
||||
- [Octopus Process Templates](octopus-process-templates.md) — OCL syntax, step template references, channel scoping, parameters, versioning, Platform Hub patterns
|
||||
- [LLM Code Security](llm-code-security.md) — Security vulnerabilities in AI-generated code: injection flaws, hardcoded secrets, hallucinated packages, over-permissive defaults, IaC risks, crypto mistakes, review checklist
|
||||
|
||||
463
api-design.md
Normal file
463
api-design.md
Normal file
@@ -0,0 +1,463 @@
|
||||
# API Design
|
||||
|
||||
Best practices for REST/HTTP APIs in internal microservices and platform services. Focused on practical defaults -- not aspirational ideals. Sourced from OWASP API Security Top 10 (2023), RFC 9700 (OAuth 2.0 Security BCP, January 2025), Google AIP, and production experience.
|
||||
|
||||
Cross-references: [Security Architecture](security-architecture.md) covers the server boundary rule and proxy patterns. [Secrets Management](secrets-management.md) covers credential storage and rotation.
|
||||
|
||||
---
|
||||
|
||||
## 1. Transport Security
|
||||
|
||||
### 1.1 HTTPS everywhere, no exceptions
|
||||
|
||||
**Principle:** Every API endpoint -- internal or external -- must serve over TLS. Plaintext HTTP must not be available, even on internal networks.
|
||||
|
||||
**Why it matters:** Without TLS, any network hop (load balancer, sidecar, switch) can observe or modify traffic. Internal networks are not trusted in a zero-trust model -- a compromised pod can sniff adjacent traffic.
|
||||
|
||||
**How to implement:**
|
||||
- Terminate TLS at the ingress controller (e.g., Traefik, NGINX) with certificates from cert-manager / Let's Encrypt.
|
||||
- For service-to-service within the cluster, use a service mesh (Istio, Linkerd) or cert-manager CSI driver to issue per-pod certificates.
|
||||
- Set `Strict-Transport-Security` headers on all responses.
|
||||
- Redirect HTTP to HTTPS at the ingress layer.
|
||||
|
||||
**Anti-patterns:**
|
||||
- "Internal traffic doesn't need encryption" -- it does under zero-trust.
|
||||
- Self-signed certificates with verification disabled (`--insecure`, `verify=False`) -- defeats the purpose of TLS.
|
||||
- Long-lived certificates (years) with no rotation -- use short-lived certs (days to weeks) with automated renewal.
|
||||
|
||||
### 1.2 mTLS between services
|
||||
|
||||
**Principle:** Service-to-service communication must use mutual TLS -- both sides present and verify certificates.
|
||||
|
||||
**Why it matters:** Server-only TLS authenticates the server to the client, but any client can connect. mTLS ensures both parties have a cryptographically verified identity, which is the foundation of zero-trust networking.
|
||||
|
||||
**How to implement:**
|
||||
- Service mesh (Istio strict mode, Linkerd) handles mTLS transparently via sidecar proxies -- no application code changes.
|
||||
- Use SPIFFE/SPIRE for standardized workload identity (SVID certificates).
|
||||
- Default certificate lifetime should be short (24 hours) with automatic rotation.
|
||||
- Start in permissive mode (allow both plain and mTLS), migrate to strict mode once all services are enrolled.
|
||||
|
||||
**Anti-patterns:**
|
||||
- Permissive mode as a permanent state -- it must be a migration step, not the end state.
|
||||
- Disabling mTLS verification for "debugging" and forgetting to re-enable it.
|
||||
- Using a single shared certificate for all services -- each workload needs its own identity.
|
||||
|
||||
### 1.3 Certificate management
|
||||
|
||||
**Principle:** Certificate issuance and rotation must be fully automated. No manual certificate management in production.
|
||||
|
||||
**Why it matters:** Manual certificate management leads to expired certificates, which cause outages. It also leads to long-lived certificates, which increase blast radius if compromised.
|
||||
|
||||
**How to implement:**
|
||||
- cert-manager in Kubernetes with ClusterIssuer for ingress certificates.
|
||||
- Service mesh control plane for workload certificates (Istio Citadel, Linkerd identity).
|
||||
- Monitor certificate expiry with alerts at 30/14/7 days before expiry.
|
||||
- Store CA keys in HSM or sealed secrets -- never in plaintext ConfigMaps.
|
||||
|
||||
**Anti-patterns:**
|
||||
- Certificates stored in Git repos (even encrypted, they need rotation).
|
||||
- Wildcard certificates shared across trust boundaries.
|
||||
- No monitoring for certificate expiry -- silent failures at 3am.
|
||||
|
||||
---
|
||||
|
||||
## 2. Authentication and Authorization
|
||||
|
||||
### 2.1 OIDC/OAuth2 for user-facing APIs (RFC 9700)
|
||||
|
||||
**Principle:** Use OAuth 2.0 Authorization Code flow with PKCE for all client types. The implicit flow and resource owner password credentials flow are deprecated per RFC 9700 (January 2025).
|
||||
|
||||
**Why it matters:** The implicit flow exposes access tokens in URLs and browser history. The password grant requires users to share credentials directly with the client, bypassing centralized identity providers.
|
||||
|
||||
**How to implement:**
|
||||
- Authorization Code + PKCE for all clients (web, mobile, CLI). PKCE is now mandatory for all client types, not just public clients.
|
||||
- Use `S256` challenge method (not `plain`).
|
||||
- Tokens issued by the authorization server, validated by the resource server.
|
||||
- Use Authorization Server Metadata (RFC 8414) for automatic discovery of endpoints and supported features.
|
||||
|
||||
**Anti-patterns:**
|
||||
- Implicit flow (`response_type=token`) -- deprecated by RFC 9700.
|
||||
- Resource Owner Password Credentials flow -- deprecated by RFC 9700.
|
||||
- Storing tokens in localStorage (accessible to XSS) -- use httpOnly cookies or in-memory storage with refresh token rotation.
|
||||
- Long-lived access tokens without refresh -- use short-lived access tokens (5-15 minutes) with refresh token rotation.
|
||||
|
||||
### 2.2 JWT best practices
|
||||
|
||||
**Principle:** JWTs must be validated completely on every request -- signature, expiry, issuer, audience, and algorithm.
|
||||
|
||||
**Why it matters:** Incomplete JWT validation is a top attack vector. Accepting expired tokens, wrong audiences, or `alg: none` enables token forgery and replay.
|
||||
|
||||
**How to implement:**
|
||||
- Validate: signature (asymmetric preferred -- RS256/ES256), `exp`, `iat`, `iss`, `aud`, `nbf`.
|
||||
- Use asymmetric signing (RS256/ES256) so that only the auth server holds the private key. Resource servers only need the public key.
|
||||
- Set `aud` claim to the specific API audience -- reject tokens intended for other services.
|
||||
- Keep tokens small -- put only identity and authorization claims in the token, fetch additional data from a userinfo endpoint.
|
||||
- Use `jti` (JWT ID) claim for token revocation checks when needed.
|
||||
|
||||
**Anti-patterns:**
|
||||
- Accepting `alg: none` or allowing algorithm switching -- pin the expected algorithm server-side.
|
||||
- Not validating `aud` -- allows tokens from one service to be replayed against another.
|
||||
- Symmetric signing (HS256) with a shared secret across services -- if one service is compromised, all are.
|
||||
- Treating JWTs as sessions -- JWTs are not revocable by default. Combine with short expiry and token introspection for revocation.
|
||||
|
||||
### 2.3 Service-to-service authentication
|
||||
|
||||
**Principle:** Services authenticate to each other using mTLS identities (SPIFFE) or short-lived JWTs from a token exchange. Never shared static API keys.
|
||||
|
||||
**Why it matters:** Shared API keys have no expiry, no rotation path, no per-service identity, and no audit trail. If one service is compromised, the key works for everything.
|
||||
|
||||
**How to implement:**
|
||||
- **Preferred: mTLS with SPIFFE.** The service mesh provides identity automatically. Authorization policies reference SPIFFE IDs (e.g., `spiffe://cluster.local/ns/payments/sa/payment-svc`).
|
||||
- **Alternative: OAuth2 Client Credentials flow.** Each service has its own `client_id` and `client_secret` (or asymmetric key pair). Tokens are short-lived and scoped to specific audiences.
|
||||
- Use asymmetric client authentication (private_key_jwt per RFC 7523) rather than client secrets where possible.
|
||||
- Implement audience restriction -- tokens minted for service A must not be accepted by service B.
|
||||
|
||||
**Anti-patterns:**
|
||||
- Shared static API keys passed in headers or query strings.
|
||||
- One "admin" service account used by all services.
|
||||
- Service-to-service tokens with no audience claim -- replayable across any internal API.
|
||||
- Bearer tokens without mTLS -- if the network is compromised, the token can be stolen and replayed from anywhere.
|
||||
|
||||
### 2.4 Authorization: object-level and function-level
|
||||
|
||||
**Principle:** Check authorization at every API endpoint, for every object access, based on the authenticated identity. Never rely on "the client won't send that request."
|
||||
|
||||
**Why it matters:** Broken Object-Level Authorization (BOLA) is the #1 risk in the OWASP API Security Top 10. Broken Function-Level Authorization is #5. These are the most common API vulnerabilities found in penetration tests.
|
||||
|
||||
**How to implement:**
|
||||
- Every endpoint that accesses a specific resource must verify the caller owns or has access to that resource.
|
||||
- Use middleware/decorators that enforce authorization before the handler runs.
|
||||
- Use random UUIDs for resource identifiers, not sequential integers (which are trivially enumerable).
|
||||
- Separate authorization for data access (BOLA) and function access (admin endpoints, bulk operations).
|
||||
- Automated tests that verify: user A cannot access user B's resources, non-admin cannot call admin endpoints.
|
||||
|
||||
**Anti-patterns:**
|
||||
- Authorization only at the API gateway -- must also be enforced at the service level.
|
||||
- Relying on obscurity of endpoint URLs for access control.
|
||||
- Sequential/predictable resource IDs without authorization checks.
|
||||
- Missing authorization on secondary endpoints (e.g., `/users/{id}/orders` checks user but not order ownership).
|
||||
|
||||
---
|
||||
|
||||
## 3. API Design Patterns
|
||||
|
||||
### 3.1 Versioning
|
||||
|
||||
**Principle:** Version your API from day one using URL path versioning (`/v1/`). Support at most two versions simultaneously.
|
||||
|
||||
**Why it matters:** Breaking changes without versioning cause cascading failures across all consumers simultaneously. Supporting too many versions creates maintenance burden and security risk (old versions may lack patches).
|
||||
|
||||
**How to implement:**
|
||||
- URL path: `/api/v1/resources` -- simple, visible, cacheable.
|
||||
- Deprecation policy: announce deprecation in response headers (`Deprecation: true`, `Sunset: <date>`).
|
||||
- Maximum two active versions. When v3 launches, v1 is removed.
|
||||
- Internal services can use header-based versioning (`Accept: application/vnd.myapi.v2+json`) if URL versioning is too rigid for rapid iteration.
|
||||
|
||||
**Anti-patterns:**
|
||||
- No versioning ("we'll be careful") -- you will break consumers.
|
||||
- Unlimited version support -- v1 through v7 all still running, each with different bugs.
|
||||
- Breaking changes in a patch version.
|
||||
- Versioning individual endpoints instead of the whole API surface.
|
||||
|
||||
### 3.2 Pagination
|
||||
|
||||
**Principle:** All list endpoints must paginate. Use cursor-based pagination for real-time data; offset-based for stable datasets.
|
||||
|
||||
**Why it matters:** Unbounded list responses cause memory exhaustion, slow responses, and database strain. Large offset values cause full table scans.
|
||||
|
||||
**How to implement:**
|
||||
- **Cursor-based (preferred):** Return an opaque `next_cursor` token. Client passes it to get the next page. Stable under concurrent writes.
|
||||
```json
|
||||
{ "data": [...], "next_cursor": "abc123", "has_more": true }
|
||||
```
|
||||
- **Offset-based (simple datasets):** `?limit=50&offset=100`. Acceptable for admin dashboards or infrequently changing data.
|
||||
- Set a maximum page size (e.g., 100) enforced server-side. Ignore client requests for larger pages.
|
||||
- Always return pagination metadata (`next_cursor`, `has_more`, or `total_count` if cheap to compute).
|
||||
|
||||
**Anti-patterns:**
|
||||
- No pagination on list endpoints -- returns 50,000 records in one response.
|
||||
- Offset-based pagination on large, frequently-changing datasets -- pages shift as records are inserted/deleted.
|
||||
- Client-controlled page size with no server-side maximum.
|
||||
- `total_count` requiring a full table scan on every request -- make it optional or cached.
|
||||
|
||||
### 3.3 Error handling
|
||||
|
||||
**Principle:** Return structured, machine-readable errors with stable error codes, human-readable messages, and consistent shape across all endpoints.
|
||||
|
||||
**Why it matters:** Inconsistent error formats force every consumer to write custom parsing logic. Missing error codes make automated retry decisions impossible. Leaking stack traces exposes internals to attackers.
|
||||
|
||||
**How to implement:**
|
||||
- Standard error envelope:
|
||||
```json
|
||||
{
|
||||
"error": {
|
||||
"code": "RESOURCE_NOT_FOUND",
|
||||
"message": "Order 7f3a... not found",
|
||||
"details": [{ "field": "order_id", "reason": "not_found" }]
|
||||
}
|
||||
}
|
||||
```
|
||||
- Use HTTP status codes correctly: 400 (bad input), 401 (unauthenticated), 403 (unauthorized), 404 (not found), 409 (conflict), 422 (validation), 429 (rate limited), 500 (server error).
|
||||
- Error codes are stable strings (not integers) that consumers can switch on.
|
||||
- Never expose stack traces, SQL errors, or internal paths in error responses.
|
||||
- Log the full error server-side with a correlation ID. Return only the correlation ID to the client.
|
||||
|
||||
**Anti-patterns:**
|
||||
- 200 OK with `{"success": false}` -- use HTTP status codes.
|
||||
- Returning raw database errors ("duplicate key violates unique constraint on...").
|
||||
- Different error shapes from different endpoints in the same API.
|
||||
- Generic "Internal Server Error" with no correlation ID -- impossible to debug.
|
||||
|
||||
### 3.4 Idempotency
|
||||
|
||||
**Principle:** All state-changing operations must be safe to retry. Use idempotency keys for POST requests; PUT and DELETE are idempotent by definition.
|
||||
|
||||
**Why it matters:** Network failures, timeouts, and retries are normal in distributed systems. Without idempotency, retried requests create duplicate orders, double payments, or inconsistent state.
|
||||
|
||||
**How to implement:**
|
||||
- Accept `Idempotency-Key` header (IETF draft: draft-ietf-httpapi-idempotency-key-header) on POST endpoints.
|
||||
- Server stores the response for a given key (TTL 24-48 hours). Duplicate requests return the stored response.
|
||||
- Use UUIDv4 for idempotency keys -- never sequential or timestamp-based (predictable/guessable).
|
||||
- Handle concurrent duplicate requests with locking: first request processes, subsequent requests wait then return cached response.
|
||||
- PUT must be truly idempotent: same request, same result, no side effects on repeat.
|
||||
|
||||
**Anti-patterns:**
|
||||
- POST endpoints with no idempotency support -- every retry creates a duplicate.
|
||||
- Idempotency keys stored forever (memory leak) or for too short a period (retries after expiry create duplicates).
|
||||
- Client-generated sequential keys (integers, timestamps) -- guessable and exploitable.
|
||||
- "Idempotent" endpoints that still send duplicate emails/webhooks on retry.
|
||||
|
||||
### 3.5 Rate limiting
|
||||
|
||||
**Principle:** Every API must enforce rate limits. Return standard headers so clients can self-throttle.
|
||||
|
||||
**Why it matters:** Without rate limits, a single misbehaving client (or attacker) can exhaust resources for all consumers. Rate limits also protect downstream dependencies.
|
||||
|
||||
**How to implement:**
|
||||
- Token bucket or sliding window algorithm (token bucket is simplest with good burst handling).
|
||||
- Return headers: `X-RateLimit-Limit`, `X-RateLimit-Remaining`, `X-RateLimit-Reset` (IETF draft still pending; `X-` prefix remains de facto standard).
|
||||
- Return `429 Too Many Requests` with `Retry-After` header (RFC 6585).
|
||||
- Rate limit checks execute before expensive operations (auth, database queries).
|
||||
- For distributed deployments, use Redis with atomic Lua scripts for counter operations -- avoid race conditions.
|
||||
- Different tiers for different consumers (internal services get higher limits than external clients).
|
||||
|
||||
**Anti-patterns:**
|
||||
- No rate limiting ("it's an internal API") -- a runaway loop in one service takes down the whole platform.
|
||||
- Rate limiting after expensive operations (database query runs, then rate limit rejects the response).
|
||||
- No `Retry-After` header -- clients retry immediately in a tight loop, making the problem worse.
|
||||
- Per-IP rate limiting only -- bypassed by distributed clients, unfair to NAT'd users.
|
||||
|
||||
---
|
||||
|
||||
## 4. Input Validation
|
||||
|
||||
### 4.1 Schema validation at the edge
|
||||
|
||||
**Principle:** Validate all request bodies against a schema (OpenAPI/JSON Schema) at the API gateway or middleware layer. Reject requests that don't conform before they reach business logic.
|
||||
|
||||
**Why it matters:** Invalid input that reaches business logic causes unpredictable behavior -- crashes, data corruption, injection attacks. Edge validation is the first line of defense.
|
||||
|
||||
**How to implement:**
|
||||
- Define request/response schemas in OpenAPI 3.x. Generate validation middleware from the spec.
|
||||
- Reject unknown fields (additionalProperties: false) -- attackers probe via unexpected fields.
|
||||
- Enforce type constraints: string lengths, integer ranges, enum values, date formats.
|
||||
- Validate `Content-Type` header -- reject requests with unexpected content types (e.g., reject `multipart/form-data` on a JSON-only endpoint).
|
||||
|
||||
**Anti-patterns:**
|
||||
- Validation only in business logic, not at the edge -- invalid data traverses the full call stack before rejection.
|
||||
- Accepting and silently ignoring unknown fields -- hides bugs and enables mass assignment attacks.
|
||||
- Validating types but not ranges -- accepting an `age` field of 99999 or -1.
|
||||
- No schema at all ("we'll validate manually") -- inconsistent validation across endpoints.
|
||||
|
||||
### 4.2 Injection prevention
|
||||
|
||||
**Principle:** Use parameterized queries for all database access. Never concatenate user input into queries, commands, or templates.
|
||||
|
||||
**Why it matters:** SQL injection remains in the OWASP Top 10 after 20+ years. NoSQL injection, LDAP injection, and command injection follow the same pattern -- unsanitized input in a query language.
|
||||
|
||||
**How to implement:**
|
||||
- Use an ORM (SQLAlchemy, Prisma, TypeORM) or parameterized queries. All major ORMs parameterize by default.
|
||||
- For raw SQL (performance-critical paths), use prepared statements exclusively.
|
||||
- Validate input with allowlists, not denylists. If a field should be a UUID, validate it's a UUID -- don't try to strip "malicious characters."
|
||||
- For template rendering, use auto-escaping (Jinja2 autoescape, React JSX auto-escaping).
|
||||
|
||||
**Anti-patterns:**
|
||||
- String concatenation in SQL: `f"SELECT * FROM users WHERE id = '{user_input}'"`.
|
||||
- Denylisting dangerous characters instead of allowlisting valid patterns.
|
||||
- Trusting input from "internal" services -- a compromised upstream service sends malicious data.
|
||||
- Disabling ORM parameterization for "performance" without understanding the security cost.
|
||||
|
||||
### 4.3 Request size and depth limits
|
||||
|
||||
**Principle:** Enforce maximum request body size, JSON nesting depth, and array length at the gateway level.
|
||||
|
||||
**Why it matters:** Deeply nested JSON or extremely large payloads cause CPU exhaustion during parsing (hash collision attacks, recursive descent parsers). This is a denial-of-service vector.
|
||||
|
||||
**How to implement:**
|
||||
- Set maximum body size at the reverse proxy/ingress (e.g., `client_max_body_size 1m` in NGINX).
|
||||
- Limit JSON nesting depth (8-16 levels is generous for any real use case).
|
||||
- Limit array sizes in request bodies (e.g., batch endpoints accept max 100 items).
|
||||
- Set request timeouts at the gateway -- don't let slow clients hold connections open.
|
||||
|
||||
**Anti-patterns:**
|
||||
- No body size limit -- 100MB JSON payload parsed by every middleware layer.
|
||||
- Accepting arbitrarily nested JSON -- `{"a":{"a":{"a":...}}}` 1000 levels deep.
|
||||
- Batch endpoints with no limit -- client sends 1 million items in one request.
|
||||
|
||||
---
|
||||
|
||||
## 5. Secrets in APIs
|
||||
|
||||
### 5.1 Never in URLs or query parameters
|
||||
|
||||
**Principle:** Authentication tokens, API keys, and any secret material must be sent in headers (Authorization, custom headers) or request bodies. Never in URLs or query parameters.
|
||||
|
||||
**Why it matters:** URLs are logged everywhere -- web server access logs, proxy logs, browser history, referrer headers, CDN logs, monitoring tools. A token in a URL is a token in every log file in the request path.
|
||||
|
||||
**How to implement:**
|
||||
- Use `Authorization: Bearer <token>` header for all token-based auth.
|
||||
- For webhook signatures, use a signature header (e.g., `X-Hub-Signature-256`).
|
||||
- If an API currently accepts tokens in query params, deprecate that path and migrate to header-based auth.
|
||||
- Configure log scrubbing to redact Authorization headers, but don't rely on it as the primary control.
|
||||
|
||||
**Anti-patterns:**
|
||||
- `GET /api/resources?api_key=sk_live_abc123` -- key in every access log.
|
||||
- OAuth redirect URIs with tokens in query params (use `response_mode=form_post` or authorization code flow).
|
||||
- Webhook URLs with embedded secrets (`/webhook?secret=abc`) -- logged, cached, shared.
|
||||
|
||||
### 5.2 Token rotation and expiry
|
||||
|
||||
**Principle:** All tokens and API keys must have expiry dates and a documented rotation procedure. No permanent credentials.
|
||||
|
||||
**Why it matters:** Leaked tokens without expiry are valid forever. Rotation limits blast radius -- even if a token is compromised, it expires soon.
|
||||
|
||||
**How to implement:**
|
||||
- Access tokens: 5-15 minute expiry, refreshed via refresh token.
|
||||
- Refresh tokens: rotate on use (each refresh issues a new refresh token and invalidates the old one).
|
||||
- API keys for external integrations: 90-day rotation policy with overlap period (new key valid before old key expires).
|
||||
- Service account tokens (OAuth2 client credentials): short-lived (1 hour), fetched on demand.
|
||||
- Track expiry dates in a credential inventory. Alert before expiry (see [Secrets Management](secrets-management.md) -- Credential Lifecycle Management).
|
||||
|
||||
**Anti-patterns:**
|
||||
- API keys that never expire ("we'll rotate them when we need to" -- you won't).
|
||||
- Refresh tokens that don't rotate -- stolen refresh token provides permanent access.
|
||||
- No overlap period during rotation -- brief outage while all consumers update.
|
||||
- Hardcoded tokens in application config deployed via CI -- rotation requires a full redeploy.
|
||||
|
||||
### 5.3 No secrets in logs or error responses
|
||||
|
||||
**Principle:** Scrub all secrets from logs, error responses, and monitoring data. Structured logging with explicit field selection is safer than serializing request objects.
|
||||
|
||||
**Why it matters:** Log aggregation systems (ELK, Loki, Datadog) are often accessible to broader teams than production systems. A token in a log entry has a much wider exposure surface than a token in a running process.
|
||||
|
||||
**How to implement:**
|
||||
- Use structured logging. Log specific fields, not entire request objects.
|
||||
- Redact `Authorization` headers and any field matching `token`, `password`, `secret`, `key` patterns in log middleware.
|
||||
- Never log request bodies for auth endpoints (login, token exchange).
|
||||
- Error responses must not include internal state -- return a correlation ID and log details server-side.
|
||||
|
||||
**Anti-patterns:**
|
||||
- `logger.info(f"Request: {request.headers}")` -- logs all headers including Authorization.
|
||||
- Error responses that include the original request (including auth headers) for "debugging convenience."
|
||||
- Logging full webhook payloads that contain signing secrets in custom headers.
|
||||
|
||||
---
|
||||
|
||||
## 6. Service Mesh and Zero Trust
|
||||
|
||||
### 6.1 Default deny with explicit allow
|
||||
|
||||
**Principle:** Network policies and authorization policies must default to deny-all. Every allowed communication path is explicitly defined.
|
||||
|
||||
**Why it matters:** Default-allow means a compromised service can reach every other service in the cluster. Default-deny contains the blast radius to only the services the compromised workload was authorized to reach.
|
||||
|
||||
**How to implement:**
|
||||
- Kubernetes NetworkPolicy: deploy a default-deny policy in every namespace, then add specific allow rules.
|
||||
```yaml
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
name: default-deny-all
|
||||
spec:
|
||||
podSelector: {}
|
||||
policyTypes: [Ingress, Egress]
|
||||
```
|
||||
- Service mesh authorization policies: deny by default, allow specific source-to-destination pairs by SPIFFE ID.
|
||||
- Audit policies periodically -- remove rules for decommissioned services.
|
||||
|
||||
**Anti-patterns:**
|
||||
- No network policies ("everything's in the cluster, it's fine").
|
||||
- Overly broad allow rules (`allow all from namespace X`) -- defeats the purpose.
|
||||
- Network policies without egress rules -- ingress-only policies still allow compromised pods to exfiltrate data.
|
||||
|
||||
### 6.2 Least-privilege service identities
|
||||
|
||||
**Principle:** Each service gets its own identity (Kubernetes ServiceAccount + SPIFFE SVID) with the minimum permissions needed. No shared service accounts.
|
||||
|
||||
**Why it matters:** Shared identities prevent granular authorization, audit trails, and revocation. If services A and B share an identity, you cannot authorize A without also authorizing B.
|
||||
|
||||
**How to implement:**
|
||||
- One Kubernetes ServiceAccount per workload (not per namespace).
|
||||
- RBAC bindings scoped to exactly what the service needs (specific API groups, resources, verbs).
|
||||
- Authorization policies reference specific service identities: "payment-svc can call order-svc on POST /orders/{id}/payment."
|
||||
- Regularly audit which identities have access to which services -- prune unused access.
|
||||
|
||||
**Anti-patterns:**
|
||||
- Default ServiceAccount used by all pods in a namespace.
|
||||
- Cluster-wide RBAC bindings for convenience.
|
||||
- Service identities with wildcard permissions ("allow all methods on all paths").
|
||||
- No audit of identity-to-service mappings.
|
||||
|
||||
### 6.3 Observability as a security control
|
||||
|
||||
**Principle:** Distributed tracing, access logs, and metrics from the service mesh are security controls, not just debugging tools. Monitor them for anomalies.
|
||||
|
||||
**Why it matters:** Zero trust assumes breach. Detection depends on visibility. If you can't see who called what, you can't detect lateral movement.
|
||||
|
||||
**How to implement:**
|
||||
- Enable access logging in the service mesh (Istio/Envoy access logs, Linkerd tap).
|
||||
- Distributed tracing (OpenTelemetry, Jaeger) with trace context propagated across all service calls.
|
||||
- Alert on anomalies: unexpected source-destination pairs, unusual request volumes, authorization denials.
|
||||
- Retain access logs long enough for incident investigation (30-90 days minimum).
|
||||
|
||||
**Anti-patterns:**
|
||||
- Disabling access logging for performance -- sample instead of disabling entirely.
|
||||
- Tracing only in development, not production.
|
||||
- No alerting on authorization policy denials -- failed access attempts are the signal.
|
||||
|
||||
---
|
||||
|
||||
## OWASP API Security Top 10 (2023) Quick Reference
|
||||
|
||||
For context, the current OWASP API Security Top 10 maps to the practices above:
|
||||
|
||||
| # | Risk | Where addressed |
|
||||
|---|------|----------------|
|
||||
| API1 | Broken Object-Level Authorization | Section 2.4 |
|
||||
| API2 | Broken Authentication | Sections 2.1, 2.2, 2.3 |
|
||||
| API3 | Broken Object Property-Level Authorization | Section 2.4, 4.1 |
|
||||
| API4 | Unrestricted Resource Consumption | Sections 3.5, 4.3 |
|
||||
| API5 | Broken Function-Level Authorization | Section 2.4 |
|
||||
| API6 | Unrestricted Access to Sensitive Business Flows | Sections 3.4, 3.5 |
|
||||
| API7 | Server-Side Request Forgery | Section 4.2 |
|
||||
| API8 | Security Misconfiguration | Sections 1.1, 6.1 |
|
||||
| API9 | Improper Inventory Management | Section 3.1 |
|
||||
| API10 | Unsafe Consumption of APIs | Section 4.2 |
|
||||
|
||||
---
|
||||
|
||||
## Sources
|
||||
|
||||
- [OWASP API Security Top 10](https://owasp.org/API-Security/)
|
||||
- [RFC 9700 - OAuth 2.0 Security Best Current Practice (January 2025)](https://datatracker.ietf.org/doc/rfc9700/)
|
||||
- [OAuth best practices: RFC 9700 summary -- WorkOS](https://workos.com/blog/oauth-best-practices)
|
||||
- [IETF Idempotency-Key Header Draft](https://datatracker.ietf.org/doc/draft-ietf-httpapi-idempotency-key-header/)
|
||||
- [Google AIP-193: Errors](https://google.aip.dev/193)
|
||||
- [ByteByteGo: REST API Design](https://blog.bytebytego.com/p/the-art-of-rest-api-design-idempotency)
|
||||
- [Zuplo: Rate Limiting Best Practices](https://zuplo.com/learning-center/10-best-practices-for-api-rate-limiting-in-2025)
|
||||
- [Zuplo: Input/Output Validation](https://zuplo.com/blog/2025/03/25/input-output-validation-best-practices)
|
||||
- [OWASP Input Validation Cheat Sheet](https://cheatsheetseries.owasp.org/cheatsheets/Input_Validation_Cheat_Sheet.html)
|
||||
- [Machine Identity: mTLS + SPIFFE Zero Trust Guide](https://petronellatech.com/blog/machine-identity-is-the-new-perimeter-mtls-spiffe-for-zero-trust/)
|
||||
- [Buoyant: Zero Trust, mTLS, and the Service Mesh](https://www.buoyant.io/blog/zero-trust-mtls-and-the-service-mesh-explained)
|
||||
- [Kong: Zero Trust with Service Mesh](https://konghq.com/blog/engineering/zero-trust-service-mesh-security)
|
||||
- [Microsoft Azure: Web API Design Best Practices](https://learn.microsoft.com/en-us/azure/architecture/best-practices/api-design)
|
||||
800
llm-code-security.md
Normal file
800
llm-code-security.md
Normal file
@@ -0,0 +1,800 @@
|
||||
# Security of LLM-Generated Code
|
||||
|
||||
Practical guide to security vulnerabilities commonly introduced by LLMs (Claude, GPT-4, Copilot) when generating Python, shell scripts, Kubernetes manifests, and Helm charts. Based on published research from 2024-2026.
|
||||
|
||||
## Key Statistics
|
||||
|
||||
- 25-75% of AI-generated code contains security vulnerabilities depending on language, model, and prompting (Endor Labs, multiple academic studies)
|
||||
- 29.5% of Copilot-generated Python snippets and 24.2% of JavaScript snippets contain security weaknesses across 43 CWE categories (ACM study, 2024)
|
||||
- 19.7% of LLM-suggested packages are hallucinations -- non-existent package names (slopsquatting study, 576,000 code samples across 16 models)
|
||||
- 80% of AI-suggested dependencies contain known risks (Endor Labs 2025 State of Dependency Management Report)
|
||||
- Repositories with Copilot active show 6.4% secret leakage rate, 40% higher than the 4.6% baseline across public repos
|
||||
|
||||
---
|
||||
|
||||
## 1. OWASP Top 10 in LLM-Generated Code
|
||||
|
||||
Missing input sanitization is the single most common security flaw in LLM-generated code across all languages and models. The most prevalent CWE categories are:
|
||||
|
||||
| CWE | Name | Frequency |
|
||||
|-----|------|-----------|
|
||||
| CWE-89 | SQL Injection | Very High |
|
||||
| CWE-79 | Cross-Site Scripting (XSS) | Very High |
|
||||
| CWE-78 | OS Command Injection | High |
|
||||
| CWE-22 | Path Traversal | High |
|
||||
| CWE-20 | Improper Input Validation | Very High |
|
||||
| CWE-259/798 | Hard-coded Credentials | High |
|
||||
| CWE-330 | Insufficiently Random Values | High |
|
||||
| CWE-94 | Code Injection | High |
|
||||
| CWE-120/787 | Buffer Overflow | Medium (C/C++) |
|
||||
| CWE-918 | SSRF | Medium |
|
||||
|
||||
### What LLMs get wrong
|
||||
|
||||
LLMs generate code that "works" for the happy path but omits defensive coding. They reproduce patterns from training data, which is full of tutorials and Stack Overflow snippets that skip security for brevity. The model optimises for functional correctness, not security.
|
||||
|
||||
### SQL Injection
|
||||
|
||||
**Vulnerable pattern (Python):**
|
||||
```python
|
||||
# LLM-generated: string interpolation in SQL
|
||||
def get_user(username):
|
||||
query = f"SELECT * FROM users WHERE username = '{username}'"
|
||||
cursor.execute(query)
|
||||
return cursor.fetchone()
|
||||
```
|
||||
|
||||
**Secure alternative:**
|
||||
```python
|
||||
def get_user(username):
|
||||
cursor.execute("SELECT * FROM users WHERE username = %s", (username,))
|
||||
return cursor.fetchone()
|
||||
```
|
||||
|
||||
### Command Injection
|
||||
|
||||
**Vulnerable pattern (Python):**
|
||||
```python
|
||||
import subprocess
|
||||
def ping_host(hostname):
|
||||
result = subprocess.run(f"ping -c 1 {hostname}", shell=True, capture_output=True)
|
||||
return result.stdout
|
||||
```
|
||||
|
||||
**Secure alternative:**
|
||||
```python
|
||||
import subprocess
|
||||
import shlex
|
||||
def ping_host(hostname):
|
||||
# Validate hostname format first
|
||||
if not re.match(r'^[a-zA-Z0-9._-]+$', hostname):
|
||||
raise ValueError("Invalid hostname")
|
||||
result = subprocess.run(["ping", "-c", "1", hostname], capture_output=True)
|
||||
return result.stdout
|
||||
```
|
||||
|
||||
### Command Injection (Shell Scripts)
|
||||
|
||||
**Vulnerable pattern:**
|
||||
```bash
|
||||
#!/bin/bash
|
||||
# LLM-generated: unquoted variable in command
|
||||
filename=$1
|
||||
cat $filename | grep "pattern"
|
||||
```
|
||||
|
||||
**Secure alternative:**
|
||||
```bash
|
||||
#!/bin/bash
|
||||
filename="$1"
|
||||
# Validate the path is within expected directory
|
||||
realpath_file="$(realpath -- "$filename")"
|
||||
if [[ "$realpath_file" != /expected/dir/* ]]; then
|
||||
echo "Error: path outside allowed directory" >&2
|
||||
exit 1
|
||||
fi
|
||||
grep "pattern" -- "$filename"
|
||||
```
|
||||
|
||||
### Path Traversal
|
||||
|
||||
**Vulnerable pattern (Python):**
|
||||
```python
|
||||
@app.route('/files/<path:filename>')
|
||||
def serve_file(filename):
|
||||
return send_file(os.path.join('/data', filename))
|
||||
```
|
||||
|
||||
**Secure alternative:**
|
||||
```python
|
||||
@app.route('/files/<path:filename>')
|
||||
def serve_file(filename):
|
||||
# send_from_directory validates the path stays within the directory
|
||||
return send_from_directory('/data', filename)
|
||||
```
|
||||
|
||||
### How to catch it in review
|
||||
|
||||
- Search for string formatting in SQL: `f"SELECT`, `f"INSERT`, `f"UPDATE`, `f"DELETE`, `"SELECT.*" %`, `"SELECT.*" +`
|
||||
- Search for `shell=True` in subprocess calls
|
||||
- Search for `os.path.join` with user-controlled input without path validation
|
||||
- Search for unquoted `$variables` in shell scripts
|
||||
- Use SAST tools: Bandit (Python), ShellCheck (bash), semgrep with security rulesets
|
||||
|
||||
---
|
||||
|
||||
## 2. Secrets and Credentials
|
||||
|
||||
### What LLMs get wrong
|
||||
|
||||
LLMs frequently hardcode secrets directly into generated code. This happens because training data is full of tutorials with placeholder credentials that look real, and the model replicates the pattern. CWE-259 (Hard-coded Password) and CWE-798 (Hard-coded Credentials) are among the most common LLM-generated vulnerabilities.
|
||||
|
||||
Copilot specifically has been shown to leak secrets from its training context -- researchers built algorithms that generate prompts designed to extract secrets by inducing Copilot to disclose original credentials from training data.
|
||||
|
||||
### Vulnerable patterns
|
||||
|
||||
**Hardcoded API key (Python):**
|
||||
```python
|
||||
API_KEY = "sk-proj-abc123def456..."
|
||||
client = openai.OpenAI(api_key=API_KEY)
|
||||
```
|
||||
|
||||
**Hardcoded database credentials (Python):**
|
||||
```python
|
||||
conn = psycopg2.connect(
|
||||
host="db.example.com",
|
||||
user="admin",
|
||||
password="supersecret123",
|
||||
database="production"
|
||||
)
|
||||
```
|
||||
|
||||
**Hardcoded token in shell script:**
|
||||
```bash
|
||||
curl -H "Authorization: Bearer ghp_abc123def456" https://api.github.com/repos
|
||||
```
|
||||
|
||||
**Secrets in Kubernetes manifests:**
|
||||
```yaml
|
||||
env:
|
||||
- name: DATABASE_PASSWORD
|
||||
value: "plaintext-password-here" # Not a Secret reference
|
||||
```
|
||||
|
||||
### Secure alternatives
|
||||
|
||||
**Python -- environment variables or file-based secrets:**
|
||||
```python
|
||||
import os
|
||||
API_KEY = os.environ["OPENAI_API_KEY"]
|
||||
# Or read from a mounted secret file
|
||||
with open("/run/secrets/api_key") as f:
|
||||
API_KEY = f.read().strip()
|
||||
```
|
||||
|
||||
**Shell -- read from file, never as CLI argument:**
|
||||
```bash
|
||||
# Read token from file (not visible in ps output)
|
||||
TOKEN="$(cat /path/to/secret/file)"
|
||||
curl -H "Authorization: Bearer ${TOKEN}" https://api.github.com/repos
|
||||
```
|
||||
|
||||
**Kubernetes -- reference a Secret object:**
|
||||
```yaml
|
||||
env:
|
||||
- name: DATABASE_PASSWORD
|
||||
valueFrom:
|
||||
secretKeyRef:
|
||||
name: db-credentials
|
||||
key: password
|
||||
```
|
||||
|
||||
### How to catch it in review
|
||||
|
||||
- Run `detect-secrets scan` or `gitleaks` on every commit (pre-commit hook)
|
||||
- Search for patterns: `password =`, `api_key =`, `token =`, `secret =` with string literal values
|
||||
- Search for `Bearer ` followed by a literal string in shell scripts
|
||||
- In Kubernetes manifests, search for `value:` under `env:` entries (should be `valueFrom:` for sensitive values)
|
||||
- Check that `.env` files are in `.gitignore`
|
||||
|
||||
---
|
||||
|
||||
## 3. Dependency Risks
|
||||
|
||||
### What LLMs get wrong
|
||||
|
||||
LLMs hallucinate package names at alarming rates. A study of 576,000 code samples across 16 LLMs found 19.7% of suggested packages were hallucinations. Open-source models hallucinate at 21.7%, commercial models at 5.2%. Critically, 43% of hallucinated package names appeared consistently across repeated prompts, making them predictable targets.
|
||||
|
||||
This enables **slopsquatting**: attackers register packages matching commonly hallucinated names and inject malicious code. 38% of hallucinated names were similar to real package names (not random strings), making them plausible-looking.
|
||||
|
||||
Beyond hallucination, LLMs also suggest:
|
||||
- **Outdated versions** with known CVEs (training data lag)
|
||||
- **Deprecated packages** that have been superseded
|
||||
- **Packages with known vulnerabilities** -- 80% of AI-suggested dependencies contain known risks
|
||||
|
||||
### Vulnerable patterns
|
||||
|
||||
**Hallucinated package (Python):**
|
||||
```python
|
||||
# LLM suggests a package that doesn't exist (or was registered by an attacker)
|
||||
from flask_security_utils import sanitize_input # Not a real package
|
||||
```
|
||||
|
||||
**Pinned to vulnerable version:**
|
||||
```
|
||||
# requirements.txt generated by LLM
|
||||
requests==2.25.1 # Known CVE in older versions
|
||||
pyjwt==1.7.1 # Known vulnerabilities
|
||||
```
|
||||
|
||||
**Overly broad dependency (shell):**
|
||||
```bash
|
||||
pip install cryptography # Without version pin -- could get a compromised version
|
||||
```
|
||||
|
||||
### Secure alternatives
|
||||
|
||||
- **Always verify packages exist** on PyPI/npm/etc. before using LLM-suggested imports
|
||||
- **Pin versions and verify them:**
|
||||
```
|
||||
requests==2.32.3 # Verified from PyPI, no known CVEs
|
||||
```
|
||||
- **Use lockfiles** (`pip freeze`, `poetry.lock`, `package-lock.json`) and audit them
|
||||
- **Run dependency scanners:** `pip-audit`, `npm audit`, `trivy fs .`
|
||||
|
||||
### How to catch it in review
|
||||
|
||||
- Run `pip install --dry-run` or equivalent to verify packages resolve before committing
|
||||
- Use `pip-audit` / `npm audit` / `trivy` in CI to catch known vulnerabilities
|
||||
- Compare LLM-suggested package names against registry search results
|
||||
- Be suspicious of packages with very few downloads or recent creation dates
|
||||
- Search for version pins and verify them against current stable releases
|
||||
|
||||
---
|
||||
|
||||
## 4. Over-Permissive Defaults
|
||||
|
||||
### What LLMs get wrong
|
||||
|
||||
LLMs default to the most permissive configuration because it "works" with the least friction. Training data is full of tutorials and quick-start guides that use wide-open settings. The model has no concept of a deployment environment or threat model.
|
||||
|
||||
### Vulnerable patterns
|
||||
|
||||
**Binding to all interfaces (Python):**
|
||||
```python
|
||||
app.run(host="0.0.0.0", port=8080, debug=True) # Exposed to network + debug mode
|
||||
```
|
||||
|
||||
**Wide-open CORS (Python/Flask):**
|
||||
```python
|
||||
CORS(app, origins="*", supports_credentials=True)
|
||||
```
|
||||
|
||||
**Permissive file permissions (shell):**
|
||||
```bash
|
||||
chmod 777 /app/data
|
||||
chmod 666 /etc/config/credentials.yaml
|
||||
```
|
||||
|
||||
**Disabled TLS verification (Python):**
|
||||
```python
|
||||
requests.get(url, verify=False)
|
||||
```
|
||||
|
||||
**Kubernetes Service exposed externally by default:**
|
||||
```yaml
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: my-app
|
||||
spec:
|
||||
type: LoadBalancer # Exposed to the network
|
||||
ports:
|
||||
- port: 80
|
||||
```
|
||||
|
||||
### Secure alternatives
|
||||
|
||||
**Bind to localhost unless external access is needed:**
|
||||
```python
|
||||
app.run(host="127.0.0.1", port=8080, debug=False)
|
||||
```
|
||||
|
||||
**Explicit CORS origins:**
|
||||
```python
|
||||
CORS(app, origins=["https://app.example.com"], supports_credentials=True)
|
||||
```
|
||||
|
||||
**Restrictive file permissions:**
|
||||
```bash
|
||||
chmod 750 /app/data # Owner rwx, group rx, others none
|
||||
chmod 640 /etc/config/credentials.yaml # Owner rw, group r, others none
|
||||
```
|
||||
|
||||
**TLS verification enabled (always):**
|
||||
```python
|
||||
requests.get(url, verify=True) # Default, but be explicit
|
||||
# If using internal CA:
|
||||
requests.get(url, verify="/etc/ssl/certs/internal-ca.pem")
|
||||
```
|
||||
|
||||
**ClusterIP by default, expose deliberately:**
|
||||
```yaml
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: my-app
|
||||
spec:
|
||||
type: ClusterIP # Internal only, use Ingress for external access
|
||||
ports:
|
||||
- port: 80
|
||||
```
|
||||
|
||||
### How to catch it in review
|
||||
|
||||
- Search for `0.0.0.0`, `host="0.0.0.0"`, `debug=True` in application code
|
||||
- Search for `origins="*"` or `Access-Control-Allow-Origin: *` in CORS config
|
||||
- Search for `chmod 777`, `chmod 666`, or any world-readable/writable permissions
|
||||
- Search for `verify=False` in HTTP client calls
|
||||
- Search for `type: LoadBalancer` or `type: NodePort` in Kubernetes manifests without explicit justification
|
||||
- Search for `GRANT ALL` in database setup scripts
|
||||
|
||||
---
|
||||
|
||||
## 5. Infrastructure-as-Code Risks
|
||||
|
||||
### What LLMs get wrong
|
||||
|
||||
LLMs generate Kubernetes manifests and Helm charts that are functionally correct but security-negligent. They omit security contexts, resource limits, network policies, and run containers as root by default. Research (GenKubeSec, KubeGuard) found that LLMs can "confidently recommend configurations that introduce new vulnerabilities" including suggesting "allow all" rules just to satisfy constraints.
|
||||
|
||||
### Vulnerable patterns
|
||||
|
||||
**Privileged container (Kubernetes):**
|
||||
```yaml
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: my-app
|
||||
spec:
|
||||
template:
|
||||
spec:
|
||||
containers:
|
||||
- name: my-app
|
||||
image: my-app:latest # No digest, mutable tag
|
||||
# No securityContext at all -- runs as root
|
||||
# No resource limits -- can consume entire node
|
||||
# No readOnlyRootFilesystem
|
||||
```
|
||||
|
||||
**Overly broad RBAC:**
|
||||
```yaml
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: ClusterRoleBinding
|
||||
metadata:
|
||||
name: my-app
|
||||
subjects:
|
||||
- kind: ServiceAccount
|
||||
name: my-app
|
||||
roleRef:
|
||||
kind: ClusterRole
|
||||
name: cluster-admin # Full cluster access
|
||||
```
|
||||
|
||||
**No NetworkPolicy (default allows all traffic):**
|
||||
```yaml
|
||||
# LLMs typically omit NetworkPolicy entirely
|
||||
# Without it, any pod can talk to any other pod
|
||||
```
|
||||
|
||||
**Helm values without security defaults:**
|
||||
```yaml
|
||||
# values.yaml generated by LLM
|
||||
replicaCount: 1
|
||||
image:
|
||||
repository: my-app
|
||||
tag: latest # Mutable, unpinned
|
||||
service:
|
||||
type: LoadBalancer # Externally exposed
|
||||
# No securityContext, no resources, no networkPolicy
|
||||
```
|
||||
|
||||
### Secure alternatives
|
||||
|
||||
**Hardened container:**
|
||||
```yaml
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: my-app
|
||||
spec:
|
||||
template:
|
||||
spec:
|
||||
automountServiceAccountToken: false
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 1000
|
||||
runAsGroup: 1000
|
||||
fsGroup: 1000
|
||||
seccompProfile:
|
||||
type: RuntimeDefault
|
||||
containers:
|
||||
- name: my-app
|
||||
image: my-app@sha256:abc123... # Pinned by digest
|
||||
securityContext:
|
||||
allowPrivilegeEscalation: false
|
||||
readOnlyRootFilesystem: true
|
||||
capabilities:
|
||||
drop: ["ALL"]
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 128Mi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 256Mi
|
||||
```
|
||||
|
||||
**Least-privilege RBAC:**
|
||||
```yaml
|
||||
apiVersion: rbac.authorization.k8s.io/v1
|
||||
kind: Role # Namespaced, not ClusterRole
|
||||
metadata:
|
||||
name: my-app
|
||||
namespace: my-namespace
|
||||
rules:
|
||||
- apiGroups: [""]
|
||||
resources: ["configmaps"]
|
||||
verbs: ["get", "list"] # Only what's needed
|
||||
```
|
||||
|
||||
**Default-deny NetworkPolicy:**
|
||||
```yaml
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
name: my-app
|
||||
spec:
|
||||
podSelector:
|
||||
matchLabels:
|
||||
app: my-app
|
||||
policyTypes: ["Ingress", "Egress"]
|
||||
ingress:
|
||||
- from:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app: frontend
|
||||
ports:
|
||||
- port: 8080
|
||||
egress:
|
||||
- to:
|
||||
- podSelector:
|
||||
matchLabels:
|
||||
app: database
|
||||
ports:
|
||||
- port: 5432
|
||||
```
|
||||
|
||||
### How to catch it in review
|
||||
|
||||
- Run `kubesec scan`, `kube-linter`, or `trivy config` against manifests
|
||||
- Search for `privileged: true`, `allowPrivilegeEscalation: true` (should almost never appear)
|
||||
- Search for `cluster-admin` in RBAC bindings
|
||||
- Check that every Deployment/StatefulSet has `resources:` limits and `securityContext:`
|
||||
- Check that every namespace has at least one NetworkPolicy
|
||||
- Search for `image:.*:latest` -- tags should be pinned to specific versions or digests
|
||||
- Check for `automountServiceAccountToken: false` on pods that don't need the K8s API
|
||||
- In Helm charts, verify `values.yaml` includes security defaults, not just functional defaults
|
||||
|
||||
---
|
||||
|
||||
## 6. Input Validation Gaps
|
||||
|
||||
### What LLMs get wrong
|
||||
|
||||
LLMs generate code that handles the happy path but skips validation of types, lengths, formats, and ranges. They omit validation unless explicitly prompted, because training data (tutorials, examples) does the same. The model has no awareness of the threat model or what inputs are user-controlled.
|
||||
|
||||
### Vulnerable patterns
|
||||
|
||||
**No type/length validation (Python API):**
|
||||
```python
|
||||
@app.route('/api/users', methods=['POST'])
|
||||
def create_user():
|
||||
data = request.get_json()
|
||||
username = data['username'] # No validation at all
|
||||
email = data['email'] # No format check
|
||||
age = data['age'] # No type or range check
|
||||
db.execute("INSERT INTO users (username, email, age) VALUES (%s, %s, %s)",
|
||||
(username, email, age))
|
||||
```
|
||||
|
||||
**No path validation (shell):**
|
||||
```bash
|
||||
#!/bin/bash
|
||||
# LLM-generated backup script
|
||||
BACKUP_DIR="$1"
|
||||
cp -r /important/data "$BACKUP_DIR" # No validation of $1
|
||||
```
|
||||
|
||||
### Secure alternatives
|
||||
|
||||
**Validated API input (Python):**
|
||||
```python
|
||||
from pydantic import BaseModel, EmailStr, Field
|
||||
|
||||
class CreateUserRequest(BaseModel):
|
||||
username: str = Field(min_length=3, max_length=50, pattern=r'^[a-zA-Z0-9_]+$')
|
||||
email: EmailStr
|
||||
age: int = Field(ge=0, le=150)
|
||||
|
||||
@app.route('/api/users', methods=['POST'])
|
||||
def create_user():
|
||||
data = CreateUserRequest(**request.get_json()) # Validates or raises 422
|
||||
db.execute("INSERT INTO users (username, email, age) VALUES (%s, %s, %s)",
|
||||
(data.username, data.email, data.age))
|
||||
```
|
||||
|
||||
**Validated shell input:**
|
||||
```bash
|
||||
#!/bin/bash
|
||||
BACKUP_DIR="$1"
|
||||
if [[ -z "$BACKUP_DIR" ]]; then
|
||||
echo "Error: backup directory required" >&2
|
||||
exit 1
|
||||
fi
|
||||
if [[ ! -d "$BACKUP_DIR" ]]; then
|
||||
echo "Error: '$BACKUP_DIR' is not a directory" >&2
|
||||
exit 1
|
||||
fi
|
||||
# Resolve and validate path
|
||||
REAL_DIR="$(realpath -- "$BACKUP_DIR")"
|
||||
if [[ "$REAL_DIR" != /allowed/backup/* ]]; then
|
||||
echo "Error: backup directory must be under /allowed/backup/" >&2
|
||||
exit 1
|
||||
fi
|
||||
cp -r /important/data "$REAL_DIR"
|
||||
```
|
||||
|
||||
### How to catch it in review
|
||||
|
||||
- Check that all API endpoints use schema validation (Pydantic, marshmallow, JSON Schema, Joi)
|
||||
- Search for `request.get_json()`, `request.args`, `request.form` usage without subsequent validation
|
||||
- In shell scripts, check that all positional parameters (`$1`, `$2`, etc.) are validated before use
|
||||
- Look for direct use of user input in file operations, database queries, or system commands
|
||||
- Verify that numeric inputs have range checks and string inputs have length/format checks
|
||||
|
||||
---
|
||||
|
||||
## 7. Error Handling That Leaks Information
|
||||
|
||||
### What LLMs get wrong
|
||||
|
||||
LLMs generate code with verbose error handling that exposes internal details -- stack traces, file paths, database schemas, SQL queries, internal hostnames. This happens because training data includes development-mode error handling, and the model doesn't distinguish between dev and production contexts.
|
||||
|
||||
### Vulnerable patterns
|
||||
|
||||
**Leaking stack traces (Python/Flask):**
|
||||
```python
|
||||
@app.errorhandler(Exception)
|
||||
def handle_error(e):
|
||||
return jsonify({
|
||||
"error": str(e),
|
||||
"traceback": traceback.format_exc(), # Full stack trace
|
||||
"query": last_query, # SQL query that failed
|
||||
}), 500
|
||||
```
|
||||
|
||||
**Leaking database details:**
|
||||
```python
|
||||
try:
|
||||
cursor.execute(query)
|
||||
except psycopg2.Error as e:
|
||||
return f"Database error: {e}" # Includes table names, column names, query
|
||||
```
|
||||
|
||||
**Leaking file paths (shell):**
|
||||
```bash
|
||||
echo "Error: failed to read config from /etc/myapp/secrets/database.yaml"
|
||||
echo "Stack: $(python3 -c 'import traceback; traceback.print_exc()')"
|
||||
```
|
||||
|
||||
### Secure alternatives
|
||||
|
||||
**Generic error response with internal logging:**
|
||||
```python
|
||||
import logging
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
@app.errorhandler(Exception)
|
||||
def handle_error(e):
|
||||
logger.exception("Unhandled exception") # Full details go to logs
|
||||
return jsonify({"error": "Internal server error"}), 500 # Generic to client
|
||||
```
|
||||
|
||||
**Safe database error handling:**
|
||||
```python
|
||||
try:
|
||||
cursor.execute(query, params)
|
||||
except psycopg2.Error as e:
|
||||
logger.exception("Database query failed")
|
||||
return jsonify({"error": "A database error occurred"}), 500
|
||||
```
|
||||
|
||||
### How to catch it in review
|
||||
|
||||
- Search for `traceback.format_exc()` or `traceback.print_exc()` in response-building code
|
||||
- Search for `str(e)` or `repr(e)` in API responses (should go to logs, not clients)
|
||||
- Check that `DEBUG = False` / `debug=False` in production config
|
||||
- Verify error handlers return generic messages and log details internally
|
||||
- Search for internal paths (`/etc/`, `/home/`, `/var/`) in user-facing error strings
|
||||
|
||||
---
|
||||
|
||||
## 8. Cryptography Mistakes
|
||||
|
||||
### What LLMs get wrong
|
||||
|
||||
LLMs reproduce cryptographic anti-patterns from training data. CWE-780 (Use of RSA without OAEP) is the most observed weakness in Java. Common failures include using ECB mode (which leaks patterns), predictable IVs, deprecated algorithms (MD5, SHA-1 for security purposes), and rolling custom crypto. Cryptography misconfiguration appears in approximately 22-24% of security vulnerabilities across leading LLM models.
|
||||
|
||||
### Vulnerable patterns
|
||||
|
||||
**ECB mode (Python):**
|
||||
```python
|
||||
from Crypto.Cipher import AES
|
||||
cipher = AES.new(key, AES.MODE_ECB) # ECB leaks patterns in ciphertext
|
||||
ciphertext = cipher.encrypt(plaintext)
|
||||
```
|
||||
|
||||
**Hardcoded IV:**
|
||||
```python
|
||||
iv = b'\x00' * 16 # Predictable IV defeats the purpose of CBC/GCM
|
||||
cipher = AES.new(key, AES.MODE_CBC, iv=iv)
|
||||
```
|
||||
|
||||
**MD5 for password hashing:**
|
||||
```python
|
||||
import hashlib
|
||||
password_hash = hashlib.md5(password.encode()).hexdigest() # Broken for security
|
||||
```
|
||||
|
||||
**Weak random for tokens:**
|
||||
```python
|
||||
import random
|
||||
token = ''.join(random.choices(string.ascii_letters, k=32)) # Not cryptographically secure
|
||||
```
|
||||
|
||||
### Secure alternatives
|
||||
|
||||
**AES-GCM with random IV:**
|
||||
```python
|
||||
from Crypto.Cipher import AES
|
||||
from Crypto.Random import get_random_bytes
|
||||
|
||||
key = get_random_bytes(32) # AES-256
|
||||
nonce = get_random_bytes(12) # Random nonce for GCM
|
||||
cipher = AES.new(key, AES.MODE_GCM, nonce=nonce)
|
||||
ciphertext, tag = cipher.encrypt_and_digest(plaintext)
|
||||
# Store nonce + tag + ciphertext together
|
||||
```
|
||||
|
||||
**Proper password hashing:**
|
||||
```python
|
||||
import bcrypt
|
||||
# Hashing
|
||||
password_hash = bcrypt.hashpw(password.encode(), bcrypt.gensalt(rounds=12))
|
||||
# Verification
|
||||
bcrypt.checkpw(password.encode(), stored_hash)
|
||||
```
|
||||
|
||||
**Cryptographically secure random:**
|
||||
```python
|
||||
import secrets
|
||||
token = secrets.token_urlsafe(32) # Cryptographically secure
|
||||
```
|
||||
|
||||
### How to catch it in review
|
||||
|
||||
- Search for `MODE_ECB` -- should almost never be used
|
||||
- Search for `md5`, `sha1` used for passwords or security tokens (fine for checksums, not for security)
|
||||
- Search for `random.` (stdlib) used for tokens, keys, or security values -- should be `secrets.`
|
||||
- Search for hardcoded IVs: `iv = b'`, `iv = bytes(`, `nonce = b'\x00`
|
||||
- Search for `hashlib` used directly for password storage -- should be `bcrypt`, `argon2`, or `scrypt`
|
||||
- Use `bandit` which has specific checks for weak crypto (B303, B304, B305)
|
||||
|
||||
---
|
||||
|
||||
## 9. Research Findings (2024-2026)
|
||||
|
||||
### ACM / TOSEM: Security Weaknesses of Copilot-Generated Code in GitHub Projects
|
||||
Analyzed real-world Copilot-generated code on GitHub. Found 29.5% of Python and 24.2% of JavaScript snippets contained security weaknesses across 43 CWE categories. Top weaknesses: CWE-330 (insufficiently random values), CWE-94 (code injection), CWE-79 (XSS).
|
||||
|
||||
### Large-Scale GitHub Analysis (October 2025)
|
||||
Analyzed 7,703 files from 4 AI tools across public GitHub repos. Found 4,241 CWE instances across 77 distinct vulnerability types. ChatGPT-generated code comprised 91.5% of the sample, Copilot 7.5%.
|
||||
|
||||
### Slopsquatting Research (2025)
|
||||
576,000 code samples across 16 LLMs: 19.7% of suggested packages were hallucinations (205,474 unique fake names). Open-source models hallucinated at 21.7%, commercial at 5.2%. 43% of hallucinated names appeared consistently (predictable, attackable).
|
||||
|
||||
### Endor Labs: State of Dependency Management (2025)
|
||||
80% of AI-suggested dependencies contain known risks. 44-49% of dependencies imported by coding agents contained known security vulnerabilities.
|
||||
|
||||
### Copilot Code Review Study (2025)
|
||||
GitHub Copilot's code review feature frequently fails to detect critical vulnerabilities (SQL injection, XSS, insecure deserialization). Primarily flags low-severity issues like coding style.
|
||||
|
||||
### Sonar: Coding Personalities of Leading LLMs (2025)
|
||||
Multi-model analysis finding that cryptography misconfiguration appears in 22-24% of vulnerabilities across leading models. Missing input sanitization is the most common flaw category.
|
||||
|
||||
### GenKubeSec / KubeGuard (2024-2025)
|
||||
Research on LLM-generated Kubernetes configurations found models confidently recommend insecure configurations and may suggest "allow all" rules to satisfy functional requirements.
|
||||
|
||||
### OWASP Top 10 for LLM Applications (2025 Update)
|
||||
Updated to reflect agentic AI risks. Key additions: System Prompt Leakage, Excessive Agency. Improper Output Handling (treating LLM output as trusted) remains a top-5 risk. Core message: treat all LLM output as untrusted data.
|
||||
|
||||
### Security Degradation in Iterative Generation (2025)
|
||||
Code security degrades with iterative LLM refinement -- each round of "fix this" prompting can introduce new vulnerabilities while fixing the original one.
|
||||
|
||||
---
|
||||
|
||||
## 10. Practical Review Checklist
|
||||
|
||||
Use this checklist when reviewing LLM-generated code:
|
||||
|
||||
### Python
|
||||
- [ ] No string formatting in SQL queries (use parameterised queries)
|
||||
- [ ] No `shell=True` in subprocess calls
|
||||
- [ ] No `verify=False` in HTTP requests
|
||||
- [ ] No `random.` for security values (use `secrets.`)
|
||||
- [ ] No `hashlib.md5/sha1` for passwords (use `bcrypt`/`argon2`)
|
||||
- [ ] No hardcoded credentials (search for `password =`, `api_key =`, `token =`)
|
||||
- [ ] Input validation on all API endpoints (Pydantic, marshmallow)
|
||||
- [ ] Error handlers return generic messages, log details internally
|
||||
- [ ] `debug=False` in production config
|
||||
- [ ] All dependencies exist on PyPI and are pinned to audited versions
|
||||
- [ ] `host="127.0.0.1"` unless external binding is explicitly required
|
||||
|
||||
### Shell Scripts
|
||||
- [ ] All variables quoted (`"$var"` not `$var`)
|
||||
- [ ] User-provided paths validated with `realpath` and boundary checks
|
||||
- [ ] No secrets as command-line arguments (use files or env vars)
|
||||
- [ ] No `chmod 777` or `chmod 666`
|
||||
- [ ] ShellCheck passes with no warnings
|
||||
|
||||
### Kubernetes Manifests
|
||||
- [ ] `securityContext` present with `runAsNonRoot: true`, `readOnlyRootFilesystem: true`, `allowPrivilegeEscalation: false`
|
||||
- [ ] `capabilities.drop: ["ALL"]`
|
||||
- [ ] `resources.requests` and `resources.limits` defined
|
||||
- [ ] No `privileged: true`
|
||||
- [ ] No `cluster-admin` RBAC bindings
|
||||
- [ ] `automountServiceAccountToken: false` where K8s API access is not needed
|
||||
- [ ] Images pinned to digest or specific version (not `:latest`)
|
||||
- [ ] Services use `ClusterIP` by default (not `LoadBalancer`/`NodePort` without justification)
|
||||
- [ ] NetworkPolicy exists for the namespace/workload
|
||||
|
||||
### Helm Charts
|
||||
- [ ] `values.yaml` includes secure defaults for securityContext, resources, service type
|
||||
- [ ] Templates don't embed secrets in plaintext
|
||||
- [ ] Chart version and appVersion pinned
|
||||
- [ ] `helm template` renders valid, secure manifests with default values
|
||||
- [ ] `values.schema.json` validates required security fields
|
||||
|
||||
---
|
||||
|
||||
## Sources
|
||||
|
||||
- [Security Weaknesses of Copilot-Generated Code in GitHub Projects (ACM TOSEM)](https://dl.acm.org/doi/10.1145/3716848)
|
||||
- [Security Vulnerabilities in AI-Generated Code: A Large-Scale Analysis (arXiv, Oct 2025)](https://arxiv.org/abs/2510.26103)
|
||||
- [The Most Common Security Vulnerabilities in AI-Generated Code (Endor Labs)](https://www.endorlabs.com/learn/the-most-common-security-vulnerabilities-in-ai-generated-code)
|
||||
- [Endor Labs 2025 State of Dependency Management Report](https://www.prnewswire.com/news-releases/endor-labs-launches-2025-state-of-dependency-management-report-finds-80-of-ai-suggested-dependencies-contain-risks-302603438.html)
|
||||
- [LLMs' AI-Generated Code Remains Wildly Insecure (Dark Reading)](https://www.darkreading.com/application-security/llms-ai-generated-code-wildly-insecure)
|
||||
- [Popular LLMs Found to Produce Vulnerable Code by Default (Infosecurity Magazine)](https://www.infosecurity-magazine.com/news/llms-vulnerable-code-default/)
|
||||
- [Slopsquatting: How AI Hallucinations Are Fueling Supply Chain Attacks (Socket.dev)](https://socket.dev/blog/slopsquatting-how-ai-hallucinations-are-fueling-a-new-class-of-supply-chain-attacks)
|
||||
- [Slopsquatting meets Dependency Confusion (Andrew Nesbitt)](https://nesbitt.io/2025/12/10/slopsquatting-meets-dependency-confusion.html)
|
||||
- [AI-Generated Code Packages Can Lead to Slopsquatting Threat (DevOps.com)](https://devops.com/ai-generated-code-packages-can-lead-to-slopsquatting-threat-2/)
|
||||
- [OWASP Top 10 for LLM Applications 2025](https://owasp.org/www-project-top-10-for-large-language-model-applications/)
|
||||
- [OWASP LLM Top 10: How it Applies to Code Generation (Sonar)](https://www.sonarsource.com/resources/library/owasp-llm-code-generation/)
|
||||
- [The Coding Personalities of Leading LLMs (SonarSource)](https://www.sonarsource.com/the-coding-personalities-of-leading-llms.pdf)
|
||||
- [GenKubeSec: LLM-Based Kubernetes Misconfiguration Detection](https://arxiv.org/html/2405.19954v1)
|
||||
- [KubeGuard: LLM-Assisted Kubernetes Hardening](https://arxiv.org/abs/2509.04191)
|
||||
- [Security Degradation in Iterative AI Code Generation (arXiv)](https://arxiv.org/pdf/2506.11022)
|
||||
- [GitHub Copilot's Code Review: Can AI Spot Security Flaws? (arXiv)](https://arxiv.org/html/2509.13650v1)
|
||||
- [Security Risks of Vibe Coding and LLM Assistants (Kaspersky)](https://www.kaspersky.com/blog/vibe-coding-2025-risks/54584/)
|
||||
- [The Risks of Hardcoding Secrets in Code Generated by LLMs (Cycode)](https://cycode.com/blog/the-risks-of-hardcoding-secrets-in-code-generated-by-language-learning-models/)
|
||||
- [Security Flaws in DeepSeek-Generated Code (CrowdStrike)](https://www.crowdstrike.com/en-us/blog/crowdstrike-researchers-identify-hidden-vulnerabilities-ai-coded-software/)
|
||||
Reference in New Issue
Block a user