Files
best-practices/api-design.md
Paul O'Reilly 6154031cf9 api-design: add API-first methodology, DX, and contract testing
Extends api-design.md beyond its security/operations focus with three
new dimensions:

- §0 API-First Design Process — OpenAPI 3.1 as single source of truth,
  Spectral governance, dogfooding (UIs consume the public API, no
  privileged backdoors), auth-required-by-default as a design stance.
- §7 Documentation and Developer Experience — Scalar/Mintlify,
  RFC 9457 Problem Details error envelope, interactive playgrounds,
  generated SDKs (Stainless, Speakeasy, Fern), RFC 9745 deprecation
  signals and changelog UX.
- §8 Contract Testing and API Quality — schema validation in the
  test suite, Pact CDC vs provider verification, Schemathesis
  property-based fuzzing, oasdiff drift detection in CI, the API
  test pyramid.

Intro, cross-refs in §3.1/§3.3/§4.1, and Sources block reorganised
by topic. Index entry in BESTPRACTICES.md updated. PLAN file included
for traceability.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-25 13:42:03 +12:00

57 KiB
Raw Permalink Blame History

API Design

Best practices for REST/HTTP APIs in internal microservices and platform services. Covers four dimensions: methodology (design-first, dogfooding, governance), security and operations (transport, auth, validation, service mesh), developer experience (docs, SDKs, deprecation signals), and quality (contract testing, drift detection). Focused on practical defaults -- not aspirational ideals. Sourced from OWASP API Security Top 10 (2023), RFC 9700 (OAuth 2.0 Security BCP, January 2025), RFC 9457 (Problem Details, 2023), RFC 9745 (Deprecation header, 2024), Google AIP, and production experience.

Cross-references: Security Architecture covers the server boundary rule and proxy patterns. Secrets Management covers credential storage and rotation. Test-Driven Development covers the testing principles that §8 extends.


0. API-First Design Process

This section covers methodology -- how APIs get designed and governed, not what goes in them. The mechanical sections (§1-§6) assume an API-first workflow. If your team is code-first, start here.

0.1 Design the contract before writing code

Principle: The OpenAPI document is authored, reviewed, and committed before any handler code is written. The spec drives mocks, SDKs, docs, validation middleware, and contract tests in parallel -- not as artifacts generated after the fact.

Why it matters: Code-first specs (annotations on handlers exporting OpenAPI) describe how the API was implemented, not how it should be used. They rebake internal types, drift the moment someone refactors, and miss design flaws because the spec inherits them. The Postman 2024 State of the API report puts API-first adoption at 74% (83% including partial adopters).

How to implement:

  • Treat the OpenAPI file as source code -- in the repo, in PRs, code-reviewed, versioned.
  • Backend, frontend, and partners build against the same spec from day one. Use Prism (or any OpenAPI mock server) to unblock parallel work before the service exists.
  • Run a brief design review before merging the spec -- focus on resource modelling, lifecycle, and breaking-change risk; let the linter (§0.3) catch mechanical issues.
  • For greenfield, consider TypeSpec (Microsoft) for spec authoring -- it compiles to OpenAPI and is faster to write than raw YAML.

Anti-patterns:

  • Spec generated from code annotations and never reviewed independently -- drifts within months.
  • "We'll document it after v1 ships" -- guarantees a v2 rewrite once you discover the design flaws.
  • OpenAPI file treated as build output (not in PRs, not reviewed).
  • Specs that mirror the database schema 1:1 instead of designing the consumer-facing contract.

0.2 OpenAPI 3.1 as the single source of truth

Principle: Standardize on OpenAPI 3.1 (not 3.0). One document drives docs, SDKs, mocks, validation, and tests across the entire API surface.

Why it matters: OpenAPI 3.1 is a superset of JSON Schema Draft 2020-12; 3.0 was a near-but-not-quite subset that forced tooling to maintain two parallel schema engines. Standardizing on 3.1 lets you use one schema language across REST APIs, AsyncAPI 3.0 events, and validation libraries -- no more "validation schema" / "docs schema" split.

How to implement:

  • Migrate from 3.0 → 3.1: nullable: true is gone (use type: ["string", "null"]); exclusiveMinimum/exclusiveMaximum take values not booleans; example becomes examples (array); file uploads use contentMediaType/contentEncoding; the spec gains first-class webhooks.
  • Adopt AsyncAPI 3.0 for event payloads -- same JSON Schema dialect, one mental model.
  • Pin to a specific minor (openapi: 3.1.2) to avoid silent tooling drift.

Anti-patterns:

  • Staying on 3.0 to avoid the migration -- locks you out of conditional schemas (if/then/else), tuple validation, and tool consolidation.
  • Maintaining separate "schemas for validation" and "schemas for docs" -- they will diverge.
  • Using nullable: true in a 3.1 doc -- silently ignored by some tools, causes subtle validation gaps.
  • Treating OpenAPI as docs only while runtime validation is implemented separately and drifts.

0.3 Governance through linting

Principle: API style guides are enforced as code. A Spectral ruleset is checked into the repo and runs in CI on every spec change, blocking merges on errors.

Why it matters: Style guides written as wiki pages get ignored. Lint rules don't. Mechanical enforcement also frees design reviews to focus on intent and edge cases instead of bikeshedding naming.

How to implement:

  • Adopt Spectral 6.x. Extend the default oas ruleset and layer the Spectral OWASP ruleset on top -- it codifies the OWASP API Security Top 10 (2023) at the spec level (e.g. flags any operation lacking security as API2:2023).
  • Enforce: resource naming (plural nouns, kebab-case paths, camelCase fields); pagination shape; canonical error envelope (RFC 9457, see §7.3); required operationId for SDK gen; mandatory security block on every operation; response schemas on every documented status code.
  • Layer rulesets by maturity -- a base ruleset for all APIs, stricter rules for partner/public APIs.
  • Pair with oasdiff for breaking-change detection (covered in §8.4).

Anti-patterns:

  • Style guide as a wiki page nobody reads -- encode every rule that can be linted.
  • One monolithic ruleset applied identically to a prototype and a public API.
  • Per-team error envelopes -- every service inventing its own error shape.
  • Treating every Spectral warning as equally severe -- tune severities or the team will start ignoring all of them.

0.4 Dogfood your own API

Principle: Internal UIs, admin tools, mobile clients, and partner integrations all consume the same public API surface as third-party developers. No privileged backdoors. No internal-only fields. No hidden endpoints.

Why it matters: If your own dashboard cannot authenticate, paginate, or recover from a 429, neither can your customers. Dogfooding is the forcing function that keeps the API actually usable -- and it surfaces auth gaps, rate-limit gaps, and missing affordances before customers find them. The Bezos 2002 mandate at Amazon is the canonical formulation: "no direct linking, no direct reads of another team's data store, no shared-memory model, no back-doors whatsoever."

How to implement:

  • Public and internal clients live in the same repo where possible -- code review catches API shortcuts.
  • Internal traffic hits the same gateway, auth, and rate limits as external traffic. No separate "internal" tier.
  • Spectral rule that flags x-internal: true operations on the public spec -- prove the absence of backdoors mechanically.
  • Stripe-style "friction logging" -- when teams build new abstractions, document every snag before the abstraction reaches GA. The snags become the next batch of API improvements.

Anti-patterns:

  • Admin/internal endpoints that bypass auth or rate limits "because it's just us."
  • Internal-only fields on shared schemas leaking sensitive data, or worse -- external consumers come to depend on them.
  • A separate /v1-internal API that diverges from the public one, doubling maintenance.
  • UI that talks directly to the database while customers go through the API -- every UI feature becomes a customer feature request the API can't satisfy.

0.5 Auth required by default -- as a design stance

Principle: Every operation in the spec has a security requirement at design time. Unauthenticated endpoints (health checks, public OIDC discovery) are the rare, deliberately-justified exception, tagged so an auditor can list them in seconds.

Why it matters: This is "default deny" applied at API design time, not at the firewall. There is no internal network in a zero-trust model -- every request, including service-to-service, proves identity. Mechanics live in §2; the design stance in §0.5 is what decides whether your spec ever reaches a security reviewer with anonymous endpoints in it.

How to implement:

  • Spec-level: every operation has a security block. Spectral OWASP rule owasp:api2 enforces it.
  • Code-level: middleware rejects any request to an unauthenticated route unless that route is on an explicit allowlist.
  • For service-to-service calls, identity is mTLS (SPIFFE) or a short-lived JWT -- see §2.3.
  • Track unauthenticated endpoints in a single inventory file. Auditors review it; security review is required to add to it.

Anti-patterns:

  • Operations with no security block ("we'll add it later" -- you won't).
  • IP allowlist or VPN as the only control between services -- collapses the moment someone runs the service in a different environment.
  • An "internal" tier with no auth because "it's behind the load balancer."
  • API keys as the only credential, shared across services, never rotated.

1. Transport Security

1.1 HTTPS everywhere, no exceptions

Principle: Every API endpoint -- internal or external -- must serve over TLS. Plaintext HTTP must not be available, even on internal networks.

Why it matters: Without TLS, any network hop (load balancer, sidecar, switch) can observe or modify traffic. Internal networks are not trusted in a zero-trust model -- a compromised pod can sniff adjacent traffic.

How to implement:

  • Terminate TLS at the ingress controller (e.g., Traefik, NGINX) with certificates from cert-manager / Let's Encrypt.
  • For service-to-service within the cluster, use a service mesh (Istio, Linkerd) or cert-manager CSI driver to issue per-pod certificates.
  • Set Strict-Transport-Security headers on all responses.
  • Redirect HTTP to HTTPS at the ingress layer.

Anti-patterns:

  • "Internal traffic doesn't need encryption" -- it does under zero-trust.
  • Self-signed certificates with verification disabled (--insecure, verify=False) -- defeats the purpose of TLS.
  • Long-lived certificates (years) with no rotation -- use short-lived certs (days to weeks) with automated renewal.

1.2 mTLS between services

Principle: Service-to-service communication must use mutual TLS -- both sides present and verify certificates.

Why it matters: Server-only TLS authenticates the server to the client, but any client can connect. mTLS ensures both parties have a cryptographically verified identity, which is the foundation of zero-trust networking.

How to implement:

  • Service mesh (Istio strict mode, Linkerd) handles mTLS transparently via sidecar proxies -- no application code changes.
  • Use SPIFFE/SPIRE for standardized workload identity (SVID certificates).
  • Default certificate lifetime should be short (24 hours) with automatic rotation.
  • Start in permissive mode (allow both plain and mTLS), migrate to strict mode once all services are enrolled.

Anti-patterns:

  • Permissive mode as a permanent state -- it must be a migration step, not the end state.
  • Disabling mTLS verification for "debugging" and forgetting to re-enable it.
  • Using a single shared certificate for all services -- each workload needs its own identity.

1.3 Certificate management

Principle: Certificate issuance and rotation must be fully automated. No manual certificate management in production.

Why it matters: Manual certificate management leads to expired certificates, which cause outages. It also leads to long-lived certificates, which increase blast radius if compromised.

How to implement:

  • cert-manager in Kubernetes with ClusterIssuer for ingress certificates.
  • Service mesh control plane for workload certificates (Istio Citadel, Linkerd identity).
  • Monitor certificate expiry with alerts at 30/14/7 days before expiry.
  • Store CA keys in HSM or sealed secrets -- never in plaintext ConfigMaps.

Anti-patterns:

  • Certificates stored in Git repos (even encrypted, they need rotation).
  • Wildcard certificates shared across trust boundaries.
  • No monitoring for certificate expiry -- silent failures at 3am.

2. Authentication and Authorization

2.1 OIDC/OAuth2 for user-facing APIs (RFC 9700)

Principle: Use OAuth 2.0 Authorization Code flow with PKCE for all client types. The implicit flow and resource owner password credentials flow are deprecated per RFC 9700 (January 2025).

Why it matters: The implicit flow exposes access tokens in URLs and browser history. The password grant requires users to share credentials directly with the client, bypassing centralized identity providers.

How to implement:

  • Authorization Code + PKCE for all clients (web, mobile, CLI). PKCE is now mandatory for all client types, not just public clients.
  • Use S256 challenge method (not plain).
  • Tokens issued by the authorization server, validated by the resource server.
  • Use Authorization Server Metadata (RFC 8414) for automatic discovery of endpoints and supported features.

Anti-patterns:

  • Implicit flow (response_type=token) -- deprecated by RFC 9700.
  • Resource Owner Password Credentials flow -- deprecated by RFC 9700.
  • Storing tokens in localStorage (accessible to XSS) -- use httpOnly cookies or in-memory storage with refresh token rotation.
  • Long-lived access tokens without refresh -- use short-lived access tokens (5-15 minutes) with refresh token rotation.

2.2 JWT best practices

Principle: JWTs must be validated completely on every request -- signature, expiry, issuer, audience, and algorithm.

Why it matters: Incomplete JWT validation is a top attack vector. Accepting expired tokens, wrong audiences, or alg: none enables token forgery and replay.

How to implement:

  • Validate: signature (asymmetric preferred -- RS256/ES256), exp, iat, iss, aud, nbf.
  • Use asymmetric signing (RS256/ES256) so that only the auth server holds the private key. Resource servers only need the public key.
  • Set aud claim to the specific API audience -- reject tokens intended for other services.
  • Keep tokens small -- put only identity and authorization claims in the token, fetch additional data from a userinfo endpoint.
  • Use jti (JWT ID) claim for token revocation checks when needed.

Anti-patterns:

  • Accepting alg: none or allowing algorithm switching -- pin the expected algorithm server-side.
  • Not validating aud -- allows tokens from one service to be replayed against another.
  • Symmetric signing (HS256) with a shared secret across services -- if one service is compromised, all are.
  • Treating JWTs as sessions -- JWTs are not revocable by default. Combine with short expiry and token introspection for revocation.

2.3 Service-to-service authentication

Principle: Services authenticate to each other using mTLS identities (SPIFFE) or short-lived JWTs from a token exchange. Never shared static API keys.

Why it matters: Shared API keys have no expiry, no rotation path, no per-service identity, and no audit trail. If one service is compromised, the key works for everything.

How to implement:

  • Preferred: mTLS with SPIFFE. The service mesh provides identity automatically. Authorization policies reference SPIFFE IDs (e.g., spiffe://cluster.local/ns/payments/sa/payment-svc).
  • Alternative: OAuth2 Client Credentials flow. Each service has its own client_id and client_secret (or asymmetric key pair). Tokens are short-lived and scoped to specific audiences.
  • Use asymmetric client authentication (private_key_jwt per RFC 7523) rather than client secrets where possible.
  • Implement audience restriction -- tokens minted for service A must not be accepted by service B.

Anti-patterns:

  • Shared static API keys passed in headers or query strings.
  • One "admin" service account used by all services.
  • Service-to-service tokens with no audience claim -- replayable across any internal API.
  • Bearer tokens without mTLS -- if the network is compromised, the token can be stolen and replayed from anywhere.

2.4 Authorization: object-level and function-level

Principle: Check authorization at every API endpoint, for every object access, based on the authenticated identity. Never rely on "the client won't send that request."

Why it matters: Broken Object-Level Authorization (BOLA) is the #1 risk in the OWASP API Security Top 10. Broken Function-Level Authorization is #5. These are the most common API vulnerabilities found in penetration tests.

How to implement:

  • Every endpoint that accesses a specific resource must verify the caller owns or has access to that resource.
  • Use middleware/decorators that enforce authorization before the handler runs.
  • Use random UUIDs for resource identifiers, not sequential integers (which are trivially enumerable).
  • Separate authorization for data access (BOLA) and function access (admin endpoints, bulk operations).
  • Automated tests that verify: user A cannot access user B's resources, non-admin cannot call admin endpoints.

Anti-patterns:

  • Authorization only at the API gateway -- must also be enforced at the service level.
  • Relying on obscurity of endpoint URLs for access control.
  • Sequential/predictable resource IDs without authorization checks.
  • Missing authorization on secondary endpoints (e.g., /users/{id}/orders checks user but not order ownership).

3. API Design Patterns

3.1 Versioning

Principle: Version your API from day one using URL path versioning (/v1/). Support at most two versions simultaneously. See §7.5 for the deprecation-comms counterpart (Deprecation and Sunset headers, changelog UX) and the alternative versioning models (Stripe-style dated, GitHub-style header).

Why it matters: Breaking changes without versioning cause cascading failures across all consumers simultaneously. Supporting too many versions creates maintenance burden and security risk (old versions may lack patches).

How to implement:

  • URL path: /api/v1/resources -- simple, visible, cacheable.
  • Deprecation policy: announce deprecation in response headers (Deprecation: true, Sunset: <date>).
  • Maximum two active versions. When v3 launches, v1 is removed.
  • Internal services can use header-based versioning (Accept: application/vnd.myapi.v2+json) if URL versioning is too rigid for rapid iteration.

Anti-patterns:

  • No versioning ("we'll be careful") -- you will break consumers.
  • Unlimited version support -- v1 through v7 all still running, each with different bugs.
  • Breaking changes in a patch version.
  • Versioning individual endpoints instead of the whole API surface.

3.2 Pagination

Principle: All list endpoints must paginate. Use cursor-based pagination for real-time data; offset-based for stable datasets.

Why it matters: Unbounded list responses cause memory exhaustion, slow responses, and database strain. Large offset values cause full table scans.

How to implement:

  • Cursor-based (preferred): Return an opaque next_cursor token. Client passes it to get the next page. Stable under concurrent writes.
    { "data": [...], "next_cursor": "abc123", "has_more": true }
    
  • Offset-based (simple datasets): ?limit=50&offset=100. Acceptable for admin dashboards or infrequently changing data.
  • Set a maximum page size (e.g., 100) enforced server-side. Ignore client requests for larger pages.
  • Always return pagination metadata (next_cursor, has_more, or total_count if cheap to compute).

Anti-patterns:

  • No pagination on list endpoints -- returns 50,000 records in one response.
  • Offset-based pagination on large, frequently-changing datasets -- pages shift as records are inserted/deleted.
  • Client-controlled page size with no server-side maximum.
  • total_count requiring a full table scan on every request -- make it optional or cached.

3.3 Error handling

Principle: Return structured, machine-readable errors with stable error codes, human-readable messages, and consistent shape across all endpoints. Standardise on RFC 9457 Problem Details (application/problem+json) -- see §7.3 for the spec-side definition and error catalog pattern.

Why it matters: Inconsistent error formats force every consumer to write custom parsing logic. Missing error codes make automated retry decisions impossible. Leaking stack traces exposes internals to attackers.

How to implement:

  • Standard error envelope:
    {
      "error": {
        "code": "RESOURCE_NOT_FOUND",
        "message": "Order 7f3a... not found",
        "details": [{ "field": "order_id", "reason": "not_found" }]
      }
    }
    
  • Use HTTP status codes correctly: 400 (bad input), 401 (unauthenticated), 403 (unauthorized), 404 (not found), 409 (conflict), 422 (validation), 429 (rate limited), 500 (server error).
  • Error codes are stable strings (not integers) that consumers can switch on.
  • Never expose stack traces, SQL errors, or internal paths in error responses.
  • Log the full error server-side with a correlation ID. Return only the correlation ID to the client.

Anti-patterns:

  • 200 OK with {"success": false} -- use HTTP status codes.
  • Returning raw database errors ("duplicate key violates unique constraint on...").
  • Different error shapes from different endpoints in the same API.
  • Generic "Internal Server Error" with no correlation ID -- impossible to debug.

3.4 Idempotency

Principle: All state-changing operations must be safe to retry. Use idempotency keys for POST requests; PUT and DELETE are idempotent by definition.

Why it matters: Network failures, timeouts, and retries are normal in distributed systems. Without idempotency, retried requests create duplicate orders, double payments, or inconsistent state.

How to implement:

  • Accept Idempotency-Key header (IETF draft: draft-ietf-httpapi-idempotency-key-header) on POST endpoints.
  • Server stores the response for a given key (TTL 24-48 hours). Duplicate requests return the stored response.
  • Use UUIDv4 for idempotency keys -- never sequential or timestamp-based (predictable/guessable).
  • Handle concurrent duplicate requests with locking: first request processes, subsequent requests wait then return cached response.
  • PUT must be truly idempotent: same request, same result, no side effects on repeat.

Anti-patterns:

  • POST endpoints with no idempotency support -- every retry creates a duplicate.
  • Idempotency keys stored forever (memory leak) or for too short a period (retries after expiry create duplicates).
  • Client-generated sequential keys (integers, timestamps) -- guessable and exploitable.
  • "Idempotent" endpoints that still send duplicate emails/webhooks on retry.

3.5 Rate limiting

Principle: Every API must enforce rate limits. Return standard headers so clients can self-throttle.

Why it matters: Without rate limits, a single misbehaving client (or attacker) can exhaust resources for all consumers. Rate limits also protect downstream dependencies.

How to implement:

  • Token bucket or sliding window algorithm (token bucket is simplest with good burst handling).
  • Return headers: X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset (IETF draft still pending; X- prefix remains de facto standard).
  • Return 429 Too Many Requests with Retry-After header (RFC 6585).
  • Rate limit checks execute before expensive operations (auth, database queries).
  • For distributed deployments, use Redis with atomic Lua scripts for counter operations -- avoid race conditions.
  • Different tiers for different consumers (internal services get higher limits than external clients).

Anti-patterns:

  • No rate limiting ("it's an internal API") -- a runaway loop in one service takes down the whole platform.
  • Rate limiting after expensive operations (database query runs, then rate limit rejects the response).
  • No Retry-After header -- clients retry immediately in a tight loop, making the problem worse.
  • Per-IP rate limiting only -- bypassed by distributed clients, unfair to NAT'd users.

4. Input Validation

4.1 Schema validation at the edge

Principle: Validate all request bodies against a schema (OpenAPI/JSON Schema) at the API gateway or middleware layer. Reject requests that don't conform before they reach business logic. The same OpenAPI document drives runtime validation here, in-test validation in §8.1, and the docs in §7.1 -- one source of truth (§0.2).

Why it matters: Invalid input that reaches business logic causes unpredictable behavior -- crashes, data corruption, injection attacks. Edge validation is the first line of defense.

How to implement:

  • Define request/response schemas in OpenAPI 3.x. Generate validation middleware from the spec.
  • Reject unknown fields (additionalProperties: false) -- attackers probe via unexpected fields.
  • Enforce type constraints: string lengths, integer ranges, enum values, date formats.
  • Validate Content-Type header -- reject requests with unexpected content types (e.g., reject multipart/form-data on a JSON-only endpoint).

Anti-patterns:

  • Validation only in business logic, not at the edge -- invalid data traverses the full call stack before rejection.
  • Accepting and silently ignoring unknown fields -- hides bugs and enables mass assignment attacks.
  • Validating types but not ranges -- accepting an age field of 99999 or -1.
  • No schema at all ("we'll validate manually") -- inconsistent validation across endpoints.

4.2 Injection prevention

Principle: Use parameterized queries for all database access. Never concatenate user input into queries, commands, or templates.

Why it matters: SQL injection remains in the OWASP Top 10 after 20+ years. NoSQL injection, LDAP injection, and command injection follow the same pattern -- unsanitized input in a query language.

How to implement:

  • Use an ORM (SQLAlchemy, Prisma, TypeORM) or parameterized queries. All major ORMs parameterize by default.
  • For raw SQL (performance-critical paths), use prepared statements exclusively.
  • Validate input with allowlists, not denylists. If a field should be a UUID, validate it's a UUID -- don't try to strip "malicious characters."
  • For template rendering, use auto-escaping (Jinja2 autoescape, React JSX auto-escaping).

Anti-patterns:

  • String concatenation in SQL: f"SELECT * FROM users WHERE id = '{user_input}'".
  • Denylisting dangerous characters instead of allowlisting valid patterns.
  • Trusting input from "internal" services -- a compromised upstream service sends malicious data.
  • Disabling ORM parameterization for "performance" without understanding the security cost.

4.3 Request size and depth limits

Principle: Enforce maximum request body size, JSON nesting depth, and array length at the gateway level.

Why it matters: Deeply nested JSON or extremely large payloads cause CPU exhaustion during parsing (hash collision attacks, recursive descent parsers). This is a denial-of-service vector.

How to implement:

  • Set maximum body size at the reverse proxy/ingress (e.g., client_max_body_size 1m in NGINX).
  • Limit JSON nesting depth (8-16 levels is generous for any real use case).
  • Limit array sizes in request bodies (e.g., batch endpoints accept max 100 items).
  • Set request timeouts at the gateway -- don't let slow clients hold connections open.

Anti-patterns:

  • No body size limit -- 100MB JSON payload parsed by every middleware layer.
  • Accepting arbitrarily nested JSON -- {"a":{"a":{"a":...}}} 1000 levels deep.
  • Batch endpoints with no limit -- client sends 1 million items in one request.

5. Secrets in APIs

5.1 Never in URLs or query parameters

Principle: Authentication tokens, API keys, and any secret material must be sent in headers (Authorization, custom headers) or request bodies. Never in URLs or query parameters.

Why it matters: URLs are logged everywhere -- web server access logs, proxy logs, browser history, referrer headers, CDN logs, monitoring tools. A token in a URL is a token in every log file in the request path.

How to implement:

  • Use Authorization: Bearer <token> header for all token-based auth.
  • For webhook signatures, use a signature header (e.g., X-Hub-Signature-256).
  • If an API currently accepts tokens in query params, deprecate that path and migrate to header-based auth.
  • Configure log scrubbing to redact Authorization headers, but don't rely on it as the primary control.

Anti-patterns:

  • GET /api/resources?api_key=sk_live_abc123 -- key in every access log.
  • OAuth redirect URIs with tokens in query params (use response_mode=form_post or authorization code flow).
  • Webhook URLs with embedded secrets (/webhook?secret=abc) -- logged, cached, shared.

5.2 Token rotation and expiry

Principle: All tokens and API keys must have expiry dates and a documented rotation procedure. No permanent credentials.

Why it matters: Leaked tokens without expiry are valid forever. Rotation limits blast radius -- even if a token is compromised, it expires soon.

How to implement:

  • Access tokens: 5-15 minute expiry, refreshed via refresh token.
  • Refresh tokens: rotate on use (each refresh issues a new refresh token and invalidates the old one).
  • API keys for external integrations: 90-day rotation policy with overlap period (new key valid before old key expires).
  • Service account tokens (OAuth2 client credentials): short-lived (1 hour), fetched on demand.
  • Track expiry dates in a credential inventory. Alert before expiry (see Secrets Management -- Credential Lifecycle Management).

Anti-patterns:

  • API keys that never expire ("we'll rotate them when we need to" -- you won't).
  • Refresh tokens that don't rotate -- stolen refresh token provides permanent access.
  • No overlap period during rotation -- brief outage while all consumers update.
  • Hardcoded tokens in application config deployed via CI -- rotation requires a full redeploy.

5.3 No secrets in logs or error responses

Principle: Scrub all secrets from logs, error responses, and monitoring data. Structured logging with explicit field selection is safer than serializing request objects.

Why it matters: Log aggregation systems (ELK, Loki, Datadog) are often accessible to broader teams than production systems. A token in a log entry has a much wider exposure surface than a token in a running process.

How to implement:

  • Use structured logging. Log specific fields, not entire request objects.
  • Redact Authorization headers and any field matching token, password, secret, key patterns in log middleware.
  • Never log request bodies for auth endpoints (login, token exchange).
  • Error responses must not include internal state -- return a correlation ID and log details server-side.

Anti-patterns:

  • logger.info(f"Request: {request.headers}") -- logs all headers including Authorization.
  • Error responses that include the original request (including auth headers) for "debugging convenience."
  • Logging full webhook payloads that contain signing secrets in custom headers.

6. Service Mesh and Zero Trust

6.1 Default deny with explicit allow

Principle: Network policies and authorization policies must default to deny-all. Every allowed communication path is explicitly defined.

Why it matters: Default-allow means a compromised service can reach every other service in the cluster. Default-deny contains the blast radius to only the services the compromised workload was authorized to reach.

How to implement:

  • Kubernetes NetworkPolicy: deploy a default-deny policy in every namespace, then add specific allow rules.
    apiVersion: networking.k8s.io/v1
    kind: NetworkPolicy
    metadata:
      name: default-deny-all
    spec:
      podSelector: {}
      policyTypes: [Ingress, Egress]
    
  • Service mesh authorization policies: deny by default, allow specific source-to-destination pairs by SPIFFE ID.
  • Audit policies periodically -- remove rules for decommissioned services.

Anti-patterns:

  • No network policies ("everything's in the cluster, it's fine").
  • Overly broad allow rules (allow all from namespace X) -- defeats the purpose.
  • Network policies without egress rules -- ingress-only policies still allow compromised pods to exfiltrate data.

6.2 Least-privilege service identities

Principle: Each service gets its own identity (Kubernetes ServiceAccount + SPIFFE SVID) with the minimum permissions needed. No shared service accounts.

Why it matters: Shared identities prevent granular authorization, audit trails, and revocation. If services A and B share an identity, you cannot authorize A without also authorizing B.

How to implement:

  • One Kubernetes ServiceAccount per workload (not per namespace).
  • RBAC bindings scoped to exactly what the service needs (specific API groups, resources, verbs).
  • Authorization policies reference specific service identities: "payment-svc can call order-svc on POST /orders/{id}/payment."
  • Regularly audit which identities have access to which services -- prune unused access.

Anti-patterns:

  • Default ServiceAccount used by all pods in a namespace.
  • Cluster-wide RBAC bindings for convenience.
  • Service identities with wildcard permissions ("allow all methods on all paths").
  • No audit of identity-to-service mappings.

6.3 Observability as a security control

Principle: Distributed tracing, access logs, and metrics from the service mesh are security controls, not just debugging tools. Monitor them for anomalies.

Why it matters: Zero trust assumes breach. Detection depends on visibility. If you can't see who called what, you can't detect lateral movement.

How to implement:

  • Enable access logging in the service mesh (Istio/Envoy access logs, Linkerd tap).
  • Distributed tracing (OpenTelemetry, Jaeger) with trace context propagated across all service calls.
  • Alert on anomalies: unexpected source-destination pairs, unusual request volumes, authorization denials.
  • Retain access logs long enough for incident investigation (30-90 days minimum).

Anti-patterns:

  • Disabling access logging for performance -- sample instead of disabling entirely.
  • Tracing only in development, not production.
  • No alerting on authorization policy denials -- failed access attempts are the signal.

7. API Documentation and Developer Experience

The OpenAPI document from §0.2 is also the input to docs, SDKs, and changelogs. Hand-written variants drift.

7.1 Render docs from the OpenAPI spec

Principle: Documentation is rendered automatically from the committed OpenAPI document by an open-source or SaaS tool -- never hand-written prose that drifts from code.

Why it matters: Hand-written docs are wrong within weeks of any active API. Spec-driven docs cannot drift because they are recompiled on every spec change. The 2024-2025 doc-rendering market shifted decisively: Scalar (open source, MIT, framework-native middleware) is the default new pick; Mintlify dominates the SaaS tier (markdown + OpenAPI, used by Anthropic, Microsoft, Coinbase). Swagger UI still works but is no longer the default new-project choice; Stoplight Elements development slowed sharply after the SmartBear acquisition.

How to implement:

  • New projects -- pick Scalar for OSS, Mintlify or ReadMe.com for SaaS. Use Redoc if you want a single-page reference style.
  • Docs build and deploy on every spec change; pipeline fails closed if the spec doesn't lint.
  • Host docs at a stable URL (docs.<domain> or <api-domain>/docs).
  • Validate examples against their schemas in CI -- Spectral rule oas3-valid-schema-example.

Anti-patterns:

  • Defaulting to Swagger UI in 2026 by habit -- Scalar is a near drop-in with better UX.
  • Hand-written docs maintained alongside the spec -- they always drift.
  • Picking Stoplight Elements expecting active development -- verify the roadmap.
  • SaaS doc vendor whose OpenAPI ingestion is brittle -- test with your real spec before committing.
  • Rendering docs from a build-time snapshot that is never re-validated against the deployed API.

7.2 Interactive playgrounds with real auth

Principle: Every endpoint in the docs is callable from the page. Auth is collected once at the top of the page; all subsequent requests sign automatically. Examples come from the spec, not from "string" / 0 placeholders.

Why it matters: The fastest path to a developer's first successful request is the only DX metric that matters. Code samples in their language reduce friction further. Playgrounds that proxy requests through the docs vendor leak credentials and break for CORS-restricted APIs.

How to implement:

  • Every endpoint shows code samples in cURL plus 4-6 SDK languages (TypeScript, Python, Go, Java, Ruby, PHP cover ~95% of demand). Generate samples from the same SDK pipeline so they cannot drift.
  • Persist credentials in browser session -- never proxy through the docs vendor.
  • Environment switcher (prod / staging / sandbox) is a first-class control.
  • Every operation has at least 2-3 named examples in the spec (examples: { minimal: ..., withMetadata: ... }); the playground pre-fills bodies from them.
  • Provide a Postman public workspace as a complement -- useful for forking and sharing collections, but not a substitute for the embedded portal.

Anti-patterns:

  • "Try it" that proxies requests through the docs vendor -- leaks creds, breaks CORS, hides the real network call.
  • Re-prompting for credentials per endpoint instead of persisting a session.
  • Auto-filled bodies as "string" / 0 placeholders -- developer must hand-build every payload.
  • Code samples written by hand, drifting behind the SDK; or only cURL shown, forcing developers to translate.

7.3 Standard error envelope: RFC 9457 Problem Details

Principle: Every error response uses the same envelope, defined once under components.schemas.Problem following RFC 9457 Problem Details for HTTP APIs. An error catalog page enumerates each type URI with its meaning, retryability, and remediation.

Why it matters: RFC 9457 (2023) replaced RFC 7807 with a clearer link between type URIs and the extension fields a client can expect. Per-endpoint error shapes force every consumer to write custom parsing logic. The catalog turns errors into documentation a developer can search. Cross-reference: §3.3 covers error handling generally; §7.3 is the docs-side counterpart.

How to implement:

  • Define one Problem schema with type (URI), title, status, detail, instance, plus your extensions (code, correlation_id, errors[] for validation). Use application/problem+json content type on error responses.
  • Every error response in the spec references #/components/schemas/Problem (or a refinement of it).
  • Error catalog page enumerates each type URI -- meaning, HTTP status, retryability, remediation, link to migration if deprecated.
  • Examples on every error response, with at least these named cases per endpoint where applicable: success, business-rule failure, validation failure, auth failure.

Anti-patterns:

  • Per-endpoint bespoke error shapes ({error: "..."} here, {message, code} there) -- clients can't write one error handler.
  • Single example per endpoint when the realistic case has 4+ shapes.
  • Documenting only HTTP status codes without an application-level type URI taxonomy -- clients branch on prose message strings.
  • Examples that don't validate against the schema -- run oas3-valid-schema-example in CI.

7.4 Generated SDKs

Principle: SDKs are generated from the OpenAPI spec on every change. The generator opens a PR against the SDK repo so changes can be reviewed and released on a deliberate cadence.

Why it matters: Hand-written SDKs maintained by the API team always fall behind the spec. A generated SDK cannot drift. Stripe's pipeline -- a single internal definition fanning out to ~10 SDKs on a daily release cadence -- is the reference architecture.

How to implement:

  • For public/customer-facing SDKs -- Stainless (used by Stripe, Anthropic, Cloudflare, OpenAI) or Speakeasy (10 languages, single-runtime-dependency TS output, ships as a self-contained binary). Fern is a credible alternative.
  • For internal stubs and prototypes -- openapi-generator (open source, 50+ languages); generated quality is uneven, fine for internal use but weak for public SDKs without heavy template customisation.
  • A good SDK includes: idiomatic per-language style; automatic retry with exponential backoff + jitter on 429/503; transparent pagination (iterator/async-iterator hides cursor mechanics); typed errors as a discriminated union; built-in auth helpers (OAuth refresh, key rotation); webhook signature verification; small dependency footprint.
  • Pin SDK versions to dated API versions (Stripe model) -- see §7.5.

Anti-patterns:

  • Hand-written SDKs maintained by the API team -- always behind the spec.
  • openapi-generator default templates shipped as a public SDK -- surfaces deprecated framework calls, looks non-idiomatic.
  • SDKs with no retry, no pagination helper, untyped errors -- every consumer rebuilds the same plumbing.
  • SDKs pulling 25-40 transitive dependencies for a thin HTTP wrapper.

7.5 Changelog, versioning UX, and deprecation signals

Principle: API changes are communicated both in docs and on the wire. Wire signals are RFC-defined headers (Deprecation, Sunset, Link); the docs side is a machine-readable changelog with migration guides linked from each entry.

Why it matters: Deprecating an endpoint in docs only means clients on old SDKs never learn. Sunset dates without migration guides leave developers with nowhere to go. Mature APIs use both channels. Cross-reference: §3.1 covers versioning models; §7.5 is the deprecation-comms counterpart.

How to implement:

  • Wire signals -- Deprecation: @<unix-ts> (RFC 9745, finalized 2024); Sunset: <http-date> (RFC 8594, 2019); Link: <migration-url>; rel="deprecation" plus rel="successor-version".
  • Versioning models -- pick one consciously:
    • Stripe-style dated versions (2026-04-22.<codename>) -- SDK clients pin a version; minor releases backward-compatible; breaking changes cluster into named majors. Best for fine-grained evolution.
    • GitHub-style X-API-Version header -- calendar-dated, opt-in. Lighter weight than Stripe's model.
    • URL-path versioning (/v1/, /v2/) -- coarse-grained, simple. Acceptable for small APIs; awkward for fine-grained evolution. Used by Twilio and SendGrid for major boundaries.
  • Generate the changelog from oasdiff output, filterable by API version, linking every entry to a migration guide.
  • Provide an RSS/Atom or JSON feed of changes so customers can wire alerts.
  • Deprecation lead time -- 12 months minimum for public APIs (Twilio policy), 6+ for partner APIs.

Anti-patterns:

  • Deprecating in docs only with nothing on the wire -- clients on old SDKs never learn.
  • Sunset dates without migration guides.
  • Major-version-bump-only versioning (/v1/v2) for small additive changes -- forces clients to choose between staying on a frozen API or rewriting everything.
  • Changelog as a hand-written prose blog with no machine-readable feed and no link to API version.
  • "Breaking changes" buried in release notes without explicit replacement field/endpoint pointers.

8. Contract Testing and API Quality

Schema validation, contract tests, and drift detection turn the OpenAPI document from §0.2 into a executable contract that the running service must obey.

8.1 Schema validation in the existing test suite

Principle: Every test that exercises an HTTP handler validates both the request and response against the committed OpenAPI spec, in-process. This is the foundation of contract testing -- cheap, fast, integrated with whatever test framework you already use.

Why it matters: Spec drift kills consumers. A middleware that asserts every request and response against the spec catches "spec lies, code is right" drift the moment it appears. Validating only requests (and trusting framework parsing for responses) is the most common gap -- response drift is the more common production bug.

How to implement:

  • Per language:
    • Python -- openapi-core (Flask, Django, Falcon, Starlette, Werkzeug, Requests integrations).
    • Node -- express-openapi-validator (auto-validates requests, responses, security).
    • Ruby -- committee (Rack middleware + Committee::Test::Methods test helpers).
    • Java/Kotlin -- springdoc-openapi + Spring REST Docs, or atlassian/swagger-request-validator for MockMvc/RestAssured.
    • Go -- kin-openapi (openapi3filter).
  • Validate both directions in tests, not just at runtime.
  • Reload the spec on every test run -- never cache it across runs in CI.
  • Validation failures must be build-failing; warning-level is ignored within a sprint.

Anti-patterns:

  • Validating only requests, not responses -- response drift is the more common bug.
  • Hand-written JSON Schema next to the OpenAPI doc -- they will diverge.
  • Validation failures as warnings instead of build failures.
  • Loading the spec once at app startup and never re-loading in tests, so spec edits don't reach the validator.

8.2 Provider verification vs consumer-driven contracts

Principle: Provider-side spec compliance (Schemathesis, Dredd) is the default. Consumer-driven contracts (Pact) are a deliberate add-on for APIs with a small, known set of internal consumers.

Why it matters: CDC with Pact is high-value when you have a mobile app + web SPA + a couple of internal services that talk to one provider -- each consumer publishes a contract describing what it actually calls; the provider verifies against all of them and uses can-i-deploy to gate releases. CDC is overkill (and frequently abandoned) when the provider has many unknown consumers, when teams are organisationally distant, or when an OpenAPI-first workflow already gives you provider compliance via §8.1 and §8.3.

How to implement:

  • Default -- Schemathesis (provider-side, OpenAPI-driven, see §8.3) plus the schema validation in §8.1.
  • Add Pact when -- you publish SDKs you control, you have a known set of internal consumers, or consumer teams want guarantees independent of the provider's tests.
  • For Pact -- store contracts in PactFlow / Pact Broker. Gate provider deploys on can-i-deploy. The contract is not advisory -- the broker is part of the release pipeline.
  • For mixed worlds (some consumers known, many not) -- bi-directional contract testing in PactFlow accepts the OpenAPI document as the provider contract; gives you provider compliance without per-consumer Pacts.

Anti-patterns:

  • Adopting Pact organisation-wide for an API with public/unknown consumers -- you can't enumerate the contracts.
  • Consumer Pacts using random data -- every run produces a "new" contract; the broker fills with noise.
  • Provider verification as advisory (not gating deploys) -- the contract becomes documentation only.
  • Conflating CDC with end-to-end testing -- they answer different questions, and CDC should not stand in for E2E nor vice-versa.

8.3 Property-based fuzz testing the spec

Principle: Run Schemathesis against the running service in CI. It generates conformant requests from the OpenAPI document, exercises endpoints in stateful sequences derived from links, and asserts behavioural properties -- no 5xx where 4xx is expected, response bodies match the declared response schema, status codes are documented, security boundaries hold.

Why it matters: Unit tests probe values you thought of; property-based fuzzing probes values the schema says are legal but you never tried -- Unicode, edge integers, deeply nested optional fields, missing-but-valid combinations. Schemathesis caught 1.4×-4.5× more defects than competing API fuzzers in published evaluations; production users include Spotify, JetBrains, Red Hat, WordPress.

How to implement:

  • Adopt Schemathesis 4.x. Use stateful mode (st fuzz) -- most real bugs are sequence-dependent (create → read returns wrong shape) and stateless fuzzing can't find them.
  • Run against a real service backed by a real database, not mocks -- many bugs only surface with persistent state.
  • Wire reports into CI -- Allure, JUnit, or HAR for replaying failures.
  • Use after_validate hooks for custom invariants beyond what the schema declares (e.g. "the same id must reappear on subsequent GETs").
  • For dynamic auth (OAuth, refresh tokens), Schemathesis 4.13+ has first-class config -- no Python glue code needed.

Anti-patterns:

  • Running Schemathesis only against a mocked dev server -- misses bugs that depend on real DB state and auth.
  • Excluding 5xx checks because they're "flaky" -- a 5xx on a schema-valid request is the most important signal Schemathesis produces. Fix the bug, don't suppress the check.
  • Running it once and suppressing all failures with --exclude -- hides drift forever.
  • Skipping stateful mode -- most real bugs are sequence-dependent.

8.4 Drift detection in CI

Principle: Three artefacts must stay in sync -- (a) the spec in the repo, (b) the spec the running service exposes, (c) the behaviour of the deployed production service. CI must detect drift between any pair.

Why it matters: A code-first project without a "regenerate-and-diff" CI step will drift -- the only question is when. Runtime drift between deployed-API behaviour and the documented spec is the most expensive failure because consumer SDKs were generated against the documented version.

How to implement:

  • Spec-vs-spec on PR -- oasdiff (Go CLI + GitHub Action) checks 450+ categories of breaking changes; non-zero exit gates the PR. Pair with Spectral for style/governance.
  • Code-vs-spec in CI -- for code-first stacks, regenerate the spec from code in CI and oasdiff against the committed spec -- fail the build if non-empty.
  • Runtime-vs-spec -- sample real traffic and validate against the spec. Optic captures HTTP traffic and diffs against the spec; Speakeasy offers SDK-driven runtime drift detection.
  • Breaking-change detection must be a required PR check, not a Slack notification.

Anti-patterns:

  • Code-first project with no spec regeneration step -- the committed spec rots silently.
  • Breaking-change detection as a Slack notification rather than a required PR check.
  • Linting only on a schedule rather than per-PR -- drift accumulates between runs.
  • No production sampling -- the deployed API can diverge from spec without anyone noticing until a consumer breaks.

8.5 The API test pyramid

Principle: Layer tests as -- unit → schema-validated handler tests (§8.1) → contract tests (§8.2 / §8.3) → integration (real DB, real downstreams) → smoke tests of the deployed environment. Contract tests sit between unit and integration: faster than integration, broader in scope than unit.

Why it matters: 100% mocked tests with no contract layer encode developer assumptions, not the actual provider's behaviour. The classic failure mode -- "all green in CI, broken in prod" -- comes from this gap. A contract layer breaks the dependency on slow, flaky integration environments while still catching real interface drift.

How to implement:

  • Unit tests cover pure logic with no I/O -- fastest, most numerous.
  • Schema-validated handler tests (§8.1) cover most contract-shape concerns in-process.
  • Contract tests (§8.2 / §8.3) hit the running service, no mocks. Run on every PR.
  • Integration tests cover real DB + real downstream stubs (or testcontainers). Slower; run on PR but fewer in number.
  • Post-deploy smoke tests are a tiny set of canary checks (login, list, create one resource, delete it) against the deployed environment -- catch infra/config drift (TLS, auth proxy, env var) that contract tests cannot.

Anti-patterns:

  • 100% mocked tests with no contract layer -- mocks encode assumptions, drift is invisible until prod.
  • Contract tests against a stale spec -- the suite passes but the spec doesn't match deployed behaviour. Pair with §8.4.
  • Skipping post-deploy smoke tests because "we have contract tests" -- contract tests don't catch infra drift.
  • Inverted pyramid (UI-test-heavy "ice cream cone") -- slow, flaky, brittle.
  • Treating Schemathesis as integration tests -- it's contract verification, not business-flow testing; you still need a small set of curated end-to-end scenarios.

OWASP API Security Top 10 (2023) Quick Reference

For context, the current OWASP API Security Top 10 maps to the practices above:

# Risk Where addressed
API1 Broken Object-Level Authorization Section 2.4
API2 Broken Authentication Sections 2.1, 2.2, 2.3
API3 Broken Object Property-Level Authorization Section 2.4, 4.1
API4 Unrestricted Resource Consumption Sections 3.5, 4.3
API5 Broken Function-Level Authorization Section 2.4
API6 Unrestricted Access to Sensitive Business Flows Sections 3.4, 3.5
API7 Server-Side Request Forgery Section 4.2
API8 Security Misconfiguration Sections 1.1, 6.1
API9 Improper Inventory Management Section 3.1
API10 Unsafe Consumption of APIs Section 4.2

Sources

Security and operations (§1-§6)

API-first methodology (§0)

Documentation and DX (§7)

Contract testing and quality (§8)