ADR-003: Chaos-lite resilience policy


Sep 28, 2026

ACCEPTED

Md Khaled Bin Joha

Status

Accepted

Date

2026-07-24

Context

We need confidence that staging failures (Redis down, worker crash, rate limits, hold TTL) behave predictably before PHI beta. Full Chaos Mesh / Gremlin / random prod killing is premature without mature metrics and multi-env ops.

Decision

Adopt chaos-lite:

  1. Automated tests/resilience/ suite in CI (deterministic fault helpers)
  2. Hostile cases in integration tests (refresh reuse, stolen grant, oversized upload, IDOR)
  3. Staging-only load/soak scripts (not CI-blocking)
  4. Monthly scripted game day (docs/ops/GAME_DAY.md) with scribe evidence
  5. Backup restore drill (docs/ops/BACKUP_RESTORE.md)

Out of scope: Chaos Monkey in production; FLUSHALL on shared Redis; random process killing against PHI data.

Alternatives Considered

Full chaos engineering platform now

  • Pros: Broader fault coverage
  • Cons: Ops overhead; flaky CI; no staging maturity yet
  • Deferred until staging + metrics exist

Manual-only game days

  • Pros: Cheap
  • Cons: Regressions slip between months; no PR gate
  • Rejected as sole strategy

Consequences

  • Resilience suite must stay deterministic (no random kill in CI)
  • Game-day evidence is a PHI beta gate
  • Fault classes tagged in logs/Sentry so experiments ≠ mystery 500s
  • See Phases 8b–10 in tasks/plan.md