ADR-003: Chaos-lite resilience policy
Status¶
Accepted
Date¶
2026-07-24
Context¶
We need confidence that staging failures (Redis down, worker crash, rate limits, hold TTL) behave predictably before PHI beta. Full Chaos Mesh / Gremlin / random prod killing is premature without mature metrics and multi-env ops.
Decision¶
Adopt chaos-lite:
- Automated
tests/resilience/suite in CI (deterministic fault helpers) - Hostile cases in integration tests (refresh reuse, stolen grant, oversized upload, IDOR)
- Staging-only load/soak scripts (not CI-blocking)
- Monthly scripted game day (
docs/ops/GAME_DAY.md) with scribe evidence - Backup restore drill (
docs/ops/BACKUP_RESTORE.md)
Out of scope: Chaos Monkey in production; FLUSHALL on shared Redis; random process killing against PHI data.
Alternatives Considered¶
Full chaos engineering platform now¶
- Pros: Broader fault coverage
- Cons: Ops overhead; flaky CI; no staging maturity yet
- Deferred until staging + metrics exist
Manual-only game days¶
- Pros: Cheap
- Cons: Regressions slip between months; no PR gate
- Rejected as sole strategy
Consequences¶
- Resilience suite must stay deterministic (no random kill in CI)
- Game-day evidence is a PHI beta gate
- Fault classes tagged in logs/Sentry so experiments ≠ mystery 500s
- See Phases 8b–10 in tasks/plan.md
