DR replay runbook
Disaster-recovery rehearsal procedure for AlphaSwarm. Targets:
- RPO 1 hour for
alphaswarm_admin+ control-plane services. - RTO 4 hours for the same.
- RPO 15 minutes for trading-relevant data.
- RTO 1 hour for the same.
The exercise is run quarterly (calendar reminder owned by the platform team). The first exercise is scheduled for the end of Phase 5 of the multi-account overhaul.
Pre-requisites
- AWS Organizations + Control Tower applied (Phase 4 complete).
- ArgoCD app-of-apps applied to dev + staging + prod clusters.
- Velero installed on every workload cluster (chart at alphaswarm_platform/deployments/kubernetes/helm/velero).
- ECR cross-region replication active to
us-west-2. - RDS cross-region read replica green.
- S3 CRR active on every Parquet + audit-archive bucket.
- Route 53 health-check failover record set on the
manage.alpha-swarm.aiingress.
Steps
1. Trigger the failure
Pick the rehearsal target — typically alphaswarm-dev (never prod).
Document the start time in the incident ticket.
# Disable the dev cluster's API server (simulates a control-plane outage).
aws eks update-cluster-config \
--name alphaswarm-dev \
--region us-east-1 \
--resources-vpc-config endpointPrivateAccess=false,endpointPublicAccess=false
2. Confirm impact
alphaswarm_admin should now show unreachable for the dev cluster
under /admin/kubernetes/status. The KillSwitch should still
work because it fans out to other clusters too.
3. Bring up the replay cluster
cd infrastructure/envs/dev
terraform apply -var-file=terraform.tfvars
This re-creates the EKS cluster with the same name + node groups.
ArgoCD picks up the new cluster via its Cluster generator (label
alphaswarm.io/managed=true).
4. Replay state from Velero
velero backup-location get
velero restore create dr-replay-$(date +%s) \
--from-backup daily-full-$(velero backup get | tail -1 | awk '{print $1}')
5. Restore RDS
The cross-region read replica in us-west-2 is promoted to
primary; the DR replay points the dev cluster's RDS DSN at the
new primary. The Postgres instance comes up with the audit ledger
intact so no admin actions are lost.
6. Verify
alphaswarm_adminhealth should return 200 within 4h.- The audit ledger should show the gap as a single contiguous block (no missing rows beyond the RPO window).
- Paper-trading runs that were active are stamped
status=haltedby the watchdog. - The ArgoCD app-of-apps sync should converge within 15min after the cluster comes back.
7. Document
Append to the rehearsal log at
alphaswarm_docs/docs/operations/dr-rehearsal-log.md with:
- Start / end timestamps.
- Actual RPO + RTO measured.
- Issues encountered + remediations.
- Sign-off from the security officer.
Game-day: DR replay from encrypted state and WORM ledger
Phase 4.6(c) game-day from the Unified Infrastructure Control Plane Plan (private
alphaswarm_internalplanning repo) (§4, item 4.6). This extends the Velero/RDS rehearsal above with the control-plane's own recovery story: restoring a cell from OpenTofu-encrypted state with its per-cell KMS key, and proving the hash-chained WORM audit ledger survived intact. It rehearses the mechanisms decided in ADR 028 — OpenTofu cutover and the audit story in Central Deployment Control — Next Steps.
Objective
Prove that a cell can be fully rebuilt from its encrypted OpenTofu state using
only the per-cell KMS key, and that the tamper-evident audit ledger for that
cell verifies clean end-to-end after the recovery. Success = a restored cell
whose next plan shows no diff and an audit chain that audit_verify
accepts.
Grounded mechanisms
- Encrypted state (ADR 028): the OpenTofu
encryption{}state block (AWS-KMS key provider, per-cell key) renders only when the resolved binary istofu. ADR 028 §4 specifies the exact validation this game-day performs: "validating state-encryption round-trips (encrypt → destroy state → restore with the KMS key)" and-json/exit-code parity against the terraform baseline. State is owned by exactly one binary at a time (no dual-write). - WORM ledger: the hash-chained JSONL ledger is fanned out to the monolith
Postgres ledger and archived to S3 Object Lock COMPLIANCE (SSE-KMS, ≥6y)
by
WormUploader(python -m alphaswarm_controller.terraform.audit_worm). - Integrity verifier:
python -m alphaswarm_controller.terraform.audit_verifywalks the chain (entry_hash = sha256(prev_hash ‖ canonical(row))) and exits non-zero on a chain break (sev-1) — the pass/fail oracle for this game-day.
Preconditions / scope
- A rehearsal cell whose IaC runs under
tofuwith encryption enabled (never a production silo-reg cell for the first run). - Access to the cell's per-cell KMS key (and a way to test the negative case: attempt a restore without the key and confirm it fails).
- The cell's WORM bucket and a recent archived ledger snapshot are present.
- Operator holds
manage:infrastructurefor the restore plan andadmin:clusterfor any apply; four-eyes for the apply.
Roles
| Role | Responsibility |
|---|---|
| DR lead (SRE) | Runs the restore, holds the incident ticket |
| Security officer | Witnesses the KMS-key-only restore and signs the chain-verify result |
| Approver | Distinct four-eyes approver for the restore apply |
| Scribe | Records measured RTO/RPO, the no-diff plan, and the audit_verify exit code |
Step-by-step procedure
- Record the target state. Note the cell's encrypted state object location and the tip hash of its WORM-archived ledger.
- Simulate loss. Following the state-encryption round-trip discipline, destroy/withdraw the local state so recovery must come from the encrypted backend copy.
- Restore with the KMS key. Reconfigure the backend with the cell's per-cell
KMS key and run a governed
planon the cell's workspace through the One Gate. OpenTofu decrypts the state with the KMS key provider. - Negative sub-case (key-custody proof). Repeat step 3 with the KMS key denied; confirm the restore fails (state cannot be decrypted) — this proves custody of the key is load-bearing.
- Confirm no drift. The restored
planmust show no diff versus the pre-loss cell (ADR 028's no-diff migration gate), and-json/exit-code behaviour must match the terraform baseline. - Pull the WORM ledger snapshot for the cell from the Object Lock bucket.
- Verify the chain. Run
python -m alphaswarm_controller.terraform.audit_verifyagainst the recovered ledger. Exit0= intact; exit ≠0= chain break (fail the game-day, raise sev-1). - Tamper sub-case (verifier proof). Mutate one archived row in a copy and
re-run
audit_verify; confirm it exits non-zero — proving the verifier actually catches breaks. - Document in the rehearsal log alongside the RPO/RTO section above.
Success criteria (measurable)
- The cell restored from encrypted state using only the per-cell KMS key; the key-denied sub-case failed to restore.
- The post-restore
planshowed no diff and matched-json/exit-code parity. audit_verifyexited 0 on the recovered ledger; the tamper sub-case exited non-zero.- Measured RTO/RPO within the targets at the top of this runbook.
Abort / rollback
- If the restored
planshows an unexpected diff, do not apply — the state or the encryption context is wrong; stop and reconcile (never-auto-approvea drifted restore). - If
audit_verifyfails, treat the ledger as compromised: preserve the WORM object (Object Lock prevents deletion before retention), open a sev-1, and do not overwrite the archive. - Recovery of the cluster/data layer (Velero/RDS) rolls back via the main DR procedure above; this game-day adds no destructive action beyond the disposable rehearsal cell.
Evidence capture (WORM / audit ledger)
- The restore
planoutput (no-diff proof) and the-jsonparity capture. - The
audit_verifyexit codes for both the clean and tamper sub-cases. - The negative key-denied restore failure log.
- The recovered WORM ledger snapshot itself is the immutable evidence (COMPLIANCE-locked); its tip hash is recorded in the rehearsal log.
Frequency / owner
- Frequency: quarterly, folded into the DR rehearsal cadence above; also
after any ADR 028 cutover step that flips a cell to
tofu/encrypted state. - Owner:
sre-team, security officer co-signs the chain-verify result.