Game-Day: Kill-Switch Fleet Halt + Resume
Phase 4.6(b) game-day from the Unified Infrastructure Control Plane Plan (private
alphaswarm_internalplanning repo) (relocated out of this repo on 2026-07-18; see ADR 026) (§4, item 4.6). Rehearses the controller-level kill-switch — the fail-closed, Redis-backed, cluster-wide halt that freezes both the IaC and workload mutation paths — and proves a clean resume afterwards. This is the fleet halt described in the plan's Pillar-C operating model (§5, step 6) and in ADR 026; it is distinct from, and complementary to, the bots-operator KillSwitch CRD.
Objective
Prove that engaging the fleet halt blocks every mutating operation
cluster-wide while it is set — a governed terraform apply is refused, and
in-flight workload runs abort — and that clearing the halt returns the fleet to a
clean, fully-operational state. Success requires a refused gated apply during
the halt and a successful gated apply after resume, both proven from the audit
ledger.
Grounded mechanisms
| Surface | Endpoint | Behaviour | Scope |
|---|---|---|---|
| IaC gate | POST /manage/terraform/halt | Touches the kill-switch sentinel; subsequent apply/destroy return status='rejected' without invoking the executor until cleared. DELETE clears it; GET /manage/terraform/halt/status inspects it | admin:cluster + step-up |
| Workloads | POST /manage/workloads/halt | Kill-switch fan-out: signals every in-flight WorkloadRun to abort; each affected run writes its own status=HALTED audit row; idempotent | workloads:halt |
| Status | GET /manage/workloads/halt/status | Returns in-flight WorkloadRun count + whether the global halt flag is set — used by smoke tests to verify the halt fan-out worked | read:infrastructure |
Backing invariants (all real, all fail-closed):
- Redis-backed shared store (
ALPHASWARM_CP_TERRAFORM_HALT_REDIS_URL): the halt fans across every controller replica and the plan-binding survives restarts.should_halt()halts on any error — a Redis outage fails to the halted (safe) state, not the open one. Anchor: Central Deployment Control — Next Steps §3.2, §7. - Order-gate honesty: for live-trading workloads the halt sets the order-gate key, so a halt that claims to stop trading actually stops it (plan §5, step 6).
- Re-gate at consumption: execution re-checks the fresh kill-switch state before applying a reviewed plan, so a plan approved before the halt still cannot apply during it.
Preconditions / scope
- Environment: a non-production cell first; only promote to production once the dev/qa run is clean. The controller must be pointed at the shared Redis halt store (multi-replica halt is the property under test).
- A disposable, no-op-diff terraform workspace to use as the gated-apply probe (e.g. a tag-only change), so the "refused during halt / applied after resume" assertions do not mutate anything meaningful.
- At least one benign in-flight
WorkloadRun(or the ability to start one) so theworkloads/haltfan-out has something to halt and count. - Operator holds
admin:cluster,workloads:halt, and can complete step-up MFA (max_age=180s); a distinct approver is available for the post-resume prod-tier apply (four-eyes).
Roles
| Role | Responsibility |
|---|---|
| Halt operator (SRE) | Engages/clears the halt; runs the gated-apply probe |
| Approver | Provides the distinct X-Approver-Authorization for the post-resume prod apply |
| Platform on-call | Watches replica halt propagation + Redis-outage sub-case |
| Scribe | Records triggered_at, halted_count, and audit_run_ids |
Step-by-step procedure
- Baseline.
GET /manage/terraform/halt/statusandGET /manage/workloads/halt/status— confirm not-engaged and record the in-flightWorkloadRuncount. Drive one PLAN→APPLY on the probe workspace to confirm the gate applies normally before the halt. - Engage the IaC halt.
POST /manage/terraform/haltwith areason(admin:cluster+ step-up). ConfirmGET .../halt/statusreports engaged on every replica (the Redis-backed property). - Engage the workloads halt.
POST /manage/workloads/haltwith areason. ConfirmGET /manage/workloads/halt/statusshows the halt flag set and the fan-out reached the in-flight runs. - Prove mutating ops are blocked (the core assertion). Drive a PLAN→APPLY on
the probe workspace. The apply MUST return
status='rejected'without invoking the executor. Repeat with a plan that was approved before the halt — it must still be refused (re-gate at consumption). - Prove the fan-out halted work. Confirm each affected
WorkloadRunwrote astatus=HALTEDrow and the halt/status count matches. - Redis-outage sub-case (fail-closed proof). Temporarily sever the shared
Redis; confirm the gate still refuses applies (
should_halt()fails to the halted state). Restore Redis before resuming. - Resume. Clear the IaC halt with
DELETE /manage/terraform/halt; clear the workloads halt registry. Confirm both.../halt/statusendpoints report not-engaged / flagfalse. - Prove clean resume. Drive a fresh PLAN→APPLY (with a distinct approver for
prod tier) on the probe workspace — it MUST apply successfully and land a
ledger row. New
WorkloadRuns are accepted again. - Close the ticket with the refused-then-applied evidence.
Success criteria (measurable)
- During the halt, a gated
terraform applyreturnedstatus='rejected'and the executor was never invoked (no plan output, no state lock taken). - A pre-halt-approved plan was also refused during the halt (re-gate proven).
GET /manage/workloads/halt/statusreported the halt flag set and the in-flight count dropped as runs wrotestatus=HALTED.- Halt state was visible on all controller replicas (cluster-wide), and the Redis-outage sub-case still refused applies (fail-closed).
- After
DELETE /manage/terraform/halt, both status endpoints reported not-engaged and a fresh gated apply succeeded — clean resume with no manual state surgery.
Abort / rollback
- The kill-switch is the abort mechanism, so the game-day's own rollback is simply to leave the halt engaged if anything looks wrong and escalate — a stuck-halted fleet is the safe failure mode.
- If a replica fails to observe the halt (split state), do not proceed to
resume; treat divergent
halt/statusacross replicas as a P1 and fix the shared-Redis wiring before clearing. - If
DELETE /manage/terraform/haltdoes not clear (status still engaged), escalate rather than editing the sentinel by hand; break-glass is the last resort (two named operators, 60-minute auto-expiry, HIGH-severity Security Hub finding).
Evidence capture (WORM / audit ledger)
TerraformHaltResponsefields:kill_switch_path,reason,triggered_at,user_id— captured on engage.- The refused apply run row (
status='rejected') and the post-resume applied run row — the before/after pair is the proof the halt was total and the resume was clean. - Every
status=HALTEDWorkloadRunrow with itsHaltRequest.reason. - All rows are hash-chained (
entry_hash = sha256(prev_hash ‖ canonical(row))), fanned out to the monolith Postgres ledger, and archived to S3 Object Lock COMPLIANCE WORM; verify the chain withpython -m alphaswarm_controller.terraform.audit_verify(exit≠0 on a break).
Frequency / owner
- Frequency: quarterly (aligns with the existing kill-switch drill cadence in Kill-Switch Incident Response), and after any change to the halt store, Redis wiring, or authz matrix.
- Owner:
sre-team.