Skip to main content

Game-Day: Kill-Switch Fleet Halt + Resume

Phase 4.6(b) game-day from the Unified Infrastructure Control Plane Plan (private alphaswarm_internal planning repo) (relocated out of this repo on 2026-07-18; see ADR 026) (§4, item 4.6). Rehearses the controller-level kill-switch — the fail-closed, Redis-backed, cluster-wide halt that freezes both the IaC and workload mutation paths — and proves a clean resume afterwards. This is the fleet halt described in the plan's Pillar-C operating model (§5, step 6) and in ADR 026; it is distinct from, and complementary to, the bots-operator KillSwitch CRD.

Objective​

Prove that engaging the fleet halt blocks every mutating operation cluster-wide while it is set — a governed terraform apply is refused, and in-flight workload runs abort — and that clearing the halt returns the fleet to a clean, fully-operational state. Success requires a refused gated apply during the halt and a successful gated apply after resume, both proven from the audit ledger.

Grounded mechanisms​

SurfaceEndpointBehaviourScope
IaC gatePOST /manage/terraform/haltTouches the kill-switch sentinel; subsequent apply/destroy return status='rejected' without invoking the executor until cleared. DELETE clears it; GET /manage/terraform/halt/status inspects itadmin:cluster + step-up
WorkloadsPOST /manage/workloads/haltKill-switch fan-out: signals every in-flight WorkloadRun to abort; each affected run writes its own status=HALTED audit row; idempotentworkloads:halt
StatusGET /manage/workloads/halt/statusReturns in-flight WorkloadRun count + whether the global halt flag is set — used by smoke tests to verify the halt fan-out workedread:infrastructure

Backing invariants (all real, all fail-closed):

  • Redis-backed shared store (ALPHASWARM_CP_TERRAFORM_HALT_REDIS_URL): the halt fans across every controller replica and the plan-binding survives restarts. should_halt() halts on any error — a Redis outage fails to the halted (safe) state, not the open one. Anchor: Central Deployment Control — Next Steps §3.2, §7.
  • Order-gate honesty: for live-trading workloads the halt sets the order-gate key, so a halt that claims to stop trading actually stops it (plan §5, step 6).
  • Re-gate at consumption: execution re-checks the fresh kill-switch state before applying a reviewed plan, so a plan approved before the halt still cannot apply during it.

Preconditions / scope​

  • Environment: a non-production cell first; only promote to production once the dev/qa run is clean. The controller must be pointed at the shared Redis halt store (multi-replica halt is the property under test).
  • A disposable, no-op-diff terraform workspace to use as the gated-apply probe (e.g. a tag-only change), so the "refused during halt / applied after resume" assertions do not mutate anything meaningful.
  • At least one benign in-flight WorkloadRun (or the ability to start one) so the workloads/halt fan-out has something to halt and count.
  • Operator holds admin:cluster, workloads:halt, and can complete step-up MFA (max_age=180s); a distinct approver is available for the post-resume prod-tier apply (four-eyes).

Roles​

RoleResponsibility
Halt operator (SRE)Engages/clears the halt; runs the gated-apply probe
ApproverProvides the distinct X-Approver-Authorization for the post-resume prod apply
Platform on-callWatches replica halt propagation + Redis-outage sub-case
ScribeRecords triggered_at, halted_count, and audit_run_ids

Step-by-step procedure​

  1. Baseline. GET /manage/terraform/halt/status and GET /manage/workloads/halt/status — confirm not-engaged and record the in-flight WorkloadRun count. Drive one PLAN→APPLY on the probe workspace to confirm the gate applies normally before the halt.
  2. Engage the IaC halt. POST /manage/terraform/halt with a reason (admin:cluster + step-up). Confirm GET .../halt/status reports engaged on every replica (the Redis-backed property).
  3. Engage the workloads halt. POST /manage/workloads/halt with a reason. Confirm GET /manage/workloads/halt/status shows the halt flag set and the fan-out reached the in-flight runs.
  4. Prove mutating ops are blocked (the core assertion). Drive a PLAN→APPLY on the probe workspace. The apply MUST return status='rejected' without invoking the executor. Repeat with a plan that was approved before the halt — it must still be refused (re-gate at consumption).
  5. Prove the fan-out halted work. Confirm each affected WorkloadRun wrote a status=HALTED row and the halt/status count matches.
  6. Redis-outage sub-case (fail-closed proof). Temporarily sever the shared Redis; confirm the gate still refuses applies (should_halt() fails to the halted state). Restore Redis before resuming.
  7. Resume. Clear the IaC halt with DELETE /manage/terraform/halt; clear the workloads halt registry. Confirm both .../halt/status endpoints report not-engaged / flag false.
  8. Prove clean resume. Drive a fresh PLAN→APPLY (with a distinct approver for prod tier) on the probe workspace — it MUST apply successfully and land a ledger row. New WorkloadRuns are accepted again.
  9. Close the ticket with the refused-then-applied evidence.

Success criteria (measurable)​

  • During the halt, a gated terraform apply returned status='rejected' and the executor was never invoked (no plan output, no state lock taken).
  • A pre-halt-approved plan was also refused during the halt (re-gate proven).
  • GET /manage/workloads/halt/status reported the halt flag set and the in-flight count dropped as runs wrote status=HALTED.
  • Halt state was visible on all controller replicas (cluster-wide), and the Redis-outage sub-case still refused applies (fail-closed).
  • After DELETE /manage/terraform/halt, both status endpoints reported not-engaged and a fresh gated apply succeeded — clean resume with no manual state surgery.

Abort / rollback​

  • The kill-switch is the abort mechanism, so the game-day's own rollback is simply to leave the halt engaged if anything looks wrong and escalate — a stuck-halted fleet is the safe failure mode.
  • If a replica fails to observe the halt (split state), do not proceed to resume; treat divergent halt/status across replicas as a P1 and fix the shared-Redis wiring before clearing.
  • If DELETE /manage/terraform/halt does not clear (status still engaged), escalate rather than editing the sentinel by hand; break-glass is the last resort (two named operators, 60-minute auto-expiry, HIGH-severity Security Hub finding).

Evidence capture (WORM / audit ledger)​

  • TerraformHaltResponse fields: kill_switch_path, reason, triggered_at, user_id — captured on engage.
  • The refused apply run row (status='rejected') and the post-resume applied run row — the before/after pair is the proof the halt was total and the resume was clean.
  • Every status=HALTED WorkloadRun row with its HaltRequest.reason.
  • All rows are hash-chained (entry_hash = sha256(prev_hash ‖ canonical(row))), fanned out to the monolith Postgres ledger, and archived to S3 Object Lock COMPLIANCE WORM; verify the chain with python -m alphaswarm_controller.terraform.audit_verify (exit≠0 on a break).

Frequency / owner​

  • Frequency: quarterly (aligns with the existing kill-switch drill cadence in Kill-Switch Incident Response), and after any change to the halt store, Redis wiring, or authz matrix.
  • Owner: sre-team.