Game-Day: Zero-Orphaned-Order Drain During Trading Hours
Phase 4.6(a) game-day from the Unified Infrastructure Control Plane Plan (private
alphaswarm_internalplanning repo) (relocated out of this repo on 2026-07-18; see ADR 026) (§4, item 4.6). This is the repeatable exercise that proves the drain choreography leaves no open order behind when a live-trading pod is terminated during market hours. The mechanisms it rehearses are the ones shipped for money-plane workloads — it does not introduce new commands.
Objective
Prove, with evidence, that terminating a live-trading workload pod during trading hours results in zero open orders at the venue at the moment of termination, because the drain choreography either (a) cancels/flattens every open order before the container exits, or (b) blocks termination until it can. The exit criterion mirrors the plan's Phase 4 gate: the game-day passes only with evidence archived to the WORM ledger.
Grounded mechanisms
The drain choreography is a layered fail-closed sequence. Each layer is a real mechanism this program built or mandated:
| Layer | Mechanism | Source |
|---|---|---|
| Scheduler | Karpenter karpenter.sh/do-not-disrupt annotation on live-trading pods so the autoscaler never voluntarily evicts them mid-session | Phase 4.5 progressive-delivery guardrails |
| Availability | PodDisruptionBudget (PDB) blocks voluntary evictions that would drop below the trading-critical replica floor | K8s-native drain layer |
| Container | preStop hook issues a cancel/flatten on the pod's open orders before SIGTERM (same cancel/flatten semantics as the bots KillSwitch modes) | Bots operator drain path |
| Operator | The kopf finalizer is fail-closed: with drain_fail_closed: true the finalizer is KEPT on drain timeout — termination is blocked until drain confirms | AlphaSwarmServiceSpec.drain_fail_closed; pythonic-unified-control-service §3.2, G-4 |
| Rollout | Stateful trading workloads use a partitioned rolling update, never a canary/blue-green (prohibited on StatefulSets) | Phase 4.5; ADR 026 standing prohibitions |
| Calendar | Trading-hours sync windows (ArgoCD) + change-freeze config gate when GitOps may act at all | Phase 4.5; pythonic-unified §3.2 (GitOps) |
drain_fail_closed is documented as "the operator finalizer is KEPT on drain
timeout (termination blocked until drain confirms) — mandatory posture for
live-trading workloads" and pairs with drain_timeout_seconds (default 300,
max 7200).
Preconditions / scope
- Environment: a non-production trading cell running a live-trading workload
configured with
drain_fail_closed: trueand a finitedrain_timeout_seconds. Never run first against a production silo-reg cell. - Karpenter
do-not-disrupt, the PDB, thepreStopflatten hook, and the ArgoCD trading-hours sync window are all in effect for the target workload. - A venue/paper sandbox with observable open-order state (order book queryable by the workload's account) so "zero open orders at terminate" is measurable.
- The controller halt surface is reachable for the abort path
(
/manage/workloads/halt; see Game-Day: Kill-Switch Fleet Halt + Resume). - The WORM audit uploader + integrity verifier CronJobs are running (evidence destination).
Roles
| Role | Responsibility |
|---|---|
| Exercise lead (SRE) | Runs the procedure, holds the incident ticket, calls pass/abort |
| Trading approver | Confirms the target account is inside a synthetic session, owns the order-book observation |
| Platform on-call | Watches the operator/finalizer state; drives the abort halt if invoked |
| Scribe | Captures timestamps, order counts, and the audit_run_ids into the ticket |
Step-by-step procedure
- Open the exercise ticket. Record start time, target cell, workload id,
drain_timeout_seconds, and the current ArgoCD sync-window state. - Seed open orders. In the synthetic session, place resting orders on the
target account so there is a non-zero open-order count to drain. Record the
count
N_open_beforefrom the order book. - Confirm the guards are armed. Verify the pod carries
karpenter.sh/do-not-disrupt, the PDB shows an available budget, and the workload spec reportsdrain_fail_closed: true. - Trigger a governed termination. Advance the partitioned rolling update by one partition (or, for the manual variant, delete exactly one trading pod). This is the eviction the choreography must survive. Do not use a canary — canary/blue-green on StatefulSets is prohibited.
- Observe the drain. The
preStophook fires cancel/flatten; watch the order book converge to zero for that pod's account. The kopf finalizer holds the pod object until drain confirms. - Record the terminate instant. At the moment the container actually exits
(finalizer removed), snapshot the venue open-order count
N_open_at_terminatefor that account. - Negative sub-case (fail-closed proof). Re-run once with a deliberately
unreachable venue so flatten cannot complete within
drain_timeout_seconds. Confirm the finalizer is KEPT, the pod staysTerminating, and the drain-timeout alert fires — i.e. the system refuses to strand orders. - Resolve the negative sub-case via the abort path (below), then restore the venue and let the drain complete cleanly.
- Close the ticket with the measured counts and the audit references.
Success criteria (measurable)
- Primary:
N_open_at_terminate == 0for the drained account at the instant the container exits — zero orphaned orders, proven from the venue order book, not inferred. - The
preStopcancel/flatten completed withindrain_timeout_seconds; the finalizer was removed only after confirmation. - Negative sub-case: with flatten blocked, the finalizer was KEPT, the pod never terminated, and the drain-timeout alert fired — no silent strand.
- No Karpenter voluntary eviction and no PDB-violating eviction occurred during a sync-window/freeze period.
- Every terminate/finalizer transition and the halt (if used) landed a hash-chained audit row (see Evidence capture).
Abort / rollback
- Abort trigger: order-book count is not converging to zero, PnL is bleeding, or the negative sub-case must be stopped.
- Action: engage the fleet halt —
POST /manage/workloads/halt(scopeworkloads:halt) to signal every in-flightWorkloadRunto abort, and for the bot fleet apply aKillSwitchatscope: fleet, mode: flattenper the Kill-Switch Incident Response runbook. For live-trading workloads the halt also sets the order-gate key (kill-switch honesty), so no new orders can be placed while halted. - Rollback: pause the partitioned rollout, let the finalizer complete the
drain, then resume via the
Kill-Switch Fleet Halt + Resume
procedure. Never force-delete a
Terminatingtrading pod (--grace-period=0 --force) — that is the exact orphaned-order failure the game-day exists to prevent.
Evidence capture (WORM / audit ledger)
Every mutating step lands a hash-chained workload_runs row through the gate
(JSONL locally → HTTP fan-out to the monolith Postgres ledger → S3 Object Lock
COMPLIANCE WORM). For this game-day, archive:
- The pod terminate + finalizer add/remove events with timestamps.
N_open_beforeandN_open_at_terminatefrom the order book.- The drain-timeout alert payload and finalizer-KEPT proof from the negative sub-case.
- Any
status=HALTEDWorkloadRunrows if the abort halt was used, plus theHaltRequest.reasonstring. - The
audit_run_ids, verified intact with the integrity verifier (python -m alphaswarm_controller.terraform.audit_verify, exit≠0 on a chain break — see the DR replay game-day).
Frequency / owner
- Frequency: quarterly, and before any change to the drain/finalizer path or trading-hours sync-window config.
- Owner:
sre-team(calendar reminder), with the trading approver co-signing the evidence.