RL lifecycle: statistical gates & paper release train
The agentic-RL lifecycle plan (P1–P3) turns AlphaSwarm's existing
statistical machinery into an enforced research → evaluation → paper
lifecycle. Everything below lives inside the sanctioned executor,
RLRuntime
— no new service, no second execution path.
Design doctrine
- Fail closed. A configured threshold whose evidence is missing is a
REJECT (
missing_evidence), never a silent pass. - Opt-in, hash-stable. All new sub-specs (
promotion,paper, evaluation seed battery) default toNone/off.snapshot_hash()excludesNone-valued model fields, so adding optional schema fields never shifts the hash of existing specs; setting one creates a new immutable version. - Policy proposes, controls dispose. The learned policy ends at
target weights;
WeightToOrders+ the monolith risk plane own everything after.
Seeded evaluation (WS1/WS2)
training.seed now seeds Python/NumPy/torch and the first
env.reset(seed=...) of each rollout. The evaluation battery is
configured on the spec:
training:
seed: 7
evaluation:
episodes: 4
n_seeds: 5 # or an explicit seed_list: [7, 11, 13, ...]
bootstrap: {n_resamples: 1000, alpha: 0.05}
compute_dsr: true
evaluate (and the post-train battery) returns an
evaluation_report: per-seed metric rows plus a per-metric aggregate
(point / iqm / lo / hi / n from a percentile bootstrap), a
Deflated Sharpe probability over the seed battery (per-period Sharpe
discipline), and PBO over the seed-returns matrix when computable.
Per-term reward_terms totals are summarised into
result_summary.reward_attribution.
Promotion gates (WS3)
promotion:
min_sharpe: 0.5
max_pbo: 0.5
min_dsr_probability: 0.90
min_seeds: 5
require_ci_positive: true # bootstrap CI lower bound of total_return > 0
mlflow:
register_model_as: rl-ppo-portfolio
With promotion set, mlflow.register_model fires only when every
gate passes; outcomes persist in result_summary.promotion with
per-check observed / threshold / reason, named to match the
alphaswarm_mlops.factory.eval_acceptance surface. promotion: null
preserves the legacy unconditional registration path byte-for-byte.
Paper release train (WS4)
paper: null keeps the legacy single-rollout paper probe. Setting
paper upgrades RLRuntime.paper() to a bounded release-train session:
paper:
max_episodes: 20
duration_seconds: 86400
# acceptance criteria (evaluated after the session, fail-closed)
max_hard_breaches: 0
max_reject_rate: 0.05
min_total_return: 0.0
max_drawdown: 0.10
# rollback triggers (evaluated between episodes)
rollback_max_drawdown: 0.15
rollback_max_loss_pct: 0.10
approval_expiry_days: 30
Session flow
- Episodes run one at a time; the run's Redis halt key is honoured between episodes and every 50 steps inside one.
- After each episode the rollback triggers are checked. A breach
engages the run's halt key (
rollback:<trigger>), stops the session, and finalises the run asrolled_back— fail-closed, no operator in the loop. - After the session the acceptance battery runs: no-rollback,
hard-breach budget, reject-rate, return and drawdown criteria.
Criteria that need gateway counters the env doesn't expose fail with
missing_evidence. - Everything lands in
result_summary.paper_pack: spec hash, checkpoint, seed, config, per-episode metrics, session summary, operational report, acceptance verdict + criteria, rollback events, and the approval expiry timestamp.
Operational counters
Envs that route orders through the WS5 gateway may expose
hard_breach_count, pretrade_reject_count and order_submit_count;
the release train reads them duck-typed. Absent counters leave the
corresponding evidence None — and any acceptance criterion that
depends on it rejects.
Pre-trade gateway parity (WS5)
WeightToOrders now accepts a risk_manager (canonically
alphaswarm.risk.manager.RiskManager, or "auto" to best-effort build
one). Every order is checked through check_pretrade_v2 before
submission:
- breaches with severity
block/criticalreject that order with structured reason codes (WeightToOrdersResult.rejected); - a crash of the check itself rejects fail-closed
(
pretrade_gateway_error) — an order can never pass un-gated; - every emission (accepted or rejected) produces a decision record:
client order id, symbol, side, quantity, notional, target weight, the
checks applied, and a deterministic
inputs_hashover (weights, prices, equity) for audit correlation.
The kill switch remains the outermost gate: engagement aborts the whole batch before any per-order logic runs.
Run statuses
running | completed | error | timeout | halted | cancelled | rolled_back
rolled_back dominates halted when the halt was engaged by a paper
rollback trigger.
Observability (WS9)
Every lifecycle run opens an rl span (soft dependency on the observe
SDK) carrying alphaswarm.rl.run_id / spec_hash / seed / target
at open and run_status / checkpoint_id / gate outcome at close.
alphaswarm_observe derives rl_run_success_rate and
rl_gate_pass_rate SLOs; RL spans are excluded from the generic
latency objective (lifecycle runs are hours-long by design).
Drill procedure
Before relying on the release train in anger, run the rollback drill:
- Point a spec at a deliberately lossy window (or a crash-env fixture)
with
rollback_max_drawdownset below the expected drawdown. - Run
RLRuntime.paper()and confirm: statusrolled_back, arollback_eventsentry naming the trigger, the halt key set torollback:<trigger>, andacceptance.passed == false. - Confirm the
rlspan for the run carriesalphaswarm.rl.run_status = "rolled_back".
See also: rl-framework, weight-centric-pipeline, rl-prudex-evaluation.