Skip to main content

RL statistical evaluation: the two-plane seam

AlphaSwarm deliberately carries two statistical evaluation implementations: a dependency-light one inside alphaswarm_rl and the full evidence-spine battery in alphaswarm.lab.evaluation. That is not an accident of history — it is the WS2.5 posture from the agentic-RL lifecycle plan: the RL plane keeps its own implementations; contract tests pin them to numerical agreement on shared fixtures; this page documents the seam.

The two planes​

RL planeMonolith lab plane
Packagealphaswarm_rl.evaluation.statistics + alphaswarm_rl.validationalphaswarm.lab.evaluation
Dependenciesnumpy; pandas via validation.cpcv; scipy only via DSRstdlib DSR; arch.bootstrap soft-dep for bootstrap tests
DSR APIvalidation.deflated_sharpe.deflated_sharpe_ratio(returns, *, sr_hat, sr_list, n_strategies_tested) — derives T, skew and kurtosis from the raw returns seriesdeflated_sharpe_ratio(observed_sharpe, *, n_obs, n_trials, variance_of_sharpes, skewness, excess_kurtosis) — takes precomputed moments, tied to the LabRun.total_trials_searched ledger
PBOvalidation.pbo.probability_of_backtest_overfitting (CSCV, Bailey et al. 2015)— (PBO lives on the RL plane)
CPCVvalidation.cpcv.CombinatorialPurgedKFold (sklearn-compatible) + combinatorial_paths_countcpcv.combinatorial_purged_cv + CPCVConfig/CPCVPath, safe_cpcv_path_count, CPCVPlanError hard guard (100 paths)
Model comparisoncompare_to_incumbent (bootstrap mean-difference CI)model_comparison: diebold_mariano, whites_reality_check, hansen_spa, model_confidence_set → TestResult rows in the evidence spine
ConsumersRLRuntime evaluation battery, promotion/gates.py, seed sweepsLab sweeps, UI (never raw Sharpe alone), EvidenceBundle

Why both exist. The lifecycle plan's gap inventory flags the duplication itself (G11: two DSR implementations with no convergence story) and resolves it with a SPLIT posture rather than a merge:

  1. Import isolation. RLRuntime is the sole sanctioned executor (hard rule 16) and must evaluate hermetically — the WS2 acceptance criterion is a green battery without mlflow or iceberg on the path (and in practice the monolith stays off it too: the convergence contract tests skip when it is absent). Importing alphaswarm.lab from the RL evaluation path would drag the monolith (and its evidence-schema, DB and arch surface) into every RL CI lane.
  2. Different callers, different shapes. The RL plane starts from raw per-seed returns arrays produced by rollouts, so its DSR derives the moments itself. The lab plane serves sweeps and the UI, where moments and the trial count arrive precomputed from the LabRun ledger. Forcing one signature on both callers would make one of them lie about its inputs.
  3. Drift is handled by contract, not by sharing. The risk register entry "statistical duplication drift" is mitigated by cross-implementation contract tests on shared fixtures — see below — not by a shared library.

The convergence contract​

The seam is pinned by alphaswarm_rl/tests/validation/test_stats_convergence.py:

  • test_dsr_agrees_with_monolith — parameterised over (seed, n_strategies) in (0, 10), (1, 50), (2, 200) with T = 256 synthetic per-period returns. One fixture computes the winning strategy's returns, its per-period Sharpe and the cross-trial Sharpe list; each plane is fed its native parameterisation (raw returns for the RL plane, precomputed variance/skew/kurtosis for the monolith). Agreement is asserted to abs=1e-4 — the residual is the monolith's Beasley-Springer-Moro inverse CDF vs the RL plane's scipy.stats.norm.ppf.
  • test_psr_agrees_when_single_trial — with one trial the deflation term vanishes and the monolith DSR degrades to its probabilistic_sharpe_ratio; the test pins the shared PSR core (Bailey-López de Prado eq. 9 variance) to abs=1e-6.

The module uses pytest.importorskip on scipy and alphaswarm.lab.evaluation.deflated_sharpe, so the contract runs in lanes where both planes are importable and skips cleanly in standalone RL-plane CI. If you add a statistic that exists on both planes, you must extend this contract file — an unpinned duplicate is the exact failure mode the seam exists to prevent.

The evaluation battery surface​

The battery is configured on the spec via EvaluationConfig (alphaswarm_rl/spec.py). Every lifecycle knob defaults to None/absent so existing spec snapshot_hash values are unchanged (hard rule 17); setting one creates a new immutable spec version.

evaluation:
episodes: 4
n_seeds: 5 # or an explicit seed_list: [7, 11, 13, ...]
bootstrap: {n_resamples: 1000, alpha: 0.05} # optional seed: for CI reproducibility
compute_dsr: true
dsr_n_strategies: 24 # search-space size; defaults to the seed count
  • n_seeds / seed_list — seed_list wins when both are set; n_seeds derives consecutive offsets from the resolved training.seed. None keeps the legacy single-pass rollout.
  • bootstrap — {n_resamples, alpha, seed} consumed by evaluation.statistics.aggregate_metrics (defaults: 1000 resamples, alpha = 0.05).
  • compute_dsr / dsr_n_strategies — enable the Deflated Sharpe probability over the seed battery and optionally widen the deflation to the true search-space size.
  • regime_slices — the newest addition (WS2 item 3): opt-in regime-sliced aggregates in the evaluation report, so the battery can answer "does this hold in the stress regime?" rather than only "does this hold on average?". Default-off like every other knob, so it is hash-stable for existing specs.

RLRuntime._do_evaluate runs one rollout batch per seed (global RNG + env.reset(seed=...) per seed, per-period portfolio returns captured per seed) and hands the rows to evaluation.statistics.build_evaluation_report, which produces:

{
"per_seed": [{"seed": 11, "metrics": {"sharpe": 0.61, "...": "..."}}],
"aggregate": {"sharpe": {"point": 0.58, "iqm": 0.57, "lo": 0.41, "hi": 0.74, "n": 5}},
"n_seeds": 5,
"dsr_probability": 0.97,
"pbo": 0.12
}

Bootstrap CI + IQM semantics​

aggregate_metrics summarises each metric across seeds as point (mean), iqm, lo/hi (percentile bootstrap CI of the mean) and n. Two deliberate degeneracy rules keep gate checks well-defined on small batteries: a single sample collapses the CI to a point interval, and iqm (the rliable-style interquartile mean, robust to outlier seeds) falls back to the plain mean below four samples.

DSR per-period discipline​

Every Sharpe handed to the DSR is unannualised (per-period) — statistics.per_period_sharpe per seed, never the annualised headline number. Passing an annualised Sharpe (SR × √252) is the canonical DSR bug and yields nonsensical probabilities; the warning is baked into validation/deflated_sharpe.py itself. The battery treats the best-Sharpe seed as the "winning strategy", the per-seed Sharpe list as the trial set, and dsr_n_strategies (default: the seed count) as the search-space size N for the deflation.

PBO over the seed-returns matrix​

When at least two seeds produced usable returns series, the battery stacks them column-wise (truncated to the shortest series) and runs probability_of_backtest_overfitting — CSCV with up to 16 blocks — reporting the fraction of IS/OOS splits where the in-sample best seed ranks below median out-of-sample. Series shorter than four periods, or a single-seed battery, yield pbo: null rather than a fabricated number.

Challenger comparison​

statistics.compare_to_incumbent bootstraps the mean difference (candidate − incumbent) and reports separated: true only when the CI lower bound clears zero. This is the WS7 approval discipline: a MARL (or any other) challenger must show a statistically separated improvement, not a point-estimate win.

From evidence to gates​

The report is exactly the evidence surface that alphaswarm_rl.promotion.gates.evaluate_gates(report, config) consumes (see RL lifecycle: statistical gates for the full gate catalogue):

  • scalar thresholds (min_sharpe, min_sortino, min_total_return, max_drawdown) read aggregate.<metric>.point;
  • require_ci_positive reads aggregate.total_return.lo — the bootstrap lower bound, not the mean;
  • max_pbo and min_dsr_probability read the report's top-level pbo and dsr_probability;
  • min_ope_lower_bound reads ope.headline.lo for offline candidates.

Fail-closed doctrine: a configured threshold whose evidence is missing from the report — battery not run, pbo: null, NaN, absent key — is a FAIL with reason missing_evidence, never a silent pass. This is why the statistics layer prefers null over a fabricated value on degenerate inputs: null evidence rejects at the gate, which is the correct outcome for an under-powered battery. A failing gate blocks mlflow.register_model and the outcome persists in result_summary.promotion with per-check observed / threshold / reason.

Which plane to use when​

  • Inside the RL lifecycle (the RLRuntime battery, promotion gates, seed sweeps, OPE): the RL plane, always. Lifecycle logic lives in RLRuntime or pure helpers it calls (hard rule 16), and those helpers must not import the monolith.
  • Lab sweeps, the UI, the evidence spine: the monolith plane — DSR/PSR rendered next to raw Sharpe from the LabRun trial ledger, CPCV path planning behind the CPCVPlanError hard guard, and the model-comparison family (Diebold-Mariano, White's Reality Check, Hansen's SPA, Model Confidence Set) emitting TestResult rows.
  • Cross-plane comparisons (e.g. an RL candidate vs a lab-selected incumbent): compute each side's evidence on its own plane; compare with compare_to_incumbent or the monolith model-comparison tests. The convergence contract is what makes the DSR numbers commensurable across planes.
  • Never add a runtime import of alphaswarm.lab.evaluation to alphaswarm_rl evaluation paths; the contract-test lane is the only sanctioned coupling point.

See also: rl-lifecycle-gates, rl-prudex-evaluation, rl-framework.