Promotion evidence: trial ledger + evidence bundle
In plain English: if you test enough random strategies, one of them will look brilliant by pure luck. AlphaSwarm defends against this by keeping a tamper-resistant ledger of every experiment a researcher (or research agent) runs — not just the flattering ones — and by requiring statistical evidence that accounts for all those attempts before a strategy is allowed anywhere near real money. A strategy that looked great in 1 test out of 1 is very different from one that looked great in 1 test out of 10,000, and the promotion gates now know the difference.
This page covers the three "ARP" (agentic research promotion) slices landed in August 2026. All enforcement is default-off behind feature flags; see Rollout.
Why this exists
The Promotion Gates API has always required human approval plus deterministic risk/validation gates before a strategy crosses from the research plane to the money plane. What it could not previously verify was the statistical honesty of the research behind the request:
- Selection bias: deflated Sharpe ratio (DSR) and probability of
backtest overfitting (PBO) statistics are only meaningful if the trial
count
Nreflects every attempt actually made. Self-reported N under-counts. - Evidence provenance: gate inputs were supplied by the requester, not read from a system of record.
- Execution bypass: nothing at the paper/live execution boundary re-checked that a promotion had actually been approved.
The three slices
Slice 1 — TrialLedger
A unified ledger of research trials in Postgres
(trial_families / trial_attempts, migration 0160), seeded from lab
sweep stamps. Every parameter sweep, walk-forward split, and re-run lands
as an attempt row under a trial family, so the DSR/PBO denominator N
can be derived from the ledger instead of self-reported.
- Code:
alphaswarm/lab/trial_ledger/,alphaswarm/persistence/models_trial_ledger.py - MCP tools:
data.lab.trial_*(read-only trial reads for agents and operators) - Reporting:
scripts/trial_ledger_reverdict_report.pyshows how existing promotions would re-verdict under ledger-derived N — run it before enabling enforcement.
Slice 2 — EvidenceBundle in the gate chain
The lab's EvidenceBundle (DSR, PBO, trial counts, provenance) is wired
into the PromotionGateChain. With strict enforcement on, a promotion
request fails closed if the bundle is missing or incomplete, and
paper execution refuses to start without an ApprovedPromotion token.
Plane hygiene note: promotion code reaches lab evidence through a
registered probe port
(alphaswarm/promotion/lab_evidence.py
→
alphaswarm/lab/evidence/promotion_probe.py),
never by importing alphaswarm.lab directly — the
plane_money_promotion boundary stays intact.
Slice 3 — Bounded graph tools
Agents assessing a promotion need lineage context ("what feeds this strategy?", "what else breaks if this dataset is stale?") without an open-ended graph query surface. Four bounded, point-in-time-filtered MCP tools expose exactly that:
| Tool | Question it answers |
|---|---|
data.graph.temporal_context | What did the graph neighbourhood of this entity look like as of time T? |
data.graph.dependency_impact | What downstream artifacts depend on this node? |
data.graph.risk_neighbors | Which risk-relevant entities sit near this node? |
data.graph.sync_status | How fresh is the graph projection itself? |
There is deliberately no raw Cypher surface — every tool is bounded in depth, filtered point-in-time, and returns typed payloads.
How a promotion flows now
Rollout and flags
| Flag | Default | Effect when enabled |
|---|---|---|
ALPHASWARM_TRIAL_LEDGER_N_ENFORCE | off | DSR/PBO statistics use ledger-derived N. Does not auto-demote existing promotions. |
ALPHASWARM_PROMOTION_LAB_EVIDENCE_ENFORCE | off | Strict mode fails promotion closed without a complete EvidenceBundle. |
ALPHASWARM_GRAPH_TEMPORAL_TOOLS_ENABLED | off | Registers the four bounded graph tools. |
Recommended order: run the re-verdict report, review deltas with the research owners, enable ledger N in staging, then strict evidence enforcement, then the graph tools for the approval-assistant agents.
See also
- Promotion Gates API — endpoints, statuses, approval semantics.
- Approval queues — how pending approvals expire.
- Strategy lifecycle — draft → backtested → paper → live.
- RL lifecycle gates — the RL-specific gate set.
- Graph Data Pillar — the governed graph plane these tools read from.