Saltar al contenido principal

ADR 021 — ResearchExperiment lineage aggregate

  • Status: Accepted (2026-08-03) via the ARP review — implemented by the ARP Study Design context: hash-locked design/protocol/plan versions via the SpecPersister idiom with content-hash input pinning, extended with preregistration and Deviation records. Supersedes the Proposed (2026-06-19) E8/Wave-3 gating. See alphaswarm_internal docs/architecture/agent-first-research-platform/02-context-map-and-ownership.md §6 (binding disposition table) and 10-adrs.md.
  • Implementation state (unchanged by this disposition): partially implemented: the ResearchExperimentSpec aggregate and its hashing helpers (hash_payload/hash_config/hash_factor_defs) shipped the same day in alphaswarm_models/src/alphaswarm_models/research_experiment.py (rollout steps 1-2). Persisting research_experiment_versions via a SpecPersister subclass and wrapping AlphaBacktestExperiment to emit the spec (rollout step 3) had not landed as of this review — the module's own docstring calls these "the ADR 021 rollout follow-ups." Gated on the Architecture Enhancement Guide roadmap (enhancement E8, Wave 3) The unlanded rollout step 3 is precisely the gap the ARP Study Design context closes.
  • Authors: Platform team
  • Related: Enhancement Guide §6/E8, ADR 014; Hard Rules 34 (experiment_id on every run), 43/57 (hash-locked spec versions), 48 (bipartite lineage graph)

Context​

The platform proves reproducibility for side-car specs — PredictorSpec, MLSkillSpec, RLExperimentSpec, KBCorpusSpec all hash-lock a canonical body into an immutable *_spec_versions row via the shared alphaswarm_core/runtime/persistence.py SpecPersister. It also stamps experiment_id on every run (Hard Rule 34) and maintains a bipartite lineage graph (Hard Rule 48).

But the experiment record itself is not hash-locked. AlphaBacktestExperiment (alphaswarm_models) — the nearest thing to the memo's central artifact — ties dataset_cfg + model_cfg + strategy_cfg + backtest_cfg together via FK hints and unhashed JSON params; dataset_hash is caller-supplied rather than computed inside the aggregate; factor expressions are code-only (the Alpha158/360 strings) and never hashed against a run; there is no rationale and no own snapshot_hash. Editing a factor string silently changes future runs. Reproducibility is therefore by convention, not by construction — and the proven hash-lock pattern is simply not applied here.

Decision​

Introduce a hash-locked ResearchExperimentSpec aggregate, persisted through the existing SpecPersister, that pins inputs by content hash and stores results and rationale.

  1. Pin inputs by hash. Fields: hypothesis, dataset_snapshot_hash, factor_defs_hash, model_config_hash, policy_hash (signal/order/execution), backtest_config_hash, plus result_metrics and rationale. The aggregate's snapshot_hash() is the SHA-256 of its canonical JSON — the same idiom every other spec uses.
  2. Compute hashes inside the aggregate. dataset_snapshot_hash is derived from the medallion as_of / snapshot_id (not caller-supplied); factor_defs_hash hashes the resolved factor expressions at build time, so editing a factor string changes the experiment identity.
  3. Persist research_experiment_versions. Wrap AlphaBacktestExperiment to emit a ResearchExperimentSpec and write the immutable version row; link it into the bipartite lineage graph.
  4. Unify the fragments. The KB kb_runs ledger, the RL trajectory corpus, and the graph BacktestRun nodes reference the same research_experiment_id, so a result is reproducible end-to-end from one re-runnable record.

Consequences​

Positive

  • A research result becomes reproducible from named, content-addressed inputs; re-running a stored experiment is deterministic; lineage is queryable, not just "remembered." Mirrors Qlib's recorder/qrun reproducibility and TradingAgents' persistent decision log.

Negative / risks

  • Computing dataset_snapshot_hash and factor_defs_hash adds resolve-time work; keep it incremental and cache by snapshot_id.
  • Higher effort than the other enhancements — it touches the ML data/feature path; sequence it after the Wave-1/2 correctness and unification work.

Explicitly rejected

  • Leaving experiments tied by FK hints + unhashed JSON (the status quo — reproducibility by convention).
  • A bespoke hashing scheme — reuse SpecPersister / snapshot_hash() so the experiment layer matches every other hash-locked runtime.

Rollout order​

  1. Define ResearchExperimentSpec + _ResearchExperimentPersister(SpecPersister)
    • research_experiment_versions migration.
  2. Compute dataset_snapshot_hash (from medallion as_of) and factor_defs_hash (resolved expressions) inside the aggregate.
  3. Wrap AlphaBacktestExperiment; backfill lineage links from KB/RL/graph.