AlphaSwarm First-Class Module Transformation Plan
Status: Proposed (planning only — no production runtime change in this document set).
Date: 2026-08-10
Evidence tracks: DOC (4ada4996…), A+E (8019545a…), B+C (eb4b09c8…), D+F (07eed925…).
Related ADRs: 015 · 016 · 017 · 018 · 019 · 020
Evidence index: alphaswarm-first-class-module-evidence-index.md
Evidence tag legend
| Tag | Meaning |
|---|---|
| A | Claim from attached source reports (R1/R2/R3) or orientation docs |
| B | Repository fact (path/symbol verified in this planning pass or a track) |
| C | Architecture recommendation of this plan |
| D | External technical guidance cited by sources |
| E | Open question / unsupported / needs measurement |
| VERIFIED | Spot-checked or track-read against live files |
| PARTIALLY VERIFIED | Grep/doc evidence + structural inference |
| INFERRED | Pattern inference only |
| PROPOSED | Target-state design not yet implemented |
| BLOCKED | No evidence found; do not assert existence |
1. Executive summary
Primary decision (C, justified by B): Adopt a monorepo-local compatibility facade that promotes the existing alphaswarm Python namespace as the stable public domain/application kernel. Implementation owners remain separately deployable:
| Concern | Owner today (B) |
|---|---|
| Wire/contracts / workload ABC | alphaswarm_core |
Mutation authority /manage/* | alphaswarm_controller (WorkloadRuntime, TerraformRuntime) |
| Distributed execution | alphaswarm_worker (WorkRequest, Executor, ExecutorRouter, NativeExecutor) |
| Orchestration engines | alphaswarm_orchestration (EngineKind.DAGSTER / PREFECT — not Temporal) |
| ML / RL / agents / KB specialists | alphaswarm_models, alphaswarm_rl, alphaswarm_agents, alphaswarm_kb |
| Money-plane bot adapters | alphaswarm_bots |
| Quant domain + API composition | alphaswarm monolith (facade home) |
Reject (A→C, affirmed by B): SimulationMeta as core; LLMs as live execution peers; a second scheduler/execution/MLOps/deployment control plane; duplicating DeploymentSpec as RuntimeDeploymentSpec; inventing asctl or Temporal orchestration.
Secondary alternative (C): Generated contracts package (alphaswarm_catalog-style) or PEP 420 namespace aggregation — useful for wire schemas, too thin as the application kernel (see §7).
Immediate focus: close domain dual-representation hotspots (Symbol/BarData/OrderRequest), unify ExecutionProfile translation, formalize Activity vs Simulation vs Workflow vs WorkRequest vs DeploymentSpec vocabulary, and turn on safety gates (tenancy_rls_enforce, halt propagation, MCP audience) before any live-execution expansion.
2. Source-document findings
Source shorthand (A): R1 Python Trading Simulation Architecture · R2 Quantitative Trading Architecture Guide · R3 deep-research critique · Ctx alphaswarm_internal/ARCHITECTURE.md (orientation only).
Retain (map into AlphaSwarm)
| Concept | Tag | Disposition |
|---|---|---|
| Pure kernels + externalized context | A | Strengthen strategy/engine purity |
| Multi-mode engines (batch / replay / paper / live) | A+B | Already exist as backtest cascade + paper + worker profiles |
| Numba path-dependent research kernels | A+B | HFT LOB path already uses @njit |
| Canonical event log + golden replay digests | A+C | Strengthen promotion/parity |
| Independent deterministic risk / kill / fail-closed | A+B | RiskLimits, GateChain, kill switch |
| Bounded queues + typed backpressure | A+C | Forbid silent critical drops |
| Strangler Fig migration | A+C | Phases 0–8 below |
| Engine protocol + factory/registry | A(R3)+B | Prefer over metaclasses |
| LLM structured advisory above risk | A+B | OrderIntent airlock |
Reject / quarantine
| Concept | Tag | Disposition |
|---|---|---|
SimulationMeta as platform core | A→C | Reject — opacity, no semantic unification |
| Write-once identical semantics across vectorized↔live↔agent DAG | A→C | Share kernels + events, not engines |
| Marketing latency (100–500×, 1M msg/s, sub-µs) as requirements | A→E | Reject until measured harness exists |
| Broken SHM ring samples / “zero-copy” via Python dicts | A→C | Quarantine until redesigned |
| LLMs / Ray as order-path peers | A→C | Research/advisory only |
| SQLite LangGraph checkpointer as prod default | A→C | Durable checkpointer for prod |
| ROS 2 / Laguna-specific MoE hardware as platform req | A→C | Reject |
| Temporal as AlphaSwarm orchestration | A(hyp)→B | ABSENT — Dagster/Prefect only |
Unsafe source claims corrected explicitly (A→B)
- Temporal orchestration in controller — CORRECTED ABSENT (B):
EngineKind=dagster|prefectonly (alphaswarm_orchestration/.../contracts.py). asctlCLI — CORRECTED ABSENT (B): entry points arealphaswarm-controller,alphaswarm-cli-control-plane,alphaswarm-controller-operator.- Need for
RuntimeDeploymentSpec— CORRECTED (B): use existingDeploymentSpecinalphaswarm_core/.../models/deployment.py. - Terraform
cellmodule — CORRECTED ABSENT (B): cells are Kustomize overlays underalphaswarm_platform/deployments/kubernetes/cells/; ArgoCD AppSets exist. SimulationSpec/ PyO3 / SHM production path — BLOCKED in repo search; de facto simulation isGraphSpec(mode="simulation")(B).- NautilusTrader library import — domain is “Nautilus-inspired”; live bridge referenced for event-driven path — do not claim full library embedding without further verification (PARTIALLY VERIFIED / INFERRED).
3. Repository current-state architecture
Package roles (B — Track A+E, spot-verified)
Spec-driven runtimes already present (B)
| Spec | Runtime | Spec versions | Ledger |
|---|---|---|---|
AgentSpec | AgentRuntime | agent_spec_versions | agent_runs_v2 |
WorkflowSpec | WorkflowRuntime | workflow_spec_versions | workflow_runs |
BotSpec | BotRuntime | bot_versions | bot_deployments |
RLExperimentSpec | RLRuntime | rl_experiment_versions | rl_runs |
AnalysisSpec | AnalysisRuntime | analysis_spec_versions | analysis_runs |
GraphSpec | LabRuntime | content-hash on lab_graphs | lab_runs |
KBCorpusSpec | KBRuntime | kb_corpus_spec_versions | kb_runs |
MLSkillSpec | MLSkillRuntime | ml_skill_versions | ml_skill_runs |
TerraformStackSpec | TerraformRuntime | terraform_stack_spec_versions | terraform_runs |
| workload ops | WorkloadRuntime | n/a | workload_runs |
Mutation authority (B — Track D)
- Workload ops:
WorkloadRuntime(alphaswarm_core) executed via controller/embedded modes. - IaC provisioning:
TerraformRuntime(controller only sanctioned subprocess). - Surface:
/manage/*on controller; monolith brokers where flagged. - State of record: Postgres ledgers; controller JSONL + HTTP sink for terraform audit.
4. Evidence inventory
See also evidence index.
| ID | Claim | Tag | Path / symbol |
|---|---|---|---|
| E1 | WorkRequest / Executor / ExecutorRouter / NativeExecutor exist | VERIFIED | alphaswarm_worker/.../execution/{contracts,base,router,native}.py |
| E2 | DeploymentSpec exists (no RuntimeDeploymentSpec) | VERIFIED | alphaswarm_core/.../models/deployment.py |
| E3 | EngineKind = Dagster/Prefect only | VERIFIED | alphaswarm_orchestration/.../contracts.py |
| E4 | Controller entry points (no asctl) | VERIFIED | alphaswarm_controller/pyproject.toml |
| E5 | Cells via Kustomize | VERIFIED | alphaswarm_platform/deployments/kubernetes/cells/*/kustomization.yaml |
| E6 | GraphSpec simulation mode | VERIFIED | alphaswarm/lab/schema.py mode: Literal[...,"simulation"] |
| E7 | Default-OFF: RLS, MCP RFC8707, halt propagation | VERIFIED | alphaswarm/config/settings.py |
| E8 | Research→paper metadata gate | VERIFIED | alphaswarm/trading/metadata_gate.py |
| E9 | ExecutionProfile dual definition | VERIFIED | worker Enum vs orchestration PortableModel |
| E10 | Domain dual-rep hotspots | VERIFIED | core/types.py vs core/domain/ |
| E11 | Missing Forecast/Calibration/Benchmark/Scenario/Listing/Fills types | BLOCKED/GAP | Track B search |
| E12 | No PyO3/SHM/SimulationMeta/SimulationSpec | BLOCKED | Track B+C search |
| E13 | No tests/perf harness | BLOCKED | Track F |
| E14 | PromotionPolicy + GateChain | VERIFIED/PARTIALLY | lab/evidence/promotion.py, promotion/gate.py |
5. Source-to-code mapping
| Source concept (A) | Existing AlphaSwarm surface (B) | Mapping status |
|---|---|---|
| First-class Simulation | GraphSpec(mode="simulation") + backtest engines + LabRuntime | Surrogate exists; no SimulationSpec |
| VECTORIZED paradigm | VectorbtProEngine, worker VECTORIZED_BACKTEST | Mapped |
| EVENT_DRIVEN sync | EventDrivenBacktester, SimulatedBrokerage | Mapped |
| EVENT_DRIVEN async / live | paper session + NativeExecutor + bots adapters | Mapped (partial live) |
| AGENT_DIRECTED_DAG | WorkflowRuntime / LangGraph adapters | Mapped as agent plane, not engine peer |
| Deployment topology | DeploymentSpec + Kustomize cells + ArgoCD | Mapped — do not add RuntimeDeploymentSpec |
| Metaclass engine selection | @register / RLComponent / InfrastructureProviderMeta / EngineCapabilities | Prefer registry/factory (C) |
| Risk gate | GateChain, RiskLimits, worker _risk_gate, bots RTS6 | Mapped |
| SHM / PyO3 HFT | Not found as product path | Reject until redesigned (C) |
| RAG episodic journal | alphaswarm_kb + HierarchicalRAG shims | Optional KB feature |
| Distributed sweeps | worker Ray/Dask/Spark backends (research plane) | Mapped; MONEY → native only |
| Temporal workflows | — | Do not map — absent |
6. Current problems and risks
Anti-patterns to refuse (C — §26 of assignment)
- Second control plane for scheduling, execution, MLOps, or deployment.
SimulationMeta/ hidden method rebinding as public architecture.- LLM → venue order without
OrderIntent+ deterministic gates. - Duplicate
DeploymentSpec/ WorkRequest / ExecutorRouter. - Big-bang rename of
alphaswarm_core→ “the platform” (wrong layer). - Silent queue drops on risk/order events.
- Marketing latency numbers as acceptance criteria.
- Turning on money plane while RLS / halt-propagation / MCP audience remain OFF.
- Editing shipped Alembic migrations or mutating hash-locked
*_spec_versions. - Agent ORM/Iceberg direct reads (Hard Rule 22).
Highest-impact risks
| Risk | Severity | Evidence |
|---|---|---|
Domain dual-representation drift (Symbol/BarData/OrderRequest) | High | B Track B H1–H3 |
Live expansion with tenancy_rls_enforce=off | High | B settings |
orchestration_kill_propagation_enabled=False leaves sub-runs alive | High | B settings + Track D |
| ExecutionProfile translation gap worker↔orchestration | Medium | B E9 |
| No performance regression suite | High for live | B Track F Gap |
| Missing first-class Forecast/Scenario/Fills etc. | Medium | B Track B |
Credential store duplication (alphaswarm_config vs core) | Medium | Ctx / ARCHITECTURE |
7. Definition of first-class alphaswarm
Primary approach (C) — selected
Monorepo-local compatibility facade package promoting the existing alphaswarm Python namespace as the stable public domain/application kernel.
What “first-class alphaswarm” means (C, aligned to assignment §1 intents):
- Canonical public Python namespace for platform capabilities.
- Owner of stable domain vocabulary (instruments, orders, bars, intents) — not infra ABCs alone.
- Entry for defining investment/trading activities.
- Entry for creating simulations / studies / workflows / work requests (typed facades, not new engines).
- Compatibility facade over existing execution/model/orchestration/deployment owners via translation adapters.
- Home of versioned interfaces, protocols, schemas, state machines, registries that application code imports.
- Boundary preventing app code from depending on infra SDKs directly.
- Surface through which governance, authorization, provenance, and PIT constraints are expressed.
- Composable kernel — not a second monolith framework.
- Migration target preserving production behavior until cutovers are flag-gated.
Why not promote alphaswarm_core alone (C)
alphaswarm_core is correctly dependency-light contracts (DeploymentSpec, WorkloadRuntime, providers, RBAC). It is the wrong layer for quant domain types, strategy runtimes, Iceberg writes, and Celery-mounted tasks. Promoting it would either (a) bloat it into a second monolith or (b) force every app import through an incomplete kernel.
Secondary alternatives (C)
| Alternative | Pros | Cons | Verdict |
|---|---|---|---|
| PEP 420 namespace aggregation | Clean packaging story | Harder import-boundary CI; risk of circular installs | Secondary |
Generated contracts package (extend alphaswarm_catalog) | Strong for OpenAPI/JSON Schema wire | Too thin for domain kernels & runtimes | Secondary / complementary |
| Big-bang rename / merge packages | Apparent simplicity | Breaks every consumer; violates Strangler Fig | Rejected |
Expansive internal layer model mapped onto EXISTING packages (C→B)
| Layer | Responsibility | Existing home |
|---|---|---|
| L0 Presentation | UI/CLI/BFF | alphaswarm_ui, alphaswarm_admin, alphaswarm_client, alphaswarm_cli |
| L1 API composition | HTTP/WS gateway | alphaswarm_api, alphaswarm/api |
| L2 Application kernel (facade) | Domain services, activity APIs | alphaswarm (promoted) |
| L3 Domain model | Instruments, orders, positions, events | alphaswarm/core/domain (+ legacy types strangler) |
| L4 Spec/runtime services | Hash-locked runtimes | agents/bots/rl/kb/models/analysis/lab |
| L5 Execution dispatch | WorkRequest routing | alphaswarm_worker |
| L6 Orchestration adapters | Dagster/Prefect | alphaswarm_orchestration |
| L7 Mutation / infra | Workload + Terraform | alphaswarm_controller + alphaswarm_core |
| L8 Data plane | Iceberg/Hudi/Kafka/MCP | alphaswarm/data, ingest, streaming |
| L9 Identity | IdP, M2M, step-up | alphaswarm_auth + monolith security |
| L10 Observability | OTEL, progress frames | alphaswarm_core.observe, _progress |
8. Target domain model
Keep and converge (B→C)
InstrumentBasehierarchy +InstrumentId/IdentifierSet(core/domain/)DomainOrder/DomainPosition/ExecutionReportRiskLimits, kill switch,PricingContextStrategyPromotionRequest,OrderIntent- Spec family:
BotSpec,AgentSpec,RLExperimentSpec,WorkflowSpec,GraphSpec,MLSkillSpec,KBCorpusSpec
Add as first-class domain types (PROPOSED — currently BLOCKED/GAP)
| Type | Rationale | Suggested home |
|---|---|---|
Forecast | Models return ad-hoc arrays today | alphaswarm/core/domain/forecast.py |
CalibrationResult | Pricing “Calibrated” lacks entity | core/domain/calibration.py |
Benchmark | Referenced but no class | core/domain/benchmark.py |
Scenario | Overrides exist; no entity | core/domain/scenario.py |
Listing (Instrument×Venue) | Only string primary_listing_venue | core/domain/listing.py |
Fill / fill rows | Multi-partial fills need history | persist beside execution_reports |
Dual-representation consolidation (C)
Legacy (core/types.py) | Target (core/domain/) | Phase |
|---|---|---|
Symbol | InstrumentId (+ Symbol shim) | Phase 2–3 |
BarData | Bar / BarSpecification | Phase 2–3 |
OrderRequest / OrderData | DomainOrder | Phase 3–4 |
Signal | new domain Signal | Phase 3 |
PortfolioTarget | richer portfolio VO | Phase 4 |
9. Target package architecture
Rules (C):
- New public imports prefer
alphaswarm.<domain>stable paths. - Facades may wrap worker/controller/specialists; specialists must not invent parallel public APIs for the same nouns.
alphaswarm_coreremains the shared infra contract wheel; domain richness stays inalphaswarm.- No new top-level
alphaswarm_simulationruntime package unless GraphSpec/LabRuntime prove insufficient after Phase 3 review.
10. Control-plane / data-plane / money-plane architecture
| Plane | Authority (B) | Must not |
|---|---|---|
| Control | controller /manage/*, WorkloadRuntime, TerraformRuntime | Call venues; host strategy math |
| Data | Iceberg wrapper, DataMCP, lineage | Accept orders |
| Money | NativeExecutor + bots + GateChain + RiskLimits | Consult LLMs inside risk gate |
| LLM | Agent/Workflow runtimes | Bypass OrderIntent; mutate hard limits upward |
See ADR-036.
11. Activity model
Definition (C): An Activity is a short-lived, auditable unit of work with clear inputs/outputs, no long-running checkpointed graph, and optional side effects under existing gates.
Candidates mapped from repo (B→C):
| Activity | Existing mechanism |
|---|---|
| DataMCP tool invoke | DataMCPTool + audit |
| Ingestion approval step | IngestionApproval state machine |
Pricing calc() | PricingContext |
| Single cache refresh / entity list | metadata cache routes |
| Emit security audit | emit_audit_event |
Non-goals: Do not introduce Temporal Activities. Do not rename Celery tasks wholesale. Optionally add a thin alphaswarm.activity facade that documents/standardizes logging + tenancy stamps (PROPOSED Phase 2).
12. Simulation model
Definition (C): A Simulation is a deterministic or quasi-deterministic replay against pinned data with no live venue side effects.
| Mode | Existing (B) | Notes |
|---|---|---|
| Vectorized historical | VectorbtProEngine | Research throughput |
| Event-driven bar replay | EventDrivenBacktester + SimulatedBrokerage | Chronological fidelity |
| LOB/HFT | LobBacktestEngine | Numba driver |
| Lab graph simulation | GraphSpec(mode="simulation") → Dagster bridge | De facto SimulationSpec |
| RL episode (train/eval, non-paper) | RLRuntime | Trajectories to Iceberg |
Shared across modes (C): pure decision kernels + canonical market/order events + digestable state hashes.
Not shared: clocks, fill models, reconnects, ACKs — owned by each engine.
Reject: single metaclass Simulation that claims live parity. See ADR-034 / ADR-035.
13. Runtime and deployment model
Distinguish without collapsing (C)
| Concept | Question it answers | Existing type (B) |
|---|---|---|
| Activity | What atomic work just ran? | DataMCP / approval steps (facade PROPOSED) |
| Simulation | What offline replay am I running? | GraphSpec simulation / backtest engines |
| Workflow | What multi-step checkpointed graph? | WorkflowSpec + WorkflowRuntime |
| WorkRequest | What resource-bounded dispatch unit? | alphaswarm_worker.WorkRequest |
| DeploymentSpec | What infra deploy am I requesting? | alphaswarm_core.models.deployment.DeploymentSpec |
Deployment topology (B)
- Cells: Kustomize overlays (
cells/shared-*,cells/silo-*) withalphaswarm.io/cell-id. - GitOps: ArgoCD ApplicationSets for services + observability.
- Worker deploy:
deployments/kubernetes/base/alphaswarm-worker/. - Bots CRDs:
quantbot.io/v1via controller operator.
Do not invent a Terraform cell module or RuntimeDeploymentSpec.
14. Agent and RAG architecture
| Concern | Owner (B) | Rule |
|---|---|---|
| Agent execution | AgentRuntime | No direct ORM/Iceberg; DataMCP only |
| Multi-agent workflows | WorkflowRuntime + adapters | Halt + versioning |
| LLM calls | router_complete | Hard Rule 2 |
| KB remember/recall | KBRuntime / data.kb.* | Hard Rules 56–60 |
| Cross-silo recall | alphaswarm_kb_federation only | Read-only |
| Structured outputs → trade | Must become OrderIntent then gates | ADR-037 |
LLMs are advisory peers, never money-plane peers (C). They may tighten risk, never raise hard limits.
15. Multi-tenancy and security model
| Control | Default (B) | Live-expansion requirement (C) |
|---|---|---|
tenancy_rls_enforce | off | ≥ permissive, then strict with tests |
mcp_require_rfc8707 | off | permissive before external MCP exposure |
orchestration_kill_propagation_enabled | False | True before multi-runtime live |
ws_auth_required | False | True for hosted |
enable_money_plane (worker) | False | Explicit ops enable + HMAC approval |
| Step-up MFA | CI-enforced on destructive routes | Keep expanding pattern list |
| Trust zones | UI / admin / controller / data | Never merge staff & customer processes |
Identity matrix remains Entra-first for hosted UI/admin; CLI device flow; M2M via Agent Identity / CredentialResolver.
16. Public API examples
Examples use existing names (B); new facade wrappers are PROPOSED.
# Domain identity (prefer domain; Symbol shim during strangler)
from alphaswarm.core.domain.identifiers import InstrumentId
from alphaswarm.core.types import Symbol # legacy shim — deprecate
sym = Symbol.parse("AAPL.US") # Hard Rule 1 during transition
# iid = InstrumentId(...) # target
# Simulation via lab GraphSpec (de facto SimulationSpec)
from alphaswarm.lab.schema import GraphSpec
from alphaswarm.lab.runtime import LabRuntime
spec = GraphSpec(mode="simulation", ...) # hashable graph
result = LabRuntime(spec).run()
# WorkRequest dispatch (worker contracts)
from alphaswarm_worker.execution.contracts import (
WorkRequest, ExecutionProfile, Plane, DataBinding, ResourceSpec,
)
req = WorkRequest(
profile=ExecutionProfile.VECTORIZED_BACKTEST,
plane=Plane.LLM, # never MONEY without NativeExecutor path
data_binding=DataBinding(...), # Iceberg snapshot pin
resources=ResourceSpec(...),
idempotency_key="...",
)
# Bot money-plane lifecycle
from alphaswarm_bots.spec import BotSpec
from alphaswarm_bots.runtime import BotRuntime
BotRuntime(bot_spec).run_paper(...) # still subject to metadata_gate
# Infra deploy — existing DeploymentSpec, not RuntimeDeploymentSpec
from alphaswarm_core.models.deployment import DeploymentSpec
deploy = DeploymentSpec(...) # provider.deploy(deploy)
Facade aspiration (PROPOSED):
from alphaswarm import activities, simulations, workflows, work, deploy
simulations.run(graph_spec) # -> LabRuntime / backtest runner
work.submit(work_request) # -> WorkSubmissionClient
workflows.run(workflow_spec) # -> WorkflowRuntime
deploy.apply(deployment_spec) # -> controller /manage (HTTP)
17. Existing-to-target component map
| Existing package / surface | Target role | Action |
|---|---|---|
alphaswarm | Public domain/application kernel + facade | Promote; add facade modules; strangler domain |
alphaswarm_core | Infra contracts / WorkloadRuntime / DeploymentSpec | Keep thin; do not absorb quant domain |
alphaswarm_controller | Sole mutation executor | Keep; no Temporal; no asctl rename |
alphaswarm_worker | Sole WorkRequest execution plane | Keep; MONEY→NativeExecutor |
alphaswarm_orchestration | Dagster/Prefect adapters | Keep; unify ExecutionProfile translation |
alphaswarm_agents | Agent/Workflow specialists | Keep behind facade |
alphaswarm_bots | Money-plane bot unit | Keep |
alphaswarm_rl | RL specialist | Keep |
alphaswarm_models | ML specialist + interfaces | Keep; emit domain Forecast |
alphaswarm_kb (+ federation) | Cognitive memory | Keep |
alphaswarm_platform | IaC/GitOps/cells | Keep Kustomize cells |
alphaswarm_api / _auth / _cli / _ui / _admin | Edges | Unchanged boundaries |
alphaswarm_catalog | Contracts-only | Complementary generated schemas |
alphaswarm_internal | SSoT inventories | Orientation, not runtime |
Legacy core/types.py | Compatibility shims | Deprecate per ADR-038 |
| Hypothetical Temporal / asctl / SimulationMeta | — | Do not create |
18. Dependency rules
UI/CLI/Admin --HTTP--> api / auth / monolith / controller
monolith (alphaswarm) --> core, worker(contracts), agents, bots, rl, models, kb, orchestration
controller --> core ONLY (CI rg guard)
worker --> core ONLY (+ optional ray/dask/spark)
kb_federation --> NOT alphaswarm.* / alphaswarm_kb.*
agents --> DataMCP / router_complete; NOT persistence ORM
bots money adapters --> GateChain / RiskLimits; NOT raw LLM SDKs
New facade modules may import owners; owners must not import facade internals in a cycle. Translation adapters live in alphaswarm/facade/ or alphaswarm/compat/ (PROPOSED naming).
19. State-machine definitions
| Entity | States (B / C) | Notes |
|---|---|---|
IngestionApproval | pending→approved/rejected→applied/failed/expired | B |
WorkStatus | SUCCESS/FAILED/DEAD_LETTERED/DRAINED/SKIPPED_CACHED/HALTED | B worker |
CanonicalRunState | PENDING/RUNNING/SUCCESS/FAILED/CANCELLED/SKIPPED | B orch |
EntraTenantLink | pending→active (+ reject) | B |
| Paper session | start→running→stopped/halted | B |
| Order (domain) | Nautilus-inspired lifecycle on DomainOrder | B |
| Promotion | research evidence→StrategyPromotionRequest→GateChain→live eligibility | B+C |
| Kill switch | engaged/released (Redis) | B |
| Spec versions | immutable snapshots; never mutate | B hard rules |
20. Migration phases (Strangler Fig, Phase 0–8)
| Phase | Name | Goal | Acceptance | Rollback |
|---|---|---|---|---|
| 0 | Freeze vocabulary | ADRs accepted; anti-pattern list in CI docs; no new duplicate types | ADR 033–020 accepted; evidence index linked | N/A (docs) |
| 1 | Facade skeleton | alphaswarm.facade re-exports + docs; zero behavior change | Import smoke tests; no new deps cycles | Delete facade package |
| 2 | Contract alignment | Single ExecutionProfile translation module; document Activity/Simulation/Workflow/WorkRequest/DeploymentSpec | Golden translation tests worker↔orch | Feature-flag off translator |
| 3 | Domain strangler | Engine + paper paths migrate off legacy Symbol/BarData/OrderRequest | Parity tests; deprecation warnings | Shim re-enable |
| 4 | Missing domain types | Forecast, Scenario, Listing, Fill rows (additive) | Migrations + MCP/cache pickers where needed | Leave unused types; no forced cutover |
| 5 | Simulation clarity | Formal Simulation facade over GraphSpec/backtest; golden replay digests | Replay hash equality on fixture | Facade-only rollback |
| 6 | Safety gates ON | RLS permissive→strict path tested; halt propagation ON in staging; MCP audience permissive | Track F gates 3–4–8 green | Flag revert to off |
| 7 | Live-expansion readiness | Perf baselines; money-plane checklist; promotion four-eyes | Gate 9 harness + metadata_gate + GateChain e2e | Keep money plane OFF |
| 8 | Deprecation harvest | Remove unused legacy imports; update AGENTS/index | Import linter clean; curator refresh | Restore shims one release |
21. Test strategy
Build on Track F inventory (B):
| Layer | Action (C) |
|---|---|
| Unit | Expand domain shim/parity tests under tests/core/ |
| Contract | ExecutionProfile translator; WorkRequest schema freeze tests |
| Boundary CI | Keep 14 lints; add “no RuntimeDeploymentSpec / no Temporal import” guards |
| Integration | RLS-on suite (new); halt fan-out with propagation flag ON |
| Replay | Golden event-log digests for event-driven + LOB fixtures |
| Chaos | Extend beyond RL to agent + worker crash mid-WorkRequest |
| Perf | Create tests/perf/ with pytest-benchmark — do not cite unverified README latency (E) |
| E2E | Restore admin Playwright coverage; UI navigation already growing |
22. Performance strategy
Methodology (C) — reject unverified A claims:
- Define per-stage SLOs: ingest→strategy intent→risk→send (paper/live), kill-switch engage→all halt acks, WorkRequest queue wait.
- Measure p50/p95/p99/p99.9 at fixed offered load + burst; record CPU/mem.
- Separate research throughput benches (vectorized sweeps) from money-plane latency benches.
- Native/HFT paths: measure at native boundary; do not claim zero-copy across Python object handoff without buffer proofs.
- Free-threaded CPython: optional experiment (E) — not a design assumption.
No acceptance criterion may cite AlphaForge/BTQuant/blog 100–500× figures (A→E).
23. Deployment and rollout strategy
- Use existing GitOps/Kustomize cells — no new cell Terraform module.
- Roll facade by package version compatibility, not big-bang cutover.
- Hosted flags flipped per cell with dual-write only where already supported (
cell_dual_writeremains explicit).
24. Rollback strategy
| Change class | Rollback |
|---|---|
| Facade modules | Revert package; callers keep deep imports |
| Domain shims | Re-export legacy types; warning level only |
| Flag promotions (RLS/MCP/halt) | Set back to off/False via config |
| Worker/orch translator | Bypass flag to dual accept |
| Migrations | Additive only; never edit shipped; follow-up migration to undo schema |
| ArgoCD sync | Previous image digests / AppSet revision |
25. Risk register
| ID | Risk | Mitigation |
|---|---|---|
| R1 | Facade becomes a new god-object | Thin re-exports only; logic stays in owners |
| R2 | Domain migration breaks engines | Phase 3 parity + shim window |
| R3 | Live trading without RLS | Phase 6 hard gate |
| R4 | Duplicate ExecutionProfile diverges further | Phase 2 translator ownership |
| R5 | Agents invent second order path | ADR-037 + CI import/lint |
| R6 | Perf unknown at cutover | Phase 7 harness mandatory |
| R7 | Index/docs drift | Index debt note + curator pass |
| R8 | Cred store dual path | Unify config/core stores (tracked separately) |
26. Decisions requiring approval
- Accept Primary approach (facade over
alphaswarmnamespace) — ADR-033. - Reject SimulationMeta; accept Engine protocol — ADR-034/017.
- Confirm plane boundaries — ADR-036.
- Confirm research→order authorization chain — ADR-037.
- Deprecation window length for
core/types.py(propose 2 minor releases) — ADR-038. - Phase 6 flag schedule for RLS / halt propagation / MCP audience in staging→prod.
- Whether to add additive
Filltable vs enrichexecution_reportsonly. - Whether LabRuntime joins AGENTS hard-rule list as 9th canonical runtime (docs/governance).
27. Prioritized implementation backlog
Immediate next sprint (executable tickets — do not implement in this doc set)
| ID | Ticket | Outcome |
|---|---|---|
| S1 | Accept ADRs 033–020 in arch review | Governance lock |
| S2 | Add CI guard: ban new RuntimeDeploymentSpec, SimulationMeta, temporalio imports | Prevent regression |
| S3 | Draft alphaswarm/facade/ skeleton + import smoke tests (no behavior change) | Phase 1 start |
| S4 | Implement ExecutionProfile translator module + bidirectional tests | Close H4 |
| S5 | Inventory all core/types.py import sites; publish strangler spreadsheet | Phase 3 prep |
| S6 | Design RFC for Forecast + Scenario domain types (schemas only) | Phase 4 prep |
| S7 | Add tests/perf/ scaffold with kill-switch + WorkloadRuntime benchmarks | Gate 9 start |
| S8 | Staging plan: enable orchestration_kill_propagation_enabled + measure fan-out | Phase 6 prep |
| S9 | RLS-on integration test plan (fixture tenants, negative cross-tenant) | Phase 6 prep |
| S10 | Curator refresh after docs land (or keep debt note) | Index compliance |
Medium backlog
- Golden replay digests for event-driven + LOB.
- Listing + Fill persistence design.
- Admin Playwright restoration.
- Credential store unification (
alphaswarm_configvsalphaswarm_core). - Document LabRuntime in AGENTS hard-rule table.
28. Definition of done
This transformation program is done when:
- ADRs 033–020 accepted and linked from intro/architecture indexes.
- Public guidance states: import domain/application via
alphaswarm; infra contracts viaalphaswarm_core; mutate via controller; execute via worker. - No duplicate DeploymentSpec/WorkRequest/ExecutorRouter/control planes introduced.
- Domain dual-rep hotspots have a dated deprecation path with CI tracking.
- ExecutionProfile translation is single-homed with tests.
- Activity / Simulation / Workflow / WorkRequest / DeploymentSpec vocabulary appears in AGENTS/docs without collapse.
- Safety flags for RLS, halt propagation, and MCP audience have a staged enablement record.
- Perf methodology and harness exist; no marketing latency claims remain in requirements.
alphaswarm_indexrefreshed (curator) for new architecture docs.- Production behavior unchanged except behind explicit approved flags.
Appendix A — Mermaid: target request paths
Appendix B — Document control
| Field | Value |
|---|---|
| Authors | Consolidation architect (planning subagent) |
| Inputs | Tracks DOC, A+E, B+C, D+F; spot verification 2026-08-10 |
| Non-goals | No production Python/TS behavior changes in this change set |
| Next | Architecture review → sprint tickets S1–S10 |