Architecture Enhancement Guide — Agentic Quant Research & Trading Platform
How to read this guide. It operationalizes the “Enhanced Blueprint for an Agentic Quant Research and Trading Platform” memo. The memo is a strong requirements charter, but it was written against an “unspecified codebase.” AlphaSwarm is not unspecified — it is a mature, multi-repository platform (31
alphaswarm_*repositories at the time this guide was written; the estate has since grown — 41alphaswarm*working trees are checked out under the workspace root as of this update) that already enforces the majority of the blueprint as hard architecture. This guide therefore does three things the memo could not:
- Grounds every blueprint theme against real files,
alphaswarm/AGENTS.mdhard rules, and ADRs.- Corrects the memo's external citations with verified June-2026 research.
- Narrows the work to ten prioritized enhancements that close genuine gaps, each with a worked artifact and a migration path.
This supersedes the archived
archive/alphaswarm-enhancement-plan.md, which is retained as historical background only.
Status update (verified against the working trees). This guide's gap analysis reflects the platform as of 2026-06-19. Several of the ten enhancements below have since landed at least partially — ADRs 017 (E1, partially implemented:
OrderLifecycle.tla/ReplaySnapshot.tla/SpecVersion.tlanow exist and are TLC-verified), 019 (E6, Accepted and partially implemented:alphaswarm_core/registration.pyships), 020 (E7, substantially implemented), and 021 (E8, partially implemented) carry their own current implementation-status notes — check those ADRs, not this guide, for up-to-date status on those items.OrderFSM.on_fill()(E2) also now takes anexec_idand dedups on it. The narrative and worked designs below are left as originally written for historical/design-rationale context; treat "Finding" language in E1/E2 as describing the state at time of writing, not today.
1. Executive summary
The single most important finding is a reframing. The memo's central anxiety —
"turn a conceptual research repo into a governed platform with schema-first
contracts, plugin registration, agent guardrails, lineage, and replayable state"
— describes work AlphaSwarm has already largely done. The evidence is not
aspirational; it is encoded as CI-enforced hard rules (alphaswarm/AGENTS.md) and Accepted ADRs:
| Blueprint pillar | AlphaSwarm reality (already shipping) |
|---|---|
| Schema-first, versioned message contracts | Hash-locked immutable *_spec_versions rows for Agent / Bot / RL / Analysis / Workflow / Terraform / KB specs (Hard Rules 13, 15, 17, 24, 41, 43, 57); Pydantic-v2 contracts with embedded SCHEMA_VERSION and major-version fail-closed (alphaswarm_core/contracts/) |
| Plugin registration via metaclass | RLComponentMeta, InfrastructureProviderMeta, IntegrationMeta, SecretStoreMeta, KBAdapterMeta, IdentityProviderMeta, OrchestrationAdapterMeta, CanonicalSchemaMeta — definition-time auto-registration (Hard Rules 19, 25, 45, 58) |
| "Agents are advisors, not operators" | Hard Rule 22: agents must not read Postgres/Iceberg directly — only via registered DataMCPTools; role prompts forbid order/execution tools (roles.py, trader/signal_emitter.py) |
| Formal order-lifecycle state machine | Guarded OrderFSM with _VALID_FORWARD transition table, idempotent replay, over-fill→DISPUTED (alphaswarm_bots/execution/lifecycle.py) |
| Event-driven core + deterministic replay | ADR 008 event sourcing (append-only bot_events, snapshots, replay_events); CQRS projections; transactional outbox + CDC |
| Ports-and-adapters | ADR 004/005: InfrastructureProvider ABC in alphaswarm_core; ExecutionAdapter/MarketDataAdapter/ControlPlaneAdapter ABCs (alphaswarm_bots/adapters/protocol.py) |
| Bounded tool capability + auth scopes | MCP servers are RFC 9728 + RFC 8707 conformant (Hard Rule 49); DataMCPTool carries mutates/required_scopes/tenancy_posture; RFC 8693 delegated tokens (Hard Rule 54) |
| Research lineage as first-class | Hard Rule 34 (experiment_id on every run) + Hard Rule 48 (bipartite lineage graph) + per-spec hash-locking |
| Risk + execution as pluggable policy | PreTradeRiskEngine with RTS 6 / SEC 15c3-5 policies (regulatory citations); ExecutionAlgorithm ABC incl. an Almgren-Chriss IS algo (alphaswarm_bots/execution/) |
| Observability with cross-language trace parity | System of record: alphaswarm_core.observe (W3C traceparent + alphaswarm-* headers) plus the platform OpenTelemetry Collector pipeline; @alphaswarm/observe is the wire-compatible browser SDK and alphaswarm_observe is an optional ClickHouse/forwarder sink |
Against that backdrop, the genuine, high-leverage gaps are narrow and specific. Ten enhancements, grouped into four themes, close them:
- Theme A — Correctness & formal methods. (E1) Add TLA+/Apalache specs for
the dangerous state machines (order lifecycle, event-sourced replay, spec
snapshot/replay) — at the time of writing there were zero
.tlafiles in the platform (see the guide-wide status update above: this is no longer the case; ADR 017 tracks current status). (E2) Make inbound fill idempotency structural —OrderFSM.on_fill()had noexec_idand silently double-counted a duplicate sub-quantity fill (also since addressed — see the status update above). (E3) Harden replay determinism on the live bot path (a wall-clockdatetime.now()is stamped into replayable history). (E4) Introduce a design-by-contract + property-based-testing layer (hypothesisis a declared but unused dev-dependency; there is no contract library). - Theme B — Domain-model unification. (E5) Collapse the two divergent
order models (
alphaswarm_botsmsgspec vs. monolith dataclass/Pydantic) into one canonical execution aggregate with explicitPortfolioandExecutionPolicy. (E6) Extract the ~10 bespoke registration metaclasses into onealphaswarm_coremetaclass toolkit (RegisteredComponentMeta/SchemaBoundMeta/CapabilityMeta). - Theme C — Agent governance. (E7) Add
CapabilityMeta+ uniform tool-IO schema versioning + a structural advisor-only gate so the typedStrategyPromotionRequestis the only research→live write path. - Theme D — Research integrity & edges. (E8) Add a first-class hash-locked
ResearchExperimentlineage aggregate. (E9) Achieve backtest-validity parity (port the RL purged-CV suite toalphaswarm_models) and ship Freqtrade-style automated bias detectors. (E10) Close the edge contracts: publish the WebSocketLiveEventschema, maketraceparentW3C-strict, surface MCP tool annotations, and register the orphaned ADRs.
None of these is a rewrite. Every one is additive and lands along a seam the platform already exposes — exactly the migration discipline ADR 015 mandates.
2. Method and evidence base
This guide was produced by: (a) deep technical research (Tavily + web) on the
named technologies and patterns, verified against primary sources as of
2026-06; and (b) a structured analysis of all 31 alphaswarm_* repositories,
focused on the blueprint's themes (domain model, metaclasses/registries,
contracts/schemas, the order lifecycle, MCP integration, agent governance,
testing, observability). Claims below cite real files and line-level anchors
in the working tree and real URLs for external facts. Where the memo's
opaque citation tokens (turn0fileX, turnNsearchN) made a claim unverifiable,
the corrected fact and its source are called out explicitly.
3. Conformance map — blueprint → reality → status
Legend: ✅ implemented · 🟡 partial · ❌ gap.
3.1 Architecture & domain model
| Blueprint recommendation | AlphaSwarm reality | Status |
|---|---|---|
| Layered, ports-and-adapters topology | Feature-sliced monolith with a hexagonal core: alphaswarm/core/domain/ (pure model) + core/interfaces.py (~25 port ABCs: IBrokerage, IAlphaModel, IExecutionModel, …) + core/registry.py (DI). Driven adapters under providers/, data/, streaming/; driving adapters under api/, cli/, ws/ | ✅ (seams are convention/registry-enforced, not folder-enforced) |
| Modular monolith first; split only on evidence | ADR 015 runtime-decomposition: explicitly rejects per-domain microservices; adopts a cell-based modular monolith, extracting only along hash-locked spec seams / MCP surfaces / control plane | ✅ |
| Immutable value types; aggregates mutate via named commands | alphaswarm_bots/schemas/trading.py uses msgspec.Struct(frozen=True, gc=False) for NewOrder/Fill/Position; core/domain/orders.py DomainOrder.validate_flags() invariants | ✅ (in bots); 🟡 (monolith uses a second, weaker model — see E5) |
Single canonical Order/Portfolio/Fill/ExecutionPlan aggregate | Two parallel order models; no explicit Portfolio aggregate (it is an implicit projection); no ExecutionPolicy aggregate (closest is ExecutionLayerSpec) | ❌ → E5 |
3.2 Contracts, schemas, and registration
| Blueprint recommendation | AlphaSwarm reality | Status |
|---|---|---|
| Schema-first, versioned command/event contracts | Pydantic-v2 everywhere; alphaswarm_core/contracts/ DTOs embed SCHEMA_VERSION="1.0.0", fail closed on major-version mismatch; Avro .avsc + Apicurio registry for streaming | ✅ |
RegisteredComponentMeta (auto-register, unique IDs) | RLComponentMeta(ABCMeta) (alphaswarm_rl/core/base.py:68) and 8+ siblings register at class-definition time | 🟡 — no dedup-by-id (last-writer-wins), abstract-method contract unchecked; alphaswarm_models still uses a @register decorator, not a metaclass → E6 |
SchemaBoundMeta (attach JSON-Schema identity to DTOs) | Pydantic emits JSON Schema 2020-12 via model_json_schema(), but there is no metaclass binding schema identity/version onto command/event/tool-IO classes uniformly | 🟡 → E6/E7 |
CapabilityMeta (capability manifest, scopes, transport, side-effects) | DataMCPTool carries mutates/required_scopes/tenancy_posture ClassVars; AgentCapabilities/McpServerSpec carry transport metadata — but this is not a metaclass and not uniform across the CrewAI tool registry or Platform-Context MCP | ❌ → E7 |
| Hash-locked, replayable spec versions | SpecPersister ABC (alphaswarm_core/runtime/persistence.py) + snapshot_hash() (SHA-256 of canonical JSON) → immutable *_spec_versions; replay_spec_version() rehydrates | ✅ |
3.3 Agentic layer
| Blueprint recommendation | AlphaSwarm reality | Status |
|---|---|---|
| Agents emit structured proposals, never mutate state | Hard Rule 22; roles.py prompts ("you propose only… never call an order/execution tool"); trader/signal_emitter.py emits signals only | 🟡 — enforced by tool wiring + prompts, not structurally (nothing rejects a spec that binds a mutating tool) → E7 |
| Capability manifests + auth scopes + transport | ToolRef.scopes, MCPToolContext(granted_scopes=…), RFC 8693 delegated tokens (Hard Rule 54), RFC 9728/8707 conformance (Hard Rule 49) | ✅ (Data MCP) / 🟡 (CrewAI + Platform-Context MCP lack it) → E7 |
| Versioned tool input/output schemas | Data MCP versions tool descriptors (SHA-256 → mcp_tool_versions); Codebase/Platform MCP and CrewAI tools do not; Platform-Context MCP IO is raw JSON-Schema dicts | 🟡 → E7 |
| Persistent decision log + checkpoints | graph/decision_log.py (append-only), graph/checkpointer.py (RedisCheckpointer, LangGraph), RedisHybridMemory, full agent_runs_v2/agent_run_steps trace | ✅ |
| Structured/typed LLM outputs | GuardrailSpec.output_schema (JSON-Schema or Pydantic FQN) + _guardrail_check; but parsing is JSON-from-text, not provider structured-output | 🟡 |
| Role-specialized multi-agent graph | CrewAI crews + role agents (PM/QR/QD/DQE/Desk/Risk) + LangGraph orchestration | ✅ |
3.4 Execution, risk, and data integrity
| Blueprint recommendation | AlphaSwarm reality | Status |
|---|---|---|
| Formal order-lifecycle state machine | Guarded OrderFSM (alphaswarm_bots/execution/lifecycle.py) — but only as a Python dict + pytest, not model-checked | 🟡 → E1 |
| Idempotent fill handling | Fill.exec_id + dedup key documented; OrderFSM.on_fill() does not check it — duplicate sub-quantity fill double-counts | ❌ → E2 |
| Deterministic replay | Monolith backtest replay is equality-tested; live bot replay reuses handlers but stamps datetime.now() into history and EventStore.flush swallows failures | 🟡 → E3 |
| Risk as pluggable pre-trade/intra-trade policy | PreTradeRiskEngine + PreTradePolicy Protocol; RTS 6 / SEC 15c3-5 policies with regulatory citations | ✅ |
| Execution policy (Almgren-Chriss-aware) | ExecutionAlgorithm ABC: TWAPAlgo/VWAPAlgo/POVAlgo/ISAlgo (Almgren-Chriss)/IcebergAlgo; SmartOrderRouter (latency-aware) | ✅ |
| Research lineage aggregate (reproducible from inputs) | AlphaBacktestExperiment ties inputs via FK hints + unhashed JSON; dataset_hash is caller-supplied; factor defs are code-only; no rationale; no own snapshot_hash | ❌ → E8 |
| Bias/leakage tests (look-ahead, survivorship) | alphaswarm_rl ships purged CPCV/PBO/DSR/walk-forward; alphaswarm_models has none, plus a latent .ffill().bfill() leak; survivorship unaddressed everywhere | 🟡/❌ → E9 |
3.5 Specification stack & testing
| Blueprint recommendation | AlphaSwarm reality | Status |
|---|---|---|
| Runtime contracts (pre/post/invariants) | Implicit in Pydantic validators; no icontract/deal/@require/@ensure; Hypothesis property-based testing has since landed for the order FSM (alphaswarm_bots/tests/test_order_fsm_properties.py) | 🟡 → E4 |
| Schema contracts (JSON Schema / Pydantic / Zod) | Strong: Pydantic-v2 → OpenAPI; openapi-typescript codegen to the frontends; Zod at form edges | ✅ (REST) / 🟡 (no WS schema) → E10 |
| Protocol contracts (MCP, versioned HTTP) | RFC 9728/8707 conformant MCP; oasdiff OpenAPI drift gate in CI | ✅ |
| Temporal contracts (TLA+ safety/liveness) | Was none (0 .tla/.als files) at time of writing; OrderLifecycle.tla/ReplaySnapshot.tla/SpecVersion.tla now exist (ADR 017, partially implemented) | 🟡 → E1 |
| Orthogonal test planes | pytest + contract tests + replay/immutability tests + chaos tests; but hypothesis unused, no coverage gate, no bias auto-detectors | 🟡 → E4/E9 |
4. Landscape calibration (and corrections to the memo)
The memo's comparative table is directionally right but contains stale facts and undersells how much AlphaSwarm already matches. Verified June-2026 positioning:
- NautilusTrader — deterministic single-threaded
NautilusKernelover aMessageBus, with research↔live parity achieved by swapping only the data/execution client at one port (SimulatedExchange↔ venue adapter), plus anOrderEmulator+ pluggableExecAlgorithm(TWAP). AlphaSwarm'sBotKernel(single-thread uvloop, 7 coroutines) andadapters/protocol.pyalready mirror this. Still worth borrowing: the kernel-owns-the-Clock pattern (feeds E3) and the explicit emulate-then-release execution split. Source: https://nautilustrader.io/docs/latest/concepts/architecture. - QuantConnect LEAN — the five-stage Algorithm Framework
(Universe → Alpha/
Insight→ Portfolio Construction/PortfolioTarget→ Risk → Execution) cleanly separates signal from sizing from risk from execution. AlphaSwarm'sstrategy → risk → execution → reconcilekernel loop is the same shape; adopting theInsight → PortfolioTargetvocabulary would sharpen E5. Source: https://www.quantconnect.com/docs/v2/writing-algorithms/algorithm-framework/overview. - Freqtrade — ships
lookahead-analysis(chained truncated backtests) andrecursive-analysis(variable-warmup indicator variance) as first-class CLI bias detectors. AlphaSwarm has no equivalent → E9. Source: https://www.freqtrade.io/en/stable/lookahead-analysis. - Qlib — config-driven
qrun+init_instance_by_config(class/module_path/kwargs) is exactly AlphaSwarm Hard Rule 8 +core/registry.build_from_config. Qlib's nested executor (daily strategy nesting an intraday/RL sub-executor) is a pattern AlphaSwarm's RL execution envs could adopt. Qlib is actively maintained (v0.9.x, 2025) and now drives MSRA's R&D-Agent-Quant (NeurIPS 2025). Source: https://github.com/microsoft/qlib. - TradingAgents — LangGraph persona graph with bull/bear debate and a reflection node over shared structured state, plus persistent memory and checkpoints. AlphaSwarm has the memory/checkpoints; adding bounded debate/reflection sub-graphs is an incremental win. Source: https://arxiv.org/abs/2412.20138.
- OpenBB & QRAFTI — both converge on bounding agents to a typed,
validated tool interface over governed data (OpenBB's
agents.json+/querycontract; QRAFTI's MCP factor-tool servers + Pydantic-AI + reflection + computation-graph traceability). This is AlphaSwarm's Hard Rule 22 already; E7 hardens it structurally. Sources: https://docs.openbb.co/workspace/developers/agents-integration, https://arxiv.org/abs/2412.20138 (TradingAgents), QRAFTI (arXiv id pending — the framework/authors are attested but the identifier is anomalous; verify before formal citation).
Memo fact corrections (verify before quoting the memo).
| Memo claim | Verified June-2026 fact | Source |
|---|---|---|
MCP spec revision is 2025-06-18 | Current released spec is 2025-11-25; a 2026-07-28 revision is an RC, not yet stable | https://blog.modelcontextprotocol.io/posts/2025-11-25-first-mcp-anniversary |
"Python v1.x stable; TS v2 on main is pre-alpha" | Correct — Python mcp v1.x is stable (~1.28); TS SDK v1.x is the production line, v2 is pre-alpha (v2.0.0-alpha.1/.2, ~Apr 2026), GA targeted ~Q3 2026. FastMCP 1.0 is merged into the official Python SDK; FastMCP 2.x/3.x is a separate superset | https://github.com/modelcontextprotocol/typescript-sdk, https://pypistats.org/packages/mcp |
| (not in memo) MCP tool annotations | The spec defines readOnlyHint/destructiveHint/idempotentHint/openWorldHint — these are the memo's "side-effect classification," ready to adopt in E7 | https://modelcontextprotocol.io/specification/2025-11-25/server/tools |
5. The contract stack, grounded
The memo's four-layer contract stack is the right organizing idea. AlphaSwarm is strong on the middle two layers and thin on the outer two:
The enhancements below add the missing Runtime and Temporal layers without disturbing the Schema and Protocol layers that already work.
6. Prioritized enhancements
Each enhancement states the finding, the evidence (real files), the research grounding, a design with a worked artifact, and a migration note. Priority reflects architectural leverage × risk reduction, not coding cost.
Theme A — Correctness & formal methods
E1 — Model-check the dangerous state machines (TLA+ / Apalache) · Priority: Highest
Finding (as of 2026-06-19). There were zero .tla/.als/PlusCal
artifacts in any repository. The platform's most safety-critical state
machines — the order lifecycle, the event-sourced replay loop (ADR 008), and
spec snapshot/resume — were guaranteed only by Python guards and
example-based pytest.
Update. This is no longer the case: alphaswarm_bots/specs/OrderLifecycle.tla,
alphaswarm_bots/specs/ReplaySnapshot.tla, and
alphaswarm_core/specs/SpecVersion.tla now exist (verified present in the
working trees) and are TLC-verified per ADR 017, which tracks current
status — the CI gate (rollout step 2 below) was not yet wired as of ADR 017's
last review.
Evidence. alphaswarm_bots/execution/lifecycle.py encodes the FSM as
_VALID_FORWARD: dict[OrderStatus, frozenset[OrderStatus]] plus
alphaswarm/tests/bots/execution/test_order_fsm.py (8 example cases). Replay
equality is asserted only for the monolith backtest (tests/test_event_replay.py),
not the live bot path.
Research grounding. TLA+ specifies state machines as Init + next-state
actions and expresses safety (invariants) and liveness (eventually-terminal)
properties. TLC enumerates reachable states (good for liveness on small
configs); Apalache is symbolic (SMT/Z3) and proves inductive invariants
for unbounded executions. Both are complementary and well suited to an
order-lifecycle/replay machine. Sources: https://lamport.azurewebsites.net/tla/tools.html,
https://apalache-mc.org.
Design. Add specs/OrderLifecycle.tla, faithful to the real
_VALID_FORWARD table and on_fill semantics (over-fill → DISPUTED,
idempotent re-application). Model fills as a set of distinct exec_ids so the
same spec proves both the lifecycle (E1) and idempotency (E2). Check
safety with Apalache, liveness with TLC, in CI for the bots/execution module.
------------------------------ MODULE OrderLifecycle ------------------------------
EXTENDS Naturals, FiniteSets
CONSTANTS Quantity, \* target quantity (a positive Nat), e.g. 3
ExecIds \* finite set of distinct venue exec ids, e.g. {e1,e2}
ASSUME Quantity \in Nat \ {0}
VARIABLES status, \* current OrderStatus
cumQty, \* cumulative filled quantity
applied \* set of exec ids already applied (idempotency ledger)
Terminal == {"FILLED", "CANCELLED", "REJECTED", "EXPIRED"}
\* DISPUTED is the only absorbing state; Terminal states may still escalate to it.
\* Faithful transcription of _VALID_FORWARD (lifecycle.py).
Forward == [
CREATED |-> {"VALIDATED", "REJECTED", "DISPUTED"},
VALIDATED |-> {"ROUTED", "REJECTED", "CANCELLED"},
ROUTED |-> {"ACKNOWLEDGED", "REJECTED", "CANCELLED", "EXPIRED", "DISPUTED"},
ACKNOWLEDGED |-> {"PARTIALLY_FILLED", "FILLED", "CANCEL_PENDING",
"CANCELLED", "REJECTED", "EXPIRED", "DISPUTED"},
PARTIALLY_FILLED |-> {"PARTIALLY_FILLED", "FILLED", "CANCEL_PENDING",
"CANCELLED", "EXPIRED", "DISPUTED"},
CANCEL_PENDING |-> {"CANCELLED", "FILLED", "PARTIALLY_FILLED", "DISPUTED"},
FILLED |-> {"DISPUTED"},
CANCELLED |-> {"DISPUTED"},
REJECTED |-> {"DISPUTED"},
EXPIRED |-> {"DISPUTED"},
DISPUTED |-> {}
]
Init == status = "CREATED" /\ cumQty = 0 /\ applied = {}
\* Non-fill transitions (risk pass, routing, cancel, reject, expire, dispute).
Move(to) ==
/\ to \in Forward[status]
/\ to \notin {"PARTIALLY_FILLED", "FILLED"} \* fills go through Fill/Dispute below
/\ status' = to /\ UNCHANGED <<cumQty, applied>>
\* Apply a venue fill identified by exec id `e`, quantity 1 (unit fills).
\* Idempotent: a previously-applied exec id is a no-op (the SAFE behaviour E2 adds).
SafeFill(e) ==
/\ status \in {"ACKNOWLEDGED", "PARTIALLY_FILLED", "CANCEL_PENDING"}
/\ e \in ExecIds
/\ IF e \in applied
THEN UNCHANGED <<status, cumQty, applied>> \* dedup no-op
ELSE /\ applied' = applied \cup {e}
/\ cumQty' = cumQty + 1
/\ status' = IF cumQty + 1 = Quantity THEN "FILLED"
ELSE IF cumQty + 1 > Quantity THEN "DISPUTED"
ELSE "PARTIALLY_FILLED"
Next == (\E to \in {"VALIDATED","ROUTED","ACKNOWLEDGED","CANCEL_PENDING",
"CANCELLED","REJECTED","EXPIRED","DISPUTED"} : Move(to))
\/ (\E e \in ExecIds : SafeFill(e))
\* ---- Safety invariants ----
TypeOK == status \in DOMAIN Forward /\ cumQty \in 0..(Quantity + 1) /\ applied \subseteq ExecIds
QtyConservation == (cumQty <= Quantity) \/ (status = "DISPUTED")
NoResurrection == (status \in Terminal) => (status' \in (Terminal \cup {"DISPUTED"})) \* used with [Next]_vars
IdempotentLedger == Cardinality(applied) = cumQty \/ status = "DISPUTED"
\* ---- Liveness (check with TLC under weak fairness) ----
EventuallyTerminal == <>(status \in (Terminal \cup {"DISPUTED"}))
=================================================================================
QtyConservation and IdempotentLedger are exactly the properties that the
unsafe current on_fill violates: replace SafeFill with a non-dedup variant
and Apalache produces the counterexample that motivates E2.
Migration. (1) Land specs/ with OrderLifecycle.tla + a TLC .cfg for
Quantity=3, ExecIds={e1,e2,e3} and an Apalache Inv run. (2) Add a CI job
(make spec-check) gating alphaswarm_bots/execution. (3) Follow with
ReplaySnapshot.tla (ADR 008: append-only log + snapshot anchor ⇒ projection
convergence) and SpecVersion.tla (get-or-create-by-hash is idempotent and
never mutates a version). Author an ADR, “Formal specs for critical state
machines.”
E2 — Make inbound fill idempotency structural · Priority: Highest
Finding (as of 2026-06-19). OrderFSM.on_fill(self, *, fill_qty, fill_price)
had no idempotency key. A venue (or a drop-copy replay) that re-delivers
the same exec_id whose quantity keeps cumulative ≤ target would be applied
twice, silently inflating the position; the over-fill→DISPUTED guard
only fired when the duplicate pushed cumulative over target.
Evidence. alphaswarm_bots/execution/lifecycle.py:214-250 (no exec_id
parameter, at time of writing). The dedup key (trade_date, exec_id, symbol, side, exec_type) is documented on Fill in
alphaswarm_bots/schemas/trading.py but was never checked in OMS.on_fill.
Update. This has since been fixed: on_fill now takes an optional
exec_id keyword argument and dedups against an _applied_exec_ids set,
returning a no-op transition on a repeat exec_id — the docstring in
lifecycle.py explicitly tags this "enhancement E2".
Research grounding. Event-sourcing guidance is explicit that replays and at-least-once delivery must not double-apply effects; dedup belongs at the fold, keyed on a stable event id. Source: https://martinfowler.com/eaaDev/EventSourcing.html.
Design. Thread exec_id into on_fill and keep an _applied_exec_ids: set
on the FSM (and a durable projection-side ledger). This is the Python mirror of
the TLA+ applied ledger in E1.
def on_fill(self, *, exec_id: str, fill_qty: Decimal, fill_price: Decimal) -> OrderTransition:
if exec_id in self._applied_exec_ids: # idempotent: venue replay / at-least-once
return self._record_noop()
if fill_qty <= 0:
raise OrderTransitionError(f"order {self.client_order_id}: fill_qty must be > 0")
self._applied_exec_ids.add(exec_id)
... # existing VWAP + cumulative + FILLED/PARTIALLY_FILLED/DISPUTED logic
Migration. Additive signature (keyword-only exec_id); update the two
call sites in execution/oms.py + execution/reconcile.py; back it with a
fills uniqueness constraint on (bot_id, exec_id) in the event store. Verify
with the Hypothesis stateful test in E4.
E3 — Prove replay determinism on the live bot path · Priority: High
Finding. Live-path replay is probably deterministic but unproven, and
two concrete hazards undermine it: (a) OrderTransition.at_utc = datetime.now(timezone.utc) is stamped into replayable history
(lifecycle.py:202,270), so a replay produces different records than the
original; (b) EventStore.flush swallows failures, so events can be dropped and
a later replay diverges silently.
Update. Both (a) and (b) have since been addressed and are explicitly
tagged "enhancement E3" in the source: OrderFSM now takes an injectable
clock (default _wall_clock, but a deterministic clock can be substituted
so replay uses event-time, not wall-clock — alphaswarm_bots/execution/lifecycle.py),
and EventStore.flush now surfaces drops via a WARNING + dropped_events
counter and fails closed in strict=True mode
(alphaswarm_bots/state/store.py). A live-bot replay-equality test (as
opposed to the existing monolith-backtest one) was not found in the
alphaswarm_bots test tree as of this update — that part of E3 may still be
open.
Evidence. lifecycle.py (wall-clock in history); alphaswarm_bots/state/store.py
(flush exception handling); only tests/test_event_replay.py (monolith
backtest) asserts replay equality — there is no live-bot replay-equality test.
Research grounding. NautilusTrader achieves parity by having the kernel own the Clock and feeding every component the same time source in sim and live; event-sourcing replays must be byte-stable and side-effect-gated. Sources: https://nautilustrader.io/docs/latest/concepts/architecture, https://martinfowler.com/eaaDev/EventSourcing.html.
Design. Inject a Clock (the BotKernel already owns one) into OrderFSM
so at_utc is the event time under replay, not wall-clock; make
EventStore.flush fail-closed (or route through the existing transactional
outbox uniformly); add a live-path replay_events equality test
(diff_event_logs == [] + projection equality) and the ReplaySnapshot.tla
invariant from E1.
Migration. One module at a time: clock injection → flush hardening → live replay test. No interface break for callers that pass the kernel clock.
E4 — Add runtime contracts + activate property-based testing · Priority: High
Finding (as of 2026-06-19). The memo's runtime-contract layer was
absent: no icontract/deal, no @require/@ensure/@invariant. And
hypothesis>=6.100 was declared in pyproject but had zero
@given/from hypothesis usages — property-based testing was dormant
exactly where it pays most (portfolio accounting, order decomposition, fill
aggregation). There is also no enforced coverage gate.
Evidence. Grep across the tree: 0 real hypothesis imports; 0 contract-lib
imports (at time of writing). Contracts are implicit in Pydantic validators.
Update. The Hypothesis half has since landed:
alphaswarm_bots/tests/test_order_fsm_properties.py is a hypothesis
RuleBasedStateMachine fuzzing create/fill/duplicate-fill sequences against
the FSM (explicitly tagged "enhancement E4" and paired with the E1 TLA+
spec), so hypothesis is no longer an unused dependency. icontract/deal
design-by-contract usage was not found anywhere in the tree as of this
update — that half of E4 still looks open.
Research grounding. Eiffel's pre/post/invariant vocabulary maps cleanly to
icontract (@icontract.require/ensure/invariant), which bridges to Hypothesis
via icontract-hypothesis. Hypothesis RuleBasedStateMachine with @rule/
@invariant is purpose-built to fuzz a lifecycle FSM and assert global
invariants after every step. Sources:
Eiffel Design by Contract,
https://hypothesis.readthedocs.io/en/latest/stateful.html.
Design. (1) Put icontract invariants on the canonical aggregate (E5), e.g.
Portfolio.apply_fill: @require(lambda fill: fill.qty > 0),
@ensure(...) cash/position reconcile, class @invariant no-money-created.
(2) Reuse the same predicates as a Hypothesis state machine that randomizes
create/fill (incl. duplicate exec_id)/cancel/replay and proves quantity
conservation + idempotency — the executable twin of the E1 TLA+ spec.
from decimal import Decimal
from hypothesis import strategies as st
from hypothesis.stateful import RuleBasedStateMachine, rule, invariant, precondition
from alphaswarm_bots.execution.lifecycle import OrderFSM, OrderStatus
class OrderFSMMachine(RuleBasedStateMachine):
def __init__(self) -> None:
super().__init__()
self.fsm = OrderFSM(client_order_id="t", quantity=Decimal("3"))
self.fsm.transition(OrderStatus.VALIDATED); self.fsm.transition(OrderStatus.ROUTED)
self.fsm.transition(OrderStatus.ACKNOWLEDGED)
self.seen: set[str] = set()
@precondition(lambda self: not self.fsm.is_terminal())
@rule(exec_id=st.sampled_from(["e1", "e2", "e3"])) # repeats exercise idempotency
def fill_one(self, exec_id: str) -> None:
self.fsm.on_fill(exec_id=exec_id, fill_qty=Decimal("1"), fill_price=Decimal("100"))
self.seen.add(exec_id)
@invariant()
def quantity_is_conserved(self) -> None:
assert self.fsm.cumulative_qty <= self.fsm.quantity or self.fsm.state == OrderStatus.DISPUTED
@invariant()
def cumulative_equals_distinct_fills(self) -> None:
assert self.fsm.cumulative_qty == len(self.seen) or self.fsm.state == OrderStatus.DISPUTED
TestOrderFSM = OrderFSMMachine.TestCase
This test fails today (E2 not yet applied) and passes once exec_id
dedup lands — a regression guard for the whole theme. Migration: add
icontract to alphaswarm_core deps; introduce a non-blocking coverage report
first, then a per-module gate on bots/execution and core/contracts.
Theme B — Domain-model unification
E5 — One canonical execution aggregate (+ explicit Portfolio / ExecutionPolicy) · Priority: High
Finding. Two order models coexist: alphaswarm_bots (msgspec frozen
structs + guarded OrderFSM) and the monolith (core/domain/orders.py
dataclasses + trading/order_intent.py Pydantic OrderIntent/RiskPolicy/
GateChain + an unguarded trading/execution/order_state.py fold). Portfolio
is an implicit projection; there is no ExecutionPolicy aggregate.
Evidence. OrderIntent exists only in the monolith; the guarded FSM exists
only in bots; order_state.py applies reports without transition guards.
Research grounding. LEAN's Insight → PortfolioTarget and Nautilus's single
order/exec model show the value of one vocabulary from signal to fill. Sources
above.
Design. Promote the bots model as canonical (frozen, guarded, idempotent);
treat monolith OrderIntent as the pre-trade intent that the canonical
ExecutionOrder is derived from; replace order_state.py's fold with OrderFSM.
Introduce a thin Portfolio aggregate over the existing PositionProjection/
PnLProjection with apply_fill (carrying the E4 contracts) and an
ExecutionPolicy value object wrapping the existing ExecutionLayerSpec +
ExecutionAlgorithm selection. Adopt LEAN's PortfolioTarget naming for sizing.
Migration. Extract the canonical model into alphaswarm_core (or a shared
execution boundary) so both stacks import one definition; deprecate the
monolith fold behind a shim. One axis at a time: unify types before unifying
runtimes.
E6 — Unify the metaclass/registry family into a core toolkit · Priority: High
Finding. The platform has ~10 near-identical auto-registering metaclasses
(RLComponentMeta, InfrastructureProviderMeta, IntegrationMeta,
SecretStoreMeta, KBAdapterMeta, IdentityProviderMeta,
OrchestrationAdapterMeta, CanonicalSchemaMeta, …) plus a decorator registry
(core/registry.@register) and the Predictor Hub's tuple registry. They diverge
in quality: RLComponentMeta has no dedup-by-id (last-writer-wins) and does
not validate the abstract-method contract; alphaswarm_models uses a decorator,
not a metaclass.
Evidence. alphaswarm_rl/core/base.py:68 (RLComponentMeta),
alphaswarm/core/registry.py (@register), predictors/hub.py (tuple registry),
plus the *Meta siblings across alphaswarm_core.
Research grounding. The memo's RegisteredComponentMeta / SchemaBoundMeta
/ CapabilityMeta family. Pydantic v2 supplies the schema half for free:
model_json_schema() emits JSON Schema 2020-12 / OpenAPI 3.1, so a
SchemaBoundMeta can stamp schema identity onto every command/event/tool-IO
type. Source: https://docs.pydantic.dev/latest/concepts/json_schema/.
Design. Add alphaswarm_core/registration.py with a small base the existing
metaclasses subclass — preserving their behavior but fixing dedup and
abstract-skip uniformly:
from abc import ABCMeta
class RegisteredComponentMeta(ABCMeta):
"""Definition-time registry with id-uniqueness and abstract-skip."""
registry: dict[tuple[str, str], type] = {} # (kind, component_id) -> cls
def __new__(mcls, name, bases, ns, **kw):
cls = super().__new__(mcls, name, bases, ns, **kw)
if ns.get("__abstract__") or getattr(cls, "__abstractmethods__", None):
return cls # never register abstract bases
cid, kind = getattr(cls, "component_id", None), getattr(cls, "component_kind", None)
if cid and kind:
key = (kind, cid)
if key in mcls.registry and mcls.registry[key] is not cls:
raise TypeError(f"duplicate component {key}: "
f"{mcls.registry[key].__qualname__} vs {cls.__qualname__}")
mcls.registry[key] = cls
return cls
SchemaBoundMeta(RegisteredComponentMeta) additionally calls
cls.__json_schema__ = cls.model_json_schema() and records (schema_id, schema_version); CapabilityMeta (E7) adds the capability manifest. Migration:
re-base one metaclass at a time onto the shared core; add a test that the global
registry has no duplicate keys (catches the silent last-writer-wins bugs today).
Author an ADR, “Unified component registration.”
Theme C — Agent governance
E7 — CapabilityMeta + uniform tool-IO versioning + structural advisor-only gate · Priority: High
Finding. Advisor-only behavior is real but not structural: it relies on
tool wiring and prompt text. Nothing rejects an AgentSpec (e.g.
template_target="research") that binds a mutating tool. Tool-IO schema
versioning exists only for Data MCP; the CrewAI TOOL_REGISTRY and the
Platform-Context MCP (raw JSON-Schema dicts) lack it. The typed money-plane
crossing StrategyPromotionRequest (the memo's B-1) is specified but unbuilt.
Evidence. roles.py, trader/signal_emitter.py (prompt-level); tools/__init__.py
(TOOL_REGISTRY lacks side-effect metadata); alphaswarm_mcp/server.py
(untyped IO); alphaswarm_core/contracts/ (StrategyPromotionRequest DTO exists,
no enforcing gate).
Research grounding. OWASP LLM06:2025 Excessive Agency and the
2025-2026 consensus — least-privilege/default-deny tools, "agent proposes, system
disposes," plan-level governance, immutable decision logs. MCP's tool annotations
(readOnlyHint/destructiveHint/idempotentHint/openWorldHint) are a
ready-made side-effect taxonomy. Sources:
https://genai.owasp.org/llmrisk/llm062025-excessive-agency/,
https://www.anthropic.com/engineering/writing-tools-for-agents,
https://modelcontextprotocol.io/specification/2025-11-25/server/tools.
Design. A Pydantic CapabilityManifest carried by every tool via
CapabilityMeta (E6 base), aligned to MCP annotations, enforced at both
registration and dispatch:
class CapabilityManifest(BaseModel):
model_config = ConfigDict(frozen=True, extra="forbid")
tool_name: str
schema_version: str # bumped on input/output schema change
read_only: bool # == MCP readOnlyHint
destructive: bool = False # == destructiveHint
idempotent: bool = True # == idempotentHint
open_world: bool = False # == openWorldHint
required_scopes: frozenset[str]
input_schema_id: str # SchemaBoundMeta-derived
output_schema_id: str
ADVISORY_TARGETS = {"research", "selection", "analysis"}
def assert_advisor_only(spec: "AgentSpec") -> None:
"""Reject at spec-validation time: advisory agents may bind read-only tools only."""
if spec.template_target in ADVISORY_TARGETS:
for ref in spec.tools:
man = capability_for(ref) # CapabilityMeta.registry lookup
if not man.read_only:
raise ValueError(
f"advisory agent '{spec.template_target}' may not bind "
f"mutating tool '{man.tool_name}'")
This makes Hard Rule 22 structural: the only research→live write path becomes
the typed StrategyPromotionRequest, gated by ValidationVerdict +
PreTradeVerdict + out-of-band human approval. Migration: (1) backfill
CapabilityManifest for the ~55 CrewAI tools and Platform-Context MCP (Pydantic-
type its IO first); (2) wire assert_advisor_only into registry.persist_spec;
(3) extend descriptor-hash versioning (already on Data MCP) to all surfaces.
Theme D — Research integrity & edges
E8 — First-class ResearchExperiment lineage aggregate · Priority: Medium-High
Finding. The memo's central reproducibility artifact is genuinely missing in
the ML layer. AlphaBacktestExperiment is the nearest unit but ties inputs via
FK hints + unhashed JSON params, dataset_hash is caller-supplied, factor
expressions are code-only, there is no rationale, and it has no own
snapshot_hash. Editing a factor string silently changes future runs.
Evidence. alphaswarm_models/.../alpha_backtest_experiment.py
(MLAlphaBacktestRun links versions by FK; configs go to MLflow unhashed).
Hash-locking is proven for side-car specs (PredictorSpec, MLSkillSpec,
RLExperimentSpec) but decoupled from the experiment record.
Research grounding. Qlib's recorder/qrun reproducibility and TradingAgents'
persistent decision log both make the experiment a content-addressed,
re-runnable object. AlphaSwarm already has the mechanism (SpecPersister +
bipartite lineage graph, Hard Rule 48) — it just is not applied here. Sources:
https://qlib.readthedocs.io/en/stable/component/workflow.html,
https://arxiv.org/abs/2412.20138.
Design. A hash-locked ResearchExperimentSpec (via the existing
SpecPersister) that pins, by hash, the dataset snapshot, factor definition(s),
model config, signal/order policy, backtest config — and stores result metrics +
rationale:
class ResearchExperimentSpec(BaseModel): # persisted like every other spec
model_config = ConfigDict(frozen=True)
hypothesis: str
dataset_snapshot_hash: str # computed INSIDE the aggregate, not caller-supplied
factor_defs_hash: str # hash of the resolved factor expressions
model_config_hash: str
policy_hash: str # signal/order/execution policy
backtest_config_hash: str
rationale: str
def snapshot_hash(self) -> str: # SHA-256 of canonical_json — reuses the platform idiom
...
Migration. Wrap AlphaBacktestExperiment to emit a
ResearchExperimentSpec and persist a research_experiment_versions row;
compute dataset_snapshot_hash from the medallion as_of/snapshot_id; hash
factor expressions at resolve time. This unifies the fragments (KB kb_runs, RL
trajectory corpus, graph BacktestRun) under one re-runnable record. Author an
ADR, “ResearchExperiment lineage aggregate.”
E9 — Backtest-validity parity + automated bias detectors · Priority: Medium-High
Finding. Leakage defenses are asymmetric. alphaswarm_rl ships a tested
selection-bias suite (validation/: CombinatorialPurgedKFold, PBO/CSCV,
deflated Sharpe, walk-forward with purge/embargo). alphaswarm_models has
none — only an external purged splitter, with triple-barrier label horizons
not wired into its splitter's purge — and a latent .ffill().bfill() in
shared cleaning (processors.py) that pulls future values backward.
Survivorship bias is unaddressed everywhere, and there is no Freqtrade-style
auto-detector.
Evidence. alphaswarm_rl/validation/* + tests/validation/* (present);
alphaswarm_models/.../walk_forward.py/splits.py (no embargo/purge, no leakage
tests); processors.py (bfill).
Research grounding. Freqtrade's lookahead-analysis (chained truncated
backtests comparing signals) and recursive-analysis (variable-warmup variance)
are first-class, automated detectors — not careful coding. Zipline/Backtrader
enforce point-in-time delivery structurally. Sources:
https://www.freqtrade.io/en/stable/lookahead-analysis,
https://www.freqtrade.io/en/stable/recursive-analysis.
Design. (1) Promote the RL validation contracts into a shared gate used
by alphaswarm_models; wire triple-barrier t1 horizons into the splitter's
purge. (2) Build two CLI/CI detectors: alphaswarm research lookahead-analysis
(re-run on truncated windows; flag signals that shift) and recursive-analysis
(measure last-row variance across startup_candle_count). (3) Remove/guard the
shared bfill. (4) Add a survivorship-aware universe snapshot to the medallion
as_of reader. These are domain-specific correctness tests, named suites in
the E4 test matrix, not generic CI.
E10 — Close the edge contracts · Priority: Medium
Finding. The contract loop is open at the edges. (a) REST is codegen'd
(openapi-typescript) but the WebSocket LiveEvent union has no published
schema — cross-language parity stops at REST. (b) @alphaswarm/observe emits a
16-hex trace-id (Python-parity) rather than the W3C-mandated 32-hex, risking
generic OTel-collector interop. (c) Platform-Context MCP IO is untyped and lacks
tool annotations. (d) The entire ADR set (incl. ADR 015) is orphaned from the
docs sidebar.
Evidence. alphaswarm_client/src/lib/ws/types.ts (LiveEvent union, no
schema); alphaswarm_observe_js/src/propagation.ts (16-hex id);
alphaswarm_mcp/server.py (raw JSON-Schema dicts); sidebars.ts
architectureSidebar (registers only two files).
Design. (1) Publish an AsyncAPI/JSON-Schema contract for LiveEvent and
codegen the TS union from it (single source of truth, like REST). (2) Emit a
W3C-strict 32-hex trace-id in @alphaswarm/observe while keeping the
alphaswarm-* headers. (3) Pydantic-type Platform-Context MCP IO (E6
SchemaBoundMeta) and emit the four MCP tool annotations (E7). (4) Register this
guide and the ADR set in the sidebar (done alongside this document).
7. Migration sequencing
Follow ADR 015's discipline — modular monolith; extract only along seams that already have service-shaped contracts — and the memo's rule, “do not move more than one axis at a time: first wrap, then type, then extract, then specify.”
Wave 1 is pure insurance: it adds the missing Runtime and Temporal contract layers with no interface breaks and immediately retires the two latent correctness hazards (double-counted fills, non-deterministic replay). Complete Wave 1 before attempting any model or registry unification.
8. Roadmap
| Priority | Enhancement | Primary targets | New ADR? | Effort |
|---|---|---|---|---|
| Highest | E1 TLA+/Apalache specs | specs/OrderLifecycle.tla, bots/execution, CI | Yes | Medium |
| Highest | E2 fill idempotency | bots/execution/lifecycle.py, oms.py, reconcile.py | No | Low |
| High | E3 replay determinism | lifecycle.py, state/store.py, live replay test | No | Medium |
| High | E4 DbC + property tests | alphaswarm_core deps, bots/execution, core/contracts | No | Medium |
| High | E5 canonical aggregate | shared execution boundary, monolith shims | Yes | High |
| High | E6 registration toolkit | alphaswarm_core/registration.py, re-base *Meta | Yes | Medium |
| High | E7 CapabilityMeta + gate | tools/, alphaswarm_mcp, registry.persist_spec | Yes | Medium |
| Med-High | E8 ResearchExperiment | alphaswarm_models, research_experiment_versions | Yes | High |
| Med-High | E9 validity parity + detectors | alphaswarm_models/validation, CLI | No | Medium |
| Medium | E10 edge contracts | ws schema, @alphaswarm/observe, alphaswarm_mcp, sidebars.ts | No | Medium |
ADRs (drafted as Proposed alongside this guide): ADR 017 — Formal specs for critical state machines (E1); ADR 018 — Canonical execution aggregate (E5); ADR 019 — Unified component registration (E6); ADR 020 — Agent capability + advisor gate (E7); ADR 021 — ResearchExperiment lineage aggregate (E8).
9. Sources
Internal: alphaswarm/AGENTS.md
(Hard Rules 1–64); ADRs 004,
005,
008,
015; key files
alphaswarm_bots/execution/lifecycle.py, alphaswarm_core/runtime/persistence.py,
alphaswarm_core/contracts/, alphaswarm/core/registry.py,
alphaswarm_rl/core/base.py, alphaswarm_rl/validation/.
External (verified June 2026):
- NautilusTrader architecture — https://nautilustrader.io/docs/latest/concepts/architecture
- QuantConnect LEAN Algorithm Framework — https://www.quantconnect.com/docs/v2/writing-algorithms/algorithm-framework/overview
- Freqtrade lookahead / recursive analysis — https://www.freqtrade.io/en/stable/lookahead-analysis
- Microsoft Qlib — https://github.com/microsoft/qlib · https://qlib.readthedocs.io/en/stable/component/workflow.html
- TradingAgents (arXiv 2412.20138) — https://arxiv.org/abs/2412.20138
- OpenBB agents integration — https://docs.openbb.co/workspace/developers/agents-integration
- Event Sourcing (Fowler) — https://martinfowler.com/eaaDev/EventSourcing.html
- Almgren & Chriss, Optimal Execution of Portfolio Transactions — https://www.smallake.kr/wp-content/uploads/2016/03/optliq.pdf
- MCP spec 2025-11-25 (tools + annotations) — https://modelcontextprotocol.io/specification/2025-11-25/server/tools
- MCP first-anniversary spec note — https://blog.modelcontextprotocol.io/posts/2025-11-25-first-mcp-anniversary
- MCP TypeScript SDK (v2 pre-alpha) — https://github.com/modelcontextprotocol/typescript-sdk
- Pydantic JSON Schema — https://docs.pydantic.dev/latest/concepts/json_schema/
- Design by Contract (Eiffel) — Eiffel Design by Contract
- TLA+ tools (TLC) — https://lamport.azurewebsites.net/tla/tools.html
- Apalache (symbolic model checker) — https://apalache-mc.org
- Hypothesis stateful testing — https://hypothesis.readthedocs.io/en/latest/stateful.html
- OWASP LLM06:2025 Excessive Agency — https://genai.owasp.org/llmrisk/llm062025-excessive-agency/
- Anthropic, writing tools for agents — https://www.anthropic.com/engineering/writing-tools-for-agents