ADR 016 — Pluggable hybrid local execution layer
- Status: Proposed (2026-06-16)
- Authors: Platform team
- Supersedes: Nothing; extends the workload-runtime story
- Related: ADR 004 — provider abstraction, ADR 005 — separated control plane, ADR 015 — runtime decomposition (cells), ADR 003 — Auth0 zero-trust, ADR 013 — Entra as first pool. Full design:
design/hybrid-local-execution-layer.md.
Context
Regulatory, data-gravity, and privacy requirements demand that some AlphaSwarm workloads run on the customer's own machine while the control plane stays in the cloud — running ingestion next to a private dataset, executing agent-emitted strategy code without uploading it, or using a customer's GPU. An external research report ("Architectural Paradigms for Pluggable Hybrid Cloud Execution") plus independent 2025–2026 research recommend a five-part blueprint: outbound reverse-tunnel connectivity, a node abstraction, tiered execution isolation (gVisor/Firecracker), per-tenant identity/isolation, and lease/heartbeat/reaper + idempotency fault tolerance.
A repo-wide audit (the companion design doc) shows AlphaSwarm has already built or scaffolded four of those five:
- An outbound, NAT-traversing device reverse-tunnel (
/tunnel/agent,tunnel_registry) that a local connector (alphaswarm-cli connect up) dials out to, authenticated by a distinctdevice_credentialtoken — described in code as being "for reverse-tunnel agents" — but today it only relays HTTP to a user's backend; it does not carry workloads. - The
InfrastructureProviderABC +WorkloadRuntimeseam (rule 45 / ADR 004) with a self-registering metaclass + registry. - The
alphaswarm_workerExecutor/ExecutorRouterruntime with a typedWorkRequest/WorkResultcontract, idempotency markers, retry/DLQ, and a money-plane risk gate — but it runs in-process on the host with no sandbox. - Per-tenant tenancy + cells:
TenantNamespaceSpec/TenantQuotas,TenancyStrategy.Hybrid, theCell/CellStateDeployment-Stamp lifecycle, a SPIFFE backend scaffold (spire_trust_domain = alphaswarm.fund), and a per-tenant token-bucket limiter.
What is missing is the integration glue plus net-new remote-node fault
tolerance and isolation: there is no local provider, no typed work
channel on the tunnel, no node registry / lease / heartbeat / reaper
(run-ledger tables have no owner/lease/attempts columns; stall detection is a
wall-clock halt with no requeue and there is zero SELECT … FOR UPDATE SKIP LOCKED), no enforced idempotency key, and no execution isolation on
the worker.
Two design forks are load-bearing. (1) The Solution Report recommends mapping
each machine to a Virtual Kubelet "virtual node" — but that abstraction
only makes sense when kube-scheduler is your control plane; AlphaSwarm's
scheduler is Celery routing + WorkloadRuntime + InfrastructureProvider, and
Virtual Kubelet remains CNCF Sandbox after 7+ years. (2) The Report recommends
gRPC reverse tunnels — but ADR 005 already rejected gRPC ("HTTP/JSON is
already understood … until hundreds of req/s") and a working WebSocket tunnel
already exists.
Decision
Adopt a pluggable, customer-hosted Local Execution Node that the cloud
control plane drives through a new local InfrastructureProvider — not a
parallel control path. Specifically:
ProviderKind.LOCAL+LocalProvider(InfrastructureProvider)inalphaswarm_controller, self-registered via the existing metaclass. All customer-hosted workload ops go throughWorkloadRuntime(rule 45): audit row written before dispatch,filter_resourceson lists, kill-switch fan-out and telemetry inherited unchanged.- Evolve the existing WebSocket tunnel (do not add gRPC): add a
work_*frame family (work_register/dispatch/accept/heartbeat/progress/ result/cancel/usage) alongside the HTTP-relay frames, reusing the socket,stream_iddemux, anddevice_credential. Every customer machine stays outbound-only and never connects to the Celery/Redis broker. gRPC and QUIC are explicitly deferred behind a documented scale trigger. - Reject Virtual Kubelet. Model the machine as a registered, single-
tenant
ExecutionNode(a dedicated silo / Azure Deployment-Stamp / cell); add anexecution_nodesregistry and a first-classexecution_targetdispatch dimension. The controller (notkube-scheduler) matches a run'sResourceSpecto a node's advertised capacity; capacity rejection lets the control plane reschedule (the GitLab-KASResourceExhaustedpattern). - Tiered execution isolation on the node —
process→container→gvisor→microvm(Firecracker/Kata) — selected by anisolation_tierfield on the work contract. Agent-/LLM-generated code runs in a microVM (T3), mandatory in multi-tenant mode; the launcher fails closed if a requested tier is unavailable. This realizes the gVisor/Kata/Firecracker + Kyvernorequire-runtime-classintent already named inRESTRUCTURING_PLAN.md§8.3. - Identity: keep the device-credential pairing flow as the bootstrap;
add short-lived, single-run scoped per-work tokens (GitHub-Actions-runner
model) and a per-node SPIFFE SVID via the existing scaffold; prefer
short-TTL + connect-time denylist over revocation lists; fix the
worker's plaintext-token-at-rest gap (OS keyring, rule 53). Identity stays in
the IAM hub (rule 27); credentials via
CredentialResolver(rule 26). - Fault tolerance: add
owner_node_id / lease_expires_at / last_heartbeat_at / attempts / idempotency_key (UNIQUE)to the run-ledger layer (a dedicatednode_run_claimstable), claimed withSELECT … FOR UPDATE SKIP LOCKED, renewed bywork_heartbeat, and reclaimed by a reaper beat task — modelled on the existingOrderOutboxRowtransactional-outbox pattern. At-least-once + enforced idempotency; no literal exactly-once (Two Generals). - Multi-tenancy & metering: bind each node to one
org_id+Cell; propagate tenant context viaRequestContext/claims and strip baggage at the trust boundary; enforce concurrency/rate limits (not CPU quotas) via the existing token-bucket limiter. Meter the control plane / orchestration — never the customer's raw compute (the GitHub/HCP/Datadog/Temporal precedent); the node emits idempotent, event-timestamped,tenant_id-taggedwork_usageevents that roll up intoBillingSummary.
The implementation lands in the alphaswarm_local repo (node agent — no
longer empty; it now carries a CLI, tunnel daemon, bootstrap/compose/k8s
scaffolding, and a staged/inert exec/laptop_runner.py, though the
controller-side LocalProvider and tunnel work_* frame family this ADR
proposes are not yet built), alphaswarm_controller (provider + tunnel
channel), alphaswarm_core (shared contracts), alphaswarm (registry +
ledger + reaper), alphaswarm_worker (isolation launcher + keyring),
alphaswarm_auth (per-work tokens + SVID), and alphaswarm_admin (cell
binding + metering) — respecting the ADR 005 decision tree and all import
boundaries.
Consequences
Positive
- Customer-hosted execution is a provider behind rule 45, inheriting audit,
filter_resources, kill-switch, and telemetry — no new control path to secure. - Reuses the existing tunnel, device credential, worker
Executorruntime, cells, tenancy quotas, and SPIFFE scaffold — roughly 60–70 % of the work is already done; the build is mostly integration plus a focused set of net-new primitives. - The customer machine is outbound-only and never sees the broker —
closing the largest multi-tenant exposure of the current
device.{id}queue path. - A microVM isolation tier finally contains untrusted/agent-emitted code on customer hardware — the platform's single largest execution-security gap.
- The new lease/heartbeat/reaper + enforced idempotency benefits the whole
platform, not just local execution (today there is no requeue-on-stall and
no
SKIP LOCKED).
Negative
- Net-new surface to own: a node agent, a tunnel work channel, a node registry, a reaper, and isolation runtimes. Mitigated by phasing (walking skeleton → fault tolerance → isolation → multi-tenancy → GA) with hard exit criteria.
- Evolving the WebSocket tunnel rather than adopting gRPC is a deliberate short-term/scale trade-off; revisit at sustained internal req/s.
- Two execution tiers (co-located broker worker vs customer-hosted node) add operator-facing conceptual surface; documented explicitly.
- Running isolation runtimes (KVM for microVMs) imposes a host dependency; mitigated by advertising available tiers at registration and refusing code-gen work on gVisor-only nodes.
Alternatives considered
- Virtual Kubelet / OpenYurt / KubeEdge — rejected. They assume Kubernetes is the scheduling control plane; AlphaSwarm's is not. They would force every local job into a Pod manifest and run an edge-K8s stack on a laptop. Retained only as references; OpenYurt's YurtHub edge-autonomy idea may inform a future disconnected-cell ADR.
- Expose the Celery/Redis broker to the customer machine (extend the
device.{id}queue path off-network) — rejected. Putting a semi-trusted, NAT'd, multi-tenant-adjacent machine on the shared broker is a security and isolation hazard. Tier B replaces broker draining with scoped tunnel dispatch. - gRPC-over-HTTP2 reverse tunnel now — deferred (consistent with ADR 005). The WebSocket tunnel already multiplexes and traverses NAT; gRPC/QUIC are documented future options.
- Run untrusted code in plain containers only — rejected for multi-tenant Tier B. A shared host kernel is not a sufficient boundary for adversarial / LLM-generated code; microVMs (or at least gVisor) are required.
- Meter customer compute (CPU/GB-hr) — rejected. There is no cost basis on hardware we don't own, and every comparable vendor bills orchestration / managed units instead.
Implementation references
- Full design & phased plan:
design/hybrid-local-execution-layer.md - Provider seam:
alphaswarm_core/src/alphaswarm_core/providers/protocol.py,runtime/workload.py; newalphaswarm_controller/src/alphaswarm_controller/providers/local.py - Tunnel:
alphaswarm_controller/src/alphaswarm_controller/api/routers/tunnel.py,services/tunnel_registry.py - Worker runtime + isolation:
alphaswarm_worker/src/alphaswarm_worker/execution/ - Node agent (new):
alphaswarm_local/ - Identity:
alphaswarm_auth/src/alphaswarm_auth/auth/device_service.py; controllerspire_backend/spire_trust_domain - Ledger + reaper:
alphaswarm/persistence/(OrderOutboxRowpattern inmodels_orders.py+trading/outbox_relay.py);alphaswarm/tasks/beat schedule - Tenancy + metering:
alphaswarm_core/src/alphaswarm_core/models/tenancy.py,topology/models.py(cells);alphaswarm_admin/.../accounts/billing.py,api/routers/tenants.py - AGENTS rules: 45 (InfrastructureProvider), 26 (CredentialResolver), 27 (IAM hub), 49 (MCP), 51 (TenancyStrategy), 52 (step-up MFA), 53 (device grant) —
AGENTS.md