Skip to main content

ADR 011 — CDN-fronted standalone container for the cloud-hosted alphaswarm_ui

Context​

When the original alphaswarm_client/ packaging was designed (ADR 002), the platform had three coexisting presentation surfaces: a Vite operator UI, a legacy Next.js webui, and a Python Solara visualisation layer. Collapsing all three behind one FastAPI proxy was the right call for a single-tenant local-first deployment where operators bookmark one URL and the proxy hides the rest.

The cloud-hosted, customer-facing PaaS at alpha-swarm.ai / app.alpha-swarm.ai (the new alphaswarm_ui/ Next.js 14+ App Router app) has different constraints:

  1. Multi-tenant scale. Hundreds-to-thousands of concurrent tenants. Static-asset throughput and SSR throughput scale at different ratios — co-located scaling triggers wasted CPU and unnecessary memory pressure on the SSR pods.
  2. CDN-friendly assets. Next.js standalone emits hashed, immutable filenames under /_next/static/*. Serving them from the SSR pods is bandwidth waste; Cloudflare can cache them for a year with zero risk of staleness.
  3. No Python / no Solara. alphaswarm_ui/ is pure TypeScript + Next.js server. The Solara stage (ADR 002 Stage 2) doesn't apply and would only bloat the image (~300 MB heavier).
  4. Independent BFF lifecycle. Every alphaswarm_ui/api/* route is a thin BFF handler that re-checks the session, forwards a tenancy header, and proxies upstream. Reverse-proxying through FastAPI adds an extra hop with no value (the BFF is already a proxy).
  5. Edge-rendered marketing. The (marketing) route group is designed for SSR + ISR cache. Routing it through an internal FastAPI proxy defeats the whole point of edge-near rendering.

Decision​

The cloud-hosted alphaswarm_ui ships as one clean Next.js standalone container built from alphaswarm_platform/build/docker/alphaswarm_ui/Dockerfile (two stages: a node:20-bookworm builder + a node:20-bookworm-slim runtime running node server.js; the Dockerfile's header comments describe a Chainguard Wolfi migration, but the current FROM lines are still the Debian-based node:20-bookworm* images). It DOES NOT use the ADR 002 three-stage Python/ASGI pattern.

Edge caching layout:

PathCache-ControlNotes
/_next/static/*public, max-age=31536000, immutableHashed filenames; year-long TTL
/public/* /fonts/* /images/*public, max-age=259200030-day TTL, hand-curated assets
/api/*no-store + Pragma: no-cacheBFF responses; user-scoped (rule 4 + management-engine.mdc)
Everything else (SSR)public, max-age=3600, stale-while-revalidate=86400Per-tenant marketing + dashboard pages

The NGINX Ingress at alphaswarm_platform/deployments/kubernetes/base/alphaswarm-ui/ingress.yaml sets these via nginx.ingress.kubernetes.io/configuration-snippet. Cloudflare in front honours them aggressively for /_next/static/* and bypasses the cache for /api/*.

Post-deploy cache purge: intended to run as a GitHub Actions deploy job that calls the Cloudflare zone-purge API immediately after kubectl rollout status succeeds, with the Cloudflare token sourced from the CredentialResolver chain (AGENTS rule 26). As of this review, the alphaswarm_ui repo's only deploy workflow is alphaswarm_ui/.github/workflows/build-deploy.yml (there is no .github/workflows/alphaswarm-ui.yml), it deploys to a self-hosted local k3d cluster ("Alpha production deploys run on local hardware"), and it contains no Cloudflare purge step or ALPHASWARM_CLOUDFLARE_API_TOKEN reference — the automated cache-purge step described here is not yet implemented.

HPA: keep the existing hpa.yaml (CPU 70%, memory 80%, 3-20 replicas). Because static assets are CDN-offloaded, SSR pod CPU usage tracks real per-tenant rendering work — autoscaling becomes meaningful instead of a noisy mix of "serving a JS bundle" and "rendering a dashboard page".

Consequences​

Positive

  • 80%+ static-asset bandwidth offloaded to Cloudflare's edge.
  • HPA triggers on real SSR work, not bandwidth.
  • Image is smaller than the ADR 002 combined Python + Solara + Node image (the current node:20-bookworm-slim runtime is not Alpine-based, so the exact size delta differs from the original Alpine-based estimate below, but the runtime still carries no Python/Solara). Faster pod cold start, faster rolling deploys.
  • The BFF + SSR + edge layers have one ownership boundary each — Cloudflare for delivery, NGINX Ingress for cache hints, node server.js for SSR + BFF. No ASGI proxy hop in between.
  • /api/* is no-store end-to-end — no risk of a CDN edge node caching a tenant's response and serving it to a different tenant.

Negative

  • Two presentation packaging stories now exist (ADR 002 for alphaswarm_client, ADR 011 for alphaswarm_ui). Mitigated by the per-surface scoping: each ADR is the source of truth for one tree only.
  • Cloudflare cache-purge is now part of the deploy critical path. A Cloudflare API outage during deploy means stale /_next/static/* for up to 1y per hashed filename — but the hashes change on every deploy, so the impact is bounded to assets whose names didn't change (rare for a real change).
  • Adds a CLOUDFLARE_API_TOKEN secret to the deploy environment. Stored in Vault + synced via ExternalSecret per AGENTS rule 26.

Alternatives considered​

  • Stay on ADR 002 (single FastAPI proxy container) — rejected. Bandwidth-CPU coupling, larger image, unnecessary Solara/Python weight, redundant proxy hop in front of the BFF.
  • Vercel hosting — rejected. ADR 003's zero-trust constraints
    • the on-cluster control plane integration argue for keeping the SSR layer inside our own K8s + CredentialResolver perimeter.
  • CloudFront in front of a single SSR pod — rejected. We already have Cloudflare as the edge for alpha-swarm.ai. Adding a second CDN would split the cache-purge story and add edge cost.

Implementation references​