Skip to main content

IaC runbook

"I want to provision X" recipes for the Terraform IaC control plane.

Quick reference​

TaskRecipe
Stand up local AlphaSwarm on a laptopLocal environment
Stand up AlphaSwarm on rpi_kubernetesrpi Kubernetes environment
Stand up paper-trading on GCPPaper environment
Stand up production on AWSLive environment
Stand up the seeded Wiley Tech home on AzureWiley Tech environment
Add a new module kind to the codegenAdd a module kind
Add a Terraform stack via the APICreate a stack via API
Plan / apply / destroy from the UILifecycle from the frontend
Configure HCP Terraform as state backendHCP Terraform
Wire OPA policy enforcementPolicy enforcement

Local environment​

cd alphaswarm_platform/terraform/environments/local
terraform init
terraform plan
terraform apply

What this provisions (per terraform/environments/local/main.tf):

  • A k3d-in-Docker cluster (via the local_cluster module, using the kreuzwerker/docker provider) + a local image registry on :5001 (registry_port var, default 5001).
  • A single Kubernetes namespace, alphaswarm-local.
  • Workloads for the enabled_services list (default: alphaswarm-core, alphaswarm-worker, alphaswarm-beat, alphaswarm-client, postgres, redis, neo4j, chromadb, mlflow, otel-collector, jaeger) deployed as Kubernetes workloads via the alphaswarm_workloads module — not raw docker_container resources, and no MinIO service in this list.
  • Traefik as the ingress class (k3d's built-in ingress controller).

The shared multi-namespace map (local / paper / live / backtest / system / terraform / bots / agents), the cert-manager / ESO / KEDA / ingress-nginx / kube-prometheus / otel-operator / istio Helm baseline, and the per-queue KEDA ScaledObjects (including the terraform queue) live in the root terraform/main.tf composition, not terraform/environments/local/ — they apply to the cloud environments, not the local target.

State is local (alphaswarm_platform/data/terraform/state/local.tfstate, per terraform/environments/local/backend.tf).

The file header also documents a CLI-first path that routes through TerraformRuntime (so applies land in the terraform_runs ledger and respect the global kill switch): alphaswarm deploy up / plan / build / down.

rpi Kubernetes environment​

alphaswarm-cli deploy publish-rpi --registry ghcr.io/<org> --tag <immutable-tag>
terraform -chdir=alphaswarm_platform/terraform/environments/rpi init
terraform -chdir=alphaswarm_platform/terraform/environments/rpi plan
terraform -chdir=alphaswarm_platform/terraform/environments/rpi apply

Recommended bootstrap sequence for first-time bring-up:

  1. CLI-first Terraform apply until base services are healthy.
  2. Verify API + Celery + Redis + Postgres are reachable.
  3. Move to control-plane actions (/control-plane/kubernetes/targets/rpi/*).

This avoids enqueue/stream confusion during cold start when broker/DB are still bootstrapping.

Provider mirror + init retries​

When provider downloads are unstable, define a Terraform CLI config file with provider_installation mirror rules and point AlphaSwarm at it:

export ALPHASWARM_TERRAFORM_CLI_CONFIG_FILE=/absolute/path/to/terraform.tfrc
export ALPHASWARM_TERRAFORM_INIT_RETRY_ATTEMPTS=5
export ALPHASWARM_TERRAFORM_INIT_RETRY_BACKOFF_SECONDS=2
export ALPHASWARM_TERRAFORM_INIT_RETRY_MAX_BACKOFF_SECONDS=30

TerraformExecutor applies bounded retries for transient terraform init failures and reuses ALPHASWARM_TERRAFORM_PLUGIN_CACHE_DIR between runs.

Paper environment​

cd alphaswarm_platform/terraform/environments/paper
export TF_VAR_gcp_project_id=<your-gcp-project>
export TF_VAR_primary_domain=paper.alphaswarm.example
terraform init -backend-config="bucket=alphaswarm-terraform-state-paper"
terraform plan
terraform apply

What this provisions:

  • GKE cluster (auto-promoted from ALPHASWARM_DEFAULT_CLOUD_PROVIDER=gcp).
  • Cloud SQL Postgres (single AZ — cost-optimised for paper).
  • GCS bucket + Memorystore Redis.
  • GCP Secret Manager ClusterSecretStore (ESO).
  • Bot Deployments with dry_run=true for paper trading.
  • 100% traffic to the Vite frontend (no canary split in paper).

Live environment​

cd alphaswarm_platform/terraform/environments/live
export TF_VAR_aws_subnet_ids='["subnet-aaaa", "subnet-bbbb", "subnet-cccc"]'
export TF_VAR_primary_domain=app.wiley.tech
terraform init # picks up backend.tf with S3 + DynamoDB locking
terraform plan
terraform apply

What this provisions:

  • EKS cluster Multi-AZ.
  • RDS Multi-AZ Postgres + S3 versioning + ElastiCache 7+ cluster mode.
  • AWS Secrets Manager ClusterSecretStore.
  • Bot Deployments live (dry_run=false); live_control=true on the actor's Membership is required to trigger orders.
  • Per-queue KEDA maxReplicaCount from terraform/modules/faas (same fixed values across environments — not live-specific sizing): 20 default / 50 ML / 100 backtest / 20 agents / 10 terraform.

Wiley Tech environment​

This is the seeded production home for the org provisioned by Alembic 0051. Pinned to the Wiley Tech Entra tenant.

cd alphaswarm_platform/terraform/environments/wiley-tech
export TF_VAR_azure_tenant_id=<wiley tenant id>
export TF_VAR_azure_subscription_id=<sub id>
export TF_VAR_azure_resource_group=alphaswarm-wiley-tech
export TF_VAR_azure_keyvault_url=https://alphaswarm-wiley-tech-kv.vault.azure.net/
terraform init # picks up backend.tf with Azure Blob state
terraform plan
terraform apply

What this provisions:

  • AKS cluster + Azure Workload Identity for ESO.
  • Azure PostgreSQL Flexible Server (single-zone B_Standard_B1ms burstable tier for non-live environments like wiley-tech; no zone-redundant HA is configured in terraform/modules/storage today — that upgrade only applies to the live environment's SKU).
  • ADLS Gen2 storage account (HNS enabled).
  • Azure Cache for Redis (Basic SKU for non-live environments — wiley-tech gets Basic, only live gets Standard — always TLS-only / no non-SSL port).
  • Azure Key Vault ClusterSecretStore synced via ESO Workload Identity.
  • ACR registry for AlphaSwarm images.

Add a module kind​

  1. Add the kind to TERRAFORM_MODULE_KINDS in alphaswarm/persistence/models_terraform.py.
  2. Create the Jinja2 template at alphaswarm/terraform/codegen/templates/<kind>_<cloud>.tf.j2 (and a _local fallback).
  3. (Optional) Mirror as a native HCL module under alphaswarm_platform/terraform/modules/<kind>/.
  4. Operators create a stack via POST /terraform/stacks with module_kind: "<kind>".

Create a stack via API​

curl -X POST http://localhost:8000/terraform/stacks \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <token>" \
-d '{
"name": "Bronze tier storage",
"slug": "bronze-storage",
"module_kind": "storage",
"cloud_provider": "aws",
"environment": "live",
"variables": {
"aws_region": "us-east-1",
"aws_subnet_ids": ["subnet-aaa", "subnet-bbb", "subnet-ccc"],
"bucket_name": "alphaswarm-bronze",
"db_storage_gb": 500
},
"backend": { "kind": "s3", "config": { "bucket": "alphaswarm-tf-state", "key": "bronze-storage.tfstate" } },
"tags": { "tier": "bronze" }
}'

Response includes spec_version_id (immutable, hash-locked).

Then create a workspace + plan:

# Workspace
curl -X POST http://localhost:8000/terraform/workspaces \
-H "Content-Type: application/json" -H "Authorization: Bearer <token>" \
-d '{ "slug": "bronze-live", "name": "Bronze (live)", "stack_spec_id": "<id>", "environment": "live", "state_backend": "s3" }'

# Plan
curl -X POST http://localhost:8000/terraform/workspaces/<workspace_id>/plan \
-H "Authorization: Bearer <token>"

Subscribe to live progress at wss://<host>/terraform/ws/runs/<run_id>.

Lifecycle from the frontend​

The routes exist (/infra/terraform, /infra/terraform/stacks, /infra/terraform/workspaces/[id], /infra/terraform/runs/[id]), but as of this writing every one of them is a scaffolded "Coming soon" stub in alphaswarm_client — none render workspace/run data yet. Use the API directly (see Create a stack via API above) or the CLI until the UI lands. The intended lifecycle, once implemented, is:

  1. Click Plan → enqueues plan task; result lands in awaiting_approval (this is a real TerraformRun status — see alphaswarm/persistence/models_terraform.py).
  2. Review the plan summary on the run detail page (live WS stream).
  3. Click Apply this plan on the plan run row.
  4. Apply executes → state version snapshotted → outputs visible in the workspace's latest state outputs.
  5. Destroy is intended to be friction-gated (e.g. typing the workspace slug to confirm) once implemented.

HCP Terraform​

  1. Create an HCP Terraform organization + workspaces in the HCP UI.
  2. Set ALPHASWARM_HCP_TOKEN (preferred: via CredentialResolver), ALPHASWARM_HCP_ORGANIZATION, ALPHASWARM_TERRAFORM_STATE_BACKEND=hcp.
  3. Set the stack spec's backend.kind="hcp" and the workspace's hcp_workspace_id.
  4. The runtime now drives runs through HcpClient instead of the local subprocess (no terraform binary required on the runner pod).

Policy enforcement​

  1. Author OPA Rego policies that target Terraform plan JSON (the runtime emits tfplan.binary.json via terraform show -json).
  2. Insert a TerraformPolicyAttachment row binding the policy file URI to a workspace.
  3. Set hard_mandatory=True to block apply on violation; hard_mandatory=False emits a warning.
  4. When opa is on PATH the runtime invokes opa eval --format json --input tfplan.json --data policy.rego "data.terraform.alphaswarm.deny". Without OPA installed, a soft_mandatory attachment no-ops cleanly (passes, marked skipped); a hard_mandatory attachment fails closed and blocks the apply instead.