Workstreams
Part of the agentic workflows enhancement plan. Gap IDs (G1–G10) reference gaps.md. Effort classes: S ≤ 1 day, M ≤ 1 week, L = multi-week. Every action lands along an existing seam — no rewrites.
Phasing at a glance
| Phase | Window | Contents | Exit criterion |
|---|---|---|---|
| P0 — Stop the bleeding | weeks 1–2 | WS0 entirely; WS7 baseline instrumentation started | Broken/dead CI fixed; every agent-visible lie removed or labeled |
| P1 — Foundation | weeks 2–8 | WS1 (reusable CI, wheels, App token), WS2 (agent-verify, pinning) | make agent-verify valid in every repo; zero-CI repo count = 0 |
| P2 — Canon & enablement | weeks 6–14 | WS3 (canon restructure + guidance CI), WS4 (multi-tool surfaces) | Guidance CI blocking; Claude Code loads real guidance in all repos |
| P3 — Leverage | weeks 12–20 | WS5 (context layer), WS6 (dev automation bots) | First governed PR-review + CI-triage workflows in production use |
| P4 — Operate & expand | ongoing | WS7 experimentation; enforce-flag ratchets; autonomy expansion per metrics.md thresholds | KPI review cadence running; expansion gated on thresholds |
Dependencies: WS2 depends on WS1's wheel publishing; WS4 depends on WS3's
canon decisions; WS6 depends on WS5's codebase.propose_edit and on WS7's
bot-eval harness. WS0 and WS7-baseline have no dependencies — start
immediately.
WS0 — Stop the bleeding (quick wins, all S unless noted)
Every item below removes something that actively lies to agents today (G2, parts of G5/G6/G9).
Status (verified against the working tree, 2026-08): items marked ✅ below are confirmed landed. This spot-check covers only the items called out; the remaining rows have not been individually re-verified here — see handoff.md for the program's own live-state tracking.
| # | Action | Restores |
|---|---|---|
| 0.1 ✅ | Fix duplicate python: job key in alphaswarm/.github/workflows/pr-validate.yml (lines 33/47) | PR-time gitleaks + trivy + alembic-immutability gates |
| 0.2 ✅ | Fix branch triggers: alphaswarm_ide build.yml + license-check-workflow.yml master→main; alphaswarm_eval/ci.yml main→development | IDE lint/build/test gating; eval workflow fires at all |
| 0.3 | Delete dead monolith workflows path-filtered on extracted dirs (alphaswarm-client.yml, e2e-eda.yml, quantbot-bots-image.yml, ml-pipeline.yml, monolith docs-ci.yml); repoint security-scan.yml bandit/npm-audit at paths that exist | Ends false coverage impressions + weekly red noise |
| 0.4 | Add actionlint (pattern already in platform/admin/worker via raven-actions/actionlint@v2) to every CI-bearing repo | Catches class-1 workflow breakage (like 0.1) at PR time |
| 0.5 (M) | Replace the 704-byte .claude.md stub in ~35 repos with a real CLAUDE.md pointer (template in WS4.1) — one scripted sweep PR per repo | Claude Code auto-loads real guidance org-wide |
| 0.6 ✅ | Fix alphaswarm_bots/AGENTS.md validation command (pytest tests/bots → pytest -q); add pytest+ruff job beside its crd-schema-drift.yml | Documented command actually runs the tests |
| 0.7 | Reconcile monolith AGENTS.md owner-repo contradictions, owners win: line 144 → Entra-only (per alphaswarm_ui's CI-enforced contract); rule 47 domains → alpha-swarm.ai; line 140 → eight IDE extensions; fix line 135–137 rename artifacts; delete triplicated "Where to look" rows (~1136–1161) | Canon stops contradicting CI-enforced reality |
| 0.8 | Delete misleading artifacts: alphaswarm_client/package-lock.json (pnpm is canonical), alphaswarm_ui/jest.config.js+jest.setup.js, empty alphaswarm_mcp/alphaswarm_core/ dir, alphaswarm_docs/sidebars.ts.bak, committed rebase output in _research, stray .DS_Store/.idea (kb, kb_federation, catalog, research) | Agents stop keying on dead config |
| 0.9 ✅ (partial) | Flip orchestration_workflow_versioning_enabled for the internal namespace; copy AgentSpec's logged-skip fix into load_workflow_specs_from_dir (alphaswarm_agents/orchestration/spec.py:169-173 currently swallows invalid YAML silently) | Workflow specs safe to depend on (prereq for WS6) — the silent-skip fix is confirmed landed (spec.py now surfaces the file + validation error instead of swallowing it); the versioning-flag flip was not independently re-verified |
| 0.10 ✅ | Fix RosterEvaluator.evaluate_spec (alphaswarm_agents/exemplars/evaluator.py passes str where AgentSpec required — guaranteed AttributeError); add the missing test | Roster eval facade works (prereq for WS7 bot evals) |
| 0.11 | Commit the missing referenced contract docs (EVALUATION_TAXONOMY.md, GOLDEN_DATASET_PLAN.md — cited by alphaswarm_eval README/schema.py/datasets.py, exist nowhere) | Eval taxonomy reviewable |
| 0.12 | Standardize claude/** + harden/** push triggers across the ~23 CI-bearing repos (today only docs/worker/admin) | Agent branches get CI feedback before PR |
| 0.13 | Move ghcr publishing off the personal julianwiley namespace (cd-staging.yml IMAGE_OWNER); fix julianwiley/julianwileymac URLs in controller/local pyprojects and architecture.md | Org-owned supply chain |
Acceptance: all 13 landed; a re-run of the CI-landscape audit finds zero workflows that cannot trigger and zero YAML-invalid workflows.
WS1 — Trustworthy verification (G1, G3, G7; parts of G4)
Goal: "CI is green" becomes a done-signal an agent can trust, in every repo, with one implementation.
Actions
- (L) Create an org-level
.githubrepo with reusable workflows (workflow_call):python-ci(setup → sibling deps via App token or CodeArtifact → ruff/mypy/pytest → license gate),node-ci(pnpm typecheck/lint/test/build), and one canonicalbuild-sign-push(adoptalphaswarm_platform's composite — it blocks on Trivy HIGH/CRITICAL and is SOC2-referenced; delete the divergent monolith twin). Standardizetimeout-minutes,concurrency, SHA-pinned actions, and Renovate presets through these templates (today: timeouts in 12/57 files, SHA pinning only in ide). - (M) Close the zero-CI gap in dependency order, each a one-file PR
calling the reusable workflow:
alphaswarm_configandalphaswarm_corefirst (editable-installed by ~15 pipelines), thenalphaswarm_client(its full pipeline already exists fully-written as the monolith's deadalphaswarm-client.yml— move it),alphaswarm_auth,alphaswarm_assistant(ships inside the prod IDE image),alphaswarm_data,alphaswarm_finops, then the rest of the 17. - (M) Replace the six cross-repo PATs with one org GitHub App
(
actions/create-github-app-token), one secret name org-wide; medium-term eliminate most sibling checkouts by publishingalphaswarm-core/-config/-agents/-catalogwheels to the CodeArtifact index already proven inalphaswarm_admin/build-publish.yml:94-100, consumed via the existing OIDCcodeartifact-logincomposite. - (M) Kill green-by-skip: calendar deadlines (tracked in
metrics.md reviews) to flip
LICENSE_GATE_ENFORCE,EVAL_GATE_ENFORCE,ALPHASWARM_WORKER_CI_FULL,RUN_CONTRACT_RESOLUTION; convert dated quarantine lists (controller 8 files, monolith 7 ignores, ~287 environmental failures) into ticketed burn-downs with a shrinking count enforced in CI; removecontinue-on-errorfromalphaswarm_uitypecheck/lint/vitest andignoreBuildErrorsfrom itsnext.config.mjsbehind a dated ratchet. - (M) Make supply-chain enforcement uniform on the paths that actually
deploy: remove
continue-on-errorfrom cosign/syft inalphaswarm_admin/build-publish.yml, restore its Trivy scan; add signing+SBOM toalphaswarm_uidocker builds and bothalphaswarm_ideimage pipelines; then enforce at admission with the Kyverno identity the workflows already reference. - (M) Merge gating as code:
merge_grouptriggers on deploy-adjacent repos (ui, admin, ide, platform, docs, website); repo rulesets managed via the platform repo's existing Terraform surface; Renovate + Dependabot(actions) org-wide from a shared preset with digest pinning (config exists today only in monolith + ui).
Worked artifact — repo-side caller (complete file)
# .github/workflows/ci.yml — any Python repo
name: ci
on:
pull_request:
merge_group:
push: { branches: [main, development, 'claude/**', 'harden/**'] }
jobs:
ci:
uses: Alpha-Swarm-ai/.github/.github/workflows/python-ci.yml@v1
with: { python-version: '3.12' }
secrets: inherit # one App token, not six PATs
Acceptance: zero repos without CI; zero continue-on-error on
verification steps without a dated ratchet ticket; all four enforce flags
flipped or formally descoped; one build-sign-push implementation org-wide;
PAT count for sibling access = 0.
WS2 — Uniform bootstrap/verify contract (G4, G10)
Goal: any agent, any repo, fresh sandbox: make agent-bootstrap && make agent-verify — same words everywhere, and CI runs exactly the same
target.
Actions
- (L) Ship the four-target contract to all repos:
agent-bootstrap(idempotent, no prompts; auto-detect../alphaswarm_core, fall back to.siblings/clone — mirroring what agents/controller CI already does),agent-lint,agent-test,agent-verify(= bootstrap+lint+typecheck+ test). Seed each from the repo's existing AGENTS.md Validation block; rewrite every Validation section to say onlymake agent-verify. - (M) Toolchain pinning wave,
alphaswarm_qapas the template:.python-version+ committeduv.lockin every Python repo (decide the 3.11 vs 3.12 floor deliberately — 27 repos say >=3.11, monolith+config require >=3.12);packageManager+.nvmrcon one Node LTS across the 8 JS repos; org-wide exact ruff pin (extend worker's==0.15.17determinism rationale); converge ruff line-length 100 and tiered mypy strictness (strict for libraries like core; gradual with explicit marker). - (M)
repos.yamlmachine registry inalphaswarm_index(respecting the curator process): per repo — language, build tool, sibling deps, bootstrap/verify commands, CI status, owner. Generatealphaswarm-monorepo-paths.mdand the monolith AGENTS.md sibling table from it (both are wrong today: the doc names repos that don't exist and omits ~15 real ones), with a CI freshness check. - (S) One shared
devcontainer.jsontemplate (py3.12 + node LTS + uv + pnpm);.env.examplerequired in every service repo; portalphaswarm_admin's PowerShell-only dev script to bash. - (S) Converge naming: boundary scripts →
scripts/ci/check_boundaries.pyalias everywhere; dev extra nameddev(agents currentlytest); workflow filenameci.yml.
Worked artifact — Makefile skeleton
# Standard agent contract — seeded from AGENTS.md "Validation".
# CI calls exactly these targets; humans and agents run the same thing.
# (Recipe lines below are shown space-indented for the docs linter;
# real Makefile recipes must be tab-indented.)
.PHONY: agent-bootstrap agent-lint agent-test agent-verify
agent-bootstrap:
@python scripts/bootstrap_siblings.py # ../repo → .siblings/ fallback
uv sync --frozen || pip install -e .[dev]
agent-lint:
ruff check . && ruff format --check .
mypy src || true # TODO(#ticket): dated ratchet to blocking
agent-test:
pytest -q
agent-verify: agent-bootstrap agent-lint agent-test
Acceptance: make agent-verify exists and passes (or fails honestly) in
all repos; repos.yaml exists with CI freshness check; lockfile coverage
= 100% of Python repos; an agent transcript sampled per repo shows zero
turns spent discovering build commands.
WS3 — Guidance canon restructure (G5, G6, G8)
Goal: the canon becomes small, true, mechanically checked, and single-sourced; numeric rule citations are retired before the next renumber bites.
Actions
-
(L) Split the monolith
AGENTS.mdto a <400-line core: hard rules as one-liners with links; Don't-list; When-in-doubt. Move the Project map, Where-to-look table, Quick reference, and Cursor Cloud instructions into the existing 29.cursor/skills/(they already cover most of it — the monolithic file is now the redundant copy). Execute the path cutover (rules 12/13/22/40/41 and Quick reference →../alphaswarm_agents/...real targets) instead of the current "resolve mentally" patch; remove machine-specific paths in favor of a$WORKSPACE_ROOTconvention; separate rule text from rollout narrative (rules 27/42/47/53/64 embed migration status — move to dated sidecar notes in.cursor/plans/). -
(L) Slug-keyed shared-rule registry in
alphaswarm_index(curator-owned SSoT): every cross-repo rule gets a stable slug; replace the 27 pasted deployment-version boilerplate copies and all "AGENTS rule 22/26/27/44/49" numeric citations with slug references. Example:# alphaswarm_index/rules/registry.yaml (excerpt)
RULE-LLM-GATEWAY:
statement: All LLM calls go through router_complete; no vendor SDKs.
enforcement: alphaswarm/scripts/ci/check_llm_gateway.py
owner: platform-team
RULE-NO-MONOLITH-IMPORT:
statement: Satellite repos never import alphaswarm.*.
enforcement: scripts/ci/check_boundaries.py (per repo)
RULE-DEPLOY-VERSION:
statement: Deployment version bumps follow the governed control path.
enforcement: ADR 022 One Gate -
(M) Mechanical guidance CI (extends monolith
scripts/ci/, converting the manualrules-drift-sentinelsubagent into an enforced gate): (a) link-checker with../<repo>sibling resolution over AGENTS.md/WORKFLOW.md/.cursor/rules/*; (b) hard-rule counter failing when prose counts disagree with the parsed list; (c).mdcfrontmatter validator (catches the 8 headless Cursor rules that never load and 2 missing referenced files); (d) lychee promoted from advisory to blocking inalphaswarm_docsdocs-ci.yml; (e)last_reviewedstaleness warning (>60 days) foraudience: bothconcept docs. -
(M) Resolve the two-world-model fork explicitly (G6): the
.cursorrules/.clinerules/llms.txtboilerplate teaches QAP/ArcticDB/three-planes, contradicting hard rules 3/46 (Iceberg-only). Decide per doctrine element: what is real (parts of the promotion gate demonstrably are) gets promoted into AGENTS.md as rules; what is aspirational moves to a dated design note in this repo. Then regenerate every per-tool file from AGENTS.md as thin pointers —.cursorrules,.clinerules,.junie,.devin,copilot-instructions.md(which should also list the repo's required status checks so agents know what blocks merge). -
(M) Docs suite reconciliation (G8): ADR 025 banners on
workflow-studio.md/orchestration-refactor-rollout.md/multi-agent-patterns.md(Celery-beat → compatibility-mode); one canonicalWorkflowSpec/AgentSpecsurface reference (today three contradictory schemas, two guardrail vocabularies, three run endpoints); re-baseline repo topology againstrepos.yaml(WS2.3); AGENTS.md for the 11 repos missing one,alphaswarm_configandalphaswarm_evalfirst; add nested AGENTS.md where subprojects diverge (now standard: most-specific-wins). -
(S) Codify the evidence bundle in WORKFLOW.md's Reflect phase for SLOW-mode changes — one format (what ran, run-ledger row links, verification commands), building on
ui-run-report-verification.mdcandadversarial-run-report-reviewer; replaces the ad-hocPHASE_*_COMPLETION.mdfiles littering the monolith root. This is the reports' "artifact-forward progress" rule, implemented on existing seams.
Acceptance: monolith AGENTS.md <400 lines with zero dead links (CI- enforced); zero numeric rule citations remain org-wide (grep gate); guidance CI blocking in monolith + docs; per-tool files are generated, not hand-drifted; every repo has an AGENTS.md; WORKFLOW.md defines the evidence bundle.
WS4 — Multi-tool enablement on open standards (G6; landscape F3/F4/F7/F8/F9)
Goal: end the Cursor monoculture without discarding the Cursor investment: every tool loads the same truth; parallel sessions have an org convention; enterprise governance uses managed scopes.
Actions
-
(S — part of WS0.5)
CLAUDE.mdat every repo root:# CLAUDE.md
Read AGENTS.md first — it is the contract for this repo.
Verify every change with: `make agent-verify`
Org context: alphaswarm_mcp serves rules/skills/contracts (see .mcp.json).
Cross-session state: .agents/state-template.md (only if work spans sessions). -
(M) Skills to the vendor-neutral location: mirror the monolith's 30
.cursor/skills/into.agents/skills/(the tool-neutral directory; Cursor also reads.claude/skills/and.codex/skills/, so one canonical copy + generated pointers covers Claude Code, Cursor, and Codex). Front-load trigger conditions in descriptions per the Agent Skills spec; add the six skill families the reports recommend where missing (platform, service-domain, change-management, security, operations, documentation) — most content already exists in skills/rules and just needs repackaging. -
(M) Worktree convention for parallel agent sessions (both tools now first-class): document in WORKFLOW.md — Claude Code
--worktree/isolation: worktreesubagents, Cursor/worktree+/best-of-n;.worktreeincludefor env propagation; naming (worktree-<task>); thebest-of-n-runnersubagent WORKFLOW.md already mentions becomes the documented pattern rather than folklore. Note the open Claude Code bug (isolation: worktreeignored viaclaude --agent, #50357) with the workaround (invoke via subagent frontmatter, not--agent). -
(M) Org Claude Code plugin: package the canon (rules pointers, skills, hooks,
.mcp.jsonwiring toalphaswarm_mcp/Data MCP/Codebase MCP, subagent definitions ported from.cursor/agents/) as a plugin in a private marketplace repo; enterprise-manage viamanagedscope +strictKnownMarketplacesso every engineer and CI sandbox gets identical, admin-controlled context. Project-scope plugin pins in.claude/settings.jsonwhere repos need extras. -
(M) Edit-time enforcement: wire the existing boundary lints (
scripts/ci/check_*) into pre-commit (only the monolith has any today) and Claude Code hooks (PostToolUse on Edit/Write) so agents get violation feedback at edit time, not CI time. Port the highest-value Cursor subagents (alphaswarm-hard-rules-reviewer,adversarial-run-report-reviewer,rules-drift-sentinel) to tool-neutral agent definitions consumable by Claude Code subagents. -
(S) Session-start hook for cloud/web sessions (repo bootstrap +
make agent-bootstrap) so remote sandboxes start correctly without human setup.
Acceptance: Claude Code in any repo auto-loads CLAUDE.md → AGENTS.md and can invoke org skills; one skill source of truth with generated mirrors; worktree section merged into WORKFLOW.md; plugin installed via managed settings in the standard engineer + CI images; boundary violations surface at edit time in both Cursor and Claude Code.
WS5 — Context layer activation (G8; dive findings)
Goal: the already-built context infrastructure starts earning its keep for coding agents: repo maps, symbols, ownership, and org knowledge in one bounded call.
Actions
- (M) Wake the dormant semantic code index: add a merge-to-main
task (pattern:
alphaswarm/alphaswarm/tasks/graph_seed_tasks.py) runningindex_workspace → chunk_symbols → upsert_chunksinto thecode_chunkspgvector corpus, and add the realsemanticbranch toCodebaseSearchTool.run(alphaswarm/alphaswarm/codebase/mcp/tools/search.py) querying HierarchicalRAG — converts existing dead code into the highest-leverage agent capability at near-zero design cost. - (M) Persisted symbol store for the Codebase MCP (mtime-keyed
SQLite or existing Postgres) so
find_definition/find_references/get_repo_graphstop re-walking the workspace per call; extend index scopes to sibling repos (the package's own AGENTS.md already demands it). - (S) CI-generate the mechanical index artifacts:
code-index/modules.md+symbols.mdare pure signature extraction — regenerate via CI job reusingast_index.py, reserving the curator for judgment work. Retires the self-documented index-debt while preserving the sole-writer invariant (CI acts as the curator's tool). - (M) Ownership surface: generate an ownership map from
.github/CODEOWNERS+ docs frontmatterowner:intoalphaswarm_index/architecture/ownership.md; expose as acodebase.ownersMCP tool so "who owns / who reviews X" is a one-hop query (today: three unreconciled ownership systems). - (M) Org-engineering KB corpus in
alphaswarm_kb(GLOBAL scope, like the eight platform reference corpora): ADRs, docs concepts, AGENTS.md files, OpenAPI specs, the curated index — permissioned remember/recall with bi-temporal provenance, federable viaalphaswarm_kb_federation. - (L)
context.packMCP tool: one bounded call composing the governance search order (nearest AGENTS.md → index boundary/budget rows → codebase symbol slices → KB recall) so agents don't need tribal knowledge of five surfaces. - (S) MCP version posture: the
2026-07-28stateless revision is backward compatible — no forced migration for Data/Codebase/Platform MCP servers. Plan a deliberate v2-SDK adoption window post-GA; keep RFC 9728/8707 conformance (hard rule 49) as the constant.
Acceptance: semantic search answers real queries in the monolith;
codebase.* p50 latency drops (no per-call re-walk); index-debt entries for
modules/symbols closed; codebase.owners resolves every path in the
monolith + top-6 satellites; context.pack used by the WS6 bots.
WS6 — Dogfooded dev automation (dive: platform runtime; reports' maker-checker)
Goal: dev automation runs on the platform's own audited runtime —
hash-locked specs, budgets, kill-switch, decision logs — not on a parallel
unaudited stack. This is also the productized answer to the reports'
"maker-checker" and "coordination pattern registry" recommendations: the
registry is multi-agent-patterns.md + the seven orchestration
adapters, governed by workflow_spec_versions.
Actions
- (M)
pr_review.dialectical_v1WorkflowSpec: port.cursor/agents/alphaswarm-hard-rules-reviewer.mdandadversarial-run-report-reviewer.mdintoconfigs/agents/*.yamlAgentSpecs (modeled oncodebase_refactorer.yaml) bound tocodebase.*tools; compose withDialecticalDebateAdapter(Critic = hard-rules reviewer; Evaluator = deterministic verdict) under WorkflowRuntime. Every review run gets hash-locked replay, cost caps, and kill-switch halt for free. - (M)
codebase.propose_edit— the single sanctioned mutating Codebase MCP tool, gated exactly like the money plane: reuse theApprovedPromotionfail-closed capability-token pattern (alphaswarm_agents/src/alphaswarm_agents/promotion_gate.py) so bot-generated patches always pass a HITL gate. This is the one missing primitive blocking end-to-end review/auto-fix bots, and the governance pattern already exists. - (M) CI-triage workflow on
AutomationScheduleAdapter+ the decision-log loop: a cron WorkflowSpec ingesting failing-run reports, writing verdicts viaappend_pending_decision(alphaswarm_agents/.../graph/decision_log.py), withresolve_pending_decisions+ reflection learning flaky-vs-real patterns into episodic memory — the outcome-resolution machinery built for trades applies verbatim to test failures. - (S) SAF as the audit-before-build gate:
graph/saf.pyalready does discover → spec → duplicate-audit → human gate → register; wire it as the WorkflowSpec that runs when a new module/package is proposed — the structural version of the anti-duplication prose in every AGENTS.md. - (M) Doc-sync agent: pair the
alphaswarm_mcpbundle builder with a scheduled workflow diffing bundle content againstcodebase.get_repo_graph/find_references, filing evidence-ref'd staleness findings throughFindingRecorder(extendEVIDENCE_REF_PREFIXESwithfile:/commit:) — realizes thedocs-reliability-orchestratorsubagent on the real runtime, feeding WS3's guidance CI. - (M) Regression-test the bots with the platform's own harness:
golden PR-review and triage traces through
alphaswarm_agents/evaluation.py(replay + judge intoagent_evaluations) in CI, so bot-quality regressions gate merges like unit tests (depends on WS7's judge hardening). - (S) Governance seat:
internal/devopsroster category + annotation for dev bots peralphaswarm_agents/AGENTS.mdcardinal rule 7; dev bots appear in the same promotion/kill-switch surfaces as product agents. For CI-side pipelines that must run outside the monolith cluster, compile fromalphaswarm_orchestrationcontracts to Prefect/Dagster (dependency-light, secret-rejecting, engine-portable).
Rollout follows the reports' four-stage control model, mapped to our
gates: dark launch (bot comments only) → shadow (bot verdicts compared to
human review outcomes, measured in WS7) → limited write
(codebase.propose_edit behind ApprovedPromotion) → production-adjacent
(auto-fix PRs, still human-merged). Expansion between stages is gated on
metrics.md thresholds.
Acceptance: pr_review.dialectical_v1 running on ≥3 repos in shadow
mode with measured agreement rates; CI-triage verdicts on 100% of red runs
in monolith + worker; zero dev-automation code paths outside
WorkflowRuntime governance; bot regression suite blocking in
alphaswarm_agents CI.
WS7 — Measurement and experimentation
Fully specified in metrics.md. Summary: fix the inert eval
gate (real producer, enforce flag, one eval engine); instrument
agent-authored PRs (CI-green-on-first-push, review iterations,
time-to-merge, revert-within-N-days) as spans through
alphaswarm_core.observe; weekly dashboard; SRM-checked, CUPED-adjusted
experimentation for workflow changes; KPI thresholds gating every autonomy
expansion in WS4/WS6. Baseline instrumentation starts in P0 — the
program must not run for a quarter before we can measure it.
WS-G — Governance and compliance sidebar
- EU AI Act: the reports assert applicability from 2026-08-02 — this did not survive verification and is 14 days out if true. Action (S, immediate): route to counsel for a determination of which obligations (if any) bind AlphaSwarm's internal coding-agent use vs. the product platform; do not encode compliance machinery into CI until scoped.
- Data classification for agent surfaces (reports' recommendation,
adopted): prompts/telemetry from dev-agent sessions follow the existing
tenancy/redaction posture (
alphaswarm_observeingest-time redaction); restricted-class content (credentials, customer data) never enters dev-bot context — enforced via the existingDataMCPToolscopes rather than new machinery. - Red-team automation: the ADLC manifesto's "recommended, not enforced"
red-team review for tool-gaining AgentSpecs becomes a WS6-pattern
WorkflowSpec (adversarial battery before promotion), closing the loop the
docs already sketch. Extend the ADLC manifesto with Layers 9 (typed
promotion gate, from
qap-agent-layer.md) and 10 (CapabilityManifest advisor gate, ADR 020) per the docs dive. - Naming hygiene (reports' warning, confirmed by the A2A/ACP landscape):
internal artifacts never use bare "ACP"/"MCP" for internal concepts; the
spec-pattern vocabulary (
AgentSpec,WorkflowSpec) already avoids this — keep it that way in new surfaces.