Measurement and experimentation program (WS7)
Part of the agentic workflows enhancement plan. Closes gaps.md G9. Principle (from the internal research reports, adopted): run the agentic-dev program as an experimentation system, not a sequence of anecdotes — and never expand autonomy past a stage whose KPIs aren't green.
1. Fix the inert eval loop first
The gate must mean something before anything is gated on it:
- (S)
alphaswarm_eval/ci.ymltriggermain→development(WS0.2 — done in P0). - (M) Add a producer job: nightly, budget-capped runs of the real
NL→query / NL→dashboard models over the goldens — replacing the current
fixture whose outputs are byte-equal to the goldens (per the repo's own
PROVENANCE.md, it can never fail). Commit scores as the rolling baseline. - (S) Flip
EVAL_GATE_ENFORCE=trueafter two weeks of stable baselines (calendar-tracked in the §5 review). - (M) One eval engine:
alphaswarm_mlopsdepends on thealphaswarm-evalpackage; delete the divergedalphaswarm_mlops/packages/alphaswarm_evalcopy. Portalphaswarm_agents/evaluation.py's replay+judge harness ontoalphaswarm_eval's Scorer/EvalRunner contract so agent evals emit the same EVAL spans and feed the same baselines. - (M) Harden the LLM judge before trusting it: version the judge
prompt as a registered scorer (
judge@<version>), pin the judge model, add a calibration golden set (known-good/known-bad pairs with expected score bands) run in CI, and make judge unavailability loud (fail the case or alert), replacing the silenttry/except: passblocks. - (M) De-brittle goldens: execution-based SQL equivalence against a
fixture DB (not string match); expand
nl_to_querybeyond 22 cases; trajectory + safety scorers over theagent_run_stepstrace shape thattrace_from_run_stepsalready parses.
2. Baseline instrumentation of agent-authored PRs (starts in P0)
Agent-authored PRs are already identifiable (claude/*, codex/*,
harden/* branch prefixes; Codex automerge committer). Emit per-PR events
as spans through alphaswarm_core.observe (same trace store as everything
else; W3C traceparent; ingest-time redaction applies):
| Event | Attributes (minimum) |
|---|---|
dev.pr.opened | repo, pr_number, actor_type (human/claude/codex/cursor), branch, task_ref |
dev.pr.ci_first_result | green_on_first_push (bool), failing_checks[] |
dev.pr.review_iteration | iteration_n, reviewer_type (human/bot) |
dev.pr.merged | time_to_merge_s, review_iterations, additions, deletions |
dev.pr.reverted | days_since_merge, revert_pr, linked_incident |
dev.workflow.bot_run | workflow_spec_version, verdict, cost_usd, human_agreement (shadow mode) |
Stable IDs (task_ref, repo, commit_sha, spec_version_id where a WS6
bot is involved) tie dev telemetry to the same ledger discipline as product
runs. Note: the OTel GenAI semantic-conventions stability claim did not
survive external verification — we key on our own span schema (above), and
map to external conventions later if/when they stabilize.
3. KPIs and readiness thresholds
Weekly dashboard (per repo and org-rollup), segmented actor_type:
| KPI | Definition | Readiness threshold (to expand autonomy a stage, per WS6 rollout model) |
|---|---|---|
| Task success rate | merged without human rewrite / agent PRs opened | Stable or improving over 4 weeks, no SRM in any live experiment |
| CI-green-on-first-push | first CI result green / agent PRs | Improving; no repo below 50% after WS1/WS2 land there |
| Median time-to-merge | open → merge, agent PRs | Improvement or neutral vs. baseline |
| Review burden | review iterations + human rework minutes per agent PR | Downward or stable after calibration |
| Revert rate | agent PRs reverted ≤14 days / merged | Release-limiting guardrail: no meaningful degradation vs. human baseline |
| Policy violations | boundary-lint failures post-merge; approval-bypass attempts on ApprovedPromotion paths | Zero high-severity; zero bypasses |
| Gate integrity | enforce-flags on; quarantine-list count | Monotone shrinking quarantine counts; all flags flipped by their P1/P2 deadlines |
| Cost | $ per merged agent PR (model + sandbox/CI minutes) | Within budget envelope; optimize cost-per-success, never raw tokens |
| Bot quality (WS6) | shadow-mode agreement with human review verdicts; bot regression suite pass rate | ≥ agreed threshold before limited-write stage; suite blocking |
These thresholds are deliberately strict: early agent platforms fail on trust before they fail on capability.
4. Experimentation discipline
For any change to guidance, skills, CI gates, or bot behavior where we want a causal answer (adopted from the reports; they survived as methodology even where their external evidence citations did not):
- Randomization unit matches the interference locus: repo or team for guidance/skills changes (shared context contaminates within a repo); PR for CI-gate changes; developer for per-user IDE behaviors.
- Pre-register the primary metric (usually task success rate or time-to-merge) before enabling; everything else is diagnostic.
- SRM check on every experiment (chi-square on assignment counts): any significant sample-ratio mismatch freezes interpretation until root-caused.
- CUPED variance reduction using pre-period covariates we already have per repo/developer (historical PR cycle time, review-turn count, CI failure probability).
- Valid sequential monitoring — no informal daily peeking with fixed-horizon p-values; fixed analysis points are acceptable at our scale.
- The
alphaswarm_evalA/B machinery does the same job for bots: two agent configs over the same golden suite via EvalRunner, compared with the already-written Diebold-Mariano/SPA scorers (ALPHASWARM_EVAL_STAT_SCORERS_ENABLED), withcbs_from_eval_reportsfor cost-aware promote/collapse decisions.
Realistic caveat on power: with one small team, many comparisons will be underpowered for formal significance. The discipline still pays — SRM and guardrails catch broken instrumentation and regressions even when the primary metric can't reach significance; treat sub-powered results as directional and lean on the KPI trend lines.
5. Test-infrastructure ratchets (feeds gate integrity)
Per the org's own TESTING_FRAMEWORK_BLUEPRINT.md (D6) — execute, don't
redesign:
- Flaky-test analytics (pytest-rerunfailures or span-based pass/fail-window
detection); tracked quarantine list with owners; the existing
--ignorelists andcontinue-on-errorslices convert to ticketed quarantine entries with a shrinking count enforced in CI. - Ratcheted coverage floors (
--cov-fail-understarting at current per-package baselines); registered + strictly-enforced pytest markers so suites slice deterministically. - Kick off
alphaswarm_testkitPhase 1 per the blueprint.
6. Cadence and ownership
- Weekly: dashboard review in the platform channel; new quarantine entries and enforce-flag deadlines checked; experiment readouts (with SRM status) recorded as dated notes in this directory.
- Monthly: KPI-vs-threshold review decides autonomy stage changes
(WS6 rollout model) and updates
last_reviewedhere. - Quarterly: re-run the cross-repo audit (the 2026-07-19 analysis is reproducible); refresh gaps.md; retire completed workstreams into the org-audit archive.
- Owner: platform-team; the dashboard and span schema live with
alphaswarm_core.observe; eval baselines live withalphaswarm_eval.