mirror of
https://github.com/boshu2/agentops.git
synced 2026-09-14 15:08:13 +08:00
main
1372 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
20f9d4a338 |
Release AgentOps 3.7.0 (#1144)
## What Release the prepared AgentOps update as **3.7.0**, the minor release after 3.6.0. Align the CLI and plugin versions, regenerate the Gemini manifest, and rename/update the curated notes and changelog links. ## Why The operator selected a minor release. No 4.0.0 tag or release was published. Migration instructions and the documented removed commands/skills remain accurate. ## How I tested - Go lint and focused version/manifest tests passed. - Full regeneration parity, changelog mirror parity, and release-note coverage from v3.6.0 passed. - The exact 3.7.0 release rehearsal passed in 143 seconds; all 73 full repository gates passed. All 12 security tools ran with zero missing/error tools, critical findings or security-high findings; existing advisories remain reported. - All nine hosted checks passed on `092e1814b6cba46cd9ac1d797dab2a5c8c7c188c`, including Go race/shuffle tests and 1,509 executed Bats passes (31 environment-dependent skips, zero failures). A new CLI wiring regression confirms `ao version --json` reports the build version. - Actual fresh native Claude/Codex 3.7.0 installs and upgrades from 3.6.0 passed with exact 34-skill inventories. Existing implementation validation from PR #1143 remains applicable to unchanged source. - Fresh author-distinct correction review passed for exact head `092e1814b6cba46cd9ac1d797dab2a5c8c7c188c`, covering all changed paths and four acceptance criteria with no unchecked scope. Verdict digest: `e7b24a297a0b232138de011473df8be6eb3eeaf47afd1f300398eecafab2fbab`. ## Checklist - [x] Version owners and generated metadata agree on 3.7.0. - [x] Migration/removal guidance is preserved. - [x] Exact-candidate release checks pass before tagging. - [x] Fresh correction review is recorded before tagging. |
||
|
|
d972fa2090 |
Prepare AgentOps 4.0.0 plugins, skills and CLI release (#1143)
## What Prepare AgentOps 4.0.0 across the Claude plugin, Codex plugin, skills and CLI. Claude writers capture the supplied check status during its original invocation, and plugin conformance verifies exact skill membership and link destinations. Full release security now scans the repository and blocks on Python collection failures that previously produced a false green result. ## Why The 3.6.0-to-current interval removes published commands and 20 skill names, so this is a major release with migration instructions. Release validation also exposed stale skill assertions and test prerequisites that need to match the current product contracts without weakening acceptance. ## How I tested - Native Claude Opus/Haiku success, failing-check and direct-writer trials: each check ran once, and the direct child returned plain JSON. - Actual fresh installs and upgrades from 3.6.0 in isolated Codex and Claude homes: 34 skills, expected agents, and exact installed package bytes. - Exact candidate `b721d02559e1495be6095ad97b820e88ceb4a049`: all 73 full repository gates, regeneration parity, and the complete local release rehearsal passed. All 12 security tools ran with zero skips, tool errors, critical findings or high-severity security findings. The unchanged advisory policy reports 35 quality-high findings on unchanged files. - Python: 327 tests and 72 subtests passed. Hosted Bats: 1,509 passed, 31 environment-dependent skips, zero failures. Go lint/build/vet/race/shuffle checks and CLI smoke/integration passed. - All 11 hosted checks passed, including Windows correctness, macOS/Linux installation, security, and the six-target no-publish GoReleaser snapshot. Local archive checksums and a real macOS CLI initialization/status/version smoke also passed. - Fresh author-distinct review passed all four acceptance criteria and all 35 changed paths with no unchecked acceptance. Canonical subject and caller-intent verification passed; verdict digest `68af2c935ed0106cd91b3950f5d168e662f4071f660fcbd113c36b7cd0f0426e` binds manifest `7affc77e25eaff69ba36c5ce05582b4f0385c954b76b62c02b97f97041f489b2`. ## Checklist - [x] Breaking changes documented in the migration guide and complete release notes. - [x] No credentials or private runtime proof included. - [x] Final full release checks pass on the exact candidate. - [x] Fresh author-distinct final PASS is recorded before merge. This prepares the release candidate; it does not publish a tag or release. Coverage limits remain explicit: native plugin tests used isolated macOS homes and local marketplaces, guard installation remains opt-in, and reader instructions do not prove sandbox confinement. Semgrep retains pre-existing warning-level parser diagnostics. Snapshot metadata follows the existing 3.6.0 tag; this is a packaging rehearsal, not a published 4.0.0 archive. |
||
|
|
3213afcf1c |
Default to native execution and report independently accepted work (#1129)
## Change Make native coding-agent execution the default AgentOps entry path with zero mandatory skills. Preserve full bundles and add repeatable `ao skills link --skill NAME` selection, validating the entire selection before writes. Align product, installation, architecture and generated command documentation. Extend the existing trial readout to separate endpoint test results, execution state and independently accepted work. Bind supplied judgments to exact content, acceptance and native evidence. Reject empty implementation subjects and require the caller's complete criterion ID set before reporting acceptance. Preserve genuine nonempty and deletion-only subjects, valid failures and missing-proof outcomes. ## Validation - Native onboarding from empty home/consumer directories produces no setup files; selective/full linking and failure boundaries are covered. - Actual RED/GREEN regressions cover empty subjects and the partial-criterion omission found by independent review. - Full Go build, vet and race/shuffle tests; affected Go lint; 88 Python readout/statistics tests passed. - All 73 gates, generated projections, strict documentation build and local aggregate passed (10 passed; one documented optional absence). - All nine PR checks succeeded at `7df0d42b12f35ffc22008cc10a40339afcfbb6a0`. - Fresh author-distinct review passed all six acceptance criteria over all 59 changed paths, with no findings or unchecked scope, after repairing the criterion-coverage finding. ## Evidence limits The real native coding repair demonstrates usability, not comparative skill uplift. The strict live-session machine replay remains NOT_PROVEN where execution/identity observations are unavailable; the source review PASS is retained separately. Existing cohort limits and the historical aggregate-enforcement gap remain unwaived. No new comparative cohort, scheduler, skill-corpus deletion, memory migration or global installation is included. |
||
|
|
5e874b55cf |
Evaluate installed skills on isolated Go work (#1125)
AgentOps previously relied on behavioral probes and retrospective summaries to assess skills. This adds a development-only evaluator that runs a frozen installed skill package on isolated Go tasks, preserves failed and interrupted attempts, and rebuilds a comparison readout from native results without another model call. The suite contains six task families, separate executable verifiers, frozen launch identities, native session accounting, and a focused `skill-eval` maintenance workflow. The readout separates passing code from completed trials, retains incomplete cost information, and reports missing evidence without claiming equivalence or uplift. `ao eval` remains retired; no new runtime controller or required core skill is introduced. Validation: Go build/vet/race checks and all repository CI passed. Focused reader/statistics, receipt integrity, verifier integrity, fixture calibration, generated projections, and the local aggregate runner passed. A real two-variant Docker preparation check verifies that frozen worker and verifier images survive later staging. The bounded coding pilot retained all 24 starts and produced eight comparable pairs across six task families, with no observed paired endpoint difference. The separate eight-start memory experiment did not demonstrate incremental benefit and does not promote another guidance rule. Individual runtime limits were enforced; aggregate desktop deadline enforcement remains unproven. Raw trial evidence and credentials stay outside Git. |
||
|
|
8a9a01a70a |
Fix Go recovery, gate routing, evidence, and handoff defects (#1123)
## What Repair 13 audited Go CLI defects across doctor recovery, gate routing, evidence ingestion, and session handoffs. Doctor preserves recoverable snapshots and reports unresolved findings; gates retain exact changed-file scope and advisory semantics; malformed evidence fails closed; handoffs report observed Git state and chronological recency. ## Why These failures could overwrite recovery data, skip required checks, admit malformed evidence, or restore stale context. Each defect has a regression witness. Existing skill contracts remain unchanged for this bounded workflow evaluation. ## How I tested - Regression witnesses failed before repair and pass after repair. - Combined Go build, vet, race/shuffle tests with coverage, coverage floor, pinned lint, and complexity checks pass. - Generated projections are current; local aggregate reports 10 passed, 0 failed, 1 optional skip. - All 52 selected gates pass. Authoritative Linux/Windows correctness, security, full registry, and installation CI pass. - Fresh author-distinct OpenAI/Codex review verified all 13 repairs and all 33 changed paths on `84029ee533be0a73ac6c882eb9cf7369d52a0055`, including independent critical race regressions; no findings or unchecked acceptance. - Directory reverse moves retain the existing advisory-lock concurrency boundary; this does not claim exhaustive hostile filesystem-race coverage. ## Checklist - [x] Go build and tests pass, including race detection - [x] No secrets or credentials added - [x] Compatibility behavior documented in the changed contract where applicable |
||
|
|
17849bbc24 |
Improve CLI checkpoint recovery, evidence status, and skill search (#1120)
A failed mining-checkpoint write could truncate the saved watermark and cause retry to replay older events. The CLI also wrote evidence to external roots that status could not inspect. This batch fixes those behaviors and removes duplicate normalization from skill search. - Mining checkpoints use the existing atomic storage writer. A real partial-write regression test proves old bytes survive and retry retains stable event IDs. Existing mode bits are preserved; new checkpoints use 0600. Symlink and special-file destinations are rejected before reading. Before replacement, an empty same-directory probe checks ownership and permission metadata, including ACLs and inherited permissions. Unverifiable or different metadata returns an error and leaves the prior checkpoint intact. This is a conservative refusal, not ACL migration. A directory-sync error after rename can leave the new state visible. - `ao status --evidence-root PATH` inspects an explicit existing non-Git store, with matching text/JSON/YAML reports, no fallback on invalid roots, and no reads through evidence symlinks. Omitted-flag behavior remains unchanged. - Skill-query normalization has one implementation, preserving repetition versus first-occurrence semantics. Nine fixed shipped-catalog queries remain byte-identical against a source-pinned baseline. Validation: Go build, vet, tests, race/shuffle with atomic coverage, repository Bats and aggregate suites, regeneration, applicable gates and lint passed. Final Linux and Windows correctness CI and all required checks passed. A fresh author-distinct review verified every acceptance criterion across all 23 changed paths with no unchecked scope. Nine production-query outputs match the source-pinned baseline. The first CI attempt exposed a test-child coverage flush under its temporary file-size limit; the test now restores that limit before exit. Fresh review then exposed ACL loss despite green CI. The permission guard and native regression tests repair that defect while preserving the original access requirement. The failure cases and repair costs are retained in the evaluation. |
||
|
|
db1a0573ea |
Add bounded source reads and pinned OKF profile checks (#1114)
Adds two explicit read-only operations for the context delivery lifecycle: bounded raw source reads with reversible bytes and integrity checks, and structural checking of the pinned AgentOps OKF page profile. Source reads require independently selected context policy and enforce a measured serialized-output bound before emitting content. Emitted bytes do not establish host delivery or understanding; restricted-source processing remains unavailable without native enforcement. The OKF checker rejects missing status and incompatible profiles, and never grants truth, disclosure, or usefulness approval. Validation: focused tests and Linux/Windows source-reader builds passed. The combined candidate is undergoing the required full repository checks and fresh independent review before landing. |
||
|
|
57ece9fb7b |
Restore private context routes and verify native judgment receipts (#1112)
Add explicit, recoverable private context routing through `ao config context`, binding native source, owner, task, model and destination to existing policy and external storage. Recovery reads the original Beads maintenance anchor; configuration reports native access enforcement as unattested. Add `ao provenance verify-judgments` to check required review profiles against exact native transcript receipts, independent subject and acceptance, distinct contexts, completion and permitted providers. Requested identity and unreported effort do not count as runtime evidence. The verdict schema is unchanged. Repair the existing cleanup test: a 0.3-second budget could expire during preparation before either fixture process started. A separate controlled-delay test now proves preparation cannot renew that deadline. The running-cleanup case requires parent/child readiness, preserved partial output, the postlaunch cleanup result and both processes stopped within its existing four-second bound. Production timeout behavior is unchanged. Validation: fresh author-distinct review passed the exact 55-path final subject and all T05/T21 acceptance. The complete local Bats run passed (1,333 passed, two existing skips), as did Go build/vet/test/race, all 72 full-mode gates, the aggregate and generated-output checks. Ubuntu/Windows CI, security and both installation jobs passed on the final commit. The final evidence scan found no new orphaned bindings; 73 historical bindings remain preserved. Earlier failed results and private evidence remain outside the PR. |
||
|
|
8061085c89 |
Ship native evidence helpers and fresh-family review defaults (#1110)
AO now performs intent snapshots, subject manifests, strict evidence verification, atomic verdict storage, and orphan inspection through the Go binary. The command handler keeps verification separate from presentation so it meets the existing complexity limit. These operations preserve the existing evidence formats, require explicit protected storage where applicable, and run outside a checkout without Python. The unchanged Python implementation remains a developer oracle; agents still provide semantic judgment. Codex and Claude skills now default to a fresh reviewer from the author’s model family. Callers can explicitly request cross-model review or pin its model. Reviewer adapters use a finite caller timeout or remaining deadline instead of a fixed ten-minute default, while retaining output limits and abnormal-termination cleanup. Validation: Go build, vet, tests and race/shuffle tests; 1,334 shell tests; aggregate runner; regeneration check; 72 full-mode gates. Independent checks exercised 84 storage-boundary rejections and 21 evidence operations with an empty PATH. Both canonical and generated RPI reference suites pass all 48 tests after updating the migrated oracle import without weakening assertions. Change-sensitive checks explicitly compare the final committed candidate with the original PR base. Linux, Windows, installer, security, and required summary checks are green. |
||
|
|
568e99d436 |
Loop restore: converge and crank as control flow under the verdict contract (ADR-0017) (#1099)
## Loop restore: converge and crank as control flow under the verdict
contract (ADR-0017)
Intent source: `docs/plans/2026-09-03-loop-restore.md` (in this PR).
Decision record:
`docs/adr/ADR-0017-loop-as-control-flow-not-knowledge.md`.
**Why.** The 2026-07-14 single-pass cut (`482307762`) removed the
iterate loop (discovery, crank, converge, evolve, the learn write-half)
together with the unproven compounding claim, although ADR-0011 demoted
only the latter. The control flow was never demoted, and its absence
showed on 2026-09-02, when a three-lane fix needed eight validators and
two stops because the contract had no repair phase. This restores the
loop as control flow and nothing else: no knowledge store, no `ao
converge`/`ao crank`, no evolve, no canary. ADR-0004 and ADR-0011 stay
in force.
**What changes.**
- **RPI gains a bounded repair phase.** On `FAIL` or `NOT_PROVEN` with
findings, repair and re-validate freshly under the convergence law:
caller-declared `repair_rounds` (default 2); open finding set keyed by
stable `findings[].id`, union across validator families, non-growing; no
closed id reopens; the subject digest changed or, for `NOT_PROVEN`, new
digest-bound evidence resolved a named gap. Converged = fresh PASS plus
cross-family PASS on risky surfaces. Plan and Implement keep their
single dispatch. `skills/rpi/scripts/run_once.py` models the law as pure
data (33 tests): rounds are validated for shape (digest required, no
duplicate ids, no PASS with findings, no FAIL without findings),
condition 4's evidence branch needs a NOT_PROVEN previous round, a
non-FAIL current round, new evidence, and a resolved finding, and a PASS
over unchanged bytes after a FAIL is a flip that reports NOT_PROVEN.
`workflows/rpi.js` runs validation as legs (spawned or external primary,
plus a caller-supplied `crossFamily.command` on risky surfaces) merged
worst-of with a union of stable ids; a risky surface without a
cross-family leg is `diversity_unsatisfied` and never converges or
enters repair; a failed repair or re-validation returns NOT_PROVEN with
no stale verdict. Validators return `subjectDigest`, stable finding ids,
and `evidenceRefs`.
- **crank returns as a thin wave executor** (113 lines): the caller
selects the wave and the repair bound, crank invokes RPI per lane
(parallel only on disjoint write and regen scopes), runs the wave
acceptance once, returns evidence, and stops. No retry, budget, queue,
claim, lease, Git, closure, or next-work ownership. Routing golden
`rq-07-wave-execution` ranks it first.
- **validate is cross-family by default on risky surfaces**
(`cli/internal/gates/**`, `scripts/check-*.sh`, `tests/**`,
`skills/*/scripts/**`, hook policies, `lib/**`, security-scanned paths)
with the LAW-0 dispatch table: Claude orchestrating uses read-only
`codex exec`; Codex orchestrating uses an interactive Claude session in
an NTM pane, never `claude -p`. No live adapter means
`diversity_unsatisfied`, which on a risky surface is `NOT_PROVEN`. The
full literal CI command set runs once on the final integrated subject;
routine rounds keep the receipt-driven freshness contract.
- **Conformance assertions flipped under ADR-0017 only:**
`scripts/check-cathedral-cut-conformance.py` (crank live; "Stop
regardless" replaced by positive canaries for the law's four conditions;
a bounded `for` loop that compares against `repair_rounds` is required
in `run_repair_phase`, and the gate executes the law's canaries against
the reference behavior), `workflows/rpi.js`,
`skills/rpi/scripts/validate.sh`,
`evals/agentops-core/rpi-behavior.json`,
`skills/rpi/references/rpi.feature`. Every single-pass public surface
(README, AGENTS.md, PRODUCT.md, CI-CD, agent-workflow-reference,
rpi-traversal, cli/README, quickstart and demo commands, the
operating-contract and product-boundary bats, the Codex-description
oracle) now states repair to convergence.
**Known approximation, disclosed.** The Claude conveyor has no
deterministic shell primitive, so changed paths are derived by the fresh
validator (git status and diff against the clean pre-run tree) and
unioned with the implementer's report; risk is classified over that
union and unreported paths are coverage findings. A validator is still a
model; runtime derivation outside every agent is a follow-up. Family
distinctness of the cross-family leg is asserted by the caller's choice
of command and not verified by the script.
**Not in scope.** Premortem stays a single advisory judge and Plan still
only names the first check (phase boundaries unchanged). No `verdict.v2`
or `rpi-report.v1` change. The loop's own effect on outcomes is
unmeasured and owed a seeded-defect probe, like the rest of the corpus.
**Evidence on the tip.** Regen check clean; full gate green with a
HEAD-built binary; CI's bats command green; Go build/vet/test green;
golangci-lint clean; security gate quick PASS; one fresh validator over
the whole diff; one cross-family read of the design before
implementation (13 findings folded) and two of the integrated diff (9
findings in round one, 11 by round two, 15 by round three, each round
repaired and re-reviewed; the fresh validator passed the tip after round
two and the final tip
|
||
|
|
8cdcb5a903 |
Train 1: measurement substrate, context diet, retrieval-eval contract (instrument-panel roadmap) (#1087)
> **Residues closed on the caller's merge instruction** (`499d916a6`): the round-2 findings were the same failure shape — round-1 repairs patched cited lines instead of sweeping the class — so this commit sweeps each file whole: every remaining SATURATED-row-append site in skill-eval now routes to RUNBOOK retirement, the human-only-skills *description* is runtime-conditional, premortem's "(MEASURED)" label is gone, SKILL-API's context table carries all 25 rows and the enforcement table gains `disable-model-invocation`, and the fixture-identity claim is stated precisely (probe id, honesty note, and control arm are the only differing fields — as the acceptance permits). Post-sweep: validators, full Go suite, 68/68 gates, goldens + headroom bats green, projections current, gemini in sync. Merging per Bo's instruction. ## What Train 1 of the accepted [instrument-panel roadmap](docs/plans/2026-08-26-instrument-panel-roadmap.md) (intent landed at `986a4feaf`): the measurement substrate, the skill-context diet, and the retrieval-eval contract. Three worktree-isolated lanes, each independently validated by a fresh context, plus one integration commit. 103 files, +5,510/−76. **L1 — measurement substrate** (`instrument/measurement-substrate`) - Gate `skill.probe-headroom` (advisory, Fast|Full): answers the question `skill.probe-coverage` cannot — not "does a probe result exist" but "could one have existed at all". The rule, ported to Go (`cli/internal/probeheadroom` + `cli/cmd/probe-headroom` behind a thin check script — the witness-crosscheck pattern, **no new `ao` root command**): control arm ≥ 0.75 with ≥ 2 usable reps at ≥ 2 effort levels ⇒ SATURATED (void row, not an honest null); UNMEASURED outranks it; treatment-silent ⇒ FLOOR; else SEPARATED. RED first: both committed fixture pairs read `INERT` to everything else in the repo; the failing separation test predates the implementation, and a bats negative-control swaps fixture bytes and asserts the gate flips. - **First reading on real data: 7 of 11 historical probe groups are SATURATED** — including both `validate-not-proven` runs. Those INERT rows were never honest nulls; they were void. The 0/12 ledger number now argues itself. - Declared denominator for probe-coverage: `scripts/.skill-probe-denominator-exclusions`, fail-closed parser (entry without an argument, stale slug, or duplicate ⇒ exit 2). One entry (`goals`, a pure alias-of `fitness`). Net effect deliberately zero (0/12 → 0/12: alias left, `one-way-door` entered) — the gain is a declared number, not a better-looking one. - Re-landed from the recovered clean-room commit (`9872483bd`), re-validated against *current* main: `skill-eval` (defers saturation to the gate id; its shell scripts dropped, not shipped — ratchet intent), `route`, `one-way-door`, premortem reversibility check, council `caller_challenge` (schema + validator, per the agent-core boundary that the panel may challenge, never overrule). **L4 — context diet** (`instrument/context-diet`) - `disable-model-invocation: true` on 4 human-only skills (key verified verbatim against Anthropic's docs). The plan guessed 35 candidates; the graph said otherwise — 23 carry `user-invocable: true`, and 19 of those are excluded on cited evidence (rpi consumes anti-ceremony/implement/plan/validate; workflow scripts reach others; `goals` is a live migration tombstone). The exclusion evidence is retained in the lane report. - One router skill (`human-only-skills`) — the single always-loaded description that replaces four; it hints, never fires. - `.out-of-scope/` formalized with this week's three refusals (checked-in knowledge corpus; ee self-improvement loops; whole-skill A/B as the measurement unit), each citing its evidence. - Deterministic proof, no model eval: before/after bytes of always-loaded description load reported in the lane summary. **L5 — retrieval-eval contract** (`instrument/retrieval-contract`, lane verdict PASS 10/10) - `AGENTS.md` federated row now names **ee (eidetic-engine)** as a concrete caller-selected memory system — consume, never build; symlink intact. - `schemas/pack-quality-expectations.v1.schema.json` + 4 routing goldens + `scripts/check-routing-probe-goldens.sh` graded against `ao skills find`, wired as an **advisory** nightly job. Zero goldens is a failing state — no new zero-denominator green. - **The instrument caught a real miss on day one — and its own prescription fixed it.** Golden `rq-04` expects `validate` for "judge whether this finished change is actually proven before I merge it"; at authoring, `ao skills find` ranked the *forbidden* `premortem` first and `validate` nowhere in six natural phrasings. The pointer-wording-first repair (validate's description gained the caller's own words: finished, proven, verdict, merge) now ranks it #1 at 0.333; grader 6/6, and the golden pins the repair — a description regression reopens it. ## Integration `regen-all.sh` once over the merged lanes (catalog 52 → 56, four new codex twins, mesh, router, manifests); `skills/route/SKILL.md` catalog/router links became prose repo-root references (the projected twin cannot resolve `../catalog.json` — this was both the portable-conformance failure and the sole broken doc link); `codex-portable-conformance.bats` pin 52 → 56. ## Evidence - `cd cli && go build ./... && go vet ./... && go test ./...` exit 0 · `ao gate check --full` **68/68** · four skill validators PASS · probe-headroom / routing-goldens / probe-coverage bats PASS · `regen-all.sh --check` all current. - Per-lane fresh validators re-ran every suite on detached content; L5 PASS; L1/L4 NOT_PROVEN solely on the projection-regen clause reserved for integration (their remaining acceptance observed green), settled above. Cross-family (Codex) review of the integrated diff recorded in the session report. - Two disclosed scope stretches accepted at integration: a one-line `.gitignore` entry mirroring the witness-crosscheck precedent, and the probe LEDGER.md fact-correction L1's own change made necessary (noted for Train 2's L2, which owns that file next). ## Cross-family review (Codex, fresh context) Round 1: **FAIL** — two blockers (the RED fixtures didn't isolate the control arm; the goldens grader was red where the plan's acceptance says green) and eight majors (contract contradictions in the re-landed skills, a converter-substitution false claim in the codex router twin, two overreaching `.out-of-scope` entries, stale SKILL-API counts). All repaired in one bounded round (`db68935a3`): fixtures now byte-identical outside the control arm, the routing miss actually fixed rather than tolerated, every cited contradiction reconciled at the source and re-projected. Post-repair: full Go suite exit 0, `gate check --full` 68/68, all validators and probe/goldens bats green, projections current, gemini byte-identity restored. Focused re-check verdict recorded in the session report. ## Follow-ups (Train 2, already planned) Seeded-defect probes for the judgment spine (every ledger row citing a passing headroom pre-screen) and the gate-hardening pair (`Gate-Loosen-Reason` tightening ratchet; mechanical grounding-validation over evidence docs). Plus, surfaced by this train: a latent `valid_keys`/schema divergence in `validate-skill-schema.sh` (two keys the schema defines are absent from the script's allowlist — pre-existing). |
||
|
|
ffb9f122af |
refactor(cli): delete the unconsumed eval/redact surfaces — the estate audit's mechanical cut (#1082)
> **Review findings closed.** The re-check's residue (app-seam family count) is applied in `9a2790ae7` along with the full-tier CI settlements: regenerated documentation index (generated file, hand-edit drifted it), regenerated CLI-surface count fixtures (top=18 sub=44 all=62), `Test-Removal-Reason` trailer for the deliberate test deletions, and the release-tag bats output list updated to the real changes-job set. 67/67 full-tier gates green locally. Merging on Bo's instruction. ## What Deletes the provably-dead 28% of the `ao` CLI and every reference to it, per the 2026-08-23 estate audit. −19.5K lines in the lane commit plus integration fixups. **Removed (each with zero live consumers, verified by consumer-grep + `go list -deps`):** - `ao eval` — 13 subcommands, ~10.9K LOC. Its would-be consumers were already tombstones (`scripts/eval-agentops.sh` printed `RETIRED`), `release.yml` hardcoded `--eval pass`, release evidence recorded `suite_count: 0`, and three of its module tests exercised subcommands that could never register (nil composition seats). - `ao redact` — its only declared caller (`skills/compile/scripts/compile.sh`) never existed. - `cli/internal/types/memrl_policy.go` + the orphan cascade it and eval left behind (`internal/scenario`, `internal/wiki`, `internal/runtimecmd`, `internal/redact`) — all with zero importers, verified before and after. - `scripts/check-memrl-health.sh` + `examples/schedules/feedback-drain-hourly.yaml` — a health check for the feedback loop amputated on 2026-07-14; it exits 1 on main today and the example instructs a verb (`ao feedback-loop`) that no longer exists. - `corpus.secret-scan` gate — vacuous: its file filter excluded the single tracked path its globs could match, so it scanned zero files; secrets are covered by the pinned gitleaks steps in nightly and release (validate's quick toolchain mode skips gitleaks). - Docs for the deleted surface: `docs/architecture/eval-architecture.md`, `docs/code-map/eval-lid-primitives.md`; `contracts/eval-baseline-ab.md` already carried a RETIRED banner and stays as history (delisted from the live index). **Kept, deliberately:** - `ao robot-docs` — the audit's "duplicate of `doctor robot-docs`" premise was false: they render different handbooks (whole-CLI vs doctor-scoped). Verified before acting. - `completion`, `demo`, `quick-start` — interactive human furniture, not dead code. - `corpus.witness-dolt-jsonl-crosscheck` gate — retargeted, not retired: its backing script is a hermetic self-test over real tracked fixtures; globs now point at the paths it actually exercises. - `cli/internal/evalsubstrate` — Go-dead but it is the declared mirror of `schemas/outcomes-rubric.v1.schema.json`; retiring it needs a paired schemas/docs/scripts decision (package doc comment records this). - `scripts/ci-local-release.sh` eval-evidence stanza — self-contained honest bookkeeping (`status: not_applicable`), invokes nothing removed. **Tombstones + migration:** `eval` and `redact` added to `removed_command_hint.go` and `docs/MIGRATION.md`; the now-false "(`ao eval` returned in 3.3 …)" parenthetical deleted; `go-cli.md` spine and the "Eval — the Learn seat" section updated; the dated research snapshot got a HISTORICAL banner via the docs-scope self-declaration mechanism (history not rewritten). ## Why v3.6.0 binary downloads: 4 darwin-arm64, 3 linux-amd64. Only 7 of 53 shipped skills invoke `ao` at all, and none of them touch this surface. The eval family was the single largest command surface in the CLI with zero live consumers — 28% of non-test Go maintained for nobody. ## Evidence - `cd cli && go build ./... && go vet ./... && go test ./...` — exit 0 (previously-failing `TestGoCLIDocSpineMatchesApprovedSpine` and `TestRemovedVerbsHaveMigrationRows` now pass) - `scripts/check-docs-cli-snippets.sh` PASS · `check-cmdao-surface-parity.sh` PASS (54 leaf commands) · `check-corpus-path-guard.sh` PASS · `check-new-scripts-use-preamble.sh` PASS · `ao gate check --dry-run` PASS - Implemented by a worktree-isolated lane, independently validated by a fresh context that re-ran the suite itself; the two failures it found were doc files outside the lane's write scope, fixed in the integration commit. Cross-family (Codex) review verdict included in the final session report. ## Cross-family review (Codex, fresh context) First pass: **FAIL** with two majors — (1) `quality.DeprecatedCommands` still mapped five rewrite entries onto the removed eval family, so `ao doctor --fix` would have introduced dead commands; (2) retained docs (formal-verification research links, applied-ood README run block, evalsubstrate hint strings) still prescribed removed commands. Both repaired in `4da85a0d4` (one bounded round), plus its two minors (types/AGENTS.md row, .gitignore unignore, family counts, gitleaks-coverage comment). Re-verified: full suite green, snippets gate PASS. Focused re-check: first-round findings confirmed closed; one new residue (the family count above) stopped the loop under the spiral rule. ## Follow-ups (not in this PR) - `cli/internal/quality/stale_refs.go` `DeprecatedCommands`: the five eval-target entries are pruned here; the older pre-existing dead targets (forge, inject, flywheel, ratchet, …) still need a map-wide reconciliation against the live registry. - `cli/internal/evalsubstrate` retirement decision (paired schemas/docs/scripts change). - `evals/scenarios/applied-ood/`, `evals/tier2-premortem/`, `evals/_stats/` retain historical `ao eval` mentions in prereg/holdout records — dated artifacts, left as history. |
||
|
|
621dbb575f |
3.6.0 release prep: version bumps, changelog, curated notes (#1071)
Everything-but-tag for **v3.6.0**. Minor, not major: the post-3.5.0 delta retires the knowledge-flywheel product surface and aligns the estate on the operations-layer identity, matching the 3.4.0 precedent where the orchestration pack was removed in a minor. ## What this carries - **Version 3.5.0 -> 3.6.0 across all seven surfaces**: Claude plugin manifest, marketplace metadata + plugin entry, Codex manifest, Gemini image manifest, Claude image verify pin, and the `ao` source fallback. - **CHANGELOG `[3.6.0]`** (root + docs mirror): operations-layer alignment, anti-ceremony enforcement, the behavioral eval program, the flywheel retirement, and the honest 0/12 measured probe coverage. - **Curated `docs/releases/2026-08-17-v3.6.0-notes.md`**: validator PASS, tier minor, full changed-path area coverage. The Breaking Changes section lists all six removals and the handoff write-path move rather than burying them in a minor. - **New regression test `cli/cmd/ao/version_manifest_parity_test.go`** binding the `version` fallback to every version-bearing release surface. - **PRODUCT.md** reviewed against the 3.6 surface and re-stamped; **`docs/reference/skill-system-evolution.md`** gains its 3.6.0 row and drops the "current unreleased tree" framing that the tag would falsify. ## Why the new test exists This cut missed `images/claude/verify.sh`. Its version guard — whose entire stated purpose is catching plugin.json drift *behind* the release — then rejected the **correct** version, so a user following the shipped `images/claude/README.md` on the v3.6.0 tag would have hit a hard FAIL. `check_manifest_version_consistency` in `ci-local-release.sh` compares only the two Claude manifests to each other, so it structurally could not see this. The test fails on the drift and passes when correct; both directions were exercised before committing. ## Honesty notes carried into the release - Measured behavioral probe coverage is stated as **0/12** under the v3 evidence contract. The earlier wave-1 classifications are retained as `LEGACY-UNVERIFIED` rather than counted, because the probe harness did not isolate the skill corpus between control and treatment arms. Skill-efficacy claims in these notes are directional, not proven. - The estate-ablation aggregate counts are labeled legacy-unverified and non-promotable. ## Verification Full `scripts/ci-local-release.sh --release-version 3.6.0 --readiness-mode official --security-mode full`: **PASSED — 72 checks, 0 failures**. | Dimension | Status | |---|---| | SIL (race suite, 75.8% coverage) | pass | | VIL (gates, regen, digital twin) | pass | | HIL (real Darwin/arm64 target) | pass | | Artifacts (CycloneDX + SPDX SBOM) | pass | | Security (full mode) | pass | Readiness **9.0** against threshold 8. HIL used a real target with **no waiver**: `ao` built from this tree reported `ao version 3.6.0` (`version_verified=true`) and ran a full `ao init` scaffold plus `ao status` in a scratch repo. Security full mode: 0 critical, 0 high, 3 medium (non-blocking). Notes validator PASS, doc-release gate PASS. ## Post-merge Readiness lap at the merged SHA, audit record in `docs/audits/`, then the tag — per the binding process rule that the record exists **before** the tag. |
||
|
|
f3c6d0ecf2 |
Converge retained WIP and harden evidence boundaries (#1065)
Summary:
- lands the audited current WIP lanes and excludes stale/process-only
material
- hardens prune path confinement, probe-v3 evidence binding,
codebase-recon identity, handoff/release/reverse-engineer behavior, and
Codex prompt handling
- truth-labels static skill scoring and regenerates all owning
projections
Validation:
- fresh independent PASS on commit
|
||
|
|
7a765cde19 |
Align AgentOps around its operations-layer identity (#1051)
Executes docs/plans/2026-08-07-agentops-operations-layer-alignment.md: AgentOps is the operations layer for agentic engineering; the federated integration graph is the topology, the semantic work-and-proof protocol is the contract, and RPI is the standard one-experiment traversal. Retires the ao flywheel command family and all knowledge-flywheel product state, tombstones the seven-move operating-loop workflow, narrows ao init and the .agents state writers to declared destinations, renames the core architecture page to rpi-traversal.md with a compatibility redirect, aligns AGENTS.md, 25 skills, public and package copy, regenerates every owned projection, and strengthens the conformance gates with planted-negative proofs. Both the alignment subject and the follow-up gate-bookkeeping commit carry fresh author-distinct validation PASS verdicts with empty not_checked scope. Test-Removal-Reason: dead knowledge-flywheel and session-store surfaces were deleted with their tests (operations-layer alignment) |
||
|
|
1c1500ce87 |
3.5.0 release prep: version bumps, changelog, curated notes (#1029)
Everything-but-tag for v3.5.0 (Bo's call: the post-3.4.0 delta carries two feature surfaces — `ao gc` and plan manifest mode — so minor, not patch). - Version 3.4.0 → 3.5.0 across all seven surfaces (plugin manifests, marketplace, image verify pin, `ao` source fallback). - CHANGELOG `[3.5.0]` section (root + docs mirror): ao gc family, manifest mode, Mayor-dispatch doctrine, honest-scoped-PASS, fresh-install fixes, init gitignore policy. - Curated `docs/releases/2026-07-31-v3.5.0-notes.md`: validator PASS, tier minor, full area coverage; upgrade notes call out the gc-maintainer-ops wrapper deprecation and the new init gitignore block. Verification: notes validator PASS · `regen-all.sh --check` all ✓ · skill-lint 0 · `go build/vet/test` 2947 passed / 73 packages. Post-merge: official-mode readiness lap at the merged SHA with real HIL, record in docs/audits/ before any tag. |
||
|
|
fd30523a75 |
docs(skills): install-agnostic loop commands and self-contained rpi examples (#1026)
## Summary
Fresh-install smoke testing found four commands/links in skills and CLI
help that fail verbatim for an installed user (only `skills/**` — not
the full repo tree — ships to an install; `.agents/skills/**` is the
installed skill root).
| # | Defect | Fix | Verified |
|---|---|---|---|
| 1 | `skills/plan/SKILL.md` step 1 told the runtime to run `python3
skills/validate/scripts/validate.py snapshot-intent ...` — a
checkout-only path. In an installed tree the real path is
`.agents/skills/validate/scripts/validate.py`. | Reworded to name both
paths explicitly (checkout: `skills/validate/scripts/validate.py`;
installed: `.agents/skills/validate/scripts/validate.py`),
install-agnostic. | Copied `skills/{validate,plan,rpi}` into a scratch
`.agents/skills/` layout and ran `echo '{"foo":"bar"}' \| python3
.agents/skills/validate/scripts/validate.py snapshot-intent --source -`
verbatim — produced a valid `intent_ref`. Re-ran the checkout-relative
form too. |
| 2 | `skills/rpi/SKILL.md` linked
`../../schemas/rpi-report.v1.schema.json` — `schemas/` isn't shipped to
installs, so the link 404s for an installed user. | Inlined the minimal
required `rpi-report.v1` shape as a fenced JSON block (with field
semantics), plus a note that the schema itself ships in a repo checkout.
No new files added to the skill package. | Validated the exact inlined
shape (with a concrete instance) against
`schemas/rpi-report.v1.schema.json` via `jsonschema.validate()` —
passes. Confirmed all 9 required keys and digest patterns match the real
schema. |
| 3 | `skills/rpi/SKILL.md`'s continuation-envelope example (~lines
115-121) cited this repo's own internal 2026-07-15 intent/verdict
digests (`26a4f2be...eb48`, `b6e759dd...cb6a`, etc.) as a normative
example — not reproducible by an installed user. | Replaced with a
generic, self-contained example (placeholder revisions/verdicts) that
illustrates the same two-stop-checkpoint behavior without citing this
repo's private history. | Reviewed the replaced prose reads correctly in
context; `bash tests/skills/run-all.sh` still passes (no broken
frontmatter/budget). |
| 4 | `ao --help` root epilog (`cli/cmd/ao/root.go`) pointed at
`docs/MIGRATION.md`, a relative path that doesn't exist for a user who
only has the `ao` binary (no `docs/` directory ships with it). | Changed
the epilog to the GitHub blob URL
(`https://github.com/boshu2/agentops/blob/main/docs/MIGRATION.md`), with
a note that a repo checkout also has it locally at `docs/MIGRATION.md`.
Left the internal `removedCommandHint()` machinery (and its
`docs/MIGRATION.md`-literal test assertions) untouched — that's a
separate, heavily-tested mechanism not covered by this defect. | `go run
./cmd/ao --help` shows the new URL. Confirmed `boshu2/agentops` is the
correct remote and `docs/MIGRATION.md` exists at that path on `main`. |
## Process / regen
- `scripts/codex-sync.sh --only plan`, `--only rpi`
- `scripts/regen-codex-hashes.sh --only plan`, `--only rpi`
- `python3 scripts/generate-skill-mesh.py`
- `scripts/generate-cli-reference.sh` (no diff — root epilog text isn't
captured in `COMMANDS.md`)
- `scripts/regen-all.sh --check` — all projections current
## Test plan
- [x] `cd cli && go build ./...` — success
- [x] `cd cli && go vet ./...` — no issues
- [x] `cd cli && go test ./...` — 2923 passed, 0 failed, 73 packages
- [x] `bash tests/skills/run-all.sh` — 54/54 skills pass, 0 failed
- [x] Simulated install (`.agents/skills/...`) and ran the exact
documented `validate.py snapshot-intent` command verbatim — works
- [x] Validated the inlined `rpi-report.v1` JSON shape against the real
schema with `jsonschema.validate()`
- [x] `go run ./cmd/ao --help` shows the corrected epilog
Write scope respected: `skills/plan/SKILL.md`, `skills/rpi/SKILL.md`,
`cli/cmd/ao/root.go` (help text only), and regenerated projections
(`skills-codex/**`, `images/gemini/skills/**`). No changes to
`skills/validate/**`, `CLAUDE.md`, `AGENTS.md`, `cli/internal/gates/**`,
or `cli/internal/doctor/**`.
|
||
|
|
efcf4879c8 |
feat(gc): port gc-maintainer-ops into the ao gc command family (#1016)
## What Ports `scripts/gc-maintainer-ops.sh` (425 lines of bash: prepare / check / recover-affinity for stock Gas City rigs) into the Go CLI as **`ao gc prepare|check|recover-affinity`**, per ADR-0016 (skill logic ships in Go via `ao`; shell stays thin glue). **Why:** skills ship via plugin/npx as SKILL.md only — a user without a repo checkout could not run the commands the shipped `using-gc` skill teaches. The skill said "From an AgentOps checkout", which was disclosed but weak. ## Changes - **`cli/internal/gcmaintainer`** — full port: rig/import pin verification, bundled pack-cache resolution, PyYAML-capable python selection, atomic runtime staging, managed check wrappers, skill links into city/rig Codex sinks, macOS LaunchAgent + doctor/status health checks, bounded affinity recovery. Output and refusal-message parity with the shell script (incl. refuse-before-mutation ordering). - **`cli/internal/commands/gc` + `cmd/ao/gc_composition.go`** — cobra module on the shared `clicontract.HostOptions` seam; global `--dry-run` always overrides `--apply`. - **Skills source resolution without a checkout**: `--skills-source` > enclosing agentops checkout > installed skills root (`~/.agents/skills`, `~/.claude/skills`). Existing rigs stay recognized: the `managed-by: agentops gc-maintainer-ops` wrapper marker is unchanged. - **Tests migrated**: `tests/python/test_gc_maintainer_ops.py` (7 cases) → Go L2 tests in `cli/internal/gcmaintainer` with the same fake-`gc` harness, plus module wiring tests. `scripts/check-gc-executor.sh` no longer runs the python suite. - **`scripts/gc-maintainer-ops.sh`** reduced to a thin wrapper exec'ing `ao gc`, pinning `--skills-source` to its checkout to preserve historical semantics (`--ao-bin` now selects the ao binary). - **Docs/projections**: `skills/using-gc/SKILL.md` now teaches `ao gc ...`; codex, gemini, and executor-pack projections regenerated via their owning generators; spine/COMMANDS.md/surface artifacts regenerated. ## Verification - `go build ./... && go vet ./... && go test ./...` — 2923 passed, 73 packages - `golangci-lint run` on new/touched packages — clean - `shellcheck -S warning` on wrapper + gate script — clean - `bash scripts/check-gc-executor.sh` — OK - Smoke: built `ao`, ran wrapper → `ao gc` delegation end-to-end |
||
|
|
9dd6e7d3f9 |
3.4.0 release prep: upstream-factories pivot, version bumps, release notes (#1013)
## Summary Everything-but-the-tag for v3.4.0, in four commits: - **docs(gc)**: the factory pivot — README and `using-gc` present the upstream [Gas City build pack](https://github.com/gastownhall/gascity-packs/tree/main/gascity) and [Agentic Coding Flywheel](https://agent-flywheel.com/) as the supported factory choices; the in-repo prototype (`deploy/gc/`) is retired in place. AgentOps' lane is the skills + evidence discipline either factory executes. - **chore(release)**: version 3.3.0 → 3.4.0 across all six surfaces (claude/codex/gemini plugin manifests, marketplace, image verify pin, `ao` source fallback). - **fix(gates)**: `check-orchestration-skill-boundaries.sh` exited 2 on every run — it probed adapter files deleted by the 3.3 single-pass refactor and three contract phrases removed by the skill-overhaul waves. The live ratchets (retired-skill absence, ATM-era naming) are kept. - **docs(release)**: 3.4.0 CHANGELOG section (root + docs mirror) and curated release notes; `validate-release-notes.sh` passes (tier minor, full area coverage). ## Verification - Full Go gate in a clean worktree: build ✓ vet ✓ test **2902 passed / 0 failed** (71 packages) - `scripts/regen-all.sh --check`: all 11 projection/doc checks ✓ (including the doc-release freeze gate) - `scripts/validate-release-notes.sh v3.4.0 --since v3.3.0`: PASS - `scripts/check-orchestration-skill-boundaries.sh`: exit 0 (was exit 2 on main) ## Notes - The earlier read that `go.cli-reference` needed unpinning from the negative-witness grandfather list was a **false positive**: gitignored session logs under `tests/claude-code/logs/` pollute the witness scan in a dirty checkout. On a clean tree the pin is correct; a follow-up task exists to make the scanner read only tracked files. - Tagging + Release Publisher run happen after merge, separately; an official-mode readiness artifact gets produced at the merged SHA **before** any tag (binding rule from the v3.3.0 record). |
||
|
|
a305de5c3e |
Go CLI audit residue: eval id containment hardened, dry-run honored, owned temp dirs (#1008)
## Summary Bead `age-skill-overhaul-reboot-sjv7v.12` — reconciliation of the 2026-07-24 Go CLI deep audit against current main. Full table with evidence: `docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/s12-go-residue.md`. **Fixed here (4):** - **Eval identifier path containment** [High] — new `evalsubstrate.ValidateID` at every identifier-to-path join, hardened through two review rounds: rejects separators, absolute/volume refs, leading/trailing space-or-dot (defeats Win32 trailing-strip renormalization), C0+C1+DEL controls, non-UTF-8, non-NFC, whitespace-only, >128 bytes; `ms:*` colons handled by injective one-way `%3A` encoding at the checked `ModelSpecPath` sink (raw `%` reserved so encoding cannot alias), both callers migrated, no unchecked join remains. - **`provenance add --dry-run`** [High] — was wired but never read; now honored with a no-write witness test. - **Live-runtime isolation dirs** — owned, cleaned on all paths, never claims a caller-supplied root. - **Stale eval help text** — corrected; COMMANDS.md regenerated via its owner. **Already landed (1):** the `--json`/`-o json` divergence for provenance/skills was resolved by the cmd/ao carve-out (probes confirm identical output). **Recorded OPEN with reproductions (3):** `gate check --dry-run` plan-only mode, bounded subprocess output streaming, and context/process-group cancellation — each a cross-package refactor (the audit's own G1/G2 programs), documented with fix sketches rather than half-fixed here. ## Validation - `go build` / `go vet` clean; `go test ./...` 2872+ pass across 70 packages; golangci-lint 0 issues on touched packages; CLI reference check current - Cross-family review two rounds: round 1 three findings (Windows renormalization traversal, canonicality bounds, unchecked sink) all fixed; round 2's one residual (non-injective colon encoding) fixed with witness cases Tracker: `age-skill-overhaul-reboot-sjv7v.12` |
||
|
|
c88a4514f9 |
W7 support wave: handoff schema truth, dcg fact corrections, honest support contracts (#1005)
## Summary Wave W7 of the skill-overhaul reboot (`age-skill-overhaul-reboot-sjv7v.8`) — the nine support skills, plus the one Go fix where the skill contract crosses the CLI boundary. - **handoff** — `ao session handoff --dry-run` output failed its own `handoff.v1.schema.json` (reproduced: 3 errors). Fixed with a consumer audit: schema keeps v1 with the doctrine-retired fields as optional deprecated read-compat properties; the generator keeps its collision-safe fractional id (schema pattern widened instead); real jsonschema validation in `TestHandoffDryRunSatisfiesSchema` + a legacy-artifact compat test; `read_clock` effect declared. - **dcg** — corrected the false "`rm -rf ./build` allowed" claim (live 0.5.6 blocks it) and a nonexistent rule id in the allowlist example (silent no-op) across six files; removed a token-splitting "workaround" that was an executable guard bypass, replaced with file/stdin handling and a never-reconstruct warning; temp-path rule live-probed and stated identically in both docs; version/path/upstream corrections. - **cc-hooks** — ships-by-default contradiction reconciled; PATH-clobbering recipe fixed; operator-private paths removed from shipped text; jq preflight added to the edit guard. - **ms** — validator no longer mechanically asserts the false `effects: []`; it extracts the frontmatter and requires the exact honest effects value. - **account-rotation / status / sbh / bootstrap** — real effects declared, both-tools-absent and destructive surfaces defined, live-output overclaims narrowed, versions pinned. - **shared** — advertising narrowed to the current no-bundled-references state; retirement NOT executed (bead `.11`). Ledger (32+8 fixed across two rounds / 8 rejected-stale / 8 deferred-with-reason): `docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/w7.md`. ## Validation - `go build` + `go vet` + 482 `cmd/ao` tests incl. the new schema-lock and legacy-compat tests; dry-run validates 0 errors - 49/49 strict frontmatter; regen clean; scenario-linkage PASS; liveness + anti-spiral + policy + edit-guard bats green; python ratchet no-growth; shellcheck clean - Cross-family review, two rounds: round 1 eight findings all fixed (schema compat, id collision, real validation, security bypass removal, live-probed temp rule, anchored greps); round 2 delta re-review **VERDICT: PASS** with zero residuals Tracker: `age-skill-overhaul-reboot-sjv7v.8` |
||
|
|
d9eeee9cb3 |
chore(release): prepare v3.3.0 changelog and version bump (#989)
## What Prepares the **v3.3.0** release cut (last shipped tag: `v3.2.0`, 2026-07-03). No tag and no GitHub release are created here — tagging is gated on the live GC canary and the operator's go. - **Changelog reframe.** The `[3.3.0]` headline now leads with the deliberate subtraction of the guardrail cathedral (strict verbose step contracts, enforcement gates, injection hooks, command tombstones, retired lifecycle surfaces, the heavy out-of-session orchestration layer) — removed because that scaffolding degrades frontier-class models rather than helping them. What remains is the judgment-boundary membrane: caller-owned intent, one fresh independent validation, durable verdicts, pinned provenance. Subtraction is git-concrete: `cli/internal` packages **79 → 43**; reference orchestrator **~29k lines → ~1,300-line thin pack** (one commit removed 29,418 lines for 1,334). - **Post-07-17 delta folded in:** Gas City factory pack (preview) — official checksummed GC/Beads binaries fetched by the installer (no compile), mayor-driven door (human attach or agent drive via mail/sling) with heartbeat shepherd, the `using-gc` mayor-orchestration skill + four-layer visibility doctrine, rig-scoped dispatch intake, three disclosed upstream v1.3.5 defects (#4586, the cross-store claim fix, gastownhall/gascity#3985) with the preview→supported promotion criteria, and plan/premortem ground-truth routing. - **Version bump:** source fallback `3.3.0-rc → 3.3.0` (`cli/cmd/ao/main.go`); plugin manifests were already `3.3.0`. Adds a release-guard test asserting the fallback carries no pre-release marker (resolves audit minor 1). - **Release-gate fix:** trimmed the `using-gc` skill description (195 → 117 prose chars) to satisfy the 180-char budget — a blocker that landed with the GC arc — and regenerated its Codex/Gemini/pack projections. ## Recommended version: **v3.3.0** `v3.3.0` was an in-progress RC (`main.go` at `3.3.0-rc`, plugin manifests at `3.3.0`, changelog `[3.3.0]` drafted) that was never tagged; the GC pack + loop doctrine landed on the RC line. This cut finalizes it. ## Go / No-Go checklist (deterministic bar: docs/audits/release-readiness-3.3-2026-07-20.md + docs/runbooks/release-process.md) | # | Criterion | Status | Evidence | |---|---|---|---| | 1 | `go build ./...` | PASS | exit 0 | | 2 | `go vet ./...` | PASS | exit 0 | | 3 | Full Go test suite (`go test ./... -count=1`) | PASS | 63 packages `ok`, 0 FAIL | | 4 | Cathedral Cut conformance (`check-cathedral-cut-conformance.py`) | PASS | "Cathedral Cut conformance: PASS" | | 5 | Skill-mesh drift (`generate-skill-mesh.py --check`) | PASS | "up to date (48 skills)" | | 6 | Derived-artifact drift (`make regen-check`) | PASS | "All generated projections are current." | | 7 | Docs + release-doc gate (`make docs-check`) | PASS | CLI ref current; 395 links, 0 broken; skill count 48; release freeze intact | | 8 | `ao gate check --full` | PASS* | 66/67 pass. Sole fail = `workflow.install-drift`, a worktree-local symlink artifact (tracked `workflows/` is byte-identical to origin/main; passes in a clean CI clone — the audit's 67/67) | | 9 | Fast release gate (`ci-local-release.sh --quick`) | PASS | "LOCAL CI QUICK SANITY PASSED": 44/44 install-surface smoke, 33/33 release smoke, init smoke | | 10 | Skill token budgets / lint | PASS | 100 budgets pass (fixed `using-gc` 195→117 prose) | | 11 | Install story smoke | PASS | README/MIGRATION document `ao skills link`; all 6 curl installers are refusing tombstones; `ao version` runs; npx-first | | 12 | Plugin manifest ↔ corpus parity | PASS | manifests carry 48 skills, 0 drift | | 13 | Version bump applied to owners | PASS | `main.go` 3.3.0-rc→3.3.0; manifests already 3.3.0 | | 14 | Changelog updated + synced | PASS | `CHANGELOG.md` == `docs/CHANGELOG.md` (changelog.sync gate) | | 15 | Command/test pairing gate | PASS | main.go bump paired with `version_test.go` release-guard test | | 16 | Full release gate (`ci-local-release.sh`: race + security + SBOM + readiness) | PASS | "LOCAL CI PASSED [102s]"; release-artifact manifest resolved `v3.3.0` | \* `workflow.install-drift` is an environment artifact of running the gate inside a nested worktree whose `.claude/workflows/` symlinks resolve to the primary checkout. It is not release content: `git diff origin/main -- workflows/` is empty, and the check passes in a fresh CI checkout (audit reported 67/67 at origin/main). ## Operator items (require a human or a live tag — NOT attempted here) - **Live GC canary** on the next official Gas City pin — the preview→supported promotion gate. - **Git tag `v3.3.0` + goreleaser publish** — gated on the canary and operator go. - **Docs-site publish** (mkdocs pipeline) — separate from this cut. No tag. No release. |
||
|
|
cbe7ddbb5f |
feat(gc): ship thin native AgentOps factory pack (#980)
## Outcome Ships AgentOps 3.3 as a thin native Gas City pack and removes the accidental second orchestration control plane. - Removes the custom GC delivery command/package, packet/schema family, feeder/program engine, reducer/order, reliability registry, Beads capability mirror, and fork-baseline runtime checks (about 29k deleted lines). - Retains six bounded adapters for pinned toolchain materialization, clean bootstrap, native sling/status/doctor invocation, bead-isolated worktrees, moving-main PR delivery, and teardown. - Uses official Gas City and Beads behavior as the state owners, with required OTEL configuration. - Preserves Fable Mayor/Refiner, Sol-high planning and fresh validation, Terra-high default implementation, Opus-medium overflow, and support-only Luna. - Supports automatic Refiner merge after hosted CI or a manual-review toggle without locking `main`. ## Evidence - Exact official GC: `8ffc009ded781a2ada2077f3a29bd712b2def0bf` - Exact official BD: `8e4e59d39f3459a43cf21a3236a13eca4dd874f7` - Full pinned native boundary: 9/9 passed in 114.047s - Exact source-bead route replay: passed in 98.051s - Release replay: manifest integrity, thin tests and both pack lints, generated projections, contract compatibility, Bats contract, ShellCheck, test-removal ratchet, and `go test ./...` all passed - Fresh independent semantic verdict: PASS (Sol-high) - Exact delivered head: `e805f0e26ce21c2eae9e720f57fff2a705f3185b` - Subject manifest: `8554d065076f70d0e639ca353409d2a75dca7ac01e4cb65cb4a118e98c0770fc` - Verdict: `3f89376b675de854832f38b46ccfa7ad7ca54079ffde578769c33e1f118c68c1` ## Release boundary After this PR merges, the release qualification runs one fresh mixed Terra/Opus canary from merged `main`, with at most one external repair and one terminal retry. No live city repairs itself. The only known local fast-gate exception is `workflow.install-drift`: the installed workflow symlink resolves to the separate dirty primary checkout, outside this candidate. All candidate-owned gates passed. |
||
|
|
faff55ac2a |
feat(gc): bind native 3.3 factory workflows (#972)
## Summary
- bind the GC 3.3 bead-native workflow: Fable Mayor, fresh Sol-high plan
and validation, Terra-high default or Opus-medium overflow
implementation, Luna support-only
- add the bounded one-shot graph feeder, strict packet/worktree
identity, semantic terminalization, and model-free protected delivery
sweep
- make official Gas City v1.3.5 plus Beads v1.1.0 bootstrap repeatable,
disable unstable event propulsion/hooks, and prove quiescent teardown
## Reliability boundaries
- no Gas City self-repair loop, daemon, private scheduler, model
fallback, or main mutex
- semantic completion is independent of delivery; protected
PR/CI/rebase/merge is deterministic and moving-main aware
- event hooks are disabled because 3.3 uses explicit program admission
and cooldown delivery; native controller bead observation remains active
## Evidence
- full local release CI PASS at
.agents/releases/local-ci/20260722T113355Z
- two pristine standalone-clone cycles with official GC v1.3.5 8ffc009d
and BD v1.1.0 8e4e59d: repeat bootstrap, exact rig Dolt context, absent
event hooks, bounded events, enabled sealed delivery, and delayed
zero-process teardown
- fresh independent Sol-high pre-delivery binding PASS over exact head
|
||
|
|
4d631207cf |
feat(gc): add bead-native crash-only delivery kernel (#970)
## Summary - replace the unreachable Python factory lifecycle with an optional typed Go delivery reducer - make Beads the delivery lifecycle authority with deterministic moving-main successor epochs - add native Git/worktree, PR, hosted-check, auto-merge, landing, and cold-replay boundaries - retire obsolete v1 role/factory schemas and preserve the pinned historical capability harness ## Validation - fresh Sol-high PASS over exact 43-path manifest ec6a2a4bf0b180782411b1ae4afd4a2ad8384129ff037ce8c119480317626b02 - go test ./... - go test -race ./internal/gcadapter/delivery - go vet ./... - scripts/check-gc-executor.sh - GC 3.3 schema, migration, provenance, factory-doctor, and bootstrap gates Bead: age-gc-scope-failclosed-release-gktia.1 |
||
|
|
26120d8394 |
feat(gc): add crash-only delivery thin slice (#969)
## Outcome Adds the bounded GC33-6 delivery slice outside core ao: a separately built crash-only reducer, strict admission certificate binding, immutable handoff and branch/PR receipts, native graph.v2 formulas, one-step disabled Order glue, and one centralized stock claim wrapper. Semantic bead: ag-agentops-33-gc-refinery-4km81.7 (closed after independent validation). Binding PASS: d3ac5a566e0f5197814e2dbcf04d0540849f58636ed8bb49a0faca7d6af94639 over manifest 155a089b4b461bda61a7f1ad1987c66f5755ac7c5628aa648e70740be2580a5a. ## Verification - Full cli Go suite: PASS (2,887 tests) - Focused race-enabled reducer suite: PASS - GC factory/executor Python contracts: PASS - Full executor gate: PASS (123 Python + 25 Bats) - Exact official GC v1.3.5 pack lint and no-start discovery: PASS - Cold replay: exactly one delivery, branch, and PR; six emitted artifacts validate against shipped schemas - Production Go surface: 922 lines, under 3,000 tripwire ## Boundaries No ao command, core port, default install, live city, real Git/forge/CI/merge adapter, daemon, scheduler, or factory.py import. Real adapters and delivery completion remain GC33-7. |
||
|
|
e5ad3cfc94 |
test(cli): guard advertised ao commands against the live cobra tree
Walks every registered command Short/Long/Example plus quick-start printed output, extracts advertised "ao ..." strings, and resolves each against the assembled rootCmd (aliases included; Example fields get full group-descent). Guards the defect class the 3.3 release audit fixed by hand: user-facing output advertising a retired command. |
||
|
|
238964fdd2 |
feat(cli): workflows are canonical product artifacts — workflows/ + ao workflows link (#945)
Workflows get the skills treatment (operator decision): canonical source in the product tree, installed by a product verb, Claude-only labeled as such. **What moves:** all seven Claude workflow scripts + README migrate from force-added exceptions inside the gitignored `.claude/` to a tracked top-level `workflows/` (sibling of `skills/`) — the four existing conveyors plus `audit-dimensions`, `verify-fixes`, `implement-wave`: three thin, args-parameterized orchestration conveyors extracted from this session's hand-rolled waves, contract-reviewed, and smoke-proven through the real Workflow runtime (the smoke caught two contract gaps static review could not: an `export default` wrapper the runtime never invokes, and args arriving as a JSON string — both fixed, string-args tolerance now built in). **New verb:** `ao workflows link` / `unlink` mirror `ao skills link` semantics — dry-run `--json`, refuse to replace real files or foreign links, unlink only checkout-owned links — targeting the project-local `.claude/workflows/` where Claude Code resolves named workflows (`--into` overrides). Checkout identity reuses the skillsapp marker discipline, fail-closed. Claude-only runtime adapter, same doctrine as the Codex-only `skills-codex/`. **Legacy surfaces repointed:** `install-workflows.sh` (user-global $HOME installer), `check-workflow-drift.sh` + gate comment, `check-bdd-foundry-markers.sh`; spine allowlist + YAML-probe excuse + go-cli.md spine region gain the workflows group; COMMANDS.md, cli-surface projections, and surface-count fixture regenerated; new tests carry per-command git-env scrubbing (test-isolation ratchet back at baseline). **Built BY the workflow being canonized** — `implement-wave` orchestrated its own canonization: two disjoint-ownership lanes plus a seam-checking verifier that ran the real binary's link → resolve → unlink cycle in the live tree (both lanes RESOLVED). The lanes correctly *refused* to self-approve their command into the spine invariants and handed integration three flagged edits instead. **Expected local gate note:** `workflow.install-drift` correctly FAILS on machines whose user-global `~/.claude/workflows` links still point at the old location — that is the transition it exists to catch. CI stays green (absent→skip). **Post-merge operator step:** `cd ~/dev/agentops && git pull && bash scripts/install-workflows.sh`. **Verified:** full suite 63/63 pkgs; golangci-lint clean; `gate check --full` over this range = 66/67 with only the documented install-drift environment finding; workflows smoke-run evidence in session logs. |
||
|
|
3d24ee0e9d |
fix(cli): residue wave — doctor dev-version coherence, diff --only, config-models removal, dual-root stragglers (#944)
Wave 4 (residue) of the new-user happy-path arc. Three scoped implementer lanes + fresh adversarial verifier (4 RESOLVED; 1 INCOMPLETE = stale generated projections, closed in integration). **Doctor**: the dev-version detector now reuses the Binary Freshness resolution — a from-source build matching its checkout is healthy (a novice building from source can finally see `ao doctor` exit 0); findings fire only on genuine drift, shadowed duplicate `ao` binaries, or an informational from-source note outside any checkout. `ao doctor diff` gains `--only` so the fix-plan preview can be scoped the way remediation text implies. **Config**: the dead `ao config models` surface is removed end-to-end (lane re-verified zero consumers before deleting; `--show` proven byte-identical before/after; removed-child hint + MIGRATION row; existing `models:` config sections still parse and are ignored). **Dual-root stragglers**: learning-coherence gate globs, `quality.CountConstraints`, and the eval sandbox corpus deny-list now cover canonical `.agents/ao/<section>` alongside legacy roots. **Doc-link hygiene**: the strict docs-link backstop's allowlist was 100% stale (53/53 entries referenced Cathedral-Cut-deleted docs) — refreshed to 10 verified accepted-class entries; ROADMAP dead links fixed; documentation-index generator emits GitHub URLs for repo-root targets; doc-skill references instruct only shipped scripts; codex twins + CLI-surface projections + surface-count fixture regenerated. **Deferred by design**: the `3.3.0-rc` fallback version bump belongs inside the v3.3.0 tag-cut commit. **Verified**: full suite 61/61 pkgs (2825 tests); golangci-lint clean; `ao gate check --full` 67/67 over this range; Test-Removal-Reason trailer covers the 6 deliberately deleted models tests. |
||
|
|
511370f5d9 |
fix(cli): doctor coherence, config truthfulness, pruned-verb hints — novice edges 4/6/7/8 (#942)
Wave 2 of the new-user happy-path fixes (fresh-eyes novice test). **Edge 4 — doctor sub-surfaces contradicted each other.** Remediation, `next_steps`, and robot-triage's `recommended_command` now instruct `--fix` only when a fixer can actually act — non-fixable findings name their real manual action (including the detect-only bridges class the verifier caught as the same lie one level up). `ao doctor health` reports every severity bucket present (it omitted P1 — the worst actually present). `ao doctor diff` renders an explicit read-only fix plan (would auto-fix vs manual action) instead of a findings copy. `ao doctor explain` is now a superset of the finding's triage entry. **Edge 6 — `ao config --show` rendered config for commands that don't exist.** `rpi.*` and `dream.*` no longer serialize in human or JSON output (struct fields retained `json:"-"` pending a follow-up removal once nothing references them). `AGENTOPS_NO_SC` had zero behavioral consumers (evidence: rg over all non-test cli/ — only config plumbing/display) — undocumented and removed from the env panel; the field stays parseable so existing config files don't error. **Edge 8 — legacy-config gaslighting.** With only `~/.agentops/config.yaml` present, `--show` used to print the deprecation warning and then claim the new path was "(not found)" and values came "(from flag)". It now shows the actually-read path labeled "(deprecated location)" with correct attribution. **Edge 7 — `ao beads` bare error.** The whole pruned 3.2 bookkeeping family gets point-of-failure migration hints; a new guard test proves no hinted verb is a live command (the `eval` trap: pruned in 3.2, returned live in 3.3). **Process:** three scoped implementers (disjoint files) + fresh adversarial verifier. The verifier refuted two lanes as incomplete (recommended-command layer, stale cmd/ao config tests asserting the old defect) — closed in integration before this PR. **Verified:** full suite 61/61 pkgs; golangci-lint clean; `ao gate check --full` 67/67 over this range; per-fix RED→GREEN sandbox repros in the workflow logs. |
||
|
|
d043390a0a |
fix(release): resolve 3.3 release-wrapper audit — blocker + 13 majors (#935)
Resolves every spellbreaking finding (the blocker + all 13 majors) from the 3.3.0 release-readiness audit ([docs/audits/release-readiness-3.3-2026-07-20.md](docs/audits/release-readiness-3.3-2026-07-20.md), included in this PR). ## Finding → fix map **CLI self-documentation (M1–M3)** - **M1** `ao robot-docs` prescribed removed `ao inject` → line removed from the canonical agent workflow; `inject` added to the removed-command hint table **and** the MIGRATION.md map (drift test `TestRemovedVerbsHaveMigrationRows` enforces the pair). - **M2** `ao config --help` documented ~14 env vars for removed subsystems (RPI/Dream/Council/tiers) → help text and the `--show` env panel pruned to the 5 vars the binary consumes; mirrored list in `internal/config` pruned identically. - **M3** `ao flywheel status` read only legacy `.agents/<section>` while `ao doctor fix` migrates learnings to canonical `.agents/ao/learnings` → new `quality.KnowledgeSectionDirs` dual-roots every knowledge reader (tier counts, new/stale artifacts, retros, health delta, utility, loop metrics, retrievable-citation stats — plus the golden-signals readers `ComputeResearchClosure`/`ComputeReuseConcentration` that the fresh verification pass caught as missed). Sandbox-proven twice: a learning existing only under `.agents/ao/learnings` appears in all metrics, and a research file only under `.agents/ao/research` flips closure from `starved/0` to `unmined/1 orphan`. **Release story (M4–M6)** - **M4** CHANGELOG `[3.3.0]` omitted post-07-17 surfaces → folded in `ao eval` (#921), default-build `ao flywheel`, the PreToolUse policy engine, and the #919 cleanup; date moved to 2026-07-20; `docs/CHANGELOG.md` re-synced (changelog.sync gate green). - **M5** MIGRATION.md attributed `ao eval` to a nonexistent "3.4" → now "returned in 3.3". - **M6** four release surfaces claimed a 50-skill corpus vs 48 everywhere real → all counts now 48 (CHANGELOG ×2, docs/3.3.md, release-notes page ×2). **Install story (M7 — product decision by Bo)** npx first (universal — installs into all coding agents), **plugins for Claude Code/Codex encouraged**, checkout + `ao skills link` as the source-tracked/contributor path; curl installers stay tombstones. Harmonized across README, UPGRADING, install-day2-ops, MIGRATION, 3.3.md, CHANGELOG, the release-notes page, all six installer tombstone messages (`install.sh` + claude/codex/agy/opencode/`codex.ps1`), and the site's CLI page. All "legacy migration-only / not the recommended path" plugin branding removed. **Docs site (B1, M8–M12)** - **B1** generated site CLI page instructed a tombstoned curl installer, nonexistent `ao rpi phased`, and wrong skills dir → `emit_index()` rewritten to the real install menu + a quickstart of commands that exist; semantic loop correctly attributed to skills. - **M8** deploy workflow's `--strict` contradicted mkdocs.yml's declared non-strict policy and aborted the build → flag dropped (lychee + validate-links.sh own link checking). - **M9** site banner said "AgentOps 2.x" → now 3.3. - **M10** ~176 internal files (audits/plans/handoffs/… + TEMP scratch doc) published and dominated search → `exclude_docs` extended; built site verified free of them; search index 4,751 → 2,444 entries. - **M11/M12** newcomer-guide skill links and all six SCHEMAS.md links 404'd on the site → absolute GitHub URLs (resolve on both GitHub and the site). The fresh verification pass found the same class on `docs/contracts/index.md` (nav-listed), `docs/contracts/corpus-learning-seam.md`, `docs/templates/slice-validation.md`, and five `docs/architecture/gas-city-factory.md` links into now-excluded `docs/audits/` — all repointed to absolute GitHub URLs. - Also from the verification pass: robot-docs exit-code table no longer says "bead claimed" (removed concept), and the docs.yml comment now cites the link checker that actually runs (`tests/docs/validate-links.sh` via doc-release checks) instead of lychee. **Skills corpus (M13)** - rch skill instructed 5 nonexistent scripts + 8 nonexistent reference files as recovery steps → pruned to the 7 real references; the wire-level `printf | rch` probe replaces the phantom `protocol_test.sh`; codex twin regenerated (parity gates green). ## Verification - `go vet` clean; **full test suite 60/60 packages pass** (includes the new-shape flywheel/quality/config tests and the inject↔MIGRATION drift test). - **`ao gate check --full --scope worktree`: 67/67 pass, 0 warnings** over this exact change set (changelog sync, shellcheck on the edited installer, skill mesh + codex parity, manifests/schema/triggers, provenance chain). - `scripts/regen-all.sh --check`: all generated projections current. - Rebuilt binary re-exercised: `robot-docs` clean, `ao inject` tombstone live, config help clean, flywheel sandbox proof above. - `mkdocs build` exit 0; warnings 84 → 46 (remainder is the accepted out-of-tree-link class per mkdocs.yml's declared policy). - Fresh-context adversarial verification workflow over all four fix groups (results in session log). ## Known residuals (deliberately out of scope) - `ao config models` subcommand still renders tier config (its two env vars ARE consumed; `COUNCIL_CLAUDE_MODEL` in its display list is not — follow-up). - Same-class single-rooted readers off the flywheel path: `learning.coherence` gate glob (`.agents/learnings/**` only), `quality.CountConstraints`, config `Paths` defaults feeding eval sandbox deny-lists. - `scripts/docs-build.sh` still uses `--strict` with its own allowlist (not wired into any workflow or gate). - Audit minors 1–8 (rc fallback version string, `config --show` legacy-fallback display, `ao init` vs doctor layout, doctor's `ao beads dir` hint, dead `tracker:` key in the example config, doc-skill phantom scripts, ROADMAP dead links, documentation-index root links). |
||
|
|
5303a816e3 | fix(cli): gate failure stderr silenced; all modules on shared seam or reasoned-exempt; guard requires coverage (age-6j9ee.1) | ||
|
|
b4aa2f0d8b | fix(cli): -o yaml is real everywhere json is — shared WriteYAML, zero silent fallbacks (age-6j9ee.2) | ||
|
|
3a54d5ff53 |
refactor(cli): unify host-seam contract, shared writeJSON/ExitError, drift guards (age-6j9ee.1)
Collapse the four drifted cmd/ao host-seam shapes into one shared clicontract.HostOptions consumed by all 17 command modules; delete the positional-func and bespoke-struct seams and the doctor GlobalOptions bundle. - clicontract/host.go: single HostOptions union (OutputMode/Verbose/Verbosef/ DryRun/ProjectRoot/GoalsPath/LedgerPath/Version/Now/EnrichFlagErr), one shared WriteJSON (byte-identical to the three deleted copies), one shared ExitError (Label field preserves doctor's stderr surfacing; gate stays silent). - config_module.go + doctor_module.go: package-level module/service singletons -> newConfigCommand()/newDoctorCommand() constructors; 7 allowlist rows removed. - carveout_regression_test.go: three AST drift guards (shared seam type / no direct host effects / no module-service singleton vars), each RED-proven. - docs/architecture/go-cli.md: document the intentional two-tier shape (5 full hexagonal, 12 app-seam); drop the stale opportunistic-extraction sentence. Zero behavior change: ao -o json capabilities byte-identical; full suite + -shuffle=on -count=2 green; 0 lint issues; gate check --full 67/67. |
||
|
|
735580d1c7 |
test(cli): durable package-var allowlist invariant + root-only harness comments
Finish line of the cmd/ao carve-out arc (age-a-plus-report-card-ieyp2.14). Changes NO command behavior — tests, guard, and allowlist only. - Add TestPackageVarsAreAllowlisted (carveout_regression_test.go): parses every cmd/ao non-test .go file with go/ast and fails on any package-level var that is not on testdata/package-var-allowlist.json, and on any stale allowlist entry. Keyed by name+file. Locks the finish-line state: after 12 family carves the only legitimate package globals are the root spine + its 5 persistent-flag targets, host-resident module/command wiring, const-like data maps, the ldflags version string, and one documented test seam (testProjectDir). - testdata/package-var-allowlist.json: 24 entries, each with a one-line reason. - Bounded-purpose comments on executeCommand and resetCommandState documenting that they save/restore ONLY still-existing root spine state — carved families own their flag state constructor-scoped in internal/commands/<family>. No dead vars found to delete: all 24 cmd/ao package vars are live. Harness was already root-only after prior waves (saves only dryRun/verbose/output/jsonFlag/ cfgFile). RED proof: a scratch `var sneaky string` in group_json.go trips the new guard; removed before commit. |
||
|
|
b390ad782a |
refactor(cli): carve flywheel into internal/commands/flywheel + internal/flywheelapp
The command module is a thin Cobra presentation seam; internal/flywheelapp owns the filesystem and clock effects and preserves the internal/evidence (ratchet.LoadCitations) dependency. flywheel status/compare surface and capabilities are byte-identical. |
||
|
|
54b462a855 | refactor(cli): carve redact into internal/commands/redact module | ||
|
|
bc4ab7970f | refactor(cli): carve quick-start into internal/commands/quickstart module | ||
|
|
c5522feebc | refactor(cli): carve robot-docs into internal/commands/robotdocs module | ||
|
|
6c938b863e | test(cli): extend AST carve-out guard to session+demo+init+version | ||
|
|
02045e6c25 |
refactor(cli): carve version into internal/commands/version module
Move the version command out of package main into internal/commands/version. version reads only build/runtime metadata (a pure effect), so it needs no app seam; the build-time version string and the global -o/--output mode are host seams injected by the composition. Unlike the other W3 families, version carries a real CommandContract that newVersionCommand attaches to the command tree, so the capabilities surface is byte-preserved. The build-time `version` string var stays in package main (the ldflags target `main.version`, shared by capabilities, doctor, and rootCmd.Version). The command-behavior tests move to the module (constructing it directly); the root-wiring tests (`version` var default, root registration, `ao --version` flag) stay in package main because the module does not own that wiring. Test-Removal-Reason: TestVersion_CommandExists removed; its Use/GroupID assertions are preserved as version module TestModule_CommandAttributes. Test-Removal-Reason: TestVersion_DevVersionDefault removed; it duplicated the version-string assertion now covered by version module TestVersion_ExecuteOutputContainsVersionString. |
||
|
|
0c8d2a4ab6 |
refactor(cli): carve init into internal/commands/init module
Move the init command out of package main into internal/commands/init, delegating the working-directory resolution and directory creation to the new internal/initapp seam. The dry-run selection is a host seam injected by the composition from the global --dry-run flag; the module performs no direct filesystem effect. |
||
|
|
1733907f24 |
refactor(cli): carve demo into internal/commands/demo module
Move the demo command out of package main into internal/commands/demo. Demo renders static explanatory text and performs no effect, so it needs no app seam; the render helpers and constructor-scoped --quick/--concepts flags live in the module. demoQuick and demoConcepts package globals die; the shared resetCommandState test helper drops its references to them (cobra flag reset still covers the live command instance). |
||
|
|
8c1076fbca |
refactor(cli): carve session into internal/commands/session module
Move the session evidence commands (bootstrap, rehydrate) out of package main into internal/commands/session, delegating filesystem effects to the new internal/sessionapp seam. The session parent command now lives in the module; the optional `ao session handoff` writer stays in package main and is attached to the module-built parent by newSessionCommand. sessionBootstrapJSON and rehydrateJSON package globals die; flag state is constructor-scoped in the module. The combined TestHandoffAndRehydratePreserveCallerTextWithoutLifecycleState is split: the handoff-writer assertions stay in cmd/ao as TestHandoffPreservesCallerTextWithoutLifecycleState; the rehydrate read side moves to the module as TestRehydrateReadsCallerAuthoredBrief plus the preserved TestRehydrateJSONEmptyStateEmitsEmptyObject and TestSessionBootstrapOnlyReportsLocalOrientation. Test-Removal-Reason: TestHandoffAndRehydratePreserveCallerTextWithoutLifecycleState removed; its handoff-writer coverage is preserved in cmd/ao and its rehydrate coverage relocated to the session module (net +2 tests). |
||
|
|
f7025acc22 | test(cli): extend AST carve-out guard to provenance+skills | ||
|
|
29948e6fee | refactor(cli): carve skills into internal/commands/skills module | ||
|
|
549dd75f85 | refactor(cli): carve provenance into internal/commands/provenance module | ||
|
|
c48b561b70 | chore(lint): enable errorlint + unused, burn backlog (age-a-plus-report-card-ieyp2.8) | ||
|
|
3069bb695f | test(cli): AST regression guard for goals+status carve-out | ||
|
|
48f1a0ce21 | refactor(cli): carve goals into internal/commands/goals module |