1372 Commits

Author SHA1 Message Date
Bo 20f9d4a338 Release AgentOps 3.7.0 (#1144)
## What

Release the prepared AgentOps update as **3.7.0**, the minor release
after 3.6.0. Align the CLI and plugin versions, regenerate the Gemini
manifest, and rename/update the curated notes and changelog links.

## Why

The operator selected a minor release. No 4.0.0 tag or release was
published. Migration instructions and the documented removed
commands/skills remain accurate.

## How I tested

- Go lint and focused version/manifest tests passed.
- Full regeneration parity, changelog mirror parity, and release-note
coverage from v3.6.0 passed.
- The exact 3.7.0 release rehearsal passed in 143 seconds; all 73 full
repository gates passed. All 12 security tools ran with zero
missing/error tools, critical findings or security-high findings;
existing advisories remain reported.
- All nine hosted checks passed on
`092e1814b6cba46cd9ac1d797dab2a5c8c7c188c`, including Go race/shuffle
tests and 1,509 executed Bats passes (31 environment-dependent skips,
zero failures). A new CLI wiring regression confirms `ao version --json`
reports the build version.
- Actual fresh native Claude/Codex 3.7.0 installs and upgrades from
3.6.0 passed with exact 34-skill inventories. Existing implementation
validation from PR #1143 remains applicable to unchanged source.
- Fresh author-distinct correction review passed for exact head
`092e1814b6cba46cd9ac1d797dab2a5c8c7c188c`, covering all changed paths
and four acceptance criteria with no unchecked scope. Verdict digest:
`e7b24a297a0b232138de011473df8be6eb3eeaf47afd1f300398eecafab2fbab`.

## Checklist

- [x] Version owners and generated metadata agree on 3.7.0.
- [x] Migration/removal guidance is preserved.
- [x] Exact-candidate release checks pass before tagging.
- [x] Fresh correction review is recorded before tagging.
2026-09-13 20:28:12 -04:00
Bo d972fa2090 Prepare AgentOps 4.0.0 plugins, skills and CLI release (#1143)
## What

Prepare AgentOps 4.0.0 across the Claude plugin, Codex plugin, skills
and CLI. Claude writers capture the supplied check status during its
original invocation, and plugin conformance verifies exact skill
membership and link destinations. Full release security now scans the
repository and blocks on Python collection failures that previously
produced a false green result.

## Why

The 3.6.0-to-current interval removes published commands and 20 skill
names, so this is a major release with migration instructions. Release
validation also exposed stale skill assertions and test prerequisites
that need to match the current product contracts without weakening
acceptance.

## How I tested

- Native Claude Opus/Haiku success, failing-check and direct-writer
trials: each check ran once, and the direct child returned plain JSON.
- Actual fresh installs and upgrades from 3.6.0 in isolated Codex and
Claude homes: 34 skills, expected agents, and exact installed package
bytes.
- Exact candidate `b721d02559e1495be6095ad97b820e88ceb4a049`: all 73
full repository gates, regeneration parity, and the complete local
release rehearsal passed. All 12 security tools ran with zero skips,
tool errors, critical findings or high-severity security findings. The
unchanged advisory policy reports 35 quality-high findings on unchanged
files.
- Python: 327 tests and 72 subtests passed. Hosted Bats: 1,509 passed,
31 environment-dependent skips, zero failures. Go
lint/build/vet/race/shuffle checks and CLI smoke/integration passed.
- All 11 hosted checks passed, including Windows correctness,
macOS/Linux installation, security, and the six-target no-publish
GoReleaser snapshot. Local archive checksums and a real macOS CLI
initialization/status/version smoke also passed.
- Fresh author-distinct review passed all four acceptance criteria and
all 35 changed paths with no unchecked acceptance. Canonical subject and
caller-intent verification passed; verdict digest
`68af2c935ed0106cd91b3950f5d168e662f4071f660fcbd113c36b7cd0f0426e` binds
manifest
`7affc77e25eaff69ba36c5ce05582b4f0385c954b76b62c02b97f97041f489b2`.

## Checklist

- [x] Breaking changes documented in the migration guide and complete
release notes.
- [x] No credentials or private runtime proof included.
- [x] Final full release checks pass on the exact candidate.
- [x] Fresh author-distinct final PASS is recorded before merge.

This prepares the release candidate; it does not publish a tag or
release.

Coverage limits remain explicit: native plugin tests used isolated macOS
homes and local marketplaces, guard installation remains opt-in, and
reader instructions do not prove sandbox confinement. Semgrep retains
pre-existing warning-level parser diagnostics. Snapshot metadata follows
the existing 3.6.0 tag; this is a packaging rehearsal, not a published
4.0.0 archive.
2026-09-13 17:21:16 -04:00
Bo 3213afcf1c Default to native execution and report independently accepted work (#1129)
## Change

Make native coding-agent execution the default AgentOps entry path with
zero mandatory skills. Preserve full bundles and add repeatable `ao
skills link --skill NAME` selection, validating the entire selection
before writes. Align product, installation, architecture and generated
command documentation.

Extend the existing trial readout to separate endpoint test results,
execution state and independently accepted work. Bind supplied judgments
to exact content, acceptance and native evidence. Reject empty
implementation subjects and require the caller's complete criterion ID
set before reporting acceptance. Preserve genuine nonempty and
deletion-only subjects, valid failures and missing-proof outcomes.

## Validation

- Native onboarding from empty home/consumer directories produces no
setup files; selective/full linking and failure boundaries are covered.
- Actual RED/GREEN regressions cover empty subjects and the
partial-criterion omission found by independent review.
- Full Go build, vet and race/shuffle tests; affected Go lint; 88 Python
readout/statistics tests passed.
- All 73 gates, generated projections, strict documentation build and
local aggregate passed (10 passed; one documented optional absence).
- All nine PR checks succeeded at
`7df0d42b12f35ffc22008cc10a40339afcfbb6a0`.
- Fresh author-distinct review passed all six acceptance criteria over
all 59 changed paths, with no findings or unchecked scope, after
repairing the criterion-coverage finding.

## Evidence limits

The real native coding repair demonstrates usability, not comparative
skill uplift. The strict live-session machine replay remains NOT_PROVEN
where execution/identity observations are unavailable; the source review
PASS is retained separately. Existing cohort limits and the historical
aggregate-enforcement gap remain unwaived. No new comparative cohort,
scheduler, skill-corpus deletion, memory migration or global
installation is included.
2026-09-10 20:28:22 +00:00
Bo 5e874b55cf Evaluate installed skills on isolated Go work (#1125)
AgentOps previously relied on behavioral probes and retrospective
summaries to assess skills. This adds a development-only evaluator that
runs a frozen installed skill package on isolated Go tasks, preserves
failed and interrupted attempts, and rebuilds a comparison readout from
native results without another model call.

The suite contains six task families, separate executable verifiers,
frozen launch identities, native session accounting, and a focused
`skill-eval` maintenance workflow. The readout separates passing code
from completed trials, retains incomplete cost information, and reports
missing evidence without claiming equivalence or uplift. `ao eval`
remains retired; no new runtime controller or required core skill is
introduced.

Validation: Go build/vet/race checks and all repository CI passed.
Focused reader/statistics, receipt integrity, verifier integrity,
fixture calibration, generated projections, and the local aggregate
runner passed. A real two-variant Docker preparation check verifies that
frozen worker and verifier images survive later staging.

The bounded coding pilot retained all 24 starts and produced eight
comparable pairs across six task families, with no observed paired
endpoint difference. The separate eight-start memory experiment did not
demonstrate incremental benefit and does not promote another guidance
rule. Individual runtime limits were enforced; aggregate desktop
deadline enforcement remains unproven. Raw trial evidence and
credentials stay outside Git.
2026-09-10 13:29:10 -04:00
Bo 8a9a01a70a Fix Go recovery, gate routing, evidence, and handoff defects (#1123)
## What

Repair 13 audited Go CLI defects across doctor recovery, gate routing,
evidence ingestion, and session handoffs. Doctor preserves recoverable
snapshots and reports unresolved findings; gates retain exact
changed-file scope and advisory semantics; malformed evidence fails
closed; handoffs report observed Git state and chronological recency.

## Why

These failures could overwrite recovery data, skip required checks,
admit malformed evidence, or restore stale context. Each defect has a
regression witness. Existing skill contracts remain unchanged for this
bounded workflow evaluation.

## How I tested

- Regression witnesses failed before repair and pass after repair.
- Combined Go build, vet, race/shuffle tests with coverage, coverage
floor, pinned lint, and complexity checks pass.
- Generated projections are current; local aggregate reports 10 passed,
0 failed, 1 optional skip.
- All 52 selected gates pass. Authoritative Linux/Windows correctness,
security, full registry, and installation CI pass.
- Fresh author-distinct OpenAI/Codex review verified all 13 repairs and
all 33 changed paths on `84029ee533be0a73ac6c882eb9cf7369d52a0055`,
including independent critical race regressions; no findings or
unchecked acceptance.
- Directory reverse moves retain the existing advisory-lock concurrency
boundary; this does not claim exhaustive hostile filesystem-race
coverage.

## Checklist

- [x] Go build and tests pass, including race detection
- [x] No secrets or credentials added
- [x] Compatibility behavior documented in the changed contract where
applicable
2026-09-10 00:13:04 -04:00
Bo 17849bbc24 Improve CLI checkpoint recovery, evidence status, and skill search (#1120)
A failed mining-checkpoint write could truncate the saved watermark and
cause retry to replay older events. The CLI also wrote evidence to
external roots that status could not inspect. This batch fixes those
behaviors and removes duplicate normalization from skill search.

- Mining checkpoints use the existing atomic storage writer. A real
partial-write regression test proves old bytes survive and retry retains
stable event IDs. Existing mode bits are preserved; new checkpoints use
0600. Symlink and special-file destinations are rejected before reading.
Before replacement, an empty same-directory probe checks ownership and
permission metadata, including ACLs and inherited permissions.
Unverifiable or different metadata returns an error and leaves the prior
checkpoint intact. This is a conservative refusal, not ACL migration. A
directory-sync error after rename can leave the new state visible.
- `ao status --evidence-root PATH` inspects an explicit existing non-Git
store, with matching text/JSON/YAML reports, no fallback on invalid
roots, and no reads through evidence symlinks. Omitted-flag behavior
remains unchanged.
- Skill-query normalization has one implementation, preserving
repetition versus first-occurrence semantics. Nine fixed shipped-catalog
queries remain byte-identical against a source-pinned baseline.

Validation: Go build, vet, tests, race/shuffle with atomic coverage,
repository Bats and aggregate suites, regeneration, applicable gates and
lint passed. Final Linux and Windows correctness CI and all required
checks passed. A fresh author-distinct review verified every acceptance
criterion across all 23 changed paths with no unchecked scope. Nine
production-query outputs match the source-pinned baseline.

The first CI attempt exposed a test-child coverage flush under its
temporary file-size limit; the test now restores that limit before exit.
Fresh review then exposed ACL loss despite green CI. The permission
guard and native regression tests repair that defect while preserving
the original access requirement. The failure cases and repair costs are
retained in the evaluation.
2026-09-09 18:48:58 -04:00
Bo db1a0573ea Add bounded source reads and pinned OKF profile checks (#1114)
Adds two explicit read-only operations for the context delivery
lifecycle: bounded raw source reads with reversible bytes and integrity
checks, and structural checking of the pinned AgentOps OKF page profile.

Source reads require independently selected context policy and enforce a
measured serialized-output bound before emitting content. Emitted bytes
do not establish host delivery or understanding; restricted-source
processing remains unavailable without native enforcement. The OKF
checker rejects missing status and incompatible profiles, and never
grants truth, disclosure, or usefulness approval.

Validation: focused tests and Linux/Windows source-reader builds passed.
The combined candidate is undergoing the required full repository checks
and fresh independent review before landing.
2026-09-09 09:37:34 -04:00
Bo 57ece9fb7b Restore private context routes and verify native judgment receipts (#1112)
Add explicit, recoverable private context routing through `ao config
context`, binding native source, owner, task, model and destination to
existing policy and external storage. Recovery reads the original Beads
maintenance anchor; configuration reports native access enforcement as
unattested.

Add `ao provenance verify-judgments` to check required review profiles
against exact native transcript receipts, independent subject and
acceptance, distinct contexts, completion and permitted providers.
Requested identity and unreported effort do not count as runtime
evidence. The verdict schema is unchanged.

Repair the existing cleanup test: a 0.3-second budget could expire
during preparation before either fixture process started. A separate
controlled-delay test now proves preparation cannot renew that deadline.
The running-cleanup case requires parent/child readiness, preserved
partial output, the postlaunch cleanup result and both processes stopped
within its existing four-second bound. Production timeout behavior is
unchanged.

Validation: fresh author-distinct review passed the exact 55-path final
subject and all T05/T21 acceptance. The complete local Bats run passed
(1,333 passed, two existing skips), as did Go build/vet/test/race, all
72 full-mode gates, the aggregate and generated-output checks.
Ubuntu/Windows CI, security and both installation jobs passed on the
final commit. The final evidence scan found no new orphaned bindings; 73
historical bindings remain preserved. Earlier failed results and private
evidence remain outside the PR.
2026-09-08 18:57:08 -04:00
Bo 8061085c89 Ship native evidence helpers and fresh-family review defaults (#1110)
AO now performs intent snapshots, subject manifests, strict evidence
verification, atomic verdict storage, and orphan inspection through the
Go binary. The command handler keeps verification separate from
presentation so it meets the existing complexity limit. These operations
preserve the existing evidence formats, require explicit protected
storage where applicable, and run outside a checkout without Python. The
unchanged Python implementation remains a developer oracle; agents still
provide semantic judgment.

Codex and Claude skills now default to a fresh reviewer from the
author’s model family. Callers can explicitly request cross-model review
or pin its model. Reviewer adapters use a finite caller timeout or
remaining deadline instead of a fixed ten-minute default, while
retaining output limits and abnormal-termination cleanup.

Validation: Go build, vet, tests and race/shuffle tests; 1,334 shell
tests; aggregate runner; regeneration check; 72 full-mode gates.
Independent checks exercised 84 storage-boundary rejections and 21
evidence operations with an empty PATH. Both canonical and generated RPI
reference suites pass all 48 tests after updating the migrated oracle
import without weakening assertions.

Change-sensitive checks explicitly compare the final committed candidate
with the original PR base. Linux, Windows, installer, security, and
required summary checks are green.
2026-09-08 16:06:33 -04:00
Bo 568e99d436 Loop restore: converge and crank as control flow under the verdict contract (ADR-0017) (#1099)
## Loop restore: converge and crank as control flow under the verdict
contract (ADR-0017)

Intent source: `docs/plans/2026-09-03-loop-restore.md` (in this PR).
Decision record:
`docs/adr/ADR-0017-loop-as-control-flow-not-knowledge.md`.

**Why.** The 2026-07-14 single-pass cut (`482307762`) removed the
iterate loop (discovery, crank, converge, evolve, the learn write-half)
together with the unproven compounding claim, although ADR-0011 demoted
only the latter. The control flow was never demoted, and its absence
showed on 2026-09-02, when a three-lane fix needed eight validators and
two stops because the contract had no repair phase. This restores the
loop as control flow and nothing else: no knowledge store, no `ao
converge`/`ao crank`, no evolve, no canary. ADR-0004 and ADR-0011 stay
in force.

**What changes.**
- **RPI gains a bounded repair phase.** On `FAIL` or `NOT_PROVEN` with
findings, repair and re-validate freshly under the convergence law:
caller-declared `repair_rounds` (default 2); open finding set keyed by
stable `findings[].id`, union across validator families, non-growing; no
closed id reopens; the subject digest changed or, for `NOT_PROVEN`, new
digest-bound evidence resolved a named gap. Converged = fresh PASS plus
cross-family PASS on risky surfaces. Plan and Implement keep their
single dispatch. `skills/rpi/scripts/run_once.py` models the law as pure
data (33 tests): rounds are validated for shape (digest required, no
duplicate ids, no PASS with findings, no FAIL without findings),
condition 4's evidence branch needs a NOT_PROVEN previous round, a
non-FAIL current round, new evidence, and a resolved finding, and a PASS
over unchanged bytes after a FAIL is a flip that reports NOT_PROVEN.
`workflows/rpi.js` runs validation as legs (spawned or external primary,
plus a caller-supplied `crossFamily.command` on risky surfaces) merged
worst-of with a union of stable ids; a risky surface without a
cross-family leg is `diversity_unsatisfied` and never converges or
enters repair; a failed repair or re-validation returns NOT_PROVEN with
no stale verdict. Validators return `subjectDigest`, stable finding ids,
and `evidenceRefs`.
- **crank returns as a thin wave executor** (113 lines): the caller
selects the wave and the repair bound, crank invokes RPI per lane
(parallel only on disjoint write and regen scopes), runs the wave
acceptance once, returns evidence, and stops. No retry, budget, queue,
claim, lease, Git, closure, or next-work ownership. Routing golden
`rq-07-wave-execution` ranks it first.
- **validate is cross-family by default on risky surfaces**
(`cli/internal/gates/**`, `scripts/check-*.sh`, `tests/**`,
`skills/*/scripts/**`, hook policies, `lib/**`, security-scanned paths)
with the LAW-0 dispatch table: Claude orchestrating uses read-only
`codex exec`; Codex orchestrating uses an interactive Claude session in
an NTM pane, never `claude -p`. No live adapter means
`diversity_unsatisfied`, which on a risky surface is `NOT_PROVEN`. The
full literal CI command set runs once on the final integrated subject;
routine rounds keep the receipt-driven freshness contract.
- **Conformance assertions flipped under ADR-0017 only:**
`scripts/check-cathedral-cut-conformance.py` (crank live; "Stop
regardless" replaced by positive canaries for the law's four conditions;
a bounded `for` loop that compares against `repair_rounds` is required
in `run_repair_phase`, and the gate executes the law's canaries against
the reference behavior), `workflows/rpi.js`,
`skills/rpi/scripts/validate.sh`,
`evals/agentops-core/rpi-behavior.json`,
`skills/rpi/references/rpi.feature`. Every single-pass public surface
(README, AGENTS.md, PRODUCT.md, CI-CD, agent-workflow-reference,
rpi-traversal, cli/README, quickstart and demo commands, the
operating-contract and product-boundary bats, the Codex-description
oracle) now states repair to convergence.

**Known approximation, disclosed.** The Claude conveyor has no
deterministic shell primitive, so changed paths are derived by the fresh
validator (git status and diff against the clean pre-run tree) and
unioned with the implementer's report; risk is classified over that
union and unreported paths are coverage findings. A validator is still a
model; runtime derivation outside every agent is a follow-up. Family
distinctness of the cross-family leg is asserted by the caller's choice
of command and not verified by the script.

**Not in scope.** Premortem stays a single advisory judge and Plan still
only names the first check (phase boundaries unchanged). No `verdict.v2`
or `rpi-report.v1` change. The loop's own effect on outcomes is
unmeasured and owed a seeded-defect probe, like the rest of the corpus.

**Evidence on the tip.** Regen check clean; full gate green with a
HEAD-built binary; CI's bats command green; Go build/vet/test green;
golangci-lint clean; security gate quick PASS; one fresh validator over
the whole diff; one cross-family read of the design before
implementation (13 findings folded) and two of the integrated diff (9
findings in round one, 11 by round two, 15 by round three, each round
repaired and re-reviewed; the fresh validator passed the tip after round
two and the final tip 1e8adb72d passed a fresh validator (14-scenario
independent harness of the law, full gate 71/71 with a HEAD-built
binary, CI bats 1164/0) and a cross-family read by Gemini 3.8 via AGY,
which closed all six remaining residues with no new findings; Codex was
unreachable at push time).

**Follow-ups filed from the final reviews, not blockers:** the JS
violation check tests growth before reopen while Python tests reopen
first (same stop, different label when both occur in one round);
`cli/testdata/compatibility-baseline/families/{demo,quickstart}/case.json`
assert help-text substrings Cobra never prints (pre-existing, no
consumer); runtime derivation of changed paths outside every agent in
the Claude conveyor.
2026-09-03 15:16:08 +00:00
Bo 8cdcb5a903 Train 1: measurement substrate, context diet, retrieval-eval contract (instrument-panel roadmap) (#1087)
> **Residues closed on the caller's merge instruction** (`499d916a6`):
the round-2 findings were the same failure shape — round-1 repairs
patched cited lines instead of sweeping the class — so this commit
sweeps each file whole: every remaining SATURATED-row-append site in
skill-eval now routes to RUNBOOK retirement, the human-only-skills
*description* is runtime-conditional, premortem's "(MEASURED)" label is
gone, SKILL-API's context table carries all 25 rows and the enforcement
table gains `disable-model-invocation`, and the fixture-identity claim
is stated precisely (probe id, honesty note, and control arm are the
only differing fields — as the acceptance permits). Post-sweep:
validators, full Go suite, 68/68 gates, goldens + headroom bats green,
projections current, gemini in sync. Merging per Bo's instruction.

## What

Train 1 of the accepted [instrument-panel
roadmap](docs/plans/2026-08-26-instrument-panel-roadmap.md) (intent
landed at `986a4feaf`): the measurement substrate, the skill-context
diet, and the retrieval-eval contract. Three worktree-isolated lanes,
each independently validated by a fresh context, plus one integration
commit. 103 files, +5,510/−76.

**L1 — measurement substrate** (`instrument/measurement-substrate`)
- Gate `skill.probe-headroom` (advisory, Fast|Full): answers the
question `skill.probe-coverage` cannot — not "does a probe result exist"
but "could one have existed at all". The rule, ported to Go
(`cli/internal/probeheadroom` + `cli/cmd/probe-headroom` behind a thin
check script — the witness-crosscheck pattern, **no new `ao` root
command**): control arm ≥ 0.75 with ≥ 2 usable reps at ≥ 2 effort levels
⇒ SATURATED (void row, not an honest null); UNMEASURED outranks it;
treatment-silent ⇒ FLOOR; else SEPARATED. RED first: both committed
fixture pairs read `INERT` to everything else in the repo; the failing
separation test predates the implementation, and a bats negative-control
swaps fixture bytes and asserts the gate flips.
- **First reading on real data: 7 of 11 historical probe groups are
SATURATED** — including both `validate-not-proven` runs. Those INERT
rows were never honest nulls; they were void. The 0/12 ledger number now
argues itself.
- Declared denominator for probe-coverage:
`scripts/.skill-probe-denominator-exclusions`, fail-closed parser (entry
without an argument, stale slug, or duplicate ⇒ exit 2). One entry
(`goals`, a pure alias-of `fitness`). Net effect deliberately zero (0/12
→ 0/12: alias left, `one-way-door` entered) — the gain is a declared
number, not a better-looking one.
- Re-landed from the recovered clean-room commit (`9872483bd`),
re-validated against *current* main: `skill-eval` (defers saturation to
the gate id; its shell scripts dropped, not shipped — ratchet intent),
`route`, `one-way-door`, premortem reversibility check, council
`caller_challenge` (schema + validator, per the agent-core boundary that
the panel may challenge, never overrule).

**L4 — context diet** (`instrument/context-diet`)
- `disable-model-invocation: true` on 4 human-only skills (key verified
verbatim against Anthropic's docs). The plan guessed 35 candidates; the
graph said otherwise — 23 carry `user-invocable: true`, and 19 of those
are excluded on cited evidence (rpi consumes
anti-ceremony/implement/plan/validate; workflow scripts reach others;
`goals` is a live migration tombstone). The exclusion evidence is
retained in the lane report.
- One router skill (`human-only-skills`) — the single always-loaded
description that replaces four; it hints, never fires.
- `.out-of-scope/` formalized with this week's three refusals
(checked-in knowledge corpus; ee self-improvement loops; whole-skill A/B
as the measurement unit), each citing its evidence.
- Deterministic proof, no model eval: before/after bytes of
always-loaded description load reported in the lane summary.

**L5 — retrieval-eval contract** (`instrument/retrieval-contract`, lane
verdict PASS 10/10)
- `AGENTS.md` federated row now names **ee (eidetic-engine)** as a
concrete caller-selected memory system — consume, never build; symlink
intact.
- `schemas/pack-quality-expectations.v1.schema.json` + 4 routing goldens
+ `scripts/check-routing-probe-goldens.sh` graded against `ao skills
find`, wired as an **advisory** nightly job. Zero goldens is a failing
state — no new zero-denominator green.
- **The instrument caught a real miss on day one — and its own
prescription fixed it.** Golden `rq-04` expects `validate` for "judge
whether this finished change is actually proven before I merge it"; at
authoring, `ao skills find` ranked the *forbidden* `premortem` first and
`validate` nowhere in six natural phrasings. The pointer-wording-first
repair (validate's description gained the caller's own words: finished,
proven, verdict, merge) now ranks it #1 at 0.333; grader 6/6, and the
golden pins the repair — a description regression reopens it.

## Integration

`regen-all.sh` once over the merged lanes (catalog 52 → 56, four new
codex twins, mesh, router, manifests); `skills/route/SKILL.md`
catalog/router links became prose repo-root references (the projected
twin cannot resolve `../catalog.json` — this was both the
portable-conformance failure and the sole broken doc link);
`codex-portable-conformance.bats` pin 52 → 56.

## Evidence

- `cd cli && go build ./... && go vet ./... && go test ./...` exit 0 ·
`ao gate check --full` **68/68** · four skill validators PASS ·
probe-headroom / routing-goldens / probe-coverage bats PASS ·
`regen-all.sh --check` all current.
- Per-lane fresh validators re-ran every suite on detached content; L5
PASS; L1/L4 NOT_PROVEN solely on the projection-regen clause reserved
for integration (their remaining acceptance observed green), settled
above. Cross-family (Codex) review of the integrated diff recorded in
the session report.
- Two disclosed scope stretches accepted at integration: a one-line
`.gitignore` entry mirroring the witness-crosscheck precedent, and the
probe LEDGER.md fact-correction L1's own change made necessary (noted
for Train 2's L2, which owns that file next).

## Cross-family review (Codex, fresh context)

Round 1: **FAIL** — two blockers (the RED fixtures didn't isolate the
control arm; the goldens grader was red where the plan's acceptance says
green) and eight majors (contract contradictions in the re-landed
skills, a converter-substitution false claim in the codex router twin,
two overreaching `.out-of-scope` entries, stale SKILL-API counts). All
repaired in one bounded round (`db68935a3`): fixtures now byte-identical
outside the control arm, the routing miss actually fixed rather than
tolerated, every cited contradiction reconciled at the source and
re-projected. Post-repair: full Go suite exit 0, `gate check --full`
68/68, all validators and probe/goldens bats green, projections current,
gemini byte-identity restored. Focused re-check verdict recorded in the
session report.

## Follow-ups (Train 2, already planned)

Seeded-defect probes for the judgment spine (every ledger row citing a
passing headroom pre-screen) and the gate-hardening pair
(`Gate-Loosen-Reason` tightening ratchet; mechanical
grounding-validation over evidence docs). Plus, surfaced by this train:
a latent `valid_keys`/schema divergence in `validate-skill-schema.sh`
(two keys the schema defines are absent from the script's allowlist —
pre-existing).
2026-08-26 23:16:51 +00:00
Bo ffb9f122af refactor(cli): delete the unconsumed eval/redact surfaces — the estate audit's mechanical cut (#1082)
> **Review findings closed.** The re-check's residue (app-seam family
count) is applied in `9a2790ae7` along with the full-tier CI
settlements: regenerated documentation index (generated file, hand-edit
drifted it), regenerated CLI-surface count fixtures (top=18 sub=44
all=62), `Test-Removal-Reason` trailer for the deliberate test
deletions, and the release-tag bats output list updated to the real
changes-job set. 67/67 full-tier gates green locally. Merging on Bo's
instruction.

## What

Deletes the provably-dead 28% of the `ao` CLI and every reference to it,
per the 2026-08-23 estate audit. −19.5K lines in the lane commit plus
integration fixups.

**Removed (each with zero live consumers, verified by consumer-grep +
`go list -deps`):**
- `ao eval` — 13 subcommands, ~10.9K LOC. Its would-be consumers were
already tombstones (`scripts/eval-agentops.sh` printed `RETIRED`),
`release.yml` hardcoded `--eval pass`, release evidence recorded
`suite_count: 0`, and three of its module tests exercised subcommands
that could never register (nil composition seats).
- `ao redact` — its only declared caller
(`skills/compile/scripts/compile.sh`) never existed.
- `cli/internal/types/memrl_policy.go` + the orphan cascade it and eval
left behind (`internal/scenario`, `internal/wiki`,
`internal/runtimecmd`, `internal/redact`) — all with zero importers,
verified before and after.
- `scripts/check-memrl-health.sh` +
`examples/schedules/feedback-drain-hourly.yaml` — a health check for the
feedback loop amputated on 2026-07-14; it exits 1 on main today and the
example instructs a verb (`ao feedback-loop`) that no longer exists.
- `corpus.secret-scan` gate — vacuous: its file filter excluded the
single tracked path its globs could match, so it scanned zero files;
secrets are covered by the pinned gitleaks steps in nightly and release
(validate's quick toolchain mode skips gitleaks).
- Docs for the deleted surface:
`docs/architecture/eval-architecture.md`,
`docs/code-map/eval-lid-primitives.md`; `contracts/eval-baseline-ab.md`
already carried a RETIRED banner and stays as history (delisted from the
live index).

**Kept, deliberately:**
- `ao robot-docs` — the audit's "duplicate of `doctor robot-docs`"
premise was false: they render different handbooks (whole-CLI vs
doctor-scoped). Verified before acting.
- `completion`, `demo`, `quick-start` — interactive human furniture, not
dead code.
- `corpus.witness-dolt-jsonl-crosscheck` gate — retargeted, not retired:
its backing script is a hermetic self-test over real tracked fixtures;
globs now point at the paths it actually exercises.
- `cli/internal/evalsubstrate` — Go-dead but it is the declared mirror
of `schemas/outcomes-rubric.v1.schema.json`; retiring it needs a paired
schemas/docs/scripts decision (package doc comment records this).
- `scripts/ci-local-release.sh` eval-evidence stanza — self-contained
honest bookkeeping (`status: not_applicable`), invokes nothing removed.

**Tombstones + migration:** `eval` and `redact` added to
`removed_command_hint.go` and `docs/MIGRATION.md`; the now-false "(`ao
eval` returned in 3.3 …)" parenthetical deleted; `go-cli.md` spine and
the "Eval — the Learn seat" section updated; the dated research snapshot
got a HISTORICAL banner via the docs-scope self-declaration mechanism
(history not rewritten).

## Why

v3.6.0 binary downloads: 4 darwin-arm64, 3 linux-amd64. Only 7 of 53
shipped skills invoke `ao` at all, and none of them touch this surface.
The eval family was the single largest command surface in the CLI with
zero live consumers — 28% of non-test Go maintained for nobody.

## Evidence

- `cd cli && go build ./... && go vet ./... && go test ./...` — exit 0
(previously-failing `TestGoCLIDocSpineMatchesApprovedSpine` and
`TestRemovedVerbsHaveMigrationRows` now pass)
- `scripts/check-docs-cli-snippets.sh` PASS ·
`check-cmdao-surface-parity.sh` PASS (54 leaf commands) ·
`check-corpus-path-guard.sh` PASS · `check-new-scripts-use-preamble.sh`
PASS · `ao gate check --dry-run` PASS
- Implemented by a worktree-isolated lane, independently validated by a
fresh context that re-ran the suite itself; the two failures it found
were doc files outside the lane's write scope, fixed in the integration
commit. Cross-family (Codex) review verdict included in the final
session report.

## Cross-family review (Codex, fresh context)

First pass: **FAIL** with two majors — (1) `quality.DeprecatedCommands`
still mapped five rewrite entries onto the removed eval family, so `ao
doctor --fix` would have introduced dead commands; (2) retained docs
(formal-verification research links, applied-ood README run block,
evalsubstrate hint strings) still prescribed removed commands. Both
repaired in `4da85a0d4` (one bounded round), plus its two minors
(types/AGENTS.md row, .gitignore unignore, family counts,
gitleaks-coverage comment). Re-verified: full suite green, snippets gate
PASS. Focused re-check: first-round findings confirmed closed; one new
residue (the family count above) stopped the loop under the spiral rule.

## Follow-ups (not in this PR)

- `cli/internal/quality/stale_refs.go` `DeprecatedCommands`: the five
eval-target entries are pruned here; the older pre-existing dead targets
(forge, inject, flywheel, ratchet, …) still need a map-wide
reconciliation against the live registry.
- `cli/internal/evalsubstrate` retirement decision (paired
schemas/docs/scripts change).
- `evals/scenarios/applied-ood/`, `evals/tier2-premortem/`,
`evals/_stats/` retain historical `ao eval` mentions in prereg/holdout
records — dated artifacts, left as history.
2026-08-25 03:34:00 +00:00
Bo 621dbb575f 3.6.0 release prep: version bumps, changelog, curated notes (#1071)
Everything-but-tag for **v3.6.0**. Minor, not major: the post-3.5.0
delta retires the knowledge-flywheel product surface and aligns the
estate on the operations-layer identity, matching the 3.4.0 precedent
where the orchestration pack was removed in a minor.

## What this carries

- **Version 3.5.0 -> 3.6.0 across all seven surfaces**: Claude plugin
manifest, marketplace metadata + plugin entry, Codex manifest, Gemini
image manifest, Claude image verify pin, and the `ao` source fallback.
- **CHANGELOG `[3.6.0]`** (root + docs mirror): operations-layer
alignment, anti-ceremony enforcement, the behavioral eval program, the
flywheel retirement, and the honest 0/12 measured probe coverage.
- **Curated `docs/releases/2026-08-17-v3.6.0-notes.md`**: validator
PASS, tier minor, full changed-path area coverage. The Breaking Changes
section lists all six removals and the handoff write-path move rather
than burying them in a minor.
- **New regression test `cli/cmd/ao/version_manifest_parity_test.go`**
binding the `version` fallback to every version-bearing release surface.
- **PRODUCT.md** reviewed against the 3.6 surface and re-stamped;
**`docs/reference/skill-system-evolution.md`** gains its 3.6.0 row and
drops the "current unreleased tree" framing that the tag would falsify.

## Why the new test exists

This cut missed `images/claude/verify.sh`. Its version guard — whose
entire stated purpose is catching plugin.json drift *behind* the release
— then rejected the **correct** version, so a user following the shipped
`images/claude/README.md` on the v3.6.0 tag would have hit a hard FAIL.
`check_manifest_version_consistency` in `ci-local-release.sh` compares
only the two Claude manifests to each other, so it structurally could
not see this.

The test fails on the drift and passes when correct; both directions
were exercised before committing.

## Honesty notes carried into the release

- Measured behavioral probe coverage is stated as **0/12** under the v3
evidence contract. The earlier wave-1 classifications are retained as
`LEGACY-UNVERIFIED` rather than counted, because the probe harness did
not isolate the skill corpus between control and treatment arms.
Skill-efficacy claims in these notes are directional, not proven.
- The estate-ablation aggregate counts are labeled legacy-unverified and
non-promotable.

## Verification

Full `scripts/ci-local-release.sh --release-version 3.6.0
--readiness-mode official --security-mode full`: **PASSED — 72 checks, 0
failures**.

| Dimension | Status |
|---|---|
| SIL (race suite, 75.8% coverage) | pass |
| VIL (gates, regen, digital twin) | pass |
| HIL (real Darwin/arm64 target) | pass |
| Artifacts (CycloneDX + SPDX SBOM) | pass |
| Security (full mode) | pass |

Readiness **9.0** against threshold 8. HIL used a real target with **no
waiver**: `ao` built from this tree reported `ao version 3.6.0`
(`version_verified=true`) and ran a full `ao init` scaffold plus `ao
status` in a scratch repo. Security full mode: 0 critical, 0 high, 3
medium (non-blocking). Notes validator PASS, doc-release gate PASS.

## Post-merge

Readiness lap at the merged SHA, audit record in `docs/audits/`, then
the tag — per the binding process rule that the record exists **before**
the tag.
2026-08-17 21:24:34 -04:00
Bo f3c6d0ecf2 Converge retained WIP and harden evidence boundaries (#1065)
Summary:
- lands the audited current WIP lanes and excludes stale/process-only
material
- hardens prune path confinement, probe-v3 evidence binding,
codebase-recon identity, handoff/release/reverse-engineer behavior, and
Codex prompt handling
- truth-labels static skill scoring and regenerates all owning
projections

Validation:
- fresh independent PASS on commit
427098ed10
- full Go suite and full Go race suite
- quick local release CI
- focused Bats, scenario/linkage, native-skill, reverse-engineer,
Cathedral, Ruff, Python ratchet, projection, and diff checks

Residual boundaries:
- final-basename ABA remains unclaimed
- live skill-probe coverage remains honestly 0/12
- release-only cross-build, SBOM, vulnerability, and release-evidence
checks are left to delivery CI
2026-08-16 18:32:26 -04:00
Bo 7a765cde19 Align AgentOps around its operations-layer identity (#1051)
Executes docs/plans/2026-08-07-agentops-operations-layer-alignment.md:
AgentOps is the operations layer for agentic engineering; the federated
integration graph is the topology, the semantic work-and-proof protocol
is the contract, and RPI is the standard one-experiment traversal.

Retires the ao flywheel command family and all knowledge-flywheel
product state, tombstones the seven-move operating-loop workflow,
narrows ao init and the .agents state writers to declared destinations,
renames the core architecture page to rpi-traversal.md with a
compatibility redirect, aligns AGENTS.md, 25 skills, public and package
copy, regenerates every owned projection, and strengthens the
conformance gates with planted-negative proofs.

Both the alignment subject and the follow-up gate-bookkeeping commit
carry fresh author-distinct validation PASS verdicts with empty
not_checked scope.

Test-Removal-Reason: dead knowledge-flywheel and session-store surfaces were deleted with their tests (operations-layer alignment)
2026-08-07 18:37:03 -04:00
Bo 1c1500ce87 3.5.0 release prep: version bumps, changelog, curated notes (#1029)
Everything-but-tag for v3.5.0 (Bo's call: the post-3.4.0 delta carries
two feature surfaces — `ao gc` and plan manifest mode — so minor, not
patch).

- Version 3.4.0 → 3.5.0 across all seven surfaces (plugin manifests,
marketplace, image verify pin, `ao` source fallback).
- CHANGELOG `[3.5.0]` section (root + docs mirror): ao gc family,
manifest mode, Mayor-dispatch doctrine, honest-scoped-PASS,
fresh-install fixes, init gitignore policy.
- Curated `docs/releases/2026-07-31-v3.5.0-notes.md`: validator PASS,
tier minor, full area coverage; upgrade notes call out the
gc-maintainer-ops wrapper deprecation and the new init gitignore block.

Verification: notes validator PASS · `regen-all.sh --check` all ✓ ·
skill-lint 0 · `go build/vet/test` 2947 passed / 73 packages.
Post-merge: official-mode readiness lap at the merged SHA with real HIL,
record in docs/audits/ before any tag.
2026-07-31 10:08:11 -04:00
Bo fd30523a75 docs(skills): install-agnostic loop commands and self-contained rpi examples (#1026)
## Summary

Fresh-install smoke testing found four commands/links in skills and CLI
help that fail verbatim for an installed user (only `skills/**` — not
the full repo tree — ships to an install; `.agents/skills/**` is the
installed skill root).

| # | Defect | Fix | Verified |
|---|---|---|---|
| 1 | `skills/plan/SKILL.md` step 1 told the runtime to run `python3
skills/validate/scripts/validate.py snapshot-intent ...` — a
checkout-only path. In an installed tree the real path is
`.agents/skills/validate/scripts/validate.py`. | Reworded to name both
paths explicitly (checkout: `skills/validate/scripts/validate.py`;
installed: `.agents/skills/validate/scripts/validate.py`),
install-agnostic. | Copied `skills/{validate,plan,rpi}` into a scratch
`.agents/skills/` layout and ran `echo '{"foo":"bar"}' \| python3
.agents/skills/validate/scripts/validate.py snapshot-intent --source -`
verbatim — produced a valid `intent_ref`. Re-ran the checkout-relative
form too. |
| 2 | `skills/rpi/SKILL.md` linked
`../../schemas/rpi-report.v1.schema.json` — `schemas/` isn't shipped to
installs, so the link 404s for an installed user. | Inlined the minimal
required `rpi-report.v1` shape as a fenced JSON block (with field
semantics), plus a note that the schema itself ships in a repo checkout.
No new files added to the skill package. | Validated the exact inlined
shape (with a concrete instance) against
`schemas/rpi-report.v1.schema.json` via `jsonschema.validate()` —
passes. Confirmed all 9 required keys and digest patterns match the real
schema. |
| 3 | `skills/rpi/SKILL.md`'s continuation-envelope example (~lines
115-121) cited this repo's own internal 2026-07-15 intent/verdict
digests (`26a4f2be...eb48`, `b6e759dd...cb6a`, etc.) as a normative
example — not reproducible by an installed user. | Replaced with a
generic, self-contained example (placeholder revisions/verdicts) that
illustrates the same two-stop-checkpoint behavior without citing this
repo's private history. | Reviewed the replaced prose reads correctly in
context; `bash tests/skills/run-all.sh` still passes (no broken
frontmatter/budget). |
| 4 | `ao --help` root epilog (`cli/cmd/ao/root.go`) pointed at
`docs/MIGRATION.md`, a relative path that doesn't exist for a user who
only has the `ao` binary (no `docs/` directory ships with it). | Changed
the epilog to the GitHub blob URL
(`https://github.com/boshu2/agentops/blob/main/docs/MIGRATION.md`), with
a note that a repo checkout also has it locally at `docs/MIGRATION.md`.
Left the internal `removedCommandHint()` machinery (and its
`docs/MIGRATION.md`-literal test assertions) untouched — that's a
separate, heavily-tested mechanism not covered by this defect. | `go run
./cmd/ao --help` shows the new URL. Confirmed `boshu2/agentops` is the
correct remote and `docs/MIGRATION.md` exists at that path on `main`. |

## Process / regen

- `scripts/codex-sync.sh --only plan`, `--only rpi`
- `scripts/regen-codex-hashes.sh --only plan`, `--only rpi`
- `python3 scripts/generate-skill-mesh.py`
- `scripts/generate-cli-reference.sh` (no diff — root epilog text isn't
captured in `COMMANDS.md`)
- `scripts/regen-all.sh --check` — all projections current

## Test plan

- [x] `cd cli && go build ./...` — success
- [x] `cd cli && go vet ./...` — no issues
- [x] `cd cli && go test ./...` — 2923 passed, 0 failed, 73 packages
- [x] `bash tests/skills/run-all.sh` — 54/54 skills pass, 0 failed
- [x] Simulated install (`.agents/skills/...`) and ran the exact
documented `validate.py snapshot-intent` command verbatim — works
- [x] Validated the inlined `rpi-report.v1` JSON shape against the real
schema with `jsonschema.validate()`
- [x] `go run ./cmd/ao --help` shows the corrected epilog

Write scope respected: `skills/plan/SKILL.md`, `skills/rpi/SKILL.md`,
`cli/cmd/ao/root.go` (help text only), and regenerated projections
(`skills-codex/**`, `images/gemini/skills/**`). No changes to
`skills/validate/**`, `CLAUDE.md`, `AGENTS.md`, `cli/internal/gates/**`,
or `cli/internal/doctor/**`.
2026-07-31 09:05:00 -04:00
Bo efcf4879c8 feat(gc): port gc-maintainer-ops into the ao gc command family (#1016)
## What

Ports `scripts/gc-maintainer-ops.sh` (425 lines of bash: prepare / check
/ recover-affinity for stock Gas City rigs) into the Go CLI as **`ao gc
prepare|check|recover-affinity`**, per ADR-0016 (skill logic ships in Go
via `ao`; shell stays thin glue).

**Why:** skills ship via plugin/npx as SKILL.md only — a user without a
repo checkout could not run the commands the shipped `using-gc` skill
teaches. The skill said "From an AgentOps checkout", which was disclosed
but weak.

## Changes

- **`cli/internal/gcmaintainer`** — full port: rig/import pin
verification, bundled pack-cache resolution, PyYAML-capable python
selection, atomic runtime staging, managed check wrappers, skill links
into city/rig Codex sinks, macOS LaunchAgent + doctor/status health
checks, bounded affinity recovery. Output and refusal-message parity
with the shell script (incl. refuse-before-mutation ordering).
- **`cli/internal/commands/gc` + `cmd/ao/gc_composition.go`** — cobra
module on the shared `clicontract.HostOptions` seam; global `--dry-run`
always overrides `--apply`.
- **Skills source resolution without a checkout**: `--skills-source` >
enclosing agentops checkout > installed skills root (`~/.agents/skills`,
`~/.claude/skills`). Existing rigs stay recognized: the `managed-by:
agentops gc-maintainer-ops` wrapper marker is unchanged.
- **Tests migrated**: `tests/python/test_gc_maintainer_ops.py` (7 cases)
→ Go L2 tests in `cli/internal/gcmaintainer` with the same fake-`gc`
harness, plus module wiring tests. `scripts/check-gc-executor.sh` no
longer runs the python suite.
- **`scripts/gc-maintainer-ops.sh`** reduced to a thin wrapper exec'ing
`ao gc`, pinning `--skills-source` to its checkout to preserve
historical semantics (`--ao-bin` now selects the ao binary).
- **Docs/projections**: `skills/using-gc/SKILL.md` now teaches `ao gc
...`; codex, gemini, and executor-pack projections regenerated via their
owning generators; spine/COMMANDS.md/surface artifacts regenerated.

## Verification

- `go build ./... && go vet ./... && go test ./...` — 2923 passed, 73
packages
- `golangci-lint run` on new/touched packages — clean
- `shellcheck -S warning` on wrapper + gate script — clean
- `bash scripts/check-gc-executor.sh` — OK
- Smoke: built `ao`, ran wrapper → `ao gc` delegation end-to-end
2026-07-30 13:16:53 -04:00
Bo 9dd6e7d3f9 3.4.0 release prep: upstream-factories pivot, version bumps, release notes (#1013)
## Summary

Everything-but-the-tag for v3.4.0, in four commits:

- **docs(gc)**: the factory pivot — README and `using-gc` present the
upstream [Gas City build
pack](https://github.com/gastownhall/gascity-packs/tree/main/gascity)
and [Agentic Coding Flywheel](https://agent-flywheel.com/) as the
supported factory choices; the in-repo prototype (`deploy/gc/`) is
retired in place. AgentOps' lane is the skills + evidence discipline
either factory executes.
- **chore(release)**: version 3.3.0 → 3.4.0 across all six surfaces
(claude/codex/gemini plugin manifests, marketplace, image verify pin,
`ao` source fallback).
- **fix(gates)**: `check-orchestration-skill-boundaries.sh` exited 2 on
every run — it probed adapter files deleted by the 3.3 single-pass
refactor and three contract phrases removed by the skill-overhaul waves.
The live ratchets (retired-skill absence, ATM-era naming) are kept.
- **docs(release)**: 3.4.0 CHANGELOG section (root + docs mirror) and
curated release notes; `validate-release-notes.sh` passes (tier minor,
full area coverage).

## Verification

- Full Go gate in a clean worktree: build ✓ vet ✓ test **2902 passed / 0
failed** (71 packages)
- `scripts/regen-all.sh --check`: all 11 projection/doc checks ✓
(including the doc-release freeze gate)
- `scripts/validate-release-notes.sh v3.4.0 --since v3.3.0`: PASS
- `scripts/check-orchestration-skill-boundaries.sh`: exit 0 (was exit 2
on main)

## Notes

- The earlier read that `go.cli-reference` needed unpinning from the
negative-witness grandfather list was a **false positive**: gitignored
session logs under `tests/claude-code/logs/` pollute the witness scan in
a dirty checkout. On a clean tree the pin is correct; a follow-up task
exists to make the scanner read only tracked files.
- Tagging + Release Publisher run happen after merge, separately; an
official-mode readiness artifact gets produced at the merged SHA
**before** any tag (binding rule from the v3.3.0 record).
2026-07-29 22:09:42 -04:00
Bo a305de5c3e Go CLI audit residue: eval id containment hardened, dry-run honored, owned temp dirs (#1008)
## Summary

Bead `age-skill-overhaul-reboot-sjv7v.12` — reconciliation of the
2026-07-24 Go CLI deep audit against current main. Full table with
evidence:
`docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/s12-go-residue.md`.

**Fixed here (4):**
- **Eval identifier path containment** [High] — new
`evalsubstrate.ValidateID` at every identifier-to-path join, hardened
through two review rounds: rejects separators, absolute/volume refs,
leading/trailing space-or-dot (defeats Win32 trailing-strip
renormalization), C0+C1+DEL controls, non-UTF-8, non-NFC,
whitespace-only, >128 bytes; `ms:*` colons handled by injective one-way
`%3A` encoding at the checked `ModelSpecPath` sink (raw `%` reserved so
encoding cannot alias), both callers migrated, no unchecked join
remains.
- **`provenance add --dry-run`** [High] — was wired but never read; now
honored with a no-write witness test.
- **Live-runtime isolation dirs** — owned, cleaned on all paths, never
claims a caller-supplied root.
- **Stale eval help text** — corrected; COMMANDS.md regenerated via its
owner.

**Already landed (1):** the `--json`/`-o json` divergence for
provenance/skills was resolved by the cmd/ao carve-out (probes confirm
identical output).

**Recorded OPEN with reproductions (3):** `gate check --dry-run`
plan-only mode, bounded subprocess output streaming, and
context/process-group cancellation — each a cross-package refactor (the
audit's own G1/G2 programs), documented with fix sketches rather than
half-fixed here.

## Validation

- `go build` / `go vet` clean; `go test ./...` 2872+ pass across 70
packages; golangci-lint 0 issues on touched packages; CLI reference
check current
- Cross-family review two rounds: round 1 three findings (Windows
renormalization traversal, canonicality bounds, unchecked sink) all
fixed; round 2's one residual (non-injective colon encoding) fixed with
witness cases

Tracker: `age-skill-overhaul-reboot-sjv7v.12`
2026-07-29 06:21:47 +00:00
Bo c88a4514f9 W7 support wave: handoff schema truth, dcg fact corrections, honest support contracts (#1005)
## Summary

Wave W7 of the skill-overhaul reboot
(`age-skill-overhaul-reboot-sjv7v.8`) — the nine support skills, plus
the one Go fix where the skill contract crosses the CLI boundary.

- **handoff** — `ao session handoff --dry-run` output failed its own
`handoff.v1.schema.json` (reproduced: 3 errors). Fixed with a consumer
audit: schema keeps v1 with the doctrine-retired fields as optional
deprecated read-compat properties; the generator keeps its
collision-safe fractional id (schema pattern widened instead); real
jsonschema validation in `TestHandoffDryRunSatisfiesSchema` + a
legacy-artifact compat test; `read_clock` effect declared.
- **dcg** — corrected the false "`rm -rf ./build` allowed" claim (live
0.5.6 blocks it) and a nonexistent rule id in the allowlist example
(silent no-op) across six files; removed a token-splitting "workaround"
that was an executable guard bypass, replaced with file/stdin handling
and a never-reconstruct warning; temp-path rule live-probed and stated
identically in both docs; version/path/upstream corrections.
- **cc-hooks** — ships-by-default contradiction reconciled;
PATH-clobbering recipe fixed; operator-private paths removed from
shipped text; jq preflight added to the edit guard.
- **ms** — validator no longer mechanically asserts the false `effects:
[]`; it extracts the frontmatter and requires the exact honest effects
value.
- **account-rotation / status / sbh / bootstrap** — real effects
declared, both-tools-absent and destructive surfaces defined,
live-output overclaims narrowed, versions pinned.
- **shared** — advertising narrowed to the current no-bundled-references
state; retirement NOT executed (bead `.11`).

Ledger (32+8 fixed across two rounds / 8 rejected-stale / 8
deferred-with-reason):
`docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/w7.md`.

## Validation

- `go build` + `go vet` + 482 `cmd/ao` tests incl. the new schema-lock
and legacy-compat tests; dry-run validates 0 errors
- 49/49 strict frontmatter; regen clean; scenario-linkage PASS; liveness
+ anti-spiral + policy + edit-guard bats green; python ratchet
no-growth; shellcheck clean
- Cross-family review, two rounds: round 1 eight findings all fixed
(schema compat, id collision, real validation, security bypass removal,
live-probed temp rule, anchored greps); round 2 delta re-review
**VERDICT: PASS** with zero residuals

Tracker: `age-skill-overhaul-reboot-sjv7v.8`
2026-07-28 16:03:12 +00:00
Bo d9eeee9cb3 chore(release): prepare v3.3.0 changelog and version bump (#989)
## What

Prepares the **v3.3.0** release cut (last shipped tag: `v3.2.0`,
2026-07-03). No tag and no GitHub release are created here — tagging is
gated on the live GC canary and the operator's go.

- **Changelog reframe.** The `[3.3.0]` headline now leads with the
deliberate subtraction of the guardrail cathedral (strict verbose step
contracts, enforcement gates, injection hooks, command tombstones,
retired lifecycle surfaces, the heavy out-of-session orchestration
layer) — removed because that scaffolding degrades frontier-class models
rather than helping them. What remains is the judgment-boundary
membrane: caller-owned intent, one fresh independent validation, durable
verdicts, pinned provenance. Subtraction is git-concrete: `cli/internal`
packages **79 → 43**; reference orchestrator **~29k lines → ~1,300-line
thin pack** (one commit removed 29,418 lines for 1,334).
- **Post-07-17 delta folded in:** Gas City factory pack (preview) —
official checksummed GC/Beads binaries fetched by the installer (no
compile), mayor-driven door (human attach or agent drive via mail/sling)
with heartbeat shepherd, the `using-gc` mayor-orchestration skill +
four-layer visibility doctrine, rig-scoped dispatch intake, three
disclosed upstream v1.3.5 defects (#4586, the cross-store claim fix,
gastownhall/gascity#3985) with the preview→supported promotion criteria,
and plan/premortem ground-truth routing.
- **Version bump:** source fallback `3.3.0-rc → 3.3.0`
(`cli/cmd/ao/main.go`); plugin manifests were already `3.3.0`. Adds a
release-guard test asserting the fallback carries no pre-release marker
(resolves audit minor 1).
- **Release-gate fix:** trimmed the `using-gc` skill description (195 →
117 prose chars) to satisfy the 180-char budget — a blocker that landed
with the GC arc — and regenerated its Codex/Gemini/pack projections.

## Recommended version: **v3.3.0**

`v3.3.0` was an in-progress RC (`main.go` at `3.3.0-rc`, plugin
manifests at `3.3.0`, changelog `[3.3.0]` drafted) that was never
tagged; the GC pack + loop doctrine landed on the RC line. This cut
finalizes it.

## Go / No-Go checklist (deterministic bar:
docs/audits/release-readiness-3.3-2026-07-20.md +
docs/runbooks/release-process.md)

| # | Criterion | Status | Evidence |
|---|---|---|---|
| 1 | `go build ./...` | PASS | exit 0 |
| 2 | `go vet ./...` | PASS | exit 0 |
| 3 | Full Go test suite (`go test ./... -count=1`) | PASS | 63 packages
`ok`, 0 FAIL |
| 4 | Cathedral Cut conformance (`check-cathedral-cut-conformance.py`) |
PASS | "Cathedral Cut conformance: PASS" |
| 5 | Skill-mesh drift (`generate-skill-mesh.py --check`) | PASS | "up
to date (48 skills)" |
| 6 | Derived-artifact drift (`make regen-check`) | PASS | "All
generated projections are current." |
| 7 | Docs + release-doc gate (`make docs-check`) | PASS | CLI ref
current; 395 links, 0 broken; skill count 48; release freeze intact |
| 8 | `ao gate check --full` | PASS* | 66/67 pass. Sole fail =
`workflow.install-drift`, a worktree-local symlink artifact (tracked
`workflows/` is byte-identical to origin/main; passes in a clean CI
clone — the audit's 67/67) |
| 9 | Fast release gate (`ci-local-release.sh --quick`) | PASS | "LOCAL
CI QUICK SANITY PASSED": 44/44 install-surface smoke, 33/33 release
smoke, init smoke |
| 10 | Skill token budgets / lint | PASS | 100 budgets pass (fixed
`using-gc` 195→117 prose) |
| 11 | Install story smoke | PASS | README/MIGRATION document `ao skills
link`; all 6 curl installers are refusing tombstones; `ao version` runs;
npx-first |
| 12 | Plugin manifest ↔ corpus parity | PASS | manifests carry 48
skills, 0 drift |
| 13 | Version bump applied to owners | PASS | `main.go` 3.3.0-rc→3.3.0;
manifests already 3.3.0 |
| 14 | Changelog updated + synced | PASS | `CHANGELOG.md` ==
`docs/CHANGELOG.md` (changelog.sync gate) |
| 15 | Command/test pairing gate | PASS | main.go bump paired with
`version_test.go` release-guard test |
| 16 | Full release gate (`ci-local-release.sh`: race + security + SBOM
+ readiness) | PASS | "LOCAL CI PASSED [102s]"; release-artifact
manifest resolved `v3.3.0` |

\* `workflow.install-drift` is an environment artifact of running the
gate inside a nested worktree whose `.claude/workflows/` symlinks
resolve to the primary checkout. It is not release content: `git diff
origin/main -- workflows/` is empty, and the check passes in a fresh CI
checkout (audit reported 67/67 at origin/main).

## Operator items (require a human or a live tag — NOT attempted here)

- **Live GC canary** on the next official Gas City pin — the
preview→supported promotion gate.
- **Git tag `v3.3.0` + goreleaser publish** — gated on the canary and
operator go.
- **Docs-site publish** (mkdocs pipeline) — separate from this cut.

No tag. No release.
2026-07-24 00:01:31 +00:00
Bo cbe7ddbb5f feat(gc): ship thin native AgentOps factory pack (#980)
## Outcome

Ships AgentOps 3.3 as a thin native Gas City pack and removes the
accidental second orchestration control plane.

- Removes the custom GC delivery command/package, packet/schema family,
feeder/program engine, reducer/order, reliability registry, Beads
capability mirror, and fork-baseline runtime checks (about 29k deleted
lines).
- Retains six bounded adapters for pinned toolchain materialization,
clean bootstrap, native sling/status/doctor invocation, bead-isolated
worktrees, moving-main PR delivery, and teardown.
- Uses official Gas City and Beads behavior as the state owners, with
required OTEL configuration.
- Preserves Fable Mayor/Refiner, Sol-high planning and fresh validation,
Terra-high default implementation, Opus-medium overflow, and
support-only Luna.
- Supports automatic Refiner merge after hosted CI or a manual-review
toggle without locking `main`.

## Evidence

- Exact official GC: `8ffc009ded781a2ada2077f3a29bd712b2def0bf`
- Exact official BD: `8e4e59d39f3459a43cf21a3236a13eca4dd874f7`
- Full pinned native boundary: 9/9 passed in 114.047s
- Exact source-bead route replay: passed in 98.051s
- Release replay: manifest integrity, thin tests and both pack lints,
generated projections, contract compatibility, Bats contract,
ShellCheck, test-removal ratchet, and `go test ./...` all passed
- Fresh independent semantic verdict: PASS (Sol-high)
- Exact delivered head: `e805f0e26ce21c2eae9e720f57fff2a705f3185b`
- Subject manifest:
`8554d065076f70d0e639ca353409d2a75dca7ac01e4cb65cb4a118e98c0770fc`
- Verdict:
`3f89376b675de854832f38b46ccfa7ad7ca54079ffde578769c33e1f118c68c1`

## Release boundary

After this PR merges, the release qualification runs one fresh mixed
Terra/Opus canary from merged `main`, with at most one external repair
and one terminal retry. No live city repairs itself.

The only known local fast-gate exception is `workflow.install-drift`:
the installed workflow symlink resolves to the separate dirty primary
checkout, outside this candidate. All candidate-owned gates passed.
2026-07-22 19:08:02 -04:00
Bo faff55ac2a feat(gc): bind native 3.3 factory workflows (#972)
## Summary

- bind the GC 3.3 bead-native workflow: Fable Mayor, fresh Sol-high plan
and validation, Terra-high default or Opus-medium overflow
implementation, Luna support-only
- add the bounded one-shot graph feeder, strict packet/worktree
identity, semantic terminalization, and model-free protected delivery
sweep
- make official Gas City v1.3.5 plus Beads v1.1.0 bootstrap repeatable,
disable unstable event propulsion/hooks, and prove quiescent teardown

## Reliability boundaries

- no Gas City self-repair loop, daemon, private scheduler, model
fallback, or main mutex
- semantic completion is independent of delivery; protected
PR/CI/rebase/merge is deterministic and moving-main aware
- event hooks are disabled because 3.3 uses explicit program admission
and cooldown delivery; native controller bead observation remains active

## Evidence

- full local release CI PASS at
.agents/releases/local-ci/20260722T113355Z
- two pristine standalone-clone cycles with official GC v1.3.5 8ffc009d
and BD v1.1.0 8e4e59d: repeat bootstrap, exact rig Dolt context, absent
event hooks, bounded events, enabled sealed delivery, and delayed
zero-process teardown
- fresh independent Sol-high pre-delivery binding PASS over exact head
e8c1253b40
- binding verdict SHA-256:
4f8d81811cd5b2eb35159c98159db2f8917d3ef54167f3c0d59d8378de14b2d1

## Tracker

Beads: age-gc-scope-failclosed-release-gktia.2

The one bounded mixed-provider live canary remains a post-merge release
gate; this PR does not claim final 3.3 release readiness before that
proof.
2026-07-22 13:49:29 +00:00
Bo 4d631207cf feat(gc): add bead-native crash-only delivery kernel (#970)
## Summary

- replace the unreachable Python factory lifecycle with an optional
typed Go delivery reducer
- make Beads the delivery lifecycle authority with deterministic
moving-main successor epochs
- add native Git/worktree, PR, hosted-check, auto-merge, landing, and
cold-replay boundaries
- retire obsolete v1 role/factory schemas and preserve the pinned
historical capability harness

## Validation

- fresh Sol-high PASS over exact 43-path manifest
ec6a2a4bf0b180782411b1ae4afd4a2ad8384129ff037ce8c119480317626b02
- go test ./...
- go test -race ./internal/gcadapter/delivery
- go vet ./...
- scripts/check-gc-executor.sh
- GC 3.3 schema, migration, provenance, factory-doctor, and bootstrap
gates

Bead: age-gc-scope-failclosed-release-gktia.1
2026-07-22 02:24:58 -04:00
Bo 26120d8394 feat(gc): add crash-only delivery thin slice (#969)
## Outcome

Adds the bounded GC33-6 delivery slice outside core ao: a separately
built crash-only reducer, strict admission certificate binding,
immutable handoff and branch/PR receipts, native graph.v2 formulas,
one-step disabled Order glue, and one centralized stock claim wrapper.

Semantic bead: ag-agentops-33-gc-refinery-4km81.7 (closed after
independent validation).
Binding PASS:
d3ac5a566e0f5197814e2dbcf04d0540849f58636ed8bb49a0faca7d6af94639 over
manifest
155a089b4b461bda61a7f1ad1987c66f5755ac7c5628aa648e70740be2580a5a.

## Verification

- Full cli Go suite: PASS (2,887 tests)
- Focused race-enabled reducer suite: PASS
- GC factory/executor Python contracts: PASS
- Full executor gate: PASS (123 Python + 25 Bats)
- Exact official GC v1.3.5 pack lint and no-start discovery: PASS
- Cold replay: exactly one delivery, branch, and PR; six emitted
artifacts validate against shipped schemas
- Production Go surface: 922 lines, under 3,000 tripwire

## Boundaries

No ao command, core port, default install, live city, real
Git/forge/CI/merge adapter, daemon, scheduler, or factory.py import.
Real adapters and delivery completion remain GC33-7.
2026-07-22 00:34:05 +00:00
Bo e5ad3cfc94 test(cli): guard advertised ao commands against the live cobra tree
Walks every registered command Short/Long/Example plus quick-start printed
output, extracts advertised "ao ..." strings, and resolves each against the
assembled rootCmd (aliases included; Example fields get full group-descent).
Guards the defect class the 3.3 release audit fixed by hand: user-facing
output advertising a retired command.
2026-07-20 19:57:01 -04:00
Bo 238964fdd2 feat(cli): workflows are canonical product artifacts — workflows/ + ao workflows link (#945)
Workflows get the skills treatment (operator decision): canonical source
in the product tree, installed by a product verb, Claude-only labeled as
such.

**What moves:** all seven Claude workflow scripts + README migrate from
force-added exceptions inside the gitignored `.claude/` to a tracked
top-level `workflows/` (sibling of `skills/`) — the four existing
conveyors plus `audit-dimensions`, `verify-fixes`, `implement-wave`:
three thin, args-parameterized orchestration conveyors extracted from
this session's hand-rolled waves, contract-reviewed, and smoke-proven
through the real Workflow runtime (the smoke caught two contract gaps
static review could not: an `export default` wrapper the runtime never
invokes, and args arriving as a JSON string — both fixed, string-args
tolerance now built in).

**New verb:** `ao workflows link` / `unlink` mirror `ao skills link`
semantics — dry-run `--json`, refuse to replace real files or foreign
links, unlink only checkout-owned links — targeting the project-local
`.claude/workflows/` where Claude Code resolves named workflows
(`--into` overrides). Checkout identity reuses the skillsapp marker
discipline, fail-closed. Claude-only runtime adapter, same doctrine as
the Codex-only `skills-codex/`.

**Legacy surfaces repointed:** `install-workflows.sh` (user-global $HOME
installer), `check-workflow-drift.sh` + gate comment,
`check-bdd-foundry-markers.sh`; spine allowlist + YAML-probe excuse +
go-cli.md spine region gain the workflows group; COMMANDS.md,
cli-surface projections, and surface-count fixture regenerated; new
tests carry per-command git-env scrubbing (test-isolation ratchet back
at baseline).

**Built BY the workflow being canonized** — `implement-wave`
orchestrated its own canonization: two disjoint-ownership lanes plus a
seam-checking verifier that ran the real binary's link → resolve →
unlink cycle in the live tree (both lanes RESOLVED). The lanes correctly
*refused* to self-approve their command into the spine invariants and
handed integration three flagged edits instead.

**Expected local gate note:** `workflow.install-drift` correctly FAILS
on machines whose user-global `~/.claude/workflows` links still point at
the old location — that is the transition it exists to catch. CI stays
green (absent→skip). **Post-merge operator step:** `cd ~/dev/agentops &&
git pull && bash scripts/install-workflows.sh`.

**Verified:** full suite 63/63 pkgs; golangci-lint clean; `gate check
--full` over this range = 66/67 with only the documented install-drift
environment finding; workflows smoke-run evidence in session logs.
2026-07-20 19:39:16 -04:00
Bo 3d24ee0e9d fix(cli): residue wave — doctor dev-version coherence, diff --only, config-models removal, dual-root stragglers (#944)
Wave 4 (residue) of the new-user happy-path arc. Three scoped
implementer lanes + fresh adversarial verifier (4 RESOLVED; 1 INCOMPLETE
= stale generated projections, closed in integration).

**Doctor**: the dev-version detector now reuses the Binary Freshness
resolution — a from-source build matching its checkout is healthy (a
novice building from source can finally see `ao doctor` exit 0);
findings fire only on genuine drift, shadowed duplicate `ao` binaries,
or an informational from-source note outside any checkout. `ao doctor
diff` gains `--only` so the fix-plan preview can be scoped the way
remediation text implies.

**Config**: the dead `ao config models` surface is removed end-to-end
(lane re-verified zero consumers before deleting; `--show` proven
byte-identical before/after; removed-child hint + MIGRATION row;
existing `models:` config sections still parse and are ignored).

**Dual-root stragglers**: learning-coherence gate globs,
`quality.CountConstraints`, and the eval sandbox corpus deny-list now
cover canonical `.agents/ao/<section>` alongside legacy roots.

**Doc-link hygiene**: the strict docs-link backstop's allowlist was 100%
stale (53/53 entries referenced Cathedral-Cut-deleted docs) — refreshed
to 10 verified accepted-class entries; ROADMAP dead links fixed;
documentation-index generator emits GitHub URLs for repo-root targets;
doc-skill references instruct only shipped scripts; codex twins +
CLI-surface projections + surface-count fixture regenerated.

**Deferred by design**: the `3.3.0-rc` fallback version bump belongs
inside the v3.3.0 tag-cut commit.

**Verified**: full suite 61/61 pkgs (2825 tests); golangci-lint clean;
`ao gate check --full` 67/67 over this range; Test-Removal-Reason
trailer covers the 6 deliberately deleted models tests.
2026-07-20 19:05:57 -04:00
Bo 511370f5d9 fix(cli): doctor coherence, config truthfulness, pruned-verb hints — novice edges 4/6/7/8 (#942)
Wave 2 of the new-user happy-path fixes (fresh-eyes novice test).

**Edge 4 — doctor sub-surfaces contradicted each other.** Remediation,
`next_steps`, and robot-triage's `recommended_command` now instruct
`--fix` only when a fixer can actually act — non-fixable findings name
their real manual action (including the detect-only bridges class the
verifier caught as the same lie one level up). `ao doctor health`
reports every severity bucket present (it omitted P1 — the worst
actually present). `ao doctor diff` renders an explicit read-only fix
plan (would auto-fix vs manual action) instead of a findings copy. `ao
doctor explain` is now a superset of the finding's triage entry.

**Edge 6 — `ao config --show` rendered config for commands that don't
exist.** `rpi.*` and `dream.*` no longer serialize in human or JSON
output (struct fields retained `json:"-"` pending a follow-up removal
once nothing references them). `AGENTOPS_NO_SC` had zero behavioral
consumers (evidence: rg over all non-test cli/ — only config
plumbing/display) — undocumented and removed from the env panel; the
field stays parseable so existing config files don't error.

**Edge 8 — legacy-config gaslighting.** With only
`~/.agentops/config.yaml` present, `--show` used to print the
deprecation warning and then claim the new path was "(not found)" and
values came "(from flag)". It now shows the actually-read path labeled
"(deprecated location)" with correct attribution.

**Edge 7 — `ao beads` bare error.** The whole pruned 3.2 bookkeeping
family gets point-of-failure migration hints; a new guard test proves no
hinted verb is a live command (the `eval` trap: pruned in 3.2, returned
live in 3.3).

**Process:** three scoped implementers (disjoint files) + fresh
adversarial verifier. The verifier refuted two lanes as incomplete
(recommended-command layer, stale cmd/ao config tests asserting the old
defect) — closed in integration before this PR.

**Verified:** full suite 61/61 pkgs; golangci-lint clean; `ao gate check
--full` 67/67 over this range; per-fix RED→GREEN sandbox repros in the
workflow logs.
2026-07-20 16:05:40 -04:00
Bo d043390a0a fix(release): resolve 3.3 release-wrapper audit — blocker + 13 majors (#935)
Resolves every spellbreaking finding (the blocker + all 13 majors) from
the 3.3.0 release-readiness audit
([docs/audits/release-readiness-3.3-2026-07-20.md](docs/audits/release-readiness-3.3-2026-07-20.md),
included in this PR).

## Finding → fix map

**CLI self-documentation (M1–M3)**
- **M1** `ao robot-docs` prescribed removed `ao inject` → line removed
from the canonical agent workflow; `inject` added to the removed-command
hint table **and** the MIGRATION.md map (drift test
`TestRemovedVerbsHaveMigrationRows` enforces the pair).
- **M2** `ao config --help` documented ~14 env vars for removed
subsystems (RPI/Dream/Council/tiers) → help text and the `--show` env
panel pruned to the 5 vars the binary consumes; mirrored list in
`internal/config` pruned identically.
- **M3** `ao flywheel status` read only legacy `.agents/<section>` while
`ao doctor fix` migrates learnings to canonical `.agents/ao/learnings` →
new `quality.KnowledgeSectionDirs` dual-roots every knowledge reader
(tier counts, new/stale artifacts, retros, health delta, utility, loop
metrics, retrievable-citation stats — plus the golden-signals readers
`ComputeResearchClosure`/`ComputeReuseConcentration` that the fresh
verification pass caught as missed). Sandbox-proven twice: a learning
existing only under `.agents/ao/learnings` appears in all metrics, and a
research file only under `.agents/ao/research` flips closure from
`starved/0` to `unmined/1 orphan`.

**Release story (M4–M6)**
- **M4** CHANGELOG `[3.3.0]` omitted post-07-17 surfaces → folded in `ao
eval` (#921), default-build `ao flywheel`, the PreToolUse policy engine,
and the #919 cleanup; date moved to 2026-07-20; `docs/CHANGELOG.md`
re-synced (changelog.sync gate green).
- **M5** MIGRATION.md attributed `ao eval` to a nonexistent "3.4" → now
"returned in 3.3".
- **M6** four release surfaces claimed a 50-skill corpus vs 48
everywhere real → all counts now 48 (CHANGELOG ×2, docs/3.3.md,
release-notes page ×2).

**Install story (M7 — product decision by Bo)**
npx first (universal — installs into all coding agents), **plugins for
Claude Code/Codex encouraged**, checkout + `ao skills link` as the
source-tracked/contributor path; curl installers stay tombstones.
Harmonized across README, UPGRADING, install-day2-ops, MIGRATION,
3.3.md, CHANGELOG, the release-notes page, all six installer tombstone
messages (`install.sh` + claude/codex/agy/opencode/`codex.ps1`), and the
site's CLI page. All "legacy migration-only / not the recommended path"
plugin branding removed.

**Docs site (B1, M8–M12)**
- **B1** generated site CLI page instructed a tombstoned curl installer,
nonexistent `ao rpi phased`, and wrong skills dir → `emit_index()`
rewritten to the real install menu + a quickstart of commands that
exist; semantic loop correctly attributed to skills.
- **M8** deploy workflow's `--strict` contradicted mkdocs.yml's declared
non-strict policy and aborted the build → flag dropped (lychee +
validate-links.sh own link checking).
- **M9** site banner said "AgentOps 2.x" → now 3.3.
- **M10** ~176 internal files (audits/plans/handoffs/… + TEMP scratch
doc) published and dominated search → `exclude_docs` extended; built
site verified free of them; search index 4,751 → 2,444 entries.
- **M11/M12** newcomer-guide skill links and all six SCHEMAS.md links
404'd on the site → absolute GitHub URLs (resolve on both GitHub and the
site). The fresh verification pass found the same class on
`docs/contracts/index.md` (nav-listed),
`docs/contracts/corpus-learning-seam.md`,
`docs/templates/slice-validation.md`, and five
`docs/architecture/gas-city-factory.md` links into now-excluded
`docs/audits/` — all repointed to absolute GitHub URLs.
- Also from the verification pass: robot-docs exit-code table no longer
says "bead claimed" (removed concept), and the docs.yml comment now
cites the link checker that actually runs
(`tests/docs/validate-links.sh` via doc-release checks) instead of
lychee.

**Skills corpus (M13)**
- rch skill instructed 5 nonexistent scripts + 8 nonexistent reference
files as recovery steps → pruned to the 7 real references; the
wire-level `printf | rch` probe replaces the phantom `protocol_test.sh`;
codex twin regenerated (parity gates green).

## Verification

- `go vet` clean; **full test suite 60/60 packages pass** (includes the
new-shape flywheel/quality/config tests and the inject↔MIGRATION drift
test).
- **`ao gate check --full --scope worktree`: 67/67 pass, 0 warnings**
over this exact change set (changelog sync, shellcheck on the edited
installer, skill mesh + codex parity, manifests/schema/triggers,
provenance chain).
- `scripts/regen-all.sh --check`: all generated projections current.
- Rebuilt binary re-exercised: `robot-docs` clean, `ao inject` tombstone
live, config help clean, flywheel sandbox proof above.
- `mkdocs build` exit 0; warnings 84 → 46 (remainder is the accepted
out-of-tree-link class per mkdocs.yml's declared policy).
- Fresh-context adversarial verification workflow over all four fix
groups (results in session log).

## Known residuals (deliberately out of scope)

- `ao config models` subcommand still renders tier config (its two env
vars ARE consumed; `COUNCIL_CLAUDE_MODEL` in its display list is not —
follow-up).
- Same-class single-rooted readers off the flywheel path:
`learning.coherence` gate glob (`.agents/learnings/**` only),
`quality.CountConstraints`, config `Paths` defaults feeding eval sandbox
deny-lists.
- `scripts/docs-build.sh` still uses `--strict` with its own allowlist
(not wired into any workflow or gate).
- Audit minors 1–8 (rc fallback version string, `config --show`
legacy-fallback display, `ao init` vs doctor layout, doctor's `ao beads
dir` hint, dead `tracker:` key in the example config, doc-skill phantom
scripts, ROADMAP dead links, documentation-index root links).
2026-07-20 12:58:06 -04:00
Bo 5303a816e3 fix(cli): gate failure stderr silenced; all modules on shared seam or reasoned-exempt; guard requires coverage (age-6j9ee.1) 2026-07-20 10:14:04 -04:00
Bo b4aa2f0d8b fix(cli): -o yaml is real everywhere json is — shared WriteYAML, zero silent fallbacks (age-6j9ee.2) 2026-07-20 10:14:04 -04:00
Bo 3a54d5ff53 refactor(cli): unify host-seam contract, shared writeJSON/ExitError, drift guards (age-6j9ee.1)
Collapse the four drifted cmd/ao host-seam shapes into one shared
clicontract.HostOptions consumed by all 17 command modules; delete the
positional-func and bespoke-struct seams and the doctor GlobalOptions bundle.

- clicontract/host.go: single HostOptions union (OutputMode/Verbose/Verbosef/
  DryRun/ProjectRoot/GoalsPath/LedgerPath/Version/Now/EnrichFlagErr), one shared
  WriteJSON (byte-identical to the three deleted copies), one shared ExitError
  (Label field preserves doctor's stderr surfacing; gate stays silent).
- config_module.go + doctor_module.go: package-level module/service singletons
  -> newConfigCommand()/newDoctorCommand() constructors; 7 allowlist rows removed.
- carveout_regression_test.go: three AST drift guards (shared seam type / no
  direct host effects / no module-service singleton vars), each RED-proven.
- docs/architecture/go-cli.md: document the intentional two-tier shape (5 full
  hexagonal, 12 app-seam); drop the stale opportunistic-extraction sentence.

Zero behavior change: ao -o json capabilities byte-identical; full suite +
-shuffle=on -count=2 green; 0 lint issues; gate check --full 67/67.
2026-07-20 10:14:04 -04:00
Bo 735580d1c7 test(cli): durable package-var allowlist invariant + root-only harness comments
Finish line of the cmd/ao carve-out arc (age-a-plus-report-card-ieyp2.14).
Changes NO command behavior — tests, guard, and allowlist only.

- Add TestPackageVarsAreAllowlisted (carveout_regression_test.go): parses every
  cmd/ao non-test .go file with go/ast and fails on any package-level var that is
  not on testdata/package-var-allowlist.json, and on any stale allowlist entry.
  Keyed by name+file. Locks the finish-line state: after 12 family carves the only
  legitimate package globals are the root spine + its 5 persistent-flag targets,
  host-resident module/command wiring, const-like data maps, the ldflags version
  string, and one documented test seam (testProjectDir).
- testdata/package-var-allowlist.json: 24 entries, each with a one-line reason.
- Bounded-purpose comments on executeCommand and resetCommandState documenting
  that they save/restore ONLY still-existing root spine state — carved families
  own their flag state constructor-scoped in internal/commands/<family>.

No dead vars found to delete: all 24 cmd/ao package vars are live. Harness was
already root-only after prior waves (saves only dryRun/verbose/output/jsonFlag/
cfgFile). RED proof: a scratch `var sneaky string` in group_json.go trips the new
guard; removed before commit.
2026-07-19 21:31:58 -04:00
Bo b390ad782a refactor(cli): carve flywheel into internal/commands/flywheel + internal/flywheelapp
The command module is a thin Cobra presentation seam; internal/flywheelapp owns
the filesystem and clock effects and preserves the internal/evidence
(ratchet.LoadCitations) dependency. flywheel status/compare surface and
capabilities are byte-identical.
2026-07-19 21:03:13 -04:00
Bo 54b462a855 refactor(cli): carve redact into internal/commands/redact module 2026-07-19 20:42:03 -04:00
Bo bc4ab7970f refactor(cli): carve quick-start into internal/commands/quickstart module 2026-07-19 20:39:24 -04:00
Bo c5522feebc refactor(cli): carve robot-docs into internal/commands/robotdocs module 2026-07-19 20:36:08 -04:00
Bo 6c938b863e test(cli): extend AST carve-out guard to session+demo+init+version 2026-07-19 20:13:58 -04:00
Bo 02045e6c25 refactor(cli): carve version into internal/commands/version module
Move the version command out of package main into internal/commands/version.
version reads only build/runtime metadata (a pure effect), so it needs no app
seam; the build-time version string and the global -o/--output mode are host
seams injected by the composition. Unlike the other W3 families, version
carries a real CommandContract that newVersionCommand attaches to the command
tree, so the capabilities surface is byte-preserved.

The build-time `version` string var stays in package main (the ldflags target
`main.version`, shared by capabilities, doctor, and rootCmd.Version). The
command-behavior tests move to the module (constructing it directly); the
root-wiring tests (`version` var default, root registration, `ao --version`
flag) stay in package main because the module does not own that wiring.

Test-Removal-Reason: TestVersion_CommandExists removed; its Use/GroupID
assertions are preserved as version module TestModule_CommandAttributes.
Test-Removal-Reason: TestVersion_DevVersionDefault removed; it duplicated the
version-string assertion now covered by version module
TestVersion_ExecuteOutputContainsVersionString.
2026-07-19 20:13:47 -04:00
Bo 0c8d2a4ab6 refactor(cli): carve init into internal/commands/init module
Move the init command out of package main into internal/commands/init,
delegating the working-directory resolution and directory creation to the new
internal/initapp seam. The dry-run selection is a host seam injected by the
composition from the global --dry-run flag; the module performs no direct
filesystem effect.
2026-07-19 20:13:46 -04:00
Bo 1733907f24 refactor(cli): carve demo into internal/commands/demo module
Move the demo command out of package main into internal/commands/demo. Demo
renders static explanatory text and performs no effect, so it needs no app
seam; the render helpers and constructor-scoped --quick/--concepts flags live
in the module. demoQuick and demoConcepts package globals die; the shared
resetCommandState test helper drops its references to them (cobra flag reset
still covers the live command instance).
2026-07-19 20:13:45 -04:00
Bo 8c1076fbca refactor(cli): carve session into internal/commands/session module
Move the session evidence commands (bootstrap, rehydrate) out of package
main into internal/commands/session, delegating filesystem effects to the
new internal/sessionapp seam. The session parent command now lives in the
module; the optional `ao session handoff` writer stays in package main and
is attached to the module-built parent by newSessionCommand.

sessionBootstrapJSON and rehydrateJSON package globals die; flag state is
constructor-scoped in the module. The combined
TestHandoffAndRehydratePreserveCallerTextWithoutLifecycleState is split: the
handoff-writer assertions stay in cmd/ao as
TestHandoffPreservesCallerTextWithoutLifecycleState; the rehydrate read side
moves to the module as TestRehydrateReadsCallerAuthoredBrief plus the
preserved TestRehydrateJSONEmptyStateEmitsEmptyObject and
TestSessionBootstrapOnlyReportsLocalOrientation.

Test-Removal-Reason: TestHandoffAndRehydratePreserveCallerTextWithoutLifecycleState removed; its handoff-writer coverage is preserved in cmd/ao and its rehydrate coverage relocated to the session module (net +2 tests).
2026-07-19 20:13:44 -04:00
Bo f7025acc22 test(cli): extend AST carve-out guard to provenance+skills 2026-07-19 18:30:07 -04:00
Bo 29948e6fee refactor(cli): carve skills into internal/commands/skills module 2026-07-19 18:29:10 -04:00
Bo 549dd75f85 refactor(cli): carve provenance into internal/commands/provenance module 2026-07-19 18:18:21 -04:00
Bo c48b561b70 chore(lint): enable errorlint + unused, burn backlog (age-a-plus-report-card-ieyp2.8) 2026-07-19 17:06:04 -04:00
Bo 3069bb695f test(cli): AST regression guard for goals+status carve-out 2026-07-19 15:54:17 -04:00
Bo 48f1a0ce21 refactor(cli): carve goals into internal/commands/goals module 2026-07-19 15:52:35 -04:00