91 Commits

Author SHA1 Message Date
Bo c6558508d1 Consolidate AgentOps into a 34-skill engineering menu (#1133)
AgentOps' 55-skill catalog contained overlapping entry points, stale
routes and descriptions that could lose meaningful guidance in the Codex
projection. This change consolidates 21 roots into existing owners,
leaving 34 distinct skills and a generated, task-oriented menu. README
documents every retired name and its replacement.

Planning now establishes observable behavior in the caller's existing
intent, using proportional Given/When/Then examples and domain language.
Implementation and final validation carry those same examples forward.
Original adaptations informed by Matt Pocock's engineering skills
strengthen existing owners rather than adding a new workflow. Routine
edits need no mandatory plan, coverage report, mutation exercise or
learning artifact.

Codex retains complete source descriptions and translates explicit-only
invocation policy. All descriptions fit the existing 180-character
limit; the root instructions retain their 250-line limit. Generated
catalogs, projections, routers, moved references/helpers and their live
consumers are updated together. RPI remains explicitly selected.

Validation passed: projection/conformance checks, the local aggregate
(10 passed; one existing optional-directory skip), and exact-commit CI
covering the complete gate registry, Bats, Go build/vet/race/coverage,
Windows and security. A fresh author-distinct reviewer passed all
acceptance criteria over the complete 573-path subject at
aa642a55d6, including the installed-link
and protected-backup changes. Review findings were repaired and
revalidated. Existing ranker goldens are regression checks, not
model-quality measurements. A fixed six-case fresh-context pilot
supplied an exact candidate menu: three of four targeted cases loaded
expected guidance, a simple refactor selected no skill, and both
no-skill controls selected none. No wrong owner was selected. This pilot
preceded final wording repairs for existing ranker/context limits; it
does not establish installed automatic activation, coding benefit or
savings. No live coding task was run in that pilot.
2026-09-10 22:18:05 -04:00
Bo 5e874b55cf Evaluate installed skills on isolated Go work (#1125)
AgentOps previously relied on behavioral probes and retrospective
summaries to assess skills. This adds a development-only evaluator that
runs a frozen installed skill package on isolated Go tasks, preserves
failed and interrupted attempts, and rebuilds a comparison readout from
native results without another model call.

The suite contains six task families, separate executable verifiers,
frozen launch identities, native session accounting, and a focused
`skill-eval` maintenance workflow. The readout separates passing code
from completed trials, retains incomplete cost information, and reports
missing evidence without claiming equivalence or uplift. `ao eval`
remains retired; no new runtime controller or required core skill is
introduced.

Validation: Go build/vet/race checks and all repository CI passed.
Focused reader/statistics, receipt integrity, verifier integrity,
fixture calibration, generated projections, and the local aggregate
runner passed. A real two-variant Docker preparation check verifies that
frozen worker and verifier images survive later staging.

The bounded coding pilot retained all 24 starts and produced eight
comparable pairs across six task families, with no observed paired
endpoint difference. The separate eight-start memory experiment did not
demonstrate incremental benefit and does not promote another guidance
rule. Individual runtime limits were enforced; aggregate desktop
deadline enforcement remains unproven. Raw trial evidence and
credentials stay outside Git.
2026-09-10 13:29:10 -04:00
Bo e1fae0dae6 Make the engineering harness lean and add optional topic memory (#1116)
RPI now owns the authorized outcome through finish, with Plan and Memory
loaded only when useful. Known defects get direct repair, and evidence
can change the approach under unchanged acceptance. Fresh exact-content
validation remains required. Memory provides optional recall, mining and
curation of reviewed topic pages; specialists and the fixed-dispatch
adapter remain optional.

The change reconciles current documentation and generated skill
projections. It preserves native budget and permission authority, BD
work ownership, protected external evidence storage and the distinction
between a supported lesson and demonstrated later benefit. It adds no
scheduler, work store, Go command or evidence schema.

Validation: required local Go/build/vet/race checks, aggregate suite,
generated-output check and 72 gates pass. The complete 44-test executor
suite passes; its shared-deadline fixture now tolerates CI scheduling
jitter while still requiring deadline exhaustion and preventing a third
launch. Fresh author-distinct review passed all 112 changed paths with
no findings; all seven exact-head CI checks passed at bfce33cce. Native
restricted-source enforcement and reduced token use are not established
by this change.
2026-09-09 10:10:29 -04:00
Bo 10f0277bdb Legible membrane, Train 2: what a stranger meets (#1100)
## Legible membrane, Train 2: what a stranger meets

Provenance: the 2026-09-02 field audit of this repo against
mattpocock/skills, compound-engineering, and the jsm corpus, findings F5
through F9. This train is sized by a consumer inventory built with `rg`
on the tip before any lane was written; the promoted-set directory move
the audit proposed is deferred because that inventory shows
skill-builder backing two blocking gates, swarm pinned by the cathedral
gate and a routing golden, using-gc required by Go code, and `ao skills
link` unable to install a second root. That inventory is the plan for a
later train.

**What changes.**
- **Archival sweep by consumer disposition.** 172 audit snapshots, 29
pawl receipts, the `evals/workbench` and `evals/membrane` trees with
their two bats consumers, four stray scratch docs, four retired eval
contracts, and nine caller-less `scripts/check-*.sh` are deleted; git
history is the archive. Every machine list that referenced them is
pruned (evidence-grounding baseline, preamble grandfather, broken-links
allowlist, `.gitattributes`, `.gitignore`, two eval fixtures, the
workflow-coverage deferred list). `docs/audits/manifests/` and
`.agents/ao/config.yaml` survive because they have live readers. About
48,000 lines.
- **Three skills retired.** `goals` (alias of fitness), `shared`
(tombstone), and `scope` (folded into plan step 3 as five write-scope
checks). Consumers edited; the probe denominator exclusion for goals
pruned; Codex package and golden count pins updated.
- **Negative routing** on research, codebase-recon, reverse-engineer,
premortem, one-way-door, and council, all within the 180-char budget,
with a teardown golden (`rq-08`). One wording was changed after the
router's prefix stemming showed "repository teardown" leaking into the
wrong skill.
- **Every promoted skill answers "It's working if"** with observable
tells in backticks, and carries a paste-ready `## Prompt` with a
concrete subject. Two fictional `ao` subcommands a draft prompt named
were caught by the body-ref validator and replaced with real commands.
- **Doctrine diet on the core five.** rpi, plan, implement, validate,
and anti-ceremony drop from about 5,100 words to 3,600 (bodies from
4,700 to 3,150) by moving the shared ownership boundary, dated
incidents, and mechanics tables into step-loaded references
(`skills/rpi/references/boundaries.md`,
`skills/validate/references/mechanics.md`,
`skills/plan/references/ground-truth-routing.md`). Every cathedral
canary and every skill validator grep survives unchanged.

- **ADR-0018** records the goals, shared, and scope retirement; the
cathedral gate tombstone and the routing goldens cite it instead of
ADR-0017.
- **Router and twins.** `ao skills find` holds a description's "Not for
X; that is <sibling>." sentence out of its haystack, so premortem no
longer ranks first for "is this live decision reversible" (golden
`rq-10` pins the reciprocal of `rq-02`); a penalty variant was tried and
reverted because it suppressed skills the caller named outright. A
declared trigger phrase of two or more words quoted whole in the query
now earns the name weight once, so "check this change" lands on validate
rather than on reality-check's name token; a live-catalog test pins
seven such queries. Single-quoted YAML descriptions unescape `''`. The
Codex catalog keeps the exclusion sentence, and a closing `>` no longer
turns `<run-id>/codebase-recon.json` into an invocation.
- **Residue the judges found.** handoff, learn, and status open a `##
Contract` heading after their tells; validate's prompt names its helper
at `skills/validate/scripts/validate.py`; the explicit-skill prompt
catalog names only live skills (five stale prompts replaced by nine,
floor 20 restored, TESTING.md names the suite); the corpus-delta receipt
binds the runner's path and SHA-256 and labels a `live_agent` claim as
an unverified caller declaration; the probe README and ledger describe
the 12-skill denominator; SKILL-API counts 30 of 54.

**Evidence on the tip.** Regen check clean; full gate green with a
HEAD-built binary; CI's bats command green; Go build/vet/test green;
lint clean; security gate quick PASS; docs-build warnings did not rise.
Fresh validation by Fable 5.1 (caller-elected) and a cross-family read
by Codex, both recorded in the PR thread.

---------

Co-authored-by: Bo <bofuller55@gmail.com>
2026-09-03 19:52:55 +00:00
Bo 568e99d436 Loop restore: converge and crank as control flow under the verdict contract (ADR-0017) (#1099)
## Loop restore: converge and crank as control flow under the verdict
contract (ADR-0017)

Intent source: `docs/plans/2026-09-03-loop-restore.md` (in this PR).
Decision record:
`docs/adr/ADR-0017-loop-as-control-flow-not-knowledge.md`.

**Why.** The 2026-07-14 single-pass cut (`482307762`) removed the
iterate loop (discovery, crank, converge, evolve, the learn write-half)
together with the unproven compounding claim, although ADR-0011 demoted
only the latter. The control flow was never demoted, and its absence
showed on 2026-09-02, when a three-lane fix needed eight validators and
two stops because the contract had no repair phase. This restores the
loop as control flow and nothing else: no knowledge store, no `ao
converge`/`ao crank`, no evolve, no canary. ADR-0004 and ADR-0011 stay
in force.

**What changes.**
- **RPI gains a bounded repair phase.** On `FAIL` or `NOT_PROVEN` with
findings, repair and re-validate freshly under the convergence law:
caller-declared `repair_rounds` (default 2); open finding set keyed by
stable `findings[].id`, union across validator families, non-growing; no
closed id reopens; the subject digest changed or, for `NOT_PROVEN`, new
digest-bound evidence resolved a named gap. Converged = fresh PASS plus
cross-family PASS on risky surfaces. Plan and Implement keep their
single dispatch. `skills/rpi/scripts/run_once.py` models the law as pure
data (33 tests): rounds are validated for shape (digest required, no
duplicate ids, no PASS with findings, no FAIL without findings),
condition 4's evidence branch needs a NOT_PROVEN previous round, a
non-FAIL current round, new evidence, and a resolved finding, and a PASS
over unchanged bytes after a FAIL is a flip that reports NOT_PROVEN.
`workflows/rpi.js` runs validation as legs (spawned or external primary,
plus a caller-supplied `crossFamily.command` on risky surfaces) merged
worst-of with a union of stable ids; a risky surface without a
cross-family leg is `diversity_unsatisfied` and never converges or
enters repair; a failed repair or re-validation returns NOT_PROVEN with
no stale verdict. Validators return `subjectDigest`, stable finding ids,
and `evidenceRefs`.
- **crank returns as a thin wave executor** (113 lines): the caller
selects the wave and the repair bound, crank invokes RPI per lane
(parallel only on disjoint write and regen scopes), runs the wave
acceptance once, returns evidence, and stops. No retry, budget, queue,
claim, lease, Git, closure, or next-work ownership. Routing golden
`rq-07-wave-execution` ranks it first.
- **validate is cross-family by default on risky surfaces**
(`cli/internal/gates/**`, `scripts/check-*.sh`, `tests/**`,
`skills/*/scripts/**`, hook policies, `lib/**`, security-scanned paths)
with the LAW-0 dispatch table: Claude orchestrating uses read-only
`codex exec`; Codex orchestrating uses an interactive Claude session in
an NTM pane, never `claude -p`. No live adapter means
`diversity_unsatisfied`, which on a risky surface is `NOT_PROVEN`. The
full literal CI command set runs once on the final integrated subject;
routine rounds keep the receipt-driven freshness contract.
- **Conformance assertions flipped under ADR-0017 only:**
`scripts/check-cathedral-cut-conformance.py` (crank live; "Stop
regardless" replaced by positive canaries for the law's four conditions;
a bounded `for` loop that compares against `repair_rounds` is required
in `run_repair_phase`, and the gate executes the law's canaries against
the reference behavior), `workflows/rpi.js`,
`skills/rpi/scripts/validate.sh`,
`evals/agentops-core/rpi-behavior.json`,
`skills/rpi/references/rpi.feature`. Every single-pass public surface
(README, AGENTS.md, PRODUCT.md, CI-CD, agent-workflow-reference,
rpi-traversal, cli/README, quickstart and demo commands, the
operating-contract and product-boundary bats, the Codex-description
oracle) now states repair to convergence.

**Known approximation, disclosed.** The Claude conveyor has no
deterministic shell primitive, so changed paths are derived by the fresh
validator (git status and diff against the clean pre-run tree) and
unioned with the implementer's report; risk is classified over that
union and unreported paths are coverage findings. A validator is still a
model; runtime derivation outside every agent is a follow-up. Family
distinctness of the cross-family leg is asserted by the caller's choice
of command and not verified by the script.

**Not in scope.** Premortem stays a single advisory judge and Plan still
only names the first check (phase boundaries unchanged). No `verdict.v2`
or `rpi-report.v1` change. The loop's own effect on outcomes is
unmeasured and owed a seeded-defect probe, like the rest of the corpus.

**Evidence on the tip.** Regen check clean; full gate green with a
HEAD-built binary; CI's bats command green; Go build/vet/test green;
golangci-lint clean; security gate quick PASS; one fresh validator over
the whole diff; one cross-family read of the design before
implementation (13 findings folded) and two of the integrated diff (9
findings in round one, 11 by round two, 15 by round three, each round
repaired and re-reviewed; the fresh validator passed the tip after round
two and the final tip 1e8adb72d passed a fresh validator (14-scenario
independent harness of the law, full gate 71/71 with a HEAD-built
binary, CI bats 1164/0) and a cross-family read by Gemini 3.8 via AGY,
which closed all six remaining residues with no new findings; Codex was
unreachable at push time).

**Follow-ups filed from the final reviews, not blockers:** the JS
violation check tests growth before reopen while Python tests reopen
first (same stop, different label when both occur in one round);
`cli/testdata/compatibility-baseline/families/{demo,quickstart}/case.json`
assert help-text substrings Cobra never prints (pre-existing, no
consumer); runtime derivation of changed paths outside every agent in
the Claude conveyor.
2026-09-03 15:16:08 +00:00
Bo 8cdcb5a903 Train 1: measurement substrate, context diet, retrieval-eval contract (instrument-panel roadmap) (#1087)
> **Residues closed on the caller's merge instruction** (`499d916a6`):
the round-2 findings were the same failure shape — round-1 repairs
patched cited lines instead of sweeping the class — so this commit
sweeps each file whole: every remaining SATURATED-row-append site in
skill-eval now routes to RUNBOOK retirement, the human-only-skills
*description* is runtime-conditional, premortem's "(MEASURED)" label is
gone, SKILL-API's context table carries all 25 rows and the enforcement
table gains `disable-model-invocation`, and the fixture-identity claim
is stated precisely (probe id, honesty note, and control arm are the
only differing fields — as the acceptance permits). Post-sweep:
validators, full Go suite, 68/68 gates, goldens + headroom bats green,
projections current, gemini in sync. Merging per Bo's instruction.

## What

Train 1 of the accepted [instrument-panel
roadmap](docs/plans/2026-08-26-instrument-panel-roadmap.md) (intent
landed at `986a4feaf`): the measurement substrate, the skill-context
diet, and the retrieval-eval contract. Three worktree-isolated lanes,
each independently validated by a fresh context, plus one integration
commit. 103 files, +5,510/−76.

**L1 — measurement substrate** (`instrument/measurement-substrate`)
- Gate `skill.probe-headroom` (advisory, Fast|Full): answers the
question `skill.probe-coverage` cannot — not "does a probe result exist"
but "could one have existed at all". The rule, ported to Go
(`cli/internal/probeheadroom` + `cli/cmd/probe-headroom` behind a thin
check script — the witness-crosscheck pattern, **no new `ao` root
command**): control arm ≥ 0.75 with ≥ 2 usable reps at ≥ 2 effort levels
⇒ SATURATED (void row, not an honest null); UNMEASURED outranks it;
treatment-silent ⇒ FLOOR; else SEPARATED. RED first: both committed
fixture pairs read `INERT` to everything else in the repo; the failing
separation test predates the implementation, and a bats negative-control
swaps fixture bytes and asserts the gate flips.
- **First reading on real data: 7 of 11 historical probe groups are
SATURATED** — including both `validate-not-proven` runs. Those INERT
rows were never honest nulls; they were void. The 0/12 ledger number now
argues itself.
- Declared denominator for probe-coverage:
`scripts/.skill-probe-denominator-exclusions`, fail-closed parser (entry
without an argument, stale slug, or duplicate ⇒ exit 2). One entry
(`goals`, a pure alias-of `fitness`). Net effect deliberately zero (0/12
→ 0/12: alias left, `one-way-door` entered) — the gain is a declared
number, not a better-looking one.
- Re-landed from the recovered clean-room commit (`9872483bd`),
re-validated against *current* main: `skill-eval` (defers saturation to
the gate id; its shell scripts dropped, not shipped — ratchet intent),
`route`, `one-way-door`, premortem reversibility check, council
`caller_challenge` (schema + validator, per the agent-core boundary that
the panel may challenge, never overrule).

**L4 — context diet** (`instrument/context-diet`)
- `disable-model-invocation: true` on 4 human-only skills (key verified
verbatim against Anthropic's docs). The plan guessed 35 candidates; the
graph said otherwise — 23 carry `user-invocable: true`, and 19 of those
are excluded on cited evidence (rpi consumes
anti-ceremony/implement/plan/validate; workflow scripts reach others;
`goals` is a live migration tombstone). The exclusion evidence is
retained in the lane report.
- One router skill (`human-only-skills`) — the single always-loaded
description that replaces four; it hints, never fires.
- `.out-of-scope/` formalized with this week's three refusals
(checked-in knowledge corpus; ee self-improvement loops; whole-skill A/B
as the measurement unit), each citing its evidence.
- Deterministic proof, no model eval: before/after bytes of
always-loaded description load reported in the lane summary.

**L5 — retrieval-eval contract** (`instrument/retrieval-contract`, lane
verdict PASS 10/10)
- `AGENTS.md` federated row now names **ee (eidetic-engine)** as a
concrete caller-selected memory system — consume, never build; symlink
intact.
- `schemas/pack-quality-expectations.v1.schema.json` + 4 routing goldens
+ `scripts/check-routing-probe-goldens.sh` graded against `ao skills
find`, wired as an **advisory** nightly job. Zero goldens is a failing
state — no new zero-denominator green.
- **The instrument caught a real miss on day one — and its own
prescription fixed it.** Golden `rq-04` expects `validate` for "judge
whether this finished change is actually proven before I merge it"; at
authoring, `ao skills find` ranked the *forbidden* `premortem` first and
`validate` nowhere in six natural phrasings. The pointer-wording-first
repair (validate's description gained the caller's own words: finished,
proven, verdict, merge) now ranks it #1 at 0.333; grader 6/6, and the
golden pins the repair — a description regression reopens it.

## Integration

`regen-all.sh` once over the merged lanes (catalog 52 → 56, four new
codex twins, mesh, router, manifests); `skills/route/SKILL.md`
catalog/router links became prose repo-root references (the projected
twin cannot resolve `../catalog.json` — this was both the
portable-conformance failure and the sole broken doc link);
`codex-portable-conformance.bats` pin 52 → 56.

## Evidence

- `cd cli && go build ./... && go vet ./... && go test ./...` exit 0 ·
`ao gate check --full` **68/68** · four skill validators PASS ·
probe-headroom / routing-goldens / probe-coverage bats PASS ·
`regen-all.sh --check` all current.
- Per-lane fresh validators re-ran every suite on detached content; L5
PASS; L1/L4 NOT_PROVEN solely on the projection-regen clause reserved
for integration (their remaining acceptance observed green), settled
above. Cross-family (Codex) review of the integrated diff recorded in
the session report.
- Two disclosed scope stretches accepted at integration: a one-line
`.gitignore` entry mirroring the witness-crosscheck precedent, and the
probe LEDGER.md fact-correction L1's own change made necessary (noted
for Train 2's L2, which owns that file next).

## Cross-family review (Codex, fresh context)

Round 1: **FAIL** — two blockers (the RED fixtures didn't isolate the
control arm; the goldens grader was red where the plan's acceptance says
green) and eight majors (contract contradictions in the re-landed
skills, a converter-substitution false claim in the codex router twin,
two overreaching `.out-of-scope` entries, stale SKILL-API counts). All
repaired in one bounded round (`db68935a3`): fixtures now byte-identical
outside the control arm, the routing miss actually fixed rather than
tolerated, every cited contradiction reconciled at the source and
re-projected. Post-repair: full Go suite exit 0, `gate check --full`
68/68, all validators and probe/goldens bats green, projections current,
gemini byte-identity restored. Focused re-check verdict recorded in the
session report.

## Follow-ups (Train 2, already planned)

Seeded-defect probes for the judgment spine (every ledger row citing a
passing headroom pre-screen) and the gate-hardening pair
(`Gate-Loosen-Reason` tightening ratchet; mechanical
grounding-validation over evidence docs). Plus, surfaced by this train:
a latent `valid_keys`/schema divergence in `validate-skill-schema.sh`
(two keys the schema defines are absent from the script's allowlist —
pre-existing).
2026-08-26 23:16:51 +00:00
Bo 621dbb575f 3.6.0 release prep: version bumps, changelog, curated notes (#1071)
Everything-but-tag for **v3.6.0**. Minor, not major: the post-3.5.0
delta retires the knowledge-flywheel product surface and aligns the
estate on the operations-layer identity, matching the 3.4.0 precedent
where the orchestration pack was removed in a minor.

## What this carries

- **Version 3.5.0 -> 3.6.0 across all seven surfaces**: Claude plugin
manifest, marketplace metadata + plugin entry, Codex manifest, Gemini
image manifest, Claude image verify pin, and the `ao` source fallback.
- **CHANGELOG `[3.6.0]`** (root + docs mirror): operations-layer
alignment, anti-ceremony enforcement, the behavioral eval program, the
flywheel retirement, and the honest 0/12 measured probe coverage.
- **Curated `docs/releases/2026-08-17-v3.6.0-notes.md`**: validator
PASS, tier minor, full changed-path area coverage. The Breaking Changes
section lists all six removals and the handoff write-path move rather
than burying them in a minor.
- **New regression test `cli/cmd/ao/version_manifest_parity_test.go`**
binding the `version` fallback to every version-bearing release surface.
- **PRODUCT.md** reviewed against the 3.6 surface and re-stamped;
**`docs/reference/skill-system-evolution.md`** gains its 3.6.0 row and
drops the "current unreleased tree" framing that the tag would falsify.

## Why the new test exists

This cut missed `images/claude/verify.sh`. Its version guard — whose
entire stated purpose is catching plugin.json drift *behind* the release
— then rejected the **correct** version, so a user following the shipped
`images/claude/README.md` on the v3.6.0 tag would have hit a hard FAIL.
`check_manifest_version_consistency` in `ci-local-release.sh` compares
only the two Claude manifests to each other, so it structurally could
not see this.

The test fails on the drift and passes when correct; both directions
were exercised before committing.

## Honesty notes carried into the release

- Measured behavioral probe coverage is stated as **0/12** under the v3
evidence contract. The earlier wave-1 classifications are retained as
`LEGACY-UNVERIFIED` rather than counted, because the probe harness did
not isolate the skill corpus between control and treatment arms.
Skill-efficacy claims in these notes are directional, not proven.
- The estate-ablation aggregate counts are labeled legacy-unverified and
non-promotable.

## Verification

Full `scripts/ci-local-release.sh --release-version 3.6.0
--readiness-mode official --security-mode full`: **PASSED — 72 checks, 0
failures**.

| Dimension | Status |
|---|---|
| SIL (race suite, 75.8% coverage) | pass |
| VIL (gates, regen, digital twin) | pass |
| HIL (real Darwin/arm64 target) | pass |
| Artifacts (CycloneDX + SPDX SBOM) | pass |
| Security (full mode) | pass |

Readiness **9.0** against threshold 8. HIL used a real target with **no
waiver**: `ao` built from this tree reported `ao version 3.6.0`
(`version_verified=true`) and ran a full `ao init` scaffold plus `ao
status` in a scratch repo. Security full mode: 0 critical, 0 high, 3
medium (non-blocking). Notes validator PASS, doc-release gate PASS.

## Post-merge

Readiness lap at the merged SHA, audit record in `docs/audits/`, then
the tag — per the binding process rule that the record exists **before**
the tag.
2026-08-17 21:24:34 -04:00
Bo f9e246d9fd fix(skills): make audit grades evidence-honest (#1066)
Audit all canonical skills against a current evidence rubric, distinguish static readiness from safety and effectiveness, strengthen scorer/schema truthfulness, and publish bounded remediation.
2026-08-16 18:55:41 -04:00
Bo f3c6d0ecf2 Converge retained WIP and harden evidence boundaries (#1065)
Summary:
- lands the audited current WIP lanes and excludes stale/process-only
material
- hardens prune path confinement, probe-v3 evidence binding,
codebase-recon identity, handoff/release/reverse-engineer behavior, and
Codex prompt handling
- truth-labels static skill scoring and regenerates all owning
projections

Validation:
- fresh independent PASS on commit
427098ed10
- full Go suite and full Go race suite
- quick local release CI
- focused Bats, scenario/linkage, native-skill, reverse-engineer,
Cathedral, Ruff, Python ratchet, projection, and diff checks

Residual boundaries:
- final-basename ABA remains unclaimed
- live skill-probe coverage remains honestly 0/12
- release-only cross-build, SBOM, vulnerability, and release-evidence
checks are left to delivery CI
2026-08-16 18:32:26 -04:00
Bo 8694b07ebf feat(skills): add anti-ceremony guard to RPI (#1064)
## Summary\n\n- add a clean-room, artifact-free anti-ceremony skill\n-
make RPI invoke its STOP/CONTINUE guard exactly once before Plan\n- STOP
before dispatching any core phase; CONTINUE preserves Plan → Implement →
fresh Validate\n- retain durable intent by reference and snapshot only
non-durable intent\n- update the canonical traversal, workflow,
standards, dependency contracts, tests, and generated projections\n\n##
Verification\n\n- RPI unit tests: 13/13\n- dependency/product-boundary
Bats: 15/15\n- RPI and Anti-Ceremony validators: PASS\n- Cathedral Cut
conformance: PASS\n- mesh, Codex/Gemini projections, parity, hashes,
manifests, and runtime formats: PASS\n- changed-scope
generation/contract checks: PASS\n- full exact-range gate: 68 pass, 0
warn, 0 fail, 0 unknown, 0 skip\n- fresh final semantic review on head
8b4a6a5ef: PASS\n\nThe branch was rebased onto the merged skill-contract
work and regenerated from canonical owners. Unrelated working-copy
changes remain excluded. Hosted Validate is running on the exact
reviewed head.
2026-08-16 16:02:53 -04:00
Bo 8e815363ab docs: add Skill-System Evolution analysis and site nav entry (#1063)
### Motivation

- Capture the full evolution of the skill system across releases and
synthesize the meta-loops, graphs, and lessons needed to treat skills as
the leverage point for validated recursive improvement.
- Provide a single, reviewable reference that maps releases to
architectural regime changes and that prescribes defensible protocols
for skill-driven compound engineering.

### Description

- Add a new research/analysis document
`docs/reference/skill-system-evolution.md` that traces each changelog
release from `2.16.0` through `3.5.0` (plus the unreleased state),
synthesizes seven evolution regimes, diagrams the present authority
graph, and enumerates surviving/retired meta-loops, design implications,
an improvement protocol, and falsifiable predictions.
- Register the new page in the site navigation by updating `mkdocs.yml`
so the analysis is discoverable under the Skills section.
- The document explicitly records evidence sources and method, clarifies
the boundary between deterministic mechanics and semantic skill
behavior, and prescribes the minimum learning record and stepwise
validation protocol for safe skill evolution.
- Small whitespace corrections were applied to the generated doc to
satisfy repository checks.

### Testing

- Ran `git diff --check` and addressed whitespace issues reported by the
check, which resolved pre-commit concerns. (passed after the fix).
- Executed a Python verification that asserts all 47 changelog release
sections are represented in the trace, which succeeded.
- Ran `make docs-check` (which exercises the docs
validation/cli-reference checks), and the invoked
doc-generation/validation script completed successfully.
- Confirmed working-tree state and that the new file is included in
documentation navigation via a local commit and status inspection.

------
[Codex
Task](https://chatgpt.com/codex/cloud/tasks/task_e_6a811cced07c832babaf3756bfe354a3)
2026-08-16 19:15:36 +00:00
Bo 5f8dab3b1d feat(skills): using-flywheel — operating manual for the second supported factory (#1017)
Closes finding 3 from the GC-work analysis: the README presents two
supported factories, but only Gas City had an operating manual
(`using-gc`). The Flywheel's "supported" claim had no surface behind it.

- **agent-flywheel.com verified live**: Jeffrey Emanuel's open-source
VPS factory (Claude Code / Codex CLI / Antigravity CLI + NTM, Agent
Mail, Beads/BV) — matches the README's description.
- **New `using-flywheel` skill** (scaffolded via skill-builder, `heal
--check --strict` clean): caller-selected adapter, upstream-wizard
provisioning (no fork/pin — upstream owns its installer),
skill-visibility check per runtime, presence≠invocation discipline,
2-round non-convergence stop, and the boundary that factory state never
becomes an AgentOps verdict.
- README factory bullet now cross-links the skill; mesh/codex/gemini
projections regenerated (51 skills); CHANGELOG Unreleased entry (move
into the 3.4.0 section if the tag lands at tip after a readiness
re-lap).

Verification: `regen-all.sh --check` all green; `go build/vet` +
gates/cmd test packages 656 passed.
2026-07-30 10:44:54 -04:00
Bo a6359795bf Make fresh validation persistence optional (#1012)
Keep fresh author-distinct validation mandatory while making verdict and report persistence consumer-driven. Align the RPI/Validate contracts, executable behavior, current guidance, regression coverage, and generated projections; preserve the Gas City 1.4 cutover.
2026-07-29 19:50:56 -04:00
Bo 25db3280ce feat(gc): integrate Gas City 1.4 packs and runs (#1011)
## Summary

- pin and fail-close on Gas City v1.4.0, paired Beads v1.1.0, and the
official `gascity` 0.1.6 registry release
- compose the official workflow pack while binding its sibling roles at
stock `defaults.rig.imports.gc`, so `do-work` resolves `gc.run-operator`
and `gc.implementation-worker`
- retire the v1.3 heartbeat workaround, use scoped
`core.control-dispatcher` propulsion, and keep the Mayor as an on-demand
operator door
- update the canonical `using-gc` skill, generated projections, and the
old-city retirement/new-city startup runbook

## Evidence

- 38 focused Python tests pass (2 opt-in native skips)
- `scripts/check-gc-executor.sh` passes: 34 tests (2 opt-in native
skips)
- `scripts/regen-all.sh --check` passes
- pinned native v1.4 bootstrap/formula/roles/doctor/teardown
qualification passes twice: author run 122.910s, independent Fable run
116.267s
- fresh Claude Fable 5 semantic verdict: **PASS**, artifact digest
`7ab1a3be2a63cb03e774936f1551176c6ab61f8214ae35e0a97d635edf4833c7`

The live mixed-provider canary remains explicitly post-merge release
qualification; this PR does not claim it already ran.
2026-07-29 23:14:55 +00:00
Bo 5908024782 Structural: goals→fitness rename with compatibility alias; shared retired to a declared contract owner (#1007)
## Summary

The two structural beads deferred from the skill-overhaul reboot waves.

**goals → fitness** (`age-skill-overhaul-reboot-sjv7v.10`, 07-24 plan
D9)
- The semantic skill is now `fitness`; the `ao goals` CLI command family
deliberately keeps its name.
- `skills/goals` remains a thin, non-advertised compatibility alias:
user-invocable, zero capabilities/effects/dependencies, an `alias-of`
context_rel edge to fitness, physical deletion only under the
observed-zero policy.
- `alias-of` added to the frontmatter schema enums (the catalog schema's
own description already listed it as valid relationship vocabulary; the
enum lagged).
- bootstrap's `goals` consumer edge migrated; the stale `goals.feature`
row in the scenario-linkage contract corrected (the feature file was
already absent).

**shared retirement** (`age-skill-overhaul-reboot-sjv7v.11`, 07-24 plan
D9)
- The runtime-neutrality contract moved to its declared owner:
`docs/contracts/runtime-neutrality.md`, scope unchanged (shared
references a consuming skill loads).
- The last consumer edge (`bootstrap consumes: shared`) removed;
`skills/shared` reduced to a non-routable tombstone pending the
observed-zero deletion window; the ports-and-adapters contract updated
to executed state.

## Validation

- Strict frontmatter 50/50 (fitness + alias); regen byte-current;
scenario-linkage PASS; liveness/anti-spiral/desc-budget bats green;
`cli/internal/skills` Go tests pass
- Cross-family review, two rounds: round 1 — one finding fixed
(unintended contract-scope broadening reverted), one rejected with
evidence (the reviewer inferred a pre-existing feature file from a stale
doc row; the pre-change tree had none); round 2 delta **VERDICT: PASS**

Tracker: `age-skill-overhaul-reboot-sjv7v.10`, `.11`
2026-07-29 05:03:12 +00:00
Bo 894c190737 W8 coherence: corpus-level trigger separation, effects vocabulary, plan closeout (#1006)
## Summary

Final wave of the skill-overhaul reboot
(`age-skill-overhaul-reboot-sjv7v.9`), run after #998–#1005 merged.

**Closing corpus-level review** (the one check no per-wave review could
do — trigger routing, effects vocabulary, authority coherence,
reachability across all 49 skills): seven findings, all applied.

- Trigger separations: rpi↔implement ("execute this plan" = the loop;
"implement this bead" = the phase), plan↔craft-goal (Mayor/goal-runner
distinction made decisive), status↔reality-check (reality-check now
requires a claim to test), research↔codebase-recon (bare "investigate"
narrowed).
- Effects vocabulary: imperative `write_*`/`regenerate_*` forms
normalized (skill-builder), `authorized_` prefix made consistent
(reverse-engineer).
- `shared` marked explicitly non-routable.
- Authority coherence verified clean: only validate persists verdicts,
only rpi dispatches the loop, no specialist starts RPI.

**Deterministic evidence**: regen byte-idempotent (two runs, zero
drift); strict frontmatter 49/49 zero warnings; scenario-linkage PASS;
description budgets green; all skill validators pass.

**Plan closeout**: reboot plan marked complete with a completion record
(8 PRs, per-wave ledgers); structural beads .10/.11/.12 remain open as
explicitly deferred tracked work.

Full report:
`docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/w8.md`

Tracker: `age-skill-overhaul-reboot-sjv7v.9`
2026-07-28 16:13:46 +00:00
Bo c88a4514f9 W7 support wave: handoff schema truth, dcg fact corrections, honest support contracts (#1005)
## Summary

Wave W7 of the skill-overhaul reboot
(`age-skill-overhaul-reboot-sjv7v.8`) — the nine support skills, plus
the one Go fix where the skill contract crosses the CLI boundary.

- **handoff** — `ao session handoff --dry-run` output failed its own
`handoff.v1.schema.json` (reproduced: 3 errors). Fixed with a consumer
audit: schema keeps v1 with the doctrine-retired fields as optional
deprecated read-compat properties; the generator keeps its
collision-safe fractional id (schema pattern widened instead); real
jsonschema validation in `TestHandoffDryRunSatisfiesSchema` + a
legacy-artifact compat test; `read_clock` effect declared.
- **dcg** — corrected the false "`rm -rf ./build` allowed" claim (live
0.5.6 blocks it) and a nonexistent rule id in the allowlist example
(silent no-op) across six files; removed a token-splitting "workaround"
that was an executable guard bypass, replaced with file/stdin handling
and a never-reconstruct warning; temp-path rule live-probed and stated
identically in both docs; version/path/upstream corrections.
- **cc-hooks** — ships-by-default contradiction reconciled;
PATH-clobbering recipe fixed; operator-private paths removed from
shipped text; jq preflight added to the edit guard.
- **ms** — validator no longer mechanically asserts the false `effects:
[]`; it extracts the frontmatter and requires the exact honest effects
value.
- **account-rotation / status / sbh / bootstrap** — real effects
declared, both-tools-absent and destructive surfaces defined,
live-output overclaims narrowed, versions pinned.
- **shared** — advertising narrowed to the current no-bundled-references
state; retirement NOT executed (bead `.11`).

Ledger (32+8 fixed across two rounds / 8 rejected-stale / 8
deferred-with-reason):
`docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/w7.md`.

## Validation

- `go build` + `go vet` + 482 `cmd/ao` tests incl. the new schema-lock
and legacy-compat tests; dry-run validates 0 errors
- 49/49 strict frontmatter; regen clean; scenario-linkage PASS; liveness
+ anti-spiral + policy + edit-guard bats green; python ratchet
no-growth; shellcheck clean
- Cross-family review, two rounds: round 1 eight findings all fixed
(schema compat, id collision, real validation, security bypass removal,
live-probed temp rule, anchored greps); round 2 delta re-review
**VERDICT: PASS** with zero residuals

Tracker: `age-skill-overhaul-reboot-sjv7v.8`
2026-07-28 16:03:12 +00:00
Bo fd7af3174a W6 runtime wave: authorize-first host mutation, honest adapter contracts, fail-closed deadlines (#1004)
## Summary

Wave W6 of the skill-overhaul reboot
(`age-skill-overhaul-reboot-sjv7v.7`) — the eight runtime-adapter
skills, every checklist item verified against the live tree first.

- **rch** [P0] — shipped reference docs instructed "always do, never
ask" for remote `sudo chown -R`/`chmod`, daemon lifecycle, and fleet
toolchain sync. Now: a strictly read-only autonomous diagnosis tier;
every mutating first-action/self-fix across all seven reference docs
marked **(authorize first)** (mechanical sweep); authority banners on
the previously-ungated catalog docs; re-running the original command is
autonomous only when that command is itself read-only;
config-threshold/force_remote/allowlist edits and local key repair all
gated.
- **agent-mail** [P0] — default coordination mode split from an
explicitly-authorized admin/DR mode (gating
`clear-and-reset-everything`); retired `python -m` CLI rewritten onto
the verified `am` CLI; typed unavailable/timeout/cleanup outcomes.
- **codex-exec** — wall-clock deadline is mandatory (declared 600s
default), expiry is fail-closed, and process-group tree reaping is a
precondition: a host that cannot guarantee it does not run (capability
unavailable) — no disclosure escape hatch.
- **using-gc** — the conditional `~/.codex/config.toml` trust write and
the codex binary update are both gated behind explicit caller
authorization.
- **agent-native** — equal author/validator context ids rejected before
any freshness attestation; reproducibility via `SOURCE_DATE_EPOCH`.
- **swarm** — disjointness check case-folded; non-empty
`write_scope.exclude` rejected instead of ignored; two new test
witnesses.
- **agy-native / ntm** — discovery commands named,
unavailable/deadline/cleanup behavior explicit, effects declared.

Ledger (31 fixed / 3 rejected-with-reason / 7 deferred-with-reason):
`docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/w6.md`.

## Validation

- 49/49 strict frontmatter; regen clean; liveness + anti-spiral bats
green; scenario-linkage PASS; swarm tests 8/8 (2 new witnesses);
agent-native fake-runner 5/5; python ratchet no-growth
- Cross-family review, two rounds (ceiling): round 1 three findings
fixed (read-only autonomous tier, gated updates, mandatory deadlines);
round 2 residuals fixed with the exact prescriptions (read-only re-run
rule, config-mutation gates, fail-closed reaping — the disclosure escape
hatch removed)

Tracker: `age-skill-overhaul-reboot-sjv7v.7`
2026-07-28 15:56:23 +00:00
Bo 2ee168902e W4 specialists wave: converter destructive path closed, authority grants stripped, honest effects (#1003)
## Summary

Wave W4 of the skill-overhaul reboot
(`age-skill-overhaul-reboot-sjv7v.5`) — seven candidate-producing
specialists, every checklist item verified against the live tree;
heavyweights reproduced before fixing.

- **converter** [P0] — `convert.sh` could `rm -rf` a caller-supplied
output dir and destroy the source package (reproduced). Now refuses when
the canonicalized output equals, contains, or is contained by the
source, or is/contains the repo root — symlinks resolved, root
special-cased. Behavioral self-test (8 assertions) proves each refusal
happens before any deletion, with source survival checked by content
digest. Default output moved to `.agents/projections/converter/`.
- **scaffold** — deleted the commit/continuation authority grants from
`generic-templates.md`; validator now sweeps `references/**` for the
full grant vocabulary (proven failing on six restored variants).
- **refactor** — feature file stripped of commit + auto-revert
authority; revert contradiction resolved; `modify_source_files`
declared.
- **doc** — `--create-issues` tracker authority and Next-Steps
work-creation removed; artifact dir into the closed set. The
`[disputed]` vacuous-validator claim was verified false and rejected.
- **skill-builder** — build-report write declared and relocated;
`context_rel` added — **strict frontmatter is now 49/49 with zero
warnings corpus-wide**. The five-Python-files ratchet claim was verified
stale (files don't exist; live ones grandfathered).
- **test** — validator replaced with load-bearing checks (fails on
removed sections); `produces` reconciled everywhere incl. the feature
file; effects now include `modify_source_files` (TDD mode edits
production code).
- **workflow-builder** — `write_workflow_script` declared.

Ledger (30 fixed / 3 rejected-stale / 15 deferred-with-reason):
`docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/w4.md`.

## Validation

- Converter self-test 8/8 (descendant, repo-ancestor,
symlink-into-source cases; digest-based survival); in-scope validators
PASS
- Strict frontmatter 49/49 zero warnings; regen clean; 23/23 bats;
scenario-linkage PASS; python ratchet no-growth; shellcheck clean
- Cross-family review, two rounds (ceiling): round 1 five findings all
fixed (bidirectional guard, self-test hardening, sweep completeness,
feature reconciliation, effects); round 2's one residual (`within()`
breaking for parent=/) fixed and witnessed

Tracker: `age-skill-overhaul-reboot-sjv7v.5`
2026-07-28 15:44:42 +00:00
Bo a0d9c3b8e1 W3b evidence wave: security suite repaired, write confinement, honest contracts across seven skills (#1002)
## Summary

Wave W3b of the skill-overhaul reboot
(`age-skill-overhaul-reboot-sjv7v.4`, evidence half) — ~40 verified
fixes across seven evidence skills. Every checklist item verified
against the live tree; heavyweights reproduced before fixing.

- **security** — deleted the unrunnable skill-local gate duplicate
(exit-1 reproduced); redteam pack refreshed against current docs — the
`context-overexposure` case was **re-pointed at the surviving control**
(its doc was deleted by `482307762`; the trust rule lives in AGENTS.md
"Authority and trust"), pack 6/6;
`tests/scripts/test-security-suite-redteam.sh` was red on `main`, now
8/8; validate.sh is behavioral (runs the redteam fail-closed);
release/merge authority stripped; all three output artifacts declared
per command surface.
- **reverse-engineer** — writes confined under `output_dir`;
FileNotFoundError canned-learning path deleted; `_TBD` gate scans the
full seven-file audit bundle fail-closed (negative self-test proves it);
ZIP extraction bounded; behavioral validator.
- **cass** — three real effects declared; TMPDIR shadow + temp-dir leak
fixed; always-failing `jq -se` extractor fixed; no-artifact-dir contract
now true (retention path removed).
- **codebase-recon** — resolves against the target repo root
(`--repo-root` + `pwd -P`); claim evidence must be a file.
- **standards** — false-`effects:[]` template annotated at its origin;
dead anchor + missing owner rows fixed; bidirectional owner-table
validator (matches table rows only, not prose).
- **domain** — new liveness-proven validator (cited contracts must
exist, term must resolve).
- **research** — `write_research_report` declared; 3/3 scenario
coverage.

Ledger (~40 fixed / 2 rejected-stale / ~13 deferred-with-reason):
`docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/w3b.md`.

## Validation

- Redteam 6/6; security suite 8/8 (previously red on main);
reverse-engineer self_test green end-to-end; per-skill validators PASS
- 49/49 frontmatter; regen-check clean; liveness+anti-spiral bats green;
python ratchet no-growth (all Python edits in grandfathered files);
shellcheck clean
- Cross-family review, two rounds (ceiling): round 1 six findings all
fixed (incl. re-pointing the silenced redteam case with commit-traced
lineage, `pwd -P` across validators, full-bundle `_TBD` scan); round 2
residual on the owner-row prose-match fixed with a witnessed table-row
check; the remaining round-2 note (the redteam case covers the trust
half of the old control; the loading-bounds half is metadata-enforced
via per-skill `context:` blocks) is recorded in the ledger.

Tracker: `age-skill-overhaul-reboot-sjv7v.4` (W3a merged in #1000; this
completes W3)
2026-07-28 15:33:40 +00:00
Bo 15206c684f W5 evolution wave: honest effects, closed-set artifact dirs, provenance decay enforced (#1001)
## Summary

Wave W5 of the skill-overhaul reboot
(`age-skill-overhaul-reboot-sjv7v.6`) — the four evolution skills, every
checklist item verified against the live tree first.

- **learn** — pruned two dead `.agents/ao/verdicts/` digest citations
per the skill's own decay rule; the harvest example now cites the living
mutating-check quarantine rule in `skills/validate/SKILL.md` instead of
an uncited paraphrase (review round 1); declared the optional write
under `.agents/scratch/learn/`.
- **operationalize** — dropped the unbacked `.v1` contract suffix;
bounded its conditional write to `.agents/scratch/operationalize/`;
boundary tightened so advisory proposals cannot self-promote.
- **pattern-mining** — `effects: [write_pattern_evidence]` declared;
artifact dir moved to the ADR-0016 closed set; jq-presence guard added
to its output validator.
- **toil-mining** — real write effect declared; 3-way output-name
mismatch reconciled to `toil-candidates-report`; artifact dir relocated;
`user-invocable` flipped to match advertised triggers; frontloaded
Constraints/Quality sections.

Ledger (11 fixed / 3 rejected-stale / 4 deferred-with-reason, zero
silent drops):
`docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/w5.md`.

## Validation

- Skill validators green; 23/23 bats; 49/49 frontmatter; `regen-all.sh
--check` clean; python ratchet PASS (no new Python)
- Cross-family (Codex) review round 1: one finding (paraphrased pruned
incident) — fixed in 40339a81d with resolvable lineage

Tracker: `age-skill-overhaul-reboot-sjv7v.6`
2026-07-28 14:05:42 +00:00
Bo 69f055f8a2 W3a judgment wave: real output contracts, honest effects, scratch-tier artifact dirs (#1000)
## Summary

Wave W3a of the skill-overhaul reboot
(`age-skill-overhaul-reboot-sjv7v.4`, judgment half) — six judgment
skills, every checklist item verified against the live tree first.

- **council** — the dangling `council-report.v1` contract is now real:
skill-local schema + jq `validate-output.sh` (rejects smuggled
verdict/readiness fields, single-judge panels, empty findings) +
validator contract guard; boundary broadened to "does not mint a verdict
of any version"; negative trigger, judge-timeout failure semantics,
output location under `.agents/scratch/council/`.
- **reality-check** — same treatment for the dangling
`reality-check-report.v1`; adds no-verdict boundary and untestable-claim
failure line.
- **postmortem** — `output_contract` now names the real markdown report
(was a `.feature` path); `effects: []` → `[write_postmortem_report]`;
artifact dir moved to the ADR-0016 closed set
(`.agents/scratch/postmortem/`).
- **idea-genie** — artifact dir moved to
`.agents/scratch/ideas/<run-id>/` with bats fixtures updated.
- **premortem** — boundary fixed to deny any verdict version.
- **scope** — zero changes: verified already-tight; its checklist
enhances dispositioned reject/defer with reasons.

Notable: the checklist's cross-cutting "live loop emits verdict.v3"
claim was **verified stale** (main is verdict.v2; no validate_v3.py
exists) and rejected — the reboot's verify-before-edit rule doing its
job.

Ledger (17 items: fixed 5 / applied 5 / rejected 4 / deferred 3, zero
silent drops):
`docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/w3a.md`.

## Validation

- 23/23 bats (liveness + anti-spiral + native-skills fixtures); output
validators smoke-tested both directions; contract guards strip-tested
- 49/49 frontmatter; `regen-all.sh --check` clean;
`check-skill-python-ratchet.sh` PASS (no new Python — validators are
jq/shell); token budgets green
- Cross-family (Codex) review in flight; two-round ceiling per plan

Tracker: `age-skill-overhaul-reboot-sjv7v.4` (closes with W3b)
2026-07-28 13:56:28 +00:00
Bo 73954c0652 W2 product/campaign wave: declare real effects, kill stale spec, tighten contracts (#999)
## Summary

Wave W2 of the skill-overhaul reboot
(`age-skill-overhaul-reboot-sjv7v.3`) — contract enhancement of the four
product/campaign skills, per the distilled audit checklist (every item
verified against the live tree first).

- **goals** — declared the real write effect (`write_goal_snapshot`;
measure/drift/export persist a derived snapshot under
`.agents/ao/goals/baselines/`), documented all 8 subcommands
(`scenarios`, `render` were missing), fixed `produces` to the real
`goal-measurement-report`, and stated the goals source itself is never
mutated.
- **product** — required `## Proven` (every claim cited) vs `##
Assumptions` headings with an explicit aspiration-laundering detector:
an uncited Proven claim is laundered and moves to Assumptions.
- **craft-goal** — authority boundary (the emitted goal prompt is inert
caller-owned text; the skill starts nothing), context block + explicit
external-tracker declaration, embedded output-shape validator for the
terminal-token contract.
- **automation-shape-routing** — deleted the stale `skill.spec.json`
sidecar that contradicted SKILL.md on 6 fields and cited an out-of-repo
spike path; SKILL.md is now the sole metadata source.

Full per-item disposition ledger (fixed 4 / applied 6 /
deferred-with-reason 5 / silent drops 0):
`docs/audits/2026-07-28-skill-overhaul-reboot/wave-reports/w2.md`.

## Validation

- `validate-skill-frontmatter.sh` 49/49; `regen-all.sh --check` clean
(all projections through owners)
- `bats` liveness + anti-spiral suites 15/15
- No new Python, no schema/frontmatter fields added, no goals→fitness
rename (separate bead), no hand-edited generated files
- Cross-family (Codex) review in flight; two-round ceiling per plan

Tracker: `age-skill-overhaul-reboot-sjv7v.3`
2026-07-28 13:43:16 +00:00
Bo 16d764b5a1 feat: add craft-goal and humanize RPI reports (#988)
## What changed

- add `craft-goal`, a bounded Mayor-style goal compiler/linter for
persistent campaigns over bead-shaped RPI experiments
- add its Codex and Gemini runtime projections plus generated catalogs,
manifests, router, graph, maps, tiers, and registry entries
- split RPI reporting into two explicit surfaces: concise human-readable
interactive output by default, with the full `rpi-report.v1` object
retained as the machine/audit artifact
- add feature and validation assertions that prevent raw schema dumps
from becoming the default interactive response again

## Why

Recent AgentOps history repeatedly showed long-running goal campaigns
confusing the outer goal with one RPI experiment. It also exposed a
presentation bug where the RPI skill instructed agents to paste the
transport-shaped YAML report into chat. This change captures the
corrected goal model and keeps machine evidence separate from the
operator-facing response.

## Validation

- `bash scripts/regen-all.sh --check`
- `bash skills/craft-goal/scripts/validate.sh`
- `bash skills/rpi/scripts/validate.sh`
- `python3 -m unittest skills/rpi/tests/test_run_once.py`
- Codex parity, generated-artifact, and override-coverage checks
- `git diff --check origin/main...HEAD`
- `bash scripts/ci-local-release.sh --fast` ran all branch-relevant
checks successfully; its aggregate exit remained nonzero only for the
pre-existing `skills/using-gc` description-budget violation, which is
unchanged from `origin/main`
2026-07-23 23:24:56 +00:00
Bo 189bc81867 feat(gc): 3.3 flow is the mayor-driven door — human or agent orchestrated (#985)
## What

The 3.3 GC pack flow becomes Gas City's native human door, drivable by
humans and agents alike (operator decision). Five commits:

1. **a903c23f8** `feat(gc)`: Mayor prompt rewritten observe-only →
**shepherd** — polls external rigs' ready step beads and dispatches each
via `gc sling <gc.run_target> <bead> --nudge`. Never claims, never
authors, never closes.
2. **e58aeb5cb** `feat(gc)`: `invoke.sh mayor start|status|tell` — the
agent control surface over native primitives (`gc session wake`, `gc
session list --json`, `gc mail send --notify`; mail chosen over
sling-text to avoid v1.3.5's inline-text auto-create). Mayor is now a
`mode=always` standing session.
3. **bd8c34df6** `docs(gc)`: README — two doors (human attaches, agent
drives via mail), heartbeat propulsion, #4586 context.
4. **2b6ffe9cf** `feat(gc)`: `skills/using-gc` rewritten as the
mayor-orchestration skill (drive loop, bead-ids-only doctrine, GC
liveness truth stack with both codex wedge classes, stall protocol) +
all regenerated projections.
5. **362f4af79** `fix(gc)` (cross-family review: 1 BLOCKER + 5 MAJORs,
all fixed): **shepherd-heartbeat order** (native `gc order`, 3m cooldown
— mode=always keeps the Mayor resident but nothing re-prompts it; the
order is the liveness engine); HQ filtered from rig enumeration; stall
recovery = `session wake` (re-sling documented as a no-op for routed
beads); `mayor status` tri-state fail-closed; mail authority limited to
dispatch/status with bodies treated as untrusted data (pause/resume
dropped — stateless under fresh wake); test stubs reject non-allowlisted
verbs.

## Why

On v1.3.5, demand-driven worker spawning never fires for rig-routed work
(upstream gastownhall/gascity#4586, ours, with a stock-pack no-LLM
repro). Shepherd-dispatch by a standing Mayor is both the
v1.3.5-functional propulsion path AND the stock-GC operating pattern
(their mayor prompt: create → sling bead-id → monitor) — so this flow
survives the upstream fix rather than fighting it. Intent stays
caller-owned: drivers hand the Mayor bead IDs, never prose work.

## Verification

- `tests.python.test_gc33_thin_pack`: 34 passed (2 integration-gated
skips)
- `shellcheck -S warning` + `bash -n` clean; `scripts/regen-all.sh
--check`: all 11 projections current
- Every gc primitive verified against v1.3.5 source (cmd_mail.go,
cmd_session_wake.go, cmd_sling.go, cmd_rig.go, decode_sessions.go,
orderdiscovery)
- Codex adversarial review round: all findings fixed
- **Not yet proven live**: heartbeat-order discovery/scheduling,
mode=always residency, wake-based recovery, end-to-end propulsion — the
live canary follows this merge.
2026-07-23 17:08:28 +00:00
boshu 0b52397d87 refactor: fold dueling-idea-genies into idea-genie as duel mode
Single root owns elicit and duel modes; validator relocated byte-identical
as validate-challenge.sh; all consumers (bats, tuning test, clean-room
LEAVES, reverse-engineer link) updated; projections regenerated via owning
commands; router 49 -> 48.

Validated fresh: verdict PASS (9ab0dba3), intent 8fde0f49.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-15 23:06:16 -04:00
boshu 9aff5df0db refactor: fold heal-skill into skill-builder
One canonical root for skill create/heal/audit. Mode routing table absorbs
all heal-skill triggers; every live consumer (gates seed, quality checks,
release scripts, bats, docs, catalogs, parity twins) updated to relocated
paths; all projections regenerated via owning commands.

Validated fresh: verdict PASS (e9b6cdb8), intent 26a4f2be, all five
acceptance scenarios re-executed by an independent context.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-15 21:57:51 -04:00
boshu 5bdbb5fd09 refactor: simplify AgentOps loop and harden CLI 2026-07-15 19:17:08 -04:00
boshu baaa9e22da feat!: prepare AgentOps 4.0 cathedral cut 2026-07-15 10:00:26 -04:00
boshu cb2a7ac398 refactor: complete the cathedral cut 2026-07-15 00:21:02 -04:00
boshu 4823077621 refactor: cut AgentOps to a single-pass evidence loop 2026-07-14 22:01:50 -04:00
boshu ecd2ee8177 Direct-cut plan readiness into Premortem
Implements: age-agentops-lean-loop-direct-cut-rp16f.3
2026-07-14 13:59:00 -04:00
boshu 43b4387dbd fix(go): seal retired-owner and lifecycle boundaries 2026-07-13 10:24:38 -04:00
boshu a2423cfaca refactor(loop): require four umbrella receipts 2026-07-13 08:23:27 -04:00
boshu 9c6a662203 feat(loop): refactor the four-umbrella operating loop
Bead: age-four-umbrella-loop-refactor-xz4ps.1
2026-07-13 06:48:50 -04:00
boshu c6e23d2036 feat(skills): mesh native genies and factory orchestration (age-o43w) 2026-07-10 10:52:06 -04:00
Boden Fuller fbba8af5ac feat: land audit hardening and goal-design workflow 2026-07-09 12:20:53 -04:00
boshu 9a23ba9cc5 refactor(skills): execute the audit retire wave — 8 skills retired/merged, 66 -> 58, spine 15 -> 13 (age-skills-audit-fable-l6ic.12)
Executes the age-e3zk decision (fresh disposition pass: docs/audits/
skills-audit-2026-07-06.md; council 2026-07-06 parked this behind age-p2c7,
which landed as d7f950ca8; operator directive 2026-07-07 authorized finishing
all filed work).

RETIRED: red-team (validate --debate absorbs), perf (frontier-generic, zero
repo bindings), flywheel (ao flywheel status CLI is the surface).
MERGED: eval-outcomes -> validate (--mode=pre-impl --target=scenario),
review -> validate (--mode=pr), compile + curate -> post-mortem (mining half;
mechanical surfaces stay ao compile / ao lookup), recover -> status
(--recover mode; deep playbook preserved at status/references/
recovery-playbook.md; delivers l6ic.7 and the l6ic.13 row reconcile).

Mechanics: ao skills retire x8 (trees incl. images/*/skills, terminal ledger
rows via --into, .agy-plugin review bundle removed); absorption tombstones in
docs/SKILLS.md + SKILL-TIERS.md (56 user-facing + 2 internal = 58, counted);
spine gate re-anchored 15 -> 13 (l6ic.11 minimal consistency: review +
red-team leave; membrane 7 + bookkeeper 6); bespoke twins post-mortem/status
hand-mirrored per AGENTS-CODEX; manifest pruned of orphan twin rows (62 -> 57);
overrides catalog pruned of 8 retired rows; gemini verify core_skills +
claude/codex image manifests updated; retired-subject eval
red-team-adversarial-validation.json removed (not canary-listed).

Gates: spine-integrity 13 PASS; wiring-closure PASS; skill-frontmatter 58/58;
codex manifest+artifacts+parity PASS; regen-check ALL GREEN; SKILL-TIERS also
fixed Opus 4.6 -> 4.8, GOALS.yaml -> GOALS.md, flywheel diagram to CLI truth
(l6ic.6 pre-work). Residual prose mentions in non-gated docs are the l6ic.2
debris sweep's scope.
2026-07-07 11:06:35 -04:00
boshu d8b840596a feat(gc-adoption): install-gc-city.sh + QUICKSTART + bd carve-out + two-ledger canon reconcile + membrane fail-open fix (age-gc-adoption-u0he arc)
- scripts/install-gc-city.sh (+15 hermetic bats): one-command membrane city standup —
  version contract fail-hard, dolt_mode=server enforced, LAW-0 print sinks, gap-A
  materialization (membrane/ + quests/_template), codex pre-trust, usage-local +
  orphan-sweep-fast fragments, sessions registered AND verified, doctor gate, city path
  canonicalized (pwd -P) + GC_CITY_PATH pinned everywhere. Live E2E green. (u0he.1, .2)
- packs/agentops-membrane: QUICKSTART.md (u0he.4); fragments; README cost section
- membrane fail-open fix (u0he.15 P0): ralph gate-bead protection + max_attempts 5 +
  honest attempt semantics; fork half = gascity 4aa582dcc
- skills/beads-br gc carve-out (8aom.2); canon reconcile (u0he.12)
(gc-outcomes-report.sh split to its own land — u0he.6)
2026-07-06 21:49:04 -04:00
boshu e7de22d1ea chore(skills): flip disposition ledger to the 2026-07-06 audit verdicts — 64/66 rows (age-skills-audit-fable-l6ic.1)
Council-decided scope A (3/3 unanimous): record docs/audits/skills-audit-2026-07-06.md
section-2 verdicts into skill-dispositions.yaml — 36 disposition changes + 28
reaffirmation stamps, EXCLUDING compile + curate (the in-flight consolidation
branch deletes/moves those rows; follow-up bead reconciles after age-p2c7).
Verdict mapping: KEEP->keep FIX->update TRIM->refactor MERGE->merge-review
RETIRE/RESOLVE->cut-review. Full-file counter reconciles the audit table
exactly (keep 32 incl curate, update 14, refactor 12 incl compile,
cut-review 6, merge-review 2). Projections regenerated (make regen-all);
regen-check ALL GREEN. No skill bodies touched; no retires executed —
cut-review rows await the age-e3zk decision.
2026-07-06 21:12:16 -04:00
boshu 0ffaa6561f feat(skills): using-gc operator skill + gc-membrane reference — gas-city factory enablement (age-gc-integrate-8aom.1, delivers age-gc-adoption-u0he.3)
The vibing-with-ntm analog for gc: standup (native store contract, LAW-0
print_args, pack wiring, pre-trust), the day-to-day quest loop, admin
(doctor cadence, backup, binary discipline, config knobs), and the
symptom-keyed stall ladder (submit-only drain recovery, trust modal,
quarantined check step, spawn livelock, diff-frame refutes). gc-membrane
carries the close-door reference (finalize semantics, pawl-verdict.v1
anatomy, RBAC). Router lines in SKILLS.md + SKILL-ROUTER.md (operator
choice, never auto-routed); AGENTS.md substrate pointer; tier + disposition
ledger entries; regen projections (64→66).

Verified: silent-novice probe — fresh agent restricted to the skill files
answered 10/10 standup/run/admin/troubleshoot scenarios correctly with
citations.
2026-07-06 08:51:28 -04:00
boshu 257be5312e feat(skills): add ms skill — meta_skill search/load doctrine (MCP consume, CLI writes) 2026-07-02 11:18:48 -04:00
Bo aae1a3af9a refactor(skills): consolidate catalog 70 → 63 — cut acfs, fold 6 skills into their canonical owners (#889)
Full-catalog skills audit + consolidation via ao skills retire: acfs cut; autodev→evolve, skill-auditor→heal-skill, beads-workflow→beads-br, continuity-loop→using-atm, inject→operationalize, forge→curate. Trigger surfaces preserved via absorbed-from notes; ao CLI surfaces unchanged; dispositions ledger flipped to historical rows; registries/codex twins/eval canaries regenerated and retargeted. Includes self-review fixes: ledger-aware DEAD_XREF in heal.sh, scenario-linkage allowlist retarget, codex-twin dead-path repairs, strengthened eval canaries, fold-survival guard for operationalize, and beads consolidation test hardening.
2026-07-02 09:23:55 -04:00
boshu c7c074e739 refactor(skills): fold substrate-skill sprawl 13->~5 via the dispositions ledger (age-focus-membrane-bookkeeper-m1wg.22)
Fold 5 substrate skills into their survivors (ao skills retire --into), after
re-litigating each keep-rationale:
- vibing-with-ntm -> ntm       (operator layer over the same NTM tools)
- agy-headless-evidence -> agy-native  (consumes agy-native; evidence is a
  listed capability; mirrors agy-mcp/rules/sidecar folds)
- codex-approval -> codex-exec (codex-runtime driving-adapter; mirrors the
  codex-* consolidation)
- dual-pane-atm -> using-atm   (a specialization of the ATM substrate)
- orchestrate -> using-atm     (route/preflight/verify instrument; the ao
  orchestrate CLI survives, only the skill doc folds)

using-atm is the surviving out-of-session-substrate consolidator. Manual cleanups
ao skills retire does NOT reach: git rm .agy-plugin/skills/vibing-with-ntm +
drop it from validate-agy-plugin.sh core_skills; drop the dual-pane-atm/orchestrate
lines from scripts/.scenario-linkage-allow; drop vibing-with-ntm from
scripts/lint/codex-cross-runtime-skills.txt; repoint 26 broken cross-ref links to
survivors + fix frozen-twin generated_hash. scenario-linkage / skill-redirects
(47 resolve) / doc-skill-refs / manifests all green.

NOTE: validate-agy-plugin.sh has a PRE-EXISTING 9-bundle drift at HEAD (agent-mail,
beads-*, cass, cc-hooks, dcg, ntm, rch — stale hand-committed .agy-plugin bundle,
no sync tooling, not a routed gate); untouched here, flagged for a follow-up.
2026-07-01 02:28:07 -04:00
boshu 0096a050c3 refactor(skills): demote evolve/autodev/acfs to a new experimental tier (age-focus-membrane-bookkeeper-m1wg.21)
Heavy legacy RPI chains with no measured uplift. Add a new `experimental` tier
and demote the three. The tier value is registered at ALL FOUR enum sites — the
authoritative schemas/skill-frontmatter.v1.schema.json (enforced by the blocking
skill.manifests gate), validate-skill-schema.sh, sku_catalog.py, and
generate-registry.sh — plus a Tier Values row in SKILL-TIERS.md. Frontmatter
metadata.tier and the SKILL-TIERS row cell flipped in lockstep for all three.

acfs is demoted, NOT cut: its cut is external (~/acfs) and a deferred operator
one-way-door decision (noted in the disposition ledger). Codex twins refreshed
by regen (not dropped — .18 owns twin cuts). registry drift, manifests schema,
sku-catalog, sync-counts all green; go build + internal tests pass.
2026-07-01 00:08:34 -04:00
Boden Fuller 86aea2431f chore(regen): reconcile generated projections post-rebase (age-revive-reverse-engineer-0ll) 2026-06-18 13:45:31 -04:00
Boden Fuller 07d9a2d6e6 feat(skills): revive + rename reverse-engineer-rpi -> reverse-engineer with the steal-map discipline
Restore the cut reverse-engineer-rpi skill (was cut in 94db74318), rename it
(drop the rpi suffix across 16 files + both trees), and upgrade it from a pure
teardown tool into teardown -> steal-map (have/gap/steal/park/reject) -> route
one-way doors to /discovery's mixed-model duel. Folds this session's discipline:
validated cross-family not self-report, park substrate we delegate (ADR-0009),
steal the pattern not the platform, probe real state not stale. 308->125 lines,
codex twin mirrored (slim frontmatter + hashes), disposition keep row added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 13:23:59 -04:00
Boden Fuller a4104bb966 feat(skill): behavior-first-planning — back bdd-foundry before thinning (age-3va.2)
Extract the behavior-first-planning discipline (intent → frozen Gherkin
behaviors → EXECUTED-red acceptance tests → derived spec → acceptance-gated
bead DAG) into skills/behavior-first-planning/SKILL.md, so the capability
survives independent of the bdd-foundry workflow. The rule it enforces: no
runnable acceptance test, no bead — closing the spec-first "beads ship with no
done-criteria" hole (the 3/10 problem) before W3 thins the workflow.

Covers the four bdd-foundry phases + the closing independent-review gate
(validate before tracker write). Codex twin mirrored (manual) + hashes regen;
dispositions ledger row (BC3 Loop, domain, planning, parity required); registry
+ skill-domain-map + SKILL-TIERS + count-propagation regenerated (73→74).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 18:48:23 -04:00
Boden Fuller daf5c4fdd3 feat(orchestrate): Phase 1 instrument lane + profiles contract (age-ueu)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-17 16:48:21 -04:00
Boden Fuller 5e6b2e2c58 fix(skills): dual-pane-atm gate admission and regen drift
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-17 11:01:15 -04:00