3 Commits

Author SHA1 Message Date
Bo d972fa2090 Prepare AgentOps 4.0.0 plugins, skills and CLI release (#1143)
## What

Prepare AgentOps 4.0.0 across the Claude plugin, Codex plugin, skills
and CLI. Claude writers capture the supplied check status during its
original invocation, and plugin conformance verifies exact skill
membership and link destinations. Full release security now scans the
repository and blocks on Python collection failures that previously
produced a false green result.

## Why

The 3.6.0-to-current interval removes published commands and 20 skill
names, so this is a major release with migration instructions. Release
validation also exposed stale skill assertions and test prerequisites
that need to match the current product contracts without weakening
acceptance.

## How I tested

- Native Claude Opus/Haiku success, failing-check and direct-writer
trials: each check ran once, and the direct child returned plain JSON.
- Actual fresh installs and upgrades from 3.6.0 in isolated Codex and
Claude homes: 34 skills, expected agents, and exact installed package
bytes.
- Exact candidate `b721d02559e1495be6095ad97b820e88ceb4a049`: all 73
full repository gates, regeneration parity, and the complete local
release rehearsal passed. All 12 security tools ran with zero skips,
tool errors, critical findings or high-severity security findings. The
unchanged advisory policy reports 35 quality-high findings on unchanged
files.
- Python: 327 tests and 72 subtests passed. Hosted Bats: 1,509 passed,
31 environment-dependent skips, zero failures. Go
lint/build/vet/race/shuffle checks and CLI smoke/integration passed.
- All 11 hosted checks passed, including Windows correctness,
macOS/Linux installation, security, and the six-target no-publish
GoReleaser snapshot. Local archive checksums and a real macOS CLI
initialization/status/version smoke also passed.
- Fresh author-distinct review passed all four acceptance criteria and
all 35 changed paths with no unchecked acceptance. Canonical subject and
caller-intent verification passed; verdict digest
`68af2c935ed0106cd91b3950f5d168e662f4071f660fcbd113c36b7cd0f0426e` binds
manifest
`7affc77e25eaff69ba36c5ce05582b4f0385c954b76b62c02b97f97041f489b2`.

## Checklist

- [x] Breaking changes documented in the migration guide and complete
release notes.
- [x] No credentials or private runtime proof included.
- [x] Final full release checks pass on the exact candidate.
- [x] Fresh author-distinct final PASS is recorded before merge.

This prepares the release candidate; it does not publish a tag or
release.

Coverage limits remain explicit: native plugin tests used isolated macOS
homes and local marketplaces, guard installation remains opt-in, and
reader instructions do not prove sandbox confinement. Semgrep retains
pre-existing warning-level parser diagnostics. Snapshot metadata follows
the existing 3.6.0 tag; this is a packaging rehearsal, not a published
4.0.0 archive.
2026-09-13 17:21:16 -04:00
Bo ab138a5418 feat(skills): add scan_descriptions --probe deterministic trigger ranker (ag-7led #trigger-probe) (#722)
## Summary

Adds a `--probe "<phrase>"` mode to
`skills/skill-builder/scripts/scan_descriptions.py` that ranks every
skill against a lexical trigger phrase **deterministically** (no live
model, byte-stable output). Introduces an optional `trigger_probes:`
skill-frontmatter field so a skill can self-declare phrases it must rank
#1 for; the probe asserts the declaring skill wins its own phrases
against the corpus.

Epic: ag-czzf.

**INFORMATIONAL-only tooling — not a blocking gate.** The probe is a
self-check authors can run; no CI gate is wired to it.

## Changes
- `skills/skill-builder/scripts/scan_descriptions.py` — `--probe` ranker
+ `trigger_probes` parser (flow + block YAML forms).
- `schemas/skill-frontmatter.v2.schema.json` — optional `trigger_probes`
array field.
- `tests/scripts/scan-descriptions-probe.bats` — 4 scenarios (rank-#1,
determinism, demotion-on-mutation, usage-error exit codes).

## Validation
- New bats suite: 4/4 pass.
- `scripts/pre-push-gate.sh` reds are pre-existing main-red
(context-packet-ab canary + its downstream pre-push-gate-governance
nested bats), not caused by this diff; CI is changed-files-scoped.

Closes-scenario: ag-7led#trigger-probe
Bounded-context: BC2-Validation
Evidence: skills/skill-builder/scripts/scan_descriptions.py

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-06-04 15:40:07 -04:00
Bo a5a0cdcb0b feat(skills): encode external skill-authoring standard + corpus trigger scanner (ag-vb3q #authoring-standard-trigger-scan) (#663)
## What

RPI cycle on **mining how AgentOps makes skills** (bead `ag-vb3q`),
using the external `meta_skill` best-practices doc as the scoring lens
(methodology only — no Rust `ms` build).

**Research:** codebase-archaeology over the 83-skill corpus + the skill
factory
(`skill-builder`/`skill-auditor`/`heal-skill`/`converter`/`forge`,
`SKILL-TIERS`, `score_agentops_skill.py`, CI gates). The corpus is
structurally excellent: **100% frontmatter conformance, 100% Codex
parity, ~94% references usage**.

**The one finding that matters:** skill selection is pure LLM reasoning
over the `description` field, yet **73 of 81 skills (90%) carry no
trigger marker**. The auditor *has* the checks
(`description-has-triggers`, `trigger-clarity`) but they were downgraded
**FAIL→WARN**, so the gap accumulated silently — a textbook
doctrine-vs-enforcement drift.

## This PR (Implement)

- **`references/skill-authoring-standard.md`** — clean-room distillation
of the external standard, cross-walked to AgentOps's *actual* enforced
rules (best-practice → gate → severity table).
- **`scripts/scan_descriptions.py`** — corpus-wide trigger scanner that
mirrors `audit.sh`'s exact three-form detection (so it never contradicts
the per-skill auditor) and adds the prioritized remediation backlog + a
suggested `Triggers:` stub the auditor never produced. `--json` robot
mode, `--strict` exit code.
- **`tests/python/test_scan_descriptions.py`** — 11 behavioral tests
(all three forms, suggestion quality, CLI exit codes, live-corpus
smoke).
- **`skill-builder/SKILL.md`** — "Corpus authoring health" section +
reference link (194/250 lines).
- Regenerated `registry.json`; patched `catalog.json` reference count.

## Plan (follow-on beads filed)

- `ag-okfu` — add trigger markers to the 73 flagged skills (wave by
tier; acceptance = scanner `--strict` exits 0)
- `ag-cx7d` — ratchet `trigger-clarity`/`description-has-triggers`
WARN→FAIL once the corpus is clean (blocked by `ag-okfu`)
- `ag-mm6q` — fix `generate-skill-catalog.sh` BSD-awk incompatibility

## Evidence

\`\`\`
python3 skills/skill-builder/scripts/scan_descriptions.py skills #
scanned 81, missing 73 (90%)
python3 -m unittest tests.python.test_scan_descriptions # 11 ok
bash scripts/check-registry-drift.sh # PASS — no drift
bash scripts/validate-context-map-drift.sh # exit 0
heal --strict skills/skill-builder # All clean
\`\`\`

Closes-scenario: ag-vb3q#authoring-standard-trigger-scan
Bounded-context: BC4-Factory
Evidence: .agents/audits/meta-skill-archaeology.md
2026-05-31 20:16:38 +00:00