Bo 3842ea0e00 docs: make README onboarding practical for Claude and Codex (#1145)
## What changed

Put Claude Code and Codex plugin installation and a complete first
research task near the top of the README. Define skills in plain
language, show the runtime-specific prompts and expected answer, and add
a short task-to-skill menu.

Clarify when the optional CLI is needed, keep the exact dependency table
in a collapsible section, and bring upgrade commands and 3.7 migration
guidance together. Advanced architecture and factory material links to
its canonical documentation. Contribution policy is preserved.

## Validation

- Documentation links, metadata counts, and release-message checks
passed.
- All 44 product-boundary/runtime-requirements Bats checks passed.
- Default local aggregate: 10 passes, zero failures, one optional
missing OL-suite skip.
- All 13 applicable AO gates passed; generated projections are current.
- GitHub GFM renders both prompt examples and the collapsed dependency
table; self-links match headings/anchors.
- Fresh independent source/readability review passed for commit
`db3e06a6bafb71696327ef20d5401e660c10692f`, covering the complete README
and all acceptance criteria.
- The full strict documentation-site check reported 99 warnings in
unchanged site pages (73 outside its allowlist). A clean
release-baseline build fails with the identical warning multiset; the
README adds zero warnings and is not a site-build input. This existing
site issue is recorded separately. No gate or allowlist was weakened.
2026-09-13 21:50:50 -04:00
2026-09-13 20:28:12 -04:00
2026-09-13 20:28:12 -04:00
2026-09-13 20:28:12 -04:00
2026-09-13 20:28:12 -04:00

AgentOps

AgentOps gives Claude Code and Codex reusable instructions, called skills, for investigating code, writing useful tests, and independently reviewing changes. Start with one task and the guidance it needs; the 34-skill library is available without adopting a new workflow.

When a coding agent says a change is done, AgentOps helps a fresh reviewer check that exact change against what you asked for. Your coding agent runs the work; your repository keeps its existing tests, tracker, and Git workflow.

Quickstart · Choose a skill · Optional CLI · Upgrade · Documentation

Quickstart

Use an installed Claude Code or Codex with plugin support. Run the commands for one runtime in your terminal. They install the full managed skill bundle; ao is not needed for the first task below.

Claude Code

claude plugin marketplace add boshu2/agentops
claude plugin install agentops@agentops-marketplace
claude plugin details agentops@agentops-marketplace

Codex

codex plugin marketplace add boshu2/agentops
codex plugin add agentops@agentops-marketplace
codex plugin list --json

The plugin should appear as agentops. Claude's inventory includes 34 skills and four agents; Codex exposes 34 skills with agentops: names. Start a new session in a project you already work on to load the installed skills.

Skills run inside your coding agent with its normal permissions. The Claude plugin also installs its policy dispatcher. Optional limits on source reads and Codex's additional agent roles have separate setup. To remove the bundle, use claude plugin uninstall agentops@agentops-marketplace or codex plugin remove agentops@agentops-marketplace in your terminal.

Try one task

Paste this into your Claude Code session, not your terminal:

/agentops:research Find how this repository validates user input. Trace one
path from the input through validation and its tests. Cite the relevant files
and line numbers, explain one edge case, and identify any missing coverage.
Answer in this conversation without changing files.

In a Codex session, use the same request with its skill prefix:

$agentops:research Find how this repository validates user input. Trace one
path from the input through validation and its tests. Cite the relevant files
and line numbers, explain one edge case, and identify any missing coverage.
Answer in this conversation without changing files.

Expect a trace you can inspect: where input enters, which checks accept or reject it, and what the tests cover. If a step cannot be established, the answer should say what is missing. You can then request a fix or a regression test using those file references. Research itself does not authorize a code change.

Choose skills by the work

Pick guidance for the task in front of you. These are independent choices, not steps you must run in order.

What you need Skill What to expect
Understand a behavior before changing it research An answer grounded in code and tests, with file references and gaps
Add tests for a behavior or regression test Tests using your repository's framework, plus the commands and results
Simplify code while preserving behavior refactor A focused change checked against the existing behavior
Clarify what a change should do plan Concrete acceptance examples and an agreed scope
Independently review a finished change validate A fresh review against the original request; requires ao

Use /agentops:test in Claude Code or $agentops:test in Codex to select Test, and substitute another skill name when needed. Ordinary language also works: "Use AgentOps Test to cover the missing edge case we just traced. Preserve the current API and run the owning package checks."

The Skill Router lists all 34 skills, including implementation, documentation, security, and memory. Installing a skill makes it available; a clear task can proceed directly in your coding agent.

Optional ao CLI

Install ao when you need its deterministic repository checks or evidence commands, or when a selected skill requires it. Research, Test, and Refactor can use your coding agent and the repository's existing tools.

With Homebrew:

brew tap boshu2/agentops https://github.com/boshu2/homebrew-agentops
brew install agentops
ao version
ao quick-start

With Go installed:

go install github.com/boshu2/agentops/cli/cmd/ao@latest

ao quick-start gives read-only guidance; ao demo prints a sample coding task. Neither writes project state. ao init is optional local evidence setup. See installation for source builds and the command reference for checks, inspection, and evidence operations.

Skill dependencies: which ones need ao, Python, or another tool?

Most skills need nothing beyond the coding agent and your repository's tools. These have additional requirements. "Conditional" means a selected task path uses the tool; "optional" means the skill can complete without it.

Skill Needs Why
rpi ao, conditional delegates exact-subject checks to Validate; only persists verdict.v2 when requested, with the fixed-dispatch adapter optional
plan ao, conditional runs ao provenance snapshot-intent with an explicit evidence root when the intent source is not durable
validate ao derives exact subject identity with the helper and uses ao provenance store-verdict when persistence is requested; Python/schema checks are developer-only
reality-check ao, conditional inspect selected goal measurements with ao goals or evidence-store facts with ao status
using-gc ao rig prep runs ao gc prepare and ao gc check
doc ao, optional a requested continuity handoff may use ao session handoff/rehydrate
reverse-engineer python3 Phase 1's mechanical teardown runs scripts/reverse_engineer.py
skill-builder python3, conditional Create mode's build.sh runs scripts/generate-skill-mesh.py; heal/check/audit modes are bash-only
ms python3, conditional, plus ms binary the MCP-search fallback runs python3 skills/ms/scripts/mcp-search.py; the ms binary is required for CLI load, write, and admin operations
memory python3, conditional a selected toil investigation can use the repository helper scripts/toil-mining/recent_human.py on cleared Codex sources
security python3, conditional the composable suite and offline redteam surfaces run security_suite.py when that scan type is selected
cass python3, optional scripts/prompt_miner.py mines repeated prompts; one of several selectable Scripts-table entries

Plugin installation and npx skills@latest add boshu2/agentops --all -g install the catalog even when these dependencies are absent. See the installation guide before using a dependent skill.

Other installation paths

Choose one source for each skill to avoid duplicate copies in your coding agent.

Path Use it when
Runtime plugin, shown above You want a bundle managed through Claude Code or Codex
npx skills@latest add boshu2/agentops --all -g You want the full library copied into supported coding agents
Checkout + ao skills link You edit skills or want to expose a selected subset from source

From an AgentOps checkout with ao installed, preview and link only the skills you want:

ao skills link --skill test --skill refactor --dry-run
ao skills link --skill test --skill refactor

Omit the selectors to link the whole catalog. Linking preserves existing real directories and foreign links. Follow the complete source checkout instructions for cloning, updating, and removing owned links.

Upgrading to 3.7

Read the migration guide before upgrading from 3.6. Version 3.7 removes CLI commands and former skill names, including learn, codebase-recon, and swarm. Their current owners are memory, research, and agent-native; the guide covers the complete mapping and preserved evidence.

Plugin updates require refreshing both the marketplace and the installed bundle:

# Claude Code
claude plugin marketplace update agentops-marketplace
claude plugin update agentops@agentops-marketplace

# Codex, for a Git-backed marketplace
codex plugin marketplace upgrade agentops-marketplace
codex plugin add agentops@agentops-marketplace

Start a new session afterward. If you also use the Homebrew CLI, run brew update followed by brew upgrade agentops. For local marketplaces, copied skills, or source links, follow the update instructions. New installs do not silently remove obsolete copied skills; inspect those separately before removing them.

See the 3.7 release notes for the full change list and validation limits.

How independent review works

AgentOps is the operations layer for agentic engineering. It connects your request, implementation, checks, and review while your existing tools keep ownership of the work.

Accepted intent -> native implementation and checks -> fresh independent judgment -> finish

Describe the outcome and acceptance examples in a conversation, issue, or your tracker. Beads is an optional tracker. Let the coding agent implement and run your repository's checks, then have a fresh context review the exact change against that same request. By default, the reviewer uses the author's model family; a different model is an explicit choice.

Result Meaning
PASS The independent reviewer checked the full accepted scope and found evidence for every criterion
FAIL A criterion failed or the change exceeded the authorized scope
NOT_PROVEN Missing evidence, coverage, or reviewer independence prevents a complete judgment

A passing test is evidence for review. The author cannot issue its own binding PASS. When you request durable proof, Validate can save a verdict.v2 record with exact content identity, checked scope, and evidence references. New proof uses caller-selected protected external non-Git storage; existing evidence is preserved.

Architecture, memory, and multi-agent workflows

AgentOps connects caller-owned tools as a federated integration graph: Git owns content and history, the tracker owns work, and the coding agent or a selected factory owns execution. AgentOps supplies guidance, deterministic checks, and independent judgment. Native execution needs zero mandatory skills. The architecture and operating contract describe these boundaries.

The RPI charter packages the workflow when explicitly selected. Memory provides on-demand recall and deliberate curation of reviewed material. Saved lessons only demonstrate benefit when later work uses them successfully; ADR-0016 covers storage, disclosure, and preservation.

One agent and one writer are the default. For selected multi-agent work, agent-native dispatches focused tasks. Gas City and Agentic Coding Flywheel are optional factory integrations. Their completion reports do not replace independent review. See model dispatch and the evidence contract for the detailed rules.

Optional admission-control hooks

The Claude Code plugin includes a PreToolUse policy dispatcher: guards against staging private tracker data, editing the provenance ledger by hand, and overwriting installed skill copies. It runs before matching tool calls and points blocked actions toward the supported command. Installing only ao does not install these hooks.

Other install paths can opt in through the CC Hooks skill. Disable the Claude plugin with /plugin disable agentops in Claude Code, or use the terminal uninstall command shown above.

Read-budget guards remain separately opt-in in both runtimes. Codex custom roles also require a separate installer and a restart; its hooks need review and trust in the native hook manager. See role and hook setup for the exact steps and update requirements.

Troubleshooting

Symptom What to check
plugin is not a recognized command Update Claude Code or Codex to a version with plugin support, then retry its install commands
A skill is missing after installation or update Check the plugin inventory with the Quickstart commands, then start a new session
ao is not found Install the optional CLI and check your shell's PATH; Go installs usually place it in $(go env GOPATH)/bin
A skill asks for Python or another tool Check the dependency table above; installing the skill does not install its dependencies
An old skill name no longer works Use the current owner in the migration guide and check for stale copies

For a reproducible problem, open an issue with the runtime version, install method, command or prompt, and observed result. Include a saved verdict only if you requested one and it is safe to share.

Some engineering guidance draws on Matt Pocock's skills, including concrete acceptance examples, domain language, and interface-focused tests. These are design influences; they do not establish measured improvements in coding outcomes.

Contributing: docs/CONTRIBUTING.md. License: Apache-2.0.

S
Description
council: Compare independent views on a consequential or contested decision. Use when: the caller selects multiple judges; evidence resolves disagreement, not voting.; swarm: Dispatch explicit disjoint packets exactly once through a caller-selected executor. Triggers: "swarm", "dispatch disjoint packets", "parallel explicit tasks".; doc: Write grounded docs, READMEs, repo instructions or continuity handoffs. Use when: these documents are requested; no reports as a routine completion ritual.
Readme 115 MiB
Languages
Go 42%
Shell 39.1%
Python 15.4%
JavaScript 1.6%
Gherkin 1.1%
Other 0.7%