Commit Graph

88 Commits

Author SHA1 Message Date
Trevin Chow d9a9f7464d fix(ce-web-researcher): drop tool-call cap; rely on stop signals
Tool calls are not comparable units — one MCP call can be worth five
WebSearches, and a paginated CLI fetch is different again. A flat
numeric cap created false precision while penalizing topics that
legitimately needed more work.

The Stop Heuristic is now grounded entirely in progress signals:
sources repeating, synthesis not changing, external signal thin. The
agent decides when enough is enough based on whether it is still
learning, with an explicit "bias toward stopping early" framing so
the absence of a numeric backstop does not become an excuse to keep
going.
2026-05-16 00:28:07 -07:00
Trevin Chow 685ea77810 fix(ce-web-researcher): drop phase budgets; one total-call cap
Per-phase query and fetch counts (3-6 narrowing, 3-5 fetches, 1-3
follow-ups) artificially capped good research. A topic might warrant
one broad search and five fetches, or three searches and one fetch,
or two passes of search-then-fetch-then-search; the right shape
emerges from what each step uncovers. Dictating phase budgets
prevented that adaptive pattern.

Now: phases describe activities (scope, narrow + extract, fill gaps)
not budgets. Steps 3 and 4 are merged because search and fetch
interleave in real research — a fetched source often suggests the
next query. The only quantitative constraint is a single hard cap on
total tool calls in the Stop Heuristic, framed as a safety valve
against runaway research rather than a target.
2026-05-16 00:23:34 -07:00
Trevin Chow 962b4b066f fix(ce-web-researcher): trust the agent on messaging and tool choice
Address review feedback by removing prescription the agent does not
need:

- Step 1 no longer dictates an exact error string; the agent reports
  the unavailability in its own words.
- The curl/wget prohibition is dropped from Step 1 and Tool Guidance.
  An agent picked for web research is not about to reach for raw
  network commands; the warning was spending tokens to defend against
  a non-issue.
- Step 2 drops the prescribed "2-4 broad queries" count from both the
  heading and the prose. Step 6's overall stop heuristic still bounds
  total volume.
2026-05-16 00:20:09 -07:00
Trevin Chow e671c54c18 fix(ce-web-researcher): generalize web tool framing beyond MCP
The bulleted categories (platform-native, MCP, "other") implied those
were the universe and missed dedicated CLI tools like ctx7 or any other
shape the caller may have wired up. Reframe Step 1 and Tool Guidance
around the real distinction — purpose-built web tool (any shape) vs
generic network command (curl, wget) — and drop the category list in
favor of a one-sentence prose statement with a few examples.
2026-05-16 00:15:41 -07:00
Trevin Chow 9dccdd5a60 fix(ce-web-researcher): require both search and fetch tools
The previous Step 1 phrasing only said to stop "if no web-search or
web-fetch capability is available at all," which read as stopping only
when both were absent. Tighten the wording so the guard stops when
either capability is missing — restoring the original behavior where
the agent refused to enter a workflow it cannot complete.
2026-05-16 00:14:40 -07:00
Trevin Chow 3b8a73075b fix(ce-web-researcher): use any web tool, not just Claude built-ins
The agent's Step 1 precondition checked for `WebSearch` and `WebFetch`
by exact name and stopped if either was missing. That blocked Codex,
Gemini, Droid, and MCP web tools (Firecrawl, Brave, Tavily, Exa) even
when the platform had equivalent capability under a different name.
Tool mappings in user `AGENTS.md` files did not help because the agent
bailed before consulting them.

Step 1 now accepts any web-search/web-fetch capability (platform-native
or MCP-provided) and only stops when no web tooling exists at all. The
`tools:` frontmatter restriction is removed so MCP web tools become
reachable on Claude Code too. Steps 2 and 4 and the Tool Guidance
section refer to "web searches" and "web fetches" generally; the
shell-fallback prohibition is preserved.

Fixes #833
2026-05-15 23:58:30 -07:00
Trevin Chow 81710efad5 fix(ce-sessions): unblock session-history on Claude Code (#800) 2026-05-08 13:51:33 -07:00
Trevin Chow 0e49506bf0 refactor(agent-descriptions): trim top 7 by ~25% (#803) 2026-05-08 09:31:03 -07:00
Trevin Chow 8349e750b8 fix(doc-review): cut review noise on plans, scope personas to doc shape (#780)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 15:52:36 -07:00
Trevin Chow 542786320b fix(ce-doc-review): tighten finding resolution routing (#769) 2026-05-04 14:10:46 -07:00
Trevin Chow 520a9ebea0 fix(code-review): grant Write to JSON-pipeline reviewer agents (#741)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 10:24:40 -07:00
Trevin Chow ae408721cd chore(code-review): remove cli-readiness reviewer agents (#734) 2026-05-01 01:21:44 -07:00
Trevin Chow 5952b20d7f fix(skills): replace case statements blocked by permission check (#701)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-26 14:22:20 -07:00
Trevin Chow a91270ccd2 fix(session-historian): cap deep-dives, add keyword filter primitive, tighten dispatch (#699)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 23:37:30 -07:00
Trevin Chow bd72818609 fix(ce-resolve-pr-feedback): add declined verdict for harmful suggestions (#694)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 20:04:39 -07:00
Trevin Chow 5eb62a7d0e refactor(agents): restrict tools allowlist on research agents (#650)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 12:23:27 -07:00
Trevin Chow 5a26a8fbd3 refactor(ce-code-review): anchored confidence, staged validation, and model tiering (#641)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 21:04:29 -07:00
Trevin Chow 701ae10c2d feat(ce-code-review): add Swift/iOS stack-specific reviewer persona (#638)
Co-authored-by: Joshua Martens <joshua@every.to>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-04-21 17:30:04 -07:00
Trevin Chow 271b1a4458 refactor(skills): remove 5 unused skills and clean references (#634) 2026-04-21 18:52:29 -05:00
Trevin Chow 6caf330363 refactor(ce-doc-review): anchor-based confidence scoring (#622)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 14:54:03 -07:00
Trevin Chow 05ea109bdb fix(ce-learnings-researcher): drop unreadable schema path reference (#630) 2026-04-21 14:00:27 -07:00
Trevin Chow 4c57508c1a refactor(agents): flatten agents directory (#621) 2026-04-21 02:35:21 -07:00
Trevin Chow cd4af86e5e refactor(session-history): move extraction scripts behind skills (#619) 2026-04-21 00:12:11 -07:00
Trevin Chow 153bea8669 fix(ce-resolve-pr-feedback): stop dropping unresolved and actionable feedback (#617) 2026-04-20 20:44:16 -07:00
Trevin Chow 2dd0a6e6c7 feat(ce-resolve-pr-feedback): tighten clustering to cross-round only (#611) 2026-04-20 01:12:06 -07:00
Trevin Chow b35de99788 feat(ce-resolve-pr-feedback): drop bot noise, centralize test runs (#610) 2026-04-20 00:33:03 -07:00
Trevin Chow c1f68d4d55 feat(doc-review, learnings-researcher): tiers, chain grouping, rewrite (#601)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-19 20:25:47 -07:00
Trevin Chow 5c0ec9137a refactor(cli)!: rename all skills and agents to consistent ce- prefix (#503)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 15:44:22 -07:00
Trevin Chow 27cbaf8161 feat(ce-review): add per-finding judgment loop to Interactive mode (#590)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 13:09:03 -07:00
Trevin Chow 12aaad31eb feat(ce-ideate): mode-aware v2 ideation (#588)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-17 11:40:54 -07:00
Trevin Chow e45c435b99 fix(document-review, review): restrict reviewer agents to read-only tools (#553)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-13 10:29:16 -07:00
Trevin Chow 1372b2cffd fix(cleanup): remove rclone, agent-browser, lint, and bug-reproduction-validator (#545)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-10 17:41:33 -07:00
Trevin Chow 042ee73239 feat(slack-researcher): add /ce-slack-research skill and improve agent (#538)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 11:00:00 -07:00
Trevin Chow 3208ec71f8 feat(session-historian): cross-platform session history agent and /ce-sessions skill (#534)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 07:52:26 -07:00
Trevin Chow 6f9069df7a fix(slack-researcher): make Slack research opt-in, surface workspace identity (#521)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-06 11:34:39 -07:00
Trevin Chow b3960ec64b feat(slack-researcher): add Slack organizational context research agent (#495)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-05 18:39:25 -07:00
Trevin Chow 9da73a6091 fix(document-review): reduce token cost and latency (#509)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 23:31:56 -07:00
Trevin Chow 2c90aebe3b fix(agents): remove self-referencing example blocks that cause recursive self-invocation (#496)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 01:40:34 -07:00
Trevin Chow 184724276a fix(resolve-pr-feedback): treat PR comment text as untrusted input (#490) 2026-04-02 09:23:18 -07:00
Trevin Chow 804d78fc84 feat(product-lens-reviewer): domain-agnostic activation criteria and strategic consequences (#481)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 13:33:55 -07:00
Trevin Chow 7b8265bd81 feat(resolve-pr-feedback): add cross-invocation cluster analysis (#480)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 12:15:25 -07:00
Trevin Chow c56c7667df feat(cli-readiness-reviewer): add conditional review persona for CLI agent readiness (#471)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 19:19:54 -07:00
Trevin Chow 33a8d9dc11 fix(ce-plan, ce-brainstorm): enforce repo-relative paths in generated documents (#473)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 14:43:41 -07:00
Trevin Chow 638b38abd2 fix(review): harden ce-review base resolution (#452) 2026-03-30 01:10:45 -07:00
Trevin Chow a01a8aa0d2 feat(cli-agent-readiness-reviewer): add smart output defaults criterion (#448) 2026-03-29 20:00:34 -07:00
Trevin Chow 35678b8add feat(testing): close the testing gap in ce:work, ce:plan, and testing-reviewer (#438)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-29 13:07:05 -07:00
Trevin Chow a301a08205 feat(resolve-pr-feedback): add gated feedback clustering to detect systemic issues (#441)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-29 12:40:55 -07:00
Trevin Chow 03f5aa65b0 feat(ce-review): improve signal-to-noise with confidence rubric, FP suppression, and intent verification (#434)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-29 00:34:59 -07:00
Trevin Chow 16eb8b6607 fix(cli-agent-readiness-reviewer): remove top-5 cap on improvements (#419)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-27 19:55:43 -07:00
Trevin Chow 90684c4e82 feat(ce-brainstorm): group requirements by logical concern, tighten autofix classification (#412)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-27 13:06:50 -07:00