Tool calls are not comparable units — one MCP call can be worth five
WebSearches, and a paginated CLI fetch is different again. A flat
numeric cap created false precision while penalizing topics that
legitimately needed more work.
The Stop Heuristic is now grounded entirely in progress signals:
sources repeating, synthesis not changing, external signal thin. The
agent decides when enough is enough based on whether it is still
learning, with an explicit "bias toward stopping early" framing so
the absence of a numeric backstop does not become an excuse to keep
going.
Per-phase query and fetch counts (3-6 narrowing, 3-5 fetches, 1-3
follow-ups) artificially capped good research. A topic might warrant
one broad search and five fetches, or three searches and one fetch,
or two passes of search-then-fetch-then-search; the right shape
emerges from what each step uncovers. Dictating phase budgets
prevented that adaptive pattern.
Now: phases describe activities (scope, narrow + extract, fill gaps)
not budgets. Steps 3 and 4 are merged because search and fetch
interleave in real research — a fetched source often suggests the
next query. The only quantitative constraint is a single hard cap on
total tool calls in the Stop Heuristic, framed as a safety valve
against runaway research rather than a target.
Address review feedback by removing prescription the agent does not
need:
- Step 1 no longer dictates an exact error string; the agent reports
the unavailability in its own words.
- The curl/wget prohibition is dropped from Step 1 and Tool Guidance.
An agent picked for web research is not about to reach for raw
network commands; the warning was spending tokens to defend against
a non-issue.
- Step 2 drops the prescribed "2-4 broad queries" count from both the
heading and the prose. Step 6's overall stop heuristic still bounds
total volume.
The bulleted categories (platform-native, MCP, "other") implied those
were the universe and missed dedicated CLI tools like ctx7 or any other
shape the caller may have wired up. Reframe Step 1 and Tool Guidance
around the real distinction — purpose-built web tool (any shape) vs
generic network command (curl, wget) — and drop the category list in
favor of a one-sentence prose statement with a few examples.
The previous Step 1 phrasing only said to stop "if no web-search or
web-fetch capability is available at all," which read as stopping only
when both were absent. Tighten the wording so the guard stops when
either capability is missing — restoring the original behavior where
the agent refused to enter a workflow it cannot complete.
The agent's Step 1 precondition checked for `WebSearch` and `WebFetch`
by exact name and stopped if either was missing. That blocked Codex,
Gemini, Droid, and MCP web tools (Firecrawl, Brave, Tavily, Exa) even
when the platform had equivalent capability under a different name.
Tool mappings in user `AGENTS.md` files did not help because the agent
bailed before consulting them.
Step 1 now accepts any web-search/web-fetch capability (platform-native
or MCP-provided) and only stops when no web tooling exists at all. The
`tools:` frontmatter restriction is removed so MCP web tools become
reachable on Claude Code too. Steps 2 and 4 and the Tool Guidance
section refer to "web searches" and "web fetches" generally; the
shell-fallback prohibition is preserved.
Fixes#833