8 Commits

Author SHA1 Message Date
Amaury Levé 77154137e8 dotnet-test: make code-testing agent tools Claude Code-compatible + add cross-host portability check (#856)
* dotnet-test: make code-testing agent tools declarations Claude Code-compatible

PR #847 added `tools: ["agent", "skill", "read", "search", "edit", "execute"]`
to the code-testing-* agents to enable VS Code / Copilot CLI subagent fan-out.
Those lowercase aliases map to real tools in VS Code and the Copilot CLI, but
Claude Code matches `tools:` against its own vocabulary (Task, Skill, Read,
Glob, Grep, Edit, Write, Bash). None of the aliases matched, so when these
agents are loaded into Claude Code via --plugin-dir and selected with
`claude --agent`, the agent was granted ZERO tools. A tool-less model asked
to generate tests emits a textual <tool_call> block and exits after one turn,
producing no file changes.

Append the Claude Code tool names to each agent's `tools:` list so the same
declaration works across all three runtimes (each honors the names it knows and
ignores the foreign ones):

- Orchestrators (generator, implementer): add Task, Skill, Read, Glob, Grep,
  Edit, Write, Bash (Task is the Claude Code equivalent of the `agent`
  fan-out tool).
- Workers (researcher, planner, builder, tester, fixer, linter): add Skill,
  Read, Glob, Grep, Edit, Write, Bash.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* skill-validator: complete built-in tools + add cross-host tool portability check

Two related follow-ups to the agent tools fix:

1. Address the skill-check review feedback. The validator's BuiltInTools set was
   missing three legitimate host tool spellings that are not case-insensitive
   matches of existing entries, so they were flagged as non-built-in:
   - "write"   — Claude Code file-creation tool (Copilot CLI / VS Code: "create")
   - "agent"   — Copilot CLI / VS Code subagent fan-out tool (Claude Code: "task")
   - "execute" — Copilot CLI / VS Code run-command tool (Claude Code: "bash")
   "agent" and "execute" were already flagged before this branch (introduced by
   the fan-out PR); adding them to BuiltInTools clears the pre-existing warnings.

2. Add a cross-host tool portability check (CheckAgentToolPortability) so an
   agent that declares a capability for only one host is flagged. Tool names are
   matched case-sensitively (hosts resolve tools by exact spelling), so an agent
   that lists e.g. only "edit" (Copilot / VS Code) without "Edit"/"Write"
   (Claude Code) is reported as working on one host and silently tool-less on the
   other. Findings are advisory (do not fail CI) and allowlistable via
   "agent-tool-portability:AGENT:capability". Wired into the agents loop in
   CheckCommand and covered by unit tests.

Also make the one existing single-host agent (optimizing-dotnet-performance)
portable by adding its Claude Code tool spellings, so the new check reports a
clean tree.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-07-07 09:00:22 +02:00
Petr Pokorny 06dd46e5f1 Enable VS Code subagent fan-out for the dotnet-test code-testing agents (#847)
* Add VS Code subagent metadata to dotnet-test code-testing agents

Enable the code-testing-* Research-Plan-Implement pipeline to fan out as
subagents in VS Code while keeping the GitHub Copilot CLI working.

Frontmatter (VS Code "coordinator/worker" pattern; portable tool aliases that
map in both VS Code and the CLI):
- code-testing-generator: tools: [agent, read, search, edit, execute] + an
  agents: list (researcher, planner, implementer, builder, tester, fixer,
  linter); softened the three runSubagent({ agent, prompt }) blocks to
  tool-agnostic delegation prose.
- code-testing-implementer: tools: [agent, read, search, edit, execute] +
  agents: (builder, tester, fixer, linter).
- Leaf agents (researcher/planner/builder/tester/fixer/linter):
  tools: [read, search, edit, execute].

Why explicit tools (not ["*"]): VS Code has no all-tools wildcard for 	ools:
and a subagent's 	ools: overrides its inherited set, so ["*"] matched nothing
and stripped subagents of file tools (they ran without read/edit/search). The
CLI treats ["*"] as all-tools, so this was VS-Code-specific. The portable
aliases agent/read/search/edit/execute map to real tools in both environments;
agents: is ignored by the CLI.

README: document the VS Code chat.subagents.allowInvocationsFromSubagents
setting (off by default) needed for the nested implementer->builder/tester/
fixer/linter layer to fan out on large scopes; the CLI has no such gate.

Validated end to end: VS Code shows researcher->planner->implementer fanning out
with real file I/O (no "subagents lack file tools" warning); CLI fan-out intact
with skills still loading and tests passing. An all-tools baseline used only
tools within this enumerated set, confirming no CLI capability is restricted.

* Include skill tool in code-testing agent tool allowlists

Address PR review: the explicit `tools:` allowlists omitted the `skill`
tool, but every code-testing-* agent's prompt instructs calling skills
(e.g. `code-testing-extensions` for per-language guidance, `test-gap-analysis`,
`assertion-quality`). Because `tools:` is an override, omitting `skill` can
prevent the agents from loading those skills in environments that gate skill
invocation by the allowlist.

Add `skill` to all eight agents:
- orchestrators (generator, implementer): [agent, skill, read, search, edit, execute]
- workers (researcher/planner/builder/tester/fixer/linter): [skill, read, search, edit, execute]

Re-verified in the Copilot CLI: full fan-out (researcher -> planner ->
implementer/tester), the `code-testing-extensions` skill is invoked, and the
generated tests pass.
2026-06-30 19:42:39 +02:00
YuliiaKovalova e1e8568375 Revert "dotnet-test: add unit-under-test + behaviors quality cue to code-test…" (#651)
This reverts commit f1b09eba79.
2026-05-14 17:00:07 +02:00
YuliiaKovalova f1b09eba79 dotnet-test: add unit-under-test + behaviors quality cue to code-testing-generator (#646)
* CTA: invocation-only baseline (dispatch mechanics, no content rules)

Experiment branch to isolate the impact of "invoke prompted subagents
more often" from the impact of "make those subagents do richer work."
Same baseline as dev/ykovalova/cta-prompt-tuning (main = 66628b6), but
strips out every content/quality rule and keeps only the dispatch
plumbing.

Comparison branch: dev/ykovalova/cta-prompt-tuning (HEAD: bd530be)
which contains both the dispatch mechanics AND content/quality rules.

Files modified (3 vs 5 in cta-prompt-tuning):
- code-testing-generator.agent.md  +110 lines
- code-testing-implementer.agent.md +11 lines
- code-testing-fixer.agent.md      +1  line
- code-testing-researcher.agent.md  UNTOUCHED (baseline)
- code-testing-planner.agent.md     UNTOUCHED (baseline)

KEPT (invocation / dispatch mechanics):

Generator:
- Rule 1: every task() call MUST use agent_type "dotnet-test:code-testing-..."
  (without this, calls dispatch generic built-ins and never reach the
  named CTA agents)
- Rule 2: routing table -- which named agent for which job
- Rule 3: prefer one named-agent dispatch over many tool calls
- Rule 4: orchestrator MUST NOT edit/create test files itself
  (forces implementer dispatch)
- Rule 5: orchestrator MUST NOT run builds/tests via terminal
  (forces builder/tester dispatch)
- Rule 6: every run MUST dispatch the planner (no exceptions for "small"
  scope; Direct still goes through planner)
- Rule 7: every build/test failure MUST dispatch the fixer
- Step 1b: mandatory initial researcher dispatch (every strategy)
- Direct strategy rewritten: dispatches planner -> implementer -> builder
  -> tester -> fixer -> linter (was "Skip Steps 3-5, write tests inline")
- All Step 3/4/5/6/7/8/9 dispatches converted from runSubagent({agent:...})
  to task({ agent_type: "dotnet-test:code-testing-...", name:..., prompt:...})
- Step 9 validator dispatch (forces builder dispatch for cleanup)
- Steps 6/7 mandatory builder/tester dispatch wrapper

Implementer:
- Section 5: "you MUST dispatch fixer for build errors" + no-inline-edit
  block (forces fixer dispatch on build failures)
- Section 6: "you MUST dispatch fixer for test failures" + no-inline-edit
  block (forces fixer dispatch on test failures)
- Section 7: "Format Code (mandatory if a lint command exists)"
  (was "Optional"; mandatory firing of linter)
- Rule 6: never declare SUCCESS while build/tests fail (gates SUCCESS on
  fixer dispatch)
- Rule 7: no inline test-file edits between failed dispatch and fixer

Fixer:
- Frontmatter description widened to advertise handling of failing tests
  (without this, the orchestrator's routing logic does not select the
  fixer for test failures, so even Rule 7's mandate produces no firing
  -- this is the change that took fixer firing from 0.00/inst to 0.39/inst
  in earlier iterations)

DROPPED (content / quality rules -- in cta-prompt-tuning, NOT here):

Generator:
- Test-strength rules embedded in implementer dispatch prompt
- Test-design rules embedded in implementer dispatch prompt (OFAT,
  mutation self-check, never mock subject under test)
- File-location rules embedded in implementer dispatch prompt
- TARGET ENTITIES / PHASE CHECKLIST / TEST TRACEABILITY blocks in
  implementer dispatch prompt
- CHECKLIST format spec in planner dispatch prompt
- Step 9 validator's detailed cleanup classification

Implementer:
- Section 4b "Verify CHECKLIST coverage" pre-completion check
- Section 8 "CHECKLIST COVERAGE" report block
- "Honor the CHECKLIST" rule

Fixer:
- "Process -- Failing Tests" section (5-step diagnosis flow)
- All anti-weakening / anti-skipping rules
- "Re-derive expected from production source" guidance

Planner:
- CHECKLIST format ("one item per TARGET BEHAVIOR, Source/Variants/
  Expected mandatory")
- "Test name from research.md conventions" rule
- "At least 2 phases" rule

Researcher:
- Section 8 "Extract Local Test Naming & Style Conventions"
- TARGET ENTITIES / TARGET BEHAVIORS / TEST INFRASTRUCTURE structure
  in research.md
- Test naming pattern extraction

WHAT THE SUBAGENTS WILL ACTUALLY DO:

The researcher / planner / implementer / fixer all operate at baseline
behavior -- they receive the same prompts they receive in the upstream
"vanilla" runs. The only difference vs vanilla is that the orchestrator
ACTUALLY DISPATCHES THEM (where vanilla often inlines the work or skips
sub-agent dispatch entirely).

EXPECTED COMPARISON:

If quality on this branch is similar to or higher than dev/ykovalova/
cta-prompt-tuning (bd530be), then "more dispatches" is the dominant
quality lever and the content/quality rules in cta-prompt-tuning are
adding marginal or noise-level value.

If quality on this branch is materially lower than cta-prompt-tuning,
then the content/quality rules are doing the heavy lifting and the
dispatch mechanics alone are insufficient.

If quality on this branch matches or exceeds vanilla but trails
cta-prompt-tuning, then the dispatch mechanics provide a baseline lift
and the content rules add an incremental quality layer on top.

Rubber-duck check passed (validated dispatch-vs-content classification;
fixer frontmatter is routing metadata not a runtime gate; surviving
dispatch prompts contain no dangling references to removed CHECKLIST /
TARGET ENTITIES / TEST STRENGTH / naming-convention concepts).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* CTA: extract unit-under-test + behaviors quality cue (no caps)

* Address PR #646 review feedback

- Step 1b prompt: explicitly request unit-under-test (file:line) and behaviors
  so the verification gate rarely needs re-dispatch.
- Step 3: rename to 'Deep Research Phase', mark skipped for Direct strategy,
  and switch from overwriting research.md to extending it (no double-research).
- Step references: update '6-9' -> '6-10' and 'Step 9' -> 'Steps 9-10' in the
  strategy table and the All-strategies-MUST line, since reporting is Step 10.
- Step 6 builder prompt: drop '*.sln' glob (could expand to multiple args);
  use 'dotnet build --no-incremental' (auto-discovers .sln) per dotnet.md.
- Step 9: stop overloading the builder agent; perform diff/cleanup directly
  in the orchestrator (Rule 5 forbids inline build/test, not git/fs hygiene).
- Fixer agent: update mission text to cover failing tests and assertion
  correction (front-matter description already mentioned this; body now
  matches), with explicit no-Ignore/no-Skip/no-production-rewrite guardrails.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Yuliia Kovalova <ykovalova@example.com>
2026-05-13 14:03:31 +02:00
Jan Krivanek 81946d2a38 Add proxy for CTA extension files (#583) 2026-04-24 12:19:17 +02:00
Jan Krivanek 05aeb657e6 Add license to agent files (#568) 2026-04-21 12:57:18 +00:00
Amaury Levé 1125fe7864 improve(dotnet-test): enhance agent descriptions for subagent discovery (#560)
- Add 'Use when:' trigger phrases to all 7 sub-agent descriptions
  so parent agents can reliably discover and delegate to them
- Add .testagent/ cleanup rule to generator agent to prevent
  ephemeral pipeline state from being committed
2026-04-20 19:51:09 +02:00
Jan Krivanek e4670b33a1 [PoC] Code testing agent + tests PoC (#433)
* Add code testing agent agents + tests PoC

* Remove AssertionEvaluator.cs change (moved to dev/jankrivanek/agents-evals)
2026-03-30 17:57:08 +02:00