* Keep test agent state out of commits
Move broad test-generation pipeline state to host scratch storage, worktree-specific Git metadata, or OS temp, and enforce the exclusion in evals.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Clarify absolute test agent state path
Use Git's explicit absolute path formatting in both test-generation entry points.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Prune Git metadata from test agent eval guards
Avoid scanning nested repositories and align the remaining TESTAGENT_DIR placeholder with the documented format.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Clarify test agent command handoff
Require callers to provide exact commands, excerpts, or absolute TESTAGENT_DIR document paths to command-running sub-agents.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Reject all repository-local testagent entries
Match .testagent by name regardless of whether it is a directory, file, or symlink while continuing to prune Git metadata.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Verify external test agent artifacts
Restore broad-run artifact checks at the Git metadata path and pass the researched lint command and state directory to the linter agent.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Make testagent eval guards constant time
Check only the forbidden workspace-root path, including broken symlinks, instead of recursively traversing dependency trees.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Broaden comprehensive test generation
Treat explicit requirements as the floor for broad suites and add mutation-relevant equivalence-partition and invariant coverage without test-count padding.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Fix external artifact grader quoting
Run state checks directly in the harness shell so TESTAGENT_DIR expands after assignment, with an isolated command probe covering valid and forbidden states.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Run broad skill evals in Git worktrees
Initialize the seven broad evaluation roots as Git repositories so TESTAGENT_DIR resolves deterministically and external artifacts remain verifiable.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Clarify non-stageable test agent state
Describe the real invariant across the pipeline: state may live under .git metadata but must never be version-controlled workspace content or appear in git status.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Standardize intermediate test state contract
Use one TESTAGENT_DIR placeholder, clearer intermediate-state terminology, and detect stageable research, plan, or status files regardless of directory name.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Use one Git root in workspace integrity eval
Baseline the fixture from the evaluation root so stageable intermediate-state files remain visible to the directory-independent guard.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Prune Vitest dependencies from state scan
Exclude node_modules through per-eval Git metadata so stageable state detection remains fast without modifying fixture content.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* Strengthen focused intermediate-state guards
Separate shell execution, reject Git-metadata files on focused runs, include ignored state files, and prune node_modules with a pathspec exclusion.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
---------
Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
* dotnet-test: make code-testing agent tools declarations Claude Code-compatible
PR #847 added `tools: ["agent", "skill", "read", "search", "edit", "execute"]`
to the code-testing-* agents to enable VS Code / Copilot CLI subagent fan-out.
Those lowercase aliases map to real tools in VS Code and the Copilot CLI, but
Claude Code matches `tools:` against its own vocabulary (Task, Skill, Read,
Glob, Grep, Edit, Write, Bash). None of the aliases matched, so when these
agents are loaded into Claude Code via --plugin-dir and selected with
`claude --agent`, the agent was granted ZERO tools. A tool-less model asked
to generate tests emits a textual <tool_call> block and exits after one turn,
producing no file changes.
Append the Claude Code tool names to each agent's `tools:` list so the same
declaration works across all three runtimes (each honors the names it knows and
ignores the foreign ones):
- Orchestrators (generator, implementer): add Task, Skill, Read, Glob, Grep,
Edit, Write, Bash (Task is the Claude Code equivalent of the `agent`
fan-out tool).
- Workers (researcher, planner, builder, tester, fixer, linter): add Skill,
Read, Glob, Grep, Edit, Write, Bash.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
* skill-validator: complete built-in tools + add cross-host tool portability check
Two related follow-ups to the agent tools fix:
1. Address the skill-check review feedback. The validator's BuiltInTools set was
missing three legitimate host tool spellings that are not case-insensitive
matches of existing entries, so they were flagged as non-built-in:
- "write" — Claude Code file-creation tool (Copilot CLI / VS Code: "create")
- "agent" — Copilot CLI / VS Code subagent fan-out tool (Claude Code: "task")
- "execute" — Copilot CLI / VS Code run-command tool (Claude Code: "bash")
"agent" and "execute" were already flagged before this branch (introduced by
the fan-out PR); adding them to BuiltInTools clears the pre-existing warnings.
2. Add a cross-host tool portability check (CheckAgentToolPortability) so an
agent that declares a capability for only one host is flagged. Tool names are
matched case-sensitively (hosts resolve tools by exact spelling), so an agent
that lists e.g. only "edit" (Copilot / VS Code) without "Edit"/"Write"
(Claude Code) is reported as working on one host and silently tool-less on the
other. Findings are advisory (do not fail CI) and allowlistable via
"agent-tool-portability:AGENT:capability". Wired into the agents loop in
CheckCommand and covered by unit tests.
Also make the one existing single-host agent (optimizing-dotnet-performance)
portable by adding its Claude Code tool spellings, so the new check reports a
clean tree.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
* Add VS Code subagent metadata to dotnet-test code-testing agents
Enable the code-testing-* Research-Plan-Implement pipeline to fan out as
subagents in VS Code while keeping the GitHub Copilot CLI working.
Frontmatter (VS Code "coordinator/worker" pattern; portable tool aliases that
map in both VS Code and the CLI):
- code-testing-generator: tools: [agent, read, search, edit, execute] + an
agents: list (researcher, planner, implementer, builder, tester, fixer,
linter); softened the three runSubagent({ agent, prompt }) blocks to
tool-agnostic delegation prose.
- code-testing-implementer: tools: [agent, read, search, edit, execute] +
agents: (builder, tester, fixer, linter).
- Leaf agents (researcher/planner/builder/tester/fixer/linter):
tools: [read, search, edit, execute].
Why explicit tools (not ["*"]): VS Code has no all-tools wildcard for ools:
and a subagent's ools: overrides its inherited set, so ["*"] matched nothing
and stripped subagents of file tools (they ran without read/edit/search). The
CLI treats ["*"] as all-tools, so this was VS-Code-specific. The portable
aliases agent/read/search/edit/execute map to real tools in both environments;
agents: is ignored by the CLI.
README: document the VS Code chat.subagents.allowInvocationsFromSubagents
setting (off by default) needed for the nested implementer->builder/tester/
fixer/linter layer to fan out on large scopes; the CLI has no such gate.
Validated end to end: VS Code shows researcher->planner->implementer fanning out
with real file I/O (no "subagents lack file tools" warning); CLI fan-out intact
with skills still loading and tests passing. An all-tools baseline used only
tools within this enumerated set, confirming no CLI capability is restricted.
* Include skill tool in code-testing agent tool allowlists
Address PR review: the explicit `tools:` allowlists omitted the `skill`
tool, but every code-testing-* agent's prompt instructs calling skills
(e.g. `code-testing-extensions` for per-language guidance, `test-gap-analysis`,
`assertion-quality`). Because `tools:` is an override, omitting `skill` can
prevent the agents from loading those skills in environments that gate skill
invocation by the allowlist.
Add `skill` to all eight agents:
- orchestrators (generator, implementer): [agent, skill, read, search, edit, execute]
- workers (researcher/planner/builder/tester/fixer/linter): [skill, read, search, edit, execute]
Re-verified in the Copilot CLI: full fan-out (researcher -> planner ->
implementer/tester), the `code-testing-extensions` skill is invoked, and the
generated tests pass.
* code-testing-agent: add polyglot pipeline examples for Python/TypeScript/Go/Java
The code-testing-agent skill family is polyglot in description but in
practice biased toward .NET because dotnet-examples.md was the only
filled-in pipeline walkthrough. The four sub-agents that participate in
the Research-Plan-Implement pipeline (researcher, planner, implementer,
generator) all pointed to dotnet-examples.md whenever they suggested a
concrete example, which made it harder for the agent to produce
idiomatic non-.NET tests (e.g. in msbench top5-* benchmarks for
Python/Flask, TypeScript/Express).
This change
* adds four new example files mirroring dotnet-examples.md format
(source → research → plan → generated test → fix cycle → final report):
- python-examples.md (pytest, unittest.mock, Mock(spec=...), parametrize)
- typescript-examples.md (Vitest with notes for Jest; it.each,
async tests, fake timers, ESM/CJS fix cycle)
- go-examples.md (standard testing package, table-driven subtests,
hand-written fake repository, injected clock)
- java-examples.md (JUnit 5 + Mockito on Maven, @ParameterizedTest +
@CsvSource, Clock.fixed, Surefire fix cycles)
* updates code-testing-extensions/SKILL.md TOC to list the new files
and clarifies usage instructions to read the matching <language>-
examples.md alongside the base extension
* makes the "Concrete example" pointers in code-testing-generator,
code-testing-implementer, code-testing-planner and
code-testing-researcher agents language-agnostic (list all available
example files instead of hard-coding dotnet-examples.md)
* expands code-testing-researcher project-structure detection list to
cover more Python (tox.ini, noxfile.py, requirements*.txt, uv.lock,
poetry.lock, pdm.lock), JS/TS (.mts/.cts/.jsx, vitest.config.*,
jest.config.*), C++ (CMakeLists.txt, BUILD.bazel, meson.build),
Java/Kotlin (pom.xml, build.gradle[.kts], wrappers), and other
ecosystem files; expands the Identify-Language section accordingly
* extends the "Language-Specific Examples" section in
code-testing-agent/SKILL.md to summarise each example file
Validated with: skill-validator check --plugin ./plugins/dotnet-test
(23 skills, 11 agents — all checks passed) and markdownlint-cli2.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Address review feedback for #708
- go-examples.md: replace the hand-rolled `contains`/`stringIndex`
helpers with `strings.Contains` from the standard library. The hand-rolled
`contains` had subtly wrong semantics — `contains(""abc"", """")` returned
`false` while `strings.Contains` returns `true` — and the helpers are
unnecessary complexity for a code-generation example.
- go-examples.md: in the `go test -run` "wrong selection regex" sample fix
cycle, quote the test name and use `single_item` (matching the underscore
that the surrounding diagnosis text refers to) instead of the unquoted
`single item` which the shell would parse as two separate CLI arguments.
- java-examples.md: the source-tree file list described `Invoice.java` as a
`record` but the `InvoiceService.markAsPaid` example mutates the invoice
via `setStatus(...)` and `setPaidDate(...)` — records are immutable, so
the description was internally inconsistent. Re-describe it as a mutable
POJO with explicit mutators to match the service code.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Add end-to-end input/output examples to anchor expected behavior for the
LLM, addressing the Example Quality (2/5) review feedback.
Changes:
- New extensions/dotnet-examples.md with .NET-specific examples: sample
source code, research output, plan output, generated test file, fix
cycle walkthroughs, and final report
- SKILL.md: replace one-liner examples with strategy selection table,
pipeline walkthrough, and pointer to language-specific extensions
- Generator agent: add strategy decision examples and sample final report
- Researcher/Planner/Implementer agents: add references to extension
examples for concrete output shapes
Architecture keeps core agents language-agnostic — all language-specific
examples live in the extensions/ folder.
- Add 'Use when:' trigger phrases to all 7 sub-agent descriptions
so parent agents can reliably discover and delegate to them
- Add .testagent/ cleanup rule to generator agent to prevent
ephemeral pipeline state from being committed