Commit Graph

9 Commits

Author SHA1 Message Date
Amaury Levé 57733bebc8 Keep test agent state out of commits (#1108)
* Keep test agent state out of commits

Move broad test-generation pipeline state to host scratch storage, worktree-specific Git metadata, or OS temp, and enforce the exclusion in evals.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify absolute test agent state path

Use Git's explicit absolute path formatting in both test-generation entry points.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Prune Git metadata from test agent eval guards

Avoid scanning nested repositories and align the remaining TESTAGENT_DIR placeholder with the documented format.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify test agent command handoff

Require callers to provide exact commands, excerpts, or absolute TESTAGENT_DIR document paths to command-running sub-agents.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Reject all repository-local testagent entries

Match .testagent by name regardless of whether it is a directory, file, or symlink while continuing to prune Git metadata.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Verify external test agent artifacts

Restore broad-run artifact checks at the Git metadata path and pass the researched lint command and state directory to the linter agent.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Make testagent eval guards constant time

Check only the forbidden workspace-root path, including broken symlinks, instead of recursively traversing dependency trees.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Broaden comprehensive test generation

Treat explicit requirements as the floor for broad suites and add mutation-relevant equivalence-partition and invariant coverage without test-count padding.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Fix external artifact grader quoting

Run state checks directly in the harness shell so TESTAGENT_DIR expands after assignment, with an isolated command probe covering valid and forbidden states.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Run broad skill evals in Git worktrees

Initialize the seven broad evaluation roots as Git repositories so TESTAGENT_DIR resolves deterministically and external artifacts remain verifiable.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify non-stageable test agent state

Describe the real invariant across the pipeline: state may live under .git metadata but must never be version-controlled workspace content or appear in git status.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Standardize intermediate test state contract

Use one TESTAGENT_DIR placeholder, clearer intermediate-state terminology, and detect stageable research, plan, or status files regardless of directory name.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Use one Git root in workspace integrity eval

Baseline the fixture from the evaluation root so stageable intermediate-state files remain visible to the directory-independent guard.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Prune Vitest dependencies from state scan

Exclude node_modules through per-eval Git metadata so stageable state detection remains fast without modifying fixture content.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Strengthen focused intermediate-state guards

Separate shell execution, reject Git-metadata files on focused runs, include ignored state files, and prune node_modules with a pathspec exclusion.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

---------

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
2026-09-03 16:36:32 -07:00
Amaury Levé 5a06b20cc9 Support classic .NET test projects in dotnet-test (#993)
* Support classic .NET test projects

Teach dotnet-test skills and agents to preserve non-SDK projects, packages.config dependencies, explicit compile registration, legacy runners, and version-compatible MSTest APIs. Add regression evals for generation, execution, coverage, and authoring.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68

* Address classic test fixture review

Tighten the MSTest version grader, make the runner fixture assertion behavioral, and use nameof for the guarded parameter.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68

---------

Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68
2026-08-12 08:38:51 -07:00
Amaury Levé 77154137e8 dotnet-test: make code-testing agent tools Claude Code-compatible + add cross-host portability check (#856)
* dotnet-test: make code-testing agent tools declarations Claude Code-compatible

PR #847 added `tools: ["agent", "skill", "read", "search", "edit", "execute"]`
to the code-testing-* agents to enable VS Code / Copilot CLI subagent fan-out.
Those lowercase aliases map to real tools in VS Code and the Copilot CLI, but
Claude Code matches `tools:` against its own vocabulary (Task, Skill, Read,
Glob, Grep, Edit, Write, Bash). None of the aliases matched, so when these
agents are loaded into Claude Code via --plugin-dir and selected with
`claude --agent`, the agent was granted ZERO tools. A tool-less model asked
to generate tests emits a textual <tool_call> block and exits after one turn,
producing no file changes.

Append the Claude Code tool names to each agent's `tools:` list so the same
declaration works across all three runtimes (each honors the names it knows and
ignores the foreign ones):

- Orchestrators (generator, implementer): add Task, Skill, Read, Glob, Grep,
  Edit, Write, Bash (Task is the Claude Code equivalent of the `agent`
  fan-out tool).
- Workers (researcher, planner, builder, tester, fixer, linter): add Skill,
  Read, Glob, Grep, Edit, Write, Bash.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* skill-validator: complete built-in tools + add cross-host tool portability check

Two related follow-ups to the agent tools fix:

1. Address the skill-check review feedback. The validator's BuiltInTools set was
   missing three legitimate host tool spellings that are not case-insensitive
   matches of existing entries, so they were flagged as non-built-in:
   - "write"   — Claude Code file-creation tool (Copilot CLI / VS Code: "create")
   - "agent"   — Copilot CLI / VS Code subagent fan-out tool (Claude Code: "task")
   - "execute" — Copilot CLI / VS Code run-command tool (Claude Code: "bash")
   "agent" and "execute" were already flagged before this branch (introduced by
   the fan-out PR); adding them to BuiltInTools clears the pre-existing warnings.

2. Add a cross-host tool portability check (CheckAgentToolPortability) so an
   agent that declares a capability for only one host is flagged. Tool names are
   matched case-sensitively (hosts resolve tools by exact spelling), so an agent
   that lists e.g. only "edit" (Copilot / VS Code) without "Edit"/"Write"
   (Claude Code) is reported as working on one host and silently tool-less on the
   other. Findings are advisory (do not fail CI) and allowlistable via
   "agent-tool-portability:AGENT:capability". Wired into the agents loop in
   CheckCommand and covered by unit tests.

Also make the one existing single-host agent (optimizing-dotnet-performance)
portable by adding its Claude Code tool spellings, so the new check reports a
clean tree.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-07-07 09:00:22 +02:00
Petr Pokorny 06dd46e5f1 Enable VS Code subagent fan-out for the dotnet-test code-testing agents (#847)
* Add VS Code subagent metadata to dotnet-test code-testing agents

Enable the code-testing-* Research-Plan-Implement pipeline to fan out as
subagents in VS Code while keeping the GitHub Copilot CLI working.

Frontmatter (VS Code "coordinator/worker" pattern; portable tool aliases that
map in both VS Code and the CLI):
- code-testing-generator: tools: [agent, read, search, edit, execute] + an
  agents: list (researcher, planner, implementer, builder, tester, fixer,
  linter); softened the three runSubagent({ agent, prompt }) blocks to
  tool-agnostic delegation prose.
- code-testing-implementer: tools: [agent, read, search, edit, execute] +
  agents: (builder, tester, fixer, linter).
- Leaf agents (researcher/planner/builder/tester/fixer/linter):
  tools: [read, search, edit, execute].

Why explicit tools (not ["*"]): VS Code has no all-tools wildcard for 	ools:
and a subagent's 	ools: overrides its inherited set, so ["*"] matched nothing
and stripped subagents of file tools (they ran without read/edit/search). The
CLI treats ["*"] as all-tools, so this was VS-Code-specific. The portable
aliases agent/read/search/edit/execute map to real tools in both environments;
agents: is ignored by the CLI.

README: document the VS Code chat.subagents.allowInvocationsFromSubagents
setting (off by default) needed for the nested implementer->builder/tester/
fixer/linter layer to fan out on large scopes; the CLI has no such gate.

Validated end to end: VS Code shows researcher->planner->implementer fanning out
with real file I/O (no "subagents lack file tools" warning); CLI fan-out intact
with skills still loading and tests passing. An all-tools baseline used only
tools within this enumerated set, confirming no CLI capability is restricted.

* Include skill tool in code-testing agent tool allowlists

Address PR review: the explicit `tools:` allowlists omitted the `skill`
tool, but every code-testing-* agent's prompt instructs calling skills
(e.g. `code-testing-extensions` for per-language guidance, `test-gap-analysis`,
`assertion-quality`). Because `tools:` is an override, omitting `skill` can
prevent the agents from loading those skills in environments that gate skill
invocation by the allowlist.

Add `skill` to all eight agents:
- orchestrators (generator, implementer): [agent, skill, read, search, edit, execute]
- workers (researcher/planner/builder/tester/fixer/linter): [skill, read, search, edit, execute]

Re-verified in the Copilot CLI: full fan-out (researcher -> planner ->
implementer/tester), the `code-testing-extensions` skill is invoked, and the
generated tests pass.
2026-06-30 19:42:39 +02:00
Amaury Levé 3d59e44c7e Fix coverage-analysis activation for plateau diagnosis prompts (#647)
* Fix coverage-analysis activation for plateau diagnosis prompts

coverage-analysis SKILL.md:
- Trim verbose implementation details (provider detection,
  ReportGenerator) that consumed description budget without
  aiding skill activation
- Add explicit USE FOR keywords: coverage stuck, coverage plateau,
  can't increase coverage, what's blocking coverage

code-testing-agent SKILL.md:
- Add 'diagnosing coverage plateaus or CRAP score computation
  (use coverage-analysis)' to DO NOT USE FOR boundary to prevent
  test-generation skill from intercepting diagnostic prompts

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* Strengthen code-testing-agent activation; harden coverage-analysis isolated mode

Address eval regressions reported on PR #647 (run 25813728646):

1. code-testing-agent: `Generate tests for ContosoUniversity ASP.NET Core MVC app`
   was NOT ACTIVATED in plugin mode (detectedSkills=[], skillEventCount=0,
   invokedAgents=[]). The model bypassed the skill system entirely.

   - SKILL.md description: restructure to use the proven `Use when user says ...`
     pattern with quoted trigger phrases (matching the run-tests skill that
     consistently activates), make the link to the code-testing-generator
     sub-agent explicit, and tighten DO NOT USE FOR clauses.
   - eval prompt (eval.yaml + eval.vally.yaml): make the request
     pipeline-shaped (`project-wide, multi-file test generation task`,
     `scaffold a new test project`) so the model recognizes it as multi-step
     work that benefits from the orchestrated pipeline. Explicitly request
     coverlet.collector + a Cobertura XML run so rubric criterion 1
     (`high line coverage as reported by the Cobertura XML in TestResults/`)
     becomes achievable without overfitting.

2. code-testing-tester agent + code-testing-extensions/dotnet.md: open a
   scoped exception to the `skip coverage tools` rule. Default behavior
   stays the same, but when the user/harness explicitly asks for a
   Cobertura/XML coverage artifact, the agent may add coverlet.collector
   to the generated test csproj so the harness's coverage command produces
   output. The agent still does not run the coverage command itself.

3. coverage-analysis SKILL.md: add a `User-visible output is mandatory`
   guard at the top of the Workflow section. The latest eval showed isolated
   mode producing literally `(no output)` in 2 of 3 scenarios — the agent
   ran Compute-CrapScores.ps1 / Extract-MethodCoverage.ps1 / ReportGenerator
   in parallel, then the session ended without ever surfacing findings.
   The guard tells the agent to always return a partial summary instead of
   ending silent, and to deprioritize ReportGenerator HTML when budget is
   tight. (Plugin-mode quality is already strong: 4.3 / 4.3 / 5.0 — no
   regression risk there.)

Aggregate dotnet-test plugin description size: 14,925 chars (limit 15,000).
skill-validator check passes (22 skills, 11 agents, 1 plugin); markdownlint
passes for all 4 modified files.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Restore MSTest modernization exclusion in code-testing-agent description

* Fix isolated-mode coverage-analysis: emit summary before optional ReportGenerator

The previous workflow encouraged the agent to run `dotnet tool install` for
ReportGenerator in parallel with the CRAP scoring scripts (Phase 2 "Steps
3 and 4 in parallel" + Phase 3 "Steps 5 and 6 in parallel"). In isolated
mode that pattern reliably crashed the session with "Failed to persist
session events: timeout while waiting for mutex to become available"
right after the scripts returned valid data, so the agent never produced
the user-facing summary.

Restructure the workflow into 5 phases:

- Phase 1 (Setup) - unchanged
- Phase 2 (Test execution) - skip when Cobertura XML already exists
- Phase 3 (Analysis) - run only the two PowerShell scripts, no RG
- Phase 4 (User-facing summary) - MANDATORY, must be the next assistant
  response after Phase 3, before any RG work; also save
  coverage-analysis.md as a secondary follow-up
- Phase 5 (ReportGenerator HTML/CSV) - strictly optional, post-summary,
  skipped by default for existing-Cobertura and plateau-diagnosis paths

Also update references/output-format.md so the Reports section marks RG
artifacts as "Not generated (optional - request HTML reports to enable)"
when Phase 5 has not run, and update references/guidelines.md so the
"show and open the markdown report" rule explicitly defers to the
user-facing assistant response.

Targets the isolated-mode regressions in PR #647 eval:
- Project-wide coverage with existing Cobertura: 1.0/5 -> expected 3+
- Coverage plateau diagnosis: 1.0/5 -> expected 3+
- Run coverage from scratch: 2.3/5 -> expected steady or up

Verified: skill-validator check --plugin ./plugins/dotnet-test passes;
markdownlint-cli2 clean on all 3 modified files.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Address PR #647 review comments

1. Prerequisites: distinguish the from-scratch path (needs NuGet for the
   coverage-provider package install + optional internet for ReportGenerator)
   from the existing-Cobertura path (needs neither).

2. Add Step 2c "Discover or accept existing Cobertura XML" so the
   existing-data path actually has $coberturaFiles populated before
   Phase 3, instead of relying on Phase 2's discovery (which it
   skips). Also clarify Step 2's destructive Remove-Item only manages
   the skill-owned coverage-analysis/ subdirectory.

3. references/output-format.md: replace the unconditional
   "Reports saved to: <coverageDir>/reports/" line with one that
   always points at <coverageDir>/ (markdown summary + raw Cobertura)
   and only mentions reports/ if Phase 5 ran.

4. Have Compute-CrapScores.ps1 emit OVERALL_LINE_COVERAGE and
   OVERALL_BRANCH_COVERAGE from the Cobertura root attributes, and
   update Phase 4 to read those values directly from the script's
   output. The Phase 4 mandatory-summary rule no longer requires a
   separate XML parse before composing the response.

Verified: skill-validator check --plugin ./plugins/dotnet-test passes;
markdownlint-cli2 clean on all modified files; Compute-CrapScores.ps1
smoke-tested on a synthetic Cobertura XML (emits OVERALL_LINE_COVERAGE:75,
OVERALL_BRANCH_COVERAGE:50 alongside HOTSPOTS).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Apply suggestion from @Evangelink

* Address unresolved coverage-analysis and eval review comments

* Refine follow-up review feedback from validation

* Tighten coverage aggregation fallback notes and counters

* Clarify pre-response save instruction wording

---------

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
2026-05-14 17:41:31 +00:00
Jan Krivanek 81946d2a38 Add proxy for CTA extension files (#583) 2026-04-24 12:19:17 +02:00
Jan Krivanek 05aeb657e6 Add license to agent files (#568) 2026-04-21 12:57:18 +00:00
Amaury Levé 1125fe7864 improve(dotnet-test): enhance agent descriptions for subagent discovery (#560)
- Add 'Use when:' trigger phrases to all 7 sub-agent descriptions
  so parent agents can reliably discover and delegate to them
- Add .testagent/ cleanup rule to generator agent to prevent
  ephemeral pipeline state from being committed
2026-04-20 19:51:09 +02:00
Jan Krivanek e4670b33a1 [PoC] Code testing agent + tests PoC (#433)
* Add code testing agent agents + tests PoC

* Remove AssertionEvaluator.cs change (moved to dev/jankrivanek/agents-evals)
2026-03-30 17:57:08 +02:00