Commit Graph

25 Commits

Author SHA1 Message Date
Amaury Levé 57733bebc8 Keep test agent state out of commits (#1108)
* Keep test agent state out of commits

Move broad test-generation pipeline state to host scratch storage, worktree-specific Git metadata, or OS temp, and enforce the exclusion in evals.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify absolute test agent state path

Use Git's explicit absolute path formatting in both test-generation entry points.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Prune Git metadata from test agent eval guards

Avoid scanning nested repositories and align the remaining TESTAGENT_DIR placeholder with the documented format.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify test agent command handoff

Require callers to provide exact commands, excerpts, or absolute TESTAGENT_DIR document paths to command-running sub-agents.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Reject all repository-local testagent entries

Match .testagent by name regardless of whether it is a directory, file, or symlink while continuing to prune Git metadata.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Verify external test agent artifacts

Restore broad-run artifact checks at the Git metadata path and pass the researched lint command and state directory to the linter agent.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Make testagent eval guards constant time

Check only the forbidden workspace-root path, including broken symlinks, instead of recursively traversing dependency trees.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Broaden comprehensive test generation

Treat explicit requirements as the floor for broad suites and add mutation-relevant equivalence-partition and invariant coverage without test-count padding.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Fix external artifact grader quoting

Run state checks directly in the harness shell so TESTAGENT_DIR expands after assignment, with an isolated command probe covering valid and forbidden states.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Run broad skill evals in Git worktrees

Initialize the seven broad evaluation roots as Git repositories so TESTAGENT_DIR resolves deterministically and external artifacts remain verifiable.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify non-stageable test agent state

Describe the real invariant across the pipeline: state may live under .git metadata but must never be version-controlled workspace content or appear in git status.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Standardize intermediate test state contract

Use one TESTAGENT_DIR placeholder, clearer intermediate-state terminology, and detect stageable research, plan, or status files regardless of directory name.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Use one Git root in workspace integrity eval

Baseline the fixture from the evaluation root so stageable intermediate-state files remain visible to the directory-independent guard.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Prune Vitest dependencies from state scan

Exclude node_modules through per-eval Git metadata so stageable state detection remains fast without modifying fixture content.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Strengthen focused intermediate-state guards

Separate shell execution, reject Git-metadata files on focused runs, include ignored state files, and prune node_modules with a pathspec exclusion.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

---------

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
2026-09-03 16:36:32 -07:00
Amaury Levé 5a06b20cc9 Support classic .NET test projects in dotnet-test (#993)
* Support classic .NET test projects

Teach dotnet-test skills and agents to preserve non-SDK projects, packages.config dependencies, explicit compile registration, legacy runners, and version-compatible MSTest APIs. Add regression evals for generation, execution, coverage, and authoring.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68

* Address classic test fixture review

Tighten the MSTest version grader, make the runner fixture assertion behavioral, and use nameof for the guarded parameter.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68

---------

Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68
2026-08-12 08:38:51 -07:00
Amaury Levé 71414ce000 Improve dotnet-test eval coverage and efficiency (#917)
* Improve dotnet-test eval coverage and efficiency

Address remaining high-confidence items from #899 by bounding the code-testing pipeline and adding eval coverage for grade-tests and find-untested-sources.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: e430fee9-d3df-4ef5-85a4-745ae4b17046

* Fix dotnet-test eval activation and quality

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 9c5c1a52-4f99-49d6-b503-1bec713a6e98

* Strengthen dotnet-test skill activation

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 9c5c1a52-4f99-49d6-b503-1bec713a6e98
2026-07-24 16:20:47 +02:00
Amaury Levé 4df4da469a Upgrade agentic workflows and fix stale PR cleanup (#916)
* chore: upgrade gh-aw runtime

* fix: paginate stale pull request cleanup

Upgrade the generated agentic workflow assets and ensure stale PR discovery includes every result page and draft pull requests.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 968a22c2-327f-4d26-8f86-1c59bcd323ea

* Improve dotnet-test eval coverage and efficiency

Address remaining high-confidence items from #899 by bounding the code-testing pipeline and adding eval coverage for grade-tests and find-untested-sources.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: e430fee9-d3df-4ef5-85a4-745ae4b17046

* fix: address agentic workflow review

Pin the Copilot setup checkout action and include the cutoff date in stale PR search results.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 968a22c2-327f-4d26-8f86-1c59bcd323ea

* Improve test migration skill guidance

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 5e19d263-02a6-45b1-9cfb-424fa4d10863

* fix: add fixture namespace imports

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: a464e6e4-3e45-41fe-b17d-887c8cb8a448

* test: assert GetOrderById grade

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: a464e6e4-3e45-41fe-b17d-887c8cb8a448

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-07-20 13:20:14 +00:00
Amaury Levé 77154137e8 dotnet-test: make code-testing agent tools Claude Code-compatible + add cross-host portability check (#856)
* dotnet-test: make code-testing agent tools declarations Claude Code-compatible

PR #847 added `tools: ["agent", "skill", "read", "search", "edit", "execute"]`
to the code-testing-* agents to enable VS Code / Copilot CLI subagent fan-out.
Those lowercase aliases map to real tools in VS Code and the Copilot CLI, but
Claude Code matches `tools:` against its own vocabulary (Task, Skill, Read,
Glob, Grep, Edit, Write, Bash). None of the aliases matched, so when these
agents are loaded into Claude Code via --plugin-dir and selected with
`claude --agent`, the agent was granted ZERO tools. A tool-less model asked
to generate tests emits a textual <tool_call> block and exits after one turn,
producing no file changes.

Append the Claude Code tool names to each agent's `tools:` list so the same
declaration works across all three runtimes (each honors the names it knows and
ignores the foreign ones):

- Orchestrators (generator, implementer): add Task, Skill, Read, Glob, Grep,
  Edit, Write, Bash (Task is the Claude Code equivalent of the `agent`
  fan-out tool).
- Workers (researcher, planner, builder, tester, fixer, linter): add Skill,
  Read, Glob, Grep, Edit, Write, Bash.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* skill-validator: complete built-in tools + add cross-host tool portability check

Two related follow-ups to the agent tools fix:

1. Address the skill-check review feedback. The validator's BuiltInTools set was
   missing three legitimate host tool spellings that are not case-insensitive
   matches of existing entries, so they were flagged as non-built-in:
   - "write"   — Claude Code file-creation tool (Copilot CLI / VS Code: "create")
   - "agent"   — Copilot CLI / VS Code subagent fan-out tool (Claude Code: "task")
   - "execute" — Copilot CLI / VS Code run-command tool (Claude Code: "bash")
   "agent" and "execute" were already flagged before this branch (introduced by
   the fan-out PR); adding them to BuiltInTools clears the pre-existing warnings.

2. Add a cross-host tool portability check (CheckAgentToolPortability) so an
   agent that declares a capability for only one host is flagged. Tool names are
   matched case-sensitively (hosts resolve tools by exact spelling), so an agent
   that lists e.g. only "edit" (Copilot / VS Code) without "Edit"/"Write"
   (Claude Code) is reported as working on one host and silently tool-less on the
   other. Findings are advisory (do not fail CI) and allowlistable via
   "agent-tool-portability:AGENT:capability". Wired into the agents loop in
   CheckCommand and covered by unit tests.

Also make the one existing single-host agent (optimizing-dotnet-performance)
portable by adding its Claude Code tool spellings, so the new check reports a
clean tree.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-07-07 09:00:22 +02:00
Petr Pokorny 06dd46e5f1 Enable VS Code subagent fan-out for the dotnet-test code-testing agents (#847)
* Add VS Code subagent metadata to dotnet-test code-testing agents

Enable the code-testing-* Research-Plan-Implement pipeline to fan out as
subagents in VS Code while keeping the GitHub Copilot CLI working.

Frontmatter (VS Code "coordinator/worker" pattern; portable tool aliases that
map in both VS Code and the CLI):
- code-testing-generator: tools: [agent, read, search, edit, execute] + an
  agents: list (researcher, planner, implementer, builder, tester, fixer,
  linter); softened the three runSubagent({ agent, prompt }) blocks to
  tool-agnostic delegation prose.
- code-testing-implementer: tools: [agent, read, search, edit, execute] +
  agents: (builder, tester, fixer, linter).
- Leaf agents (researcher/planner/builder/tester/fixer/linter):
  tools: [read, search, edit, execute].

Why explicit tools (not ["*"]): VS Code has no all-tools wildcard for 	ools:
and a subagent's 	ools: overrides its inherited set, so ["*"] matched nothing
and stripped subagents of file tools (they ran without read/edit/search). The
CLI treats ["*"] as all-tools, so this was VS-Code-specific. The portable
aliases agent/read/search/edit/execute map to real tools in both environments;
agents: is ignored by the CLI.

README: document the VS Code chat.subagents.allowInvocationsFromSubagents
setting (off by default) needed for the nested implementer->builder/tester/
fixer/linter layer to fan out on large scopes; the CLI has no such gate.

Validated end to end: VS Code shows researcher->planner->implementer fanning out
with real file I/O (no "subagents lack file tools" warning); CLI fan-out intact
with skills still loading and tests passing. An all-tools baseline used only
tools within this enumerated set, confirming no CLI capability is restricted.

* Include skill tool in code-testing agent tool allowlists

Address PR review: the explicit `tools:` allowlists omitted the `skill`
tool, but every code-testing-* agent's prompt instructs calling skills
(e.g. `code-testing-extensions` for per-language guidance, `test-gap-analysis`,
`assertion-quality`). Because `tools:` is an override, omitting `skill` can
prevent the agents from loading those skills in environments that gate skill
invocation by the allowlist.

Add `skill` to all eight agents:
- orchestrators (generator, implementer): [agent, skill, read, search, edit, execute]
- workers (researcher/planner/builder/tester/fixer/linter): [skill, read, search, edit, execute]

Re-verified in the Copilot CLI: full fan-out (researcher -> planner ->
implementer/tester), the `code-testing-extensions` skill is invoked, and the
generated tests pass.
2026-06-30 19:42:39 +02:00
Amaury Levé 0c98ce6bd9 Direct strategy must still run the Step 7 pre-completion gate (#793)
* Direct strategy must still run the Step 7 pre-completion gate

The Direct strategy correctly skips the research/plan/implement sub-agents
for small single-file tasks, but the wording let agents also skip the
Step 7 pre-completion gate (test-gap-analysis + assertion-quality +
scenario coverage) — treating a single-file task that enumerates specific
behaviors as 'trivially small'.

This is the dominant failure mode observed on behavior-enumerating tasks:
the agent writes one test file directly and finishes with no
assertion-strength or scenario-coverage check, producing weak assertions
(mutation survivors) and missing required edge/negative cases.

Clarify in both the generator Step 2 strategy table and the
code-testing-agent SKILL.md that Direct trades away only the sub-agents,
never the gate, and that a request naming a specific symbol or enumerating
scenarios is not 'trivially small' and must run the gate.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Address review: align Direct gate trigger with Step 7 threshold; clarify gate in SKILL.md

- Step 2 Direct cell no longer introduces a separate 'names a specific
  symbol' gate trigger that contradicted Step 7. It now defers to Step 7's
  own threshold (>=5 tests, or any enumerated behaviors/scenarios).
- SKILL.md now names what/where the gate is: the generator's Step 7
  (test-gap-analysis + assertion-quality).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Amaury Levé <evangelink@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-22 10:38:14 +02:00
Amaury Levé 5d717dbdd1 Add prompt-scenario coverage check to code-testing-generator gate (#789)
* Add prompt-scenario coverage check to code-testing-generator gate

The pre-completion gate already verifies assertion strength (pseudo-mutation
and assertion-depth checks), but two recurring failure modes still slip
through when the prompt enumerates specific behaviors:

- Testing an *adjacent* function/helper instead of the exact feature named
  in the objective, leaving the requested behavior uncovered.
- Covering only a single representative case when the scenario wording
  implies multiple variations or pins a condition to a specific position
  or structure.

Add a third gate item that maps each enumerated scenario to a dedicated
test, requires targeting the exact named function (preferring the canonical
existing test file), and requires honoring range/positional qualifiers
literally.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Address review: genericize example, fix gate-count consistency

- Remove benchmark-specific symbol names from the target-the-named-function
  bullet to avoid overfitting; phrase it generically.
- Fix the gate intro that said 'The two skills below' now that there are
  three numbered items (the third is a prompt self-review, not a skill).
- Update Step 8 and Rule 11 so re-running the gate includes the new
  prompt-scenario coverage check, not just test-gap-analysis + assertion-quality.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Amaury Levé <evangelink@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-18 16:20:46 +02:00
Amaury Levé f389af86c7 code-testing-generator: mandatory pre-completion self-review gate (#768)
* code-testing-generator: mandatory pre-completion self-review gate

Replaces the prose `Verify tests are implementation-specific'' bullet
in Step 7 of the code-testing-generator agent with a mandatory
pre-completion self-review gate that invokes two existing plugin
skills:

1. `test-gap-analysis'' (pseudo-mutation check) against the source
   files tested and the produced test files
2. `assertion-quality'' (trivial/tautological assertion check)
   against the produced test files

Both skills already ship in plugins/dotnet-test; this PR only wires
them into the generator's workflow as a mandatory gate before
declaring a run complete.

The two skills' `When to Use'' sections are extended to list
`called by code-testing-generator as a pre-completion self-review
step'' as a recognised use case so the model does not refuse the
invocation. The skill descriptions are unchanged (frontmatter is
already at the 1024-char limit).

Rule 11 of the generator agent is updated to list the gate alongside
final build, final test, and coverage-gap review as mandatory for ALL
strategies including Direct.

A matching rubric item is added to the ContosoUniversity scenario in
the code-testing-agent eval (yaml and vally) so the LLM judge can
verify the gate was actually invoked on the trajectory.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Apply suggestions from code review

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-06-17 10:14:26 +00:00
Amaury Levé 754011b5ee Add workspace-integrity guardrail to code-testing agents (#773)
The code-testing-generator/implementer agents could treat an unusual or
scaffolded workspace (e.g. a gutted repo with an injected synthetic module)
as corruption and "repair" it with git checkout/restore/reset/clean or rm,
restoring deleted tracked files and testing the wrong code.

- Replace generator Rule 5 ("Clean git first - stash changes") with an
  explicit "Treat the workspace as delivered" rule, and add a "Never mutate
  version control" rule. Output must be purely additive test files.
- Add a no-revert/no-clean invariant to the implementer's edit boundaries.
- Add a 'workspace integrity' eval to the code-testing-agent suite
  (eval.yaml + eval.vally.yaml). The fixture looks gutted: a metricsd project
  whose real core/io modules are committed at HEAD but deleted from the
  working tree, leaving only a synthetic 'synthstr' decoy. A git restore would
  resurrect the deleted sentinel files; graders fail if they reappear and
  require passing pytest tests for the module as delivered.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-16 13:49:34 +00:00
Amaury Levé 49a77a90bf code-testing-agent: add polyglot pipeline examples for Python/TypeScript/Go/Java (#708)
* code-testing-agent: add polyglot pipeline examples for Python/TypeScript/Go/Java

The code-testing-agent skill family is polyglot in description but in
practice biased toward .NET because dotnet-examples.md was the only
filled-in pipeline walkthrough. The four sub-agents that participate in
the Research-Plan-Implement pipeline (researcher, planner, implementer,
generator) all pointed to dotnet-examples.md whenever they suggested a
concrete example, which made it harder for the agent to produce
idiomatic non-.NET tests (e.g. in msbench top5-* benchmarks for
Python/Flask, TypeScript/Express).

This change

* adds four new example files mirroring dotnet-examples.md format
  (source → research → plan → generated test → fix cycle → final report):
  - python-examples.md (pytest, unittest.mock, Mock(spec=...), parametrize)
  - typescript-examples.md (Vitest with notes for Jest; it.each,
    async tests, fake timers, ESM/CJS fix cycle)
  - go-examples.md (standard testing package, table-driven subtests,
    hand-written fake repository, injected clock)
  - java-examples.md (JUnit 5 + Mockito on Maven, @ParameterizedTest +
    @CsvSource, Clock.fixed, Surefire fix cycles)
* updates code-testing-extensions/SKILL.md TOC to list the new files
  and clarifies usage instructions to read the matching <language>-
  examples.md alongside the base extension
* makes the "Concrete example" pointers in code-testing-generator,
  code-testing-implementer, code-testing-planner and
  code-testing-researcher agents language-agnostic (list all available
  example files instead of hard-coding dotnet-examples.md)
* expands code-testing-researcher project-structure detection list to
  cover more Python (tox.ini, noxfile.py, requirements*.txt, uv.lock,
  poetry.lock, pdm.lock), JS/TS (.mts/.cts/.jsx, vitest.config.*,
  jest.config.*), C++ (CMakeLists.txt, BUILD.bazel, meson.build),
  Java/Kotlin (pom.xml, build.gradle[.kts], wrappers), and other
  ecosystem files; expands the Identify-Language section accordingly
* extends the "Language-Specific Examples" section in
  code-testing-agent/SKILL.md to summarise each example file

Validated with: skill-validator check --plugin ./plugins/dotnet-test
(23 skills, 11 agents — all checks passed) and markdownlint-cli2.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Address review feedback for #708

- go-examples.md: replace the hand-rolled `contains`/`stringIndex`
  helpers with `strings.Contains` from the standard library. The hand-rolled
  `contains` had subtly wrong semantics — `contains(""abc"", """")` returned
  `false` while `strings.Contains` returns `true` — and the helpers are
  unnecessary complexity for a code-generation example.
- go-examples.md: in the `go test -run` "wrong selection regex" sample fix
  cycle, quote the test name and use `single_item` (matching the underscore
  that the surrounding diagnosis text refers to) instead of the unquoted
  `single item` which the shell would parse as two separate CLI arguments.
- java-examples.md: the source-tree file list described `Invoice.java` as a
  `record` but the `InvoiceService.markAsPaid` example mutates the invoice
  via `setStatus(...)` and `setPaidDate(...)` — records are immutable, so
  the description was internally inconsistent. Re-describe it as a mutable
  POJO with explicit mutators to match the service code.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-02 13:42:46 +00:00
Jan Krivanek 2fcfc099b9 Improve direct scenario (#675) 2026-05-19 19:34:29 +02:00
Amaury Levé dde9f99cf0 dotnet-test: drop tools field from agents that declared it (#665)
The four agents (code-testing-generator, test-migration, testability-migration,
test-quality-auditor) were the only dotnet-test agents declaring a 'tools:'
list. Per Jan's suggestion on #660, drop the field so they inherit the runtime's
default tool surface (matching the convention used by every other agent in the
repo).

Follow-up to #660.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-05-18 14:47:00 +02:00
Amaury Levé 2a193300ca Replace 'terminal' with 'bash' and 'powershell' in agent tools (#660)
The skill-validator only recognizes 'bash' and 'powershell' as built-in shell tools, not 'terminal'. Update the four affected agents in the dotnet-test plugin to use both for cross-platform support.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-05-18 13:39:33 +02:00
Amaury Levé 3d59e44c7e Fix coverage-analysis activation for plateau diagnosis prompts (#647)
* Fix coverage-analysis activation for plateau diagnosis prompts

coverage-analysis SKILL.md:
- Trim verbose implementation details (provider detection,
  ReportGenerator) that consumed description budget without
  aiding skill activation
- Add explicit USE FOR keywords: coverage stuck, coverage plateau,
  can't increase coverage, what's blocking coverage

code-testing-agent SKILL.md:
- Add 'diagnosing coverage plateaus or CRAP score computation
  (use coverage-analysis)' to DO NOT USE FOR boundary to prevent
  test-generation skill from intercepting diagnostic prompts

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* Strengthen code-testing-agent activation; harden coverage-analysis isolated mode

Address eval regressions reported on PR #647 (run 25813728646):

1. code-testing-agent: `Generate tests for ContosoUniversity ASP.NET Core MVC app`
   was NOT ACTIVATED in plugin mode (detectedSkills=[], skillEventCount=0,
   invokedAgents=[]). The model bypassed the skill system entirely.

   - SKILL.md description: restructure to use the proven `Use when user says ...`
     pattern with quoted trigger phrases (matching the run-tests skill that
     consistently activates), make the link to the code-testing-generator
     sub-agent explicit, and tighten DO NOT USE FOR clauses.
   - eval prompt (eval.yaml + eval.vally.yaml): make the request
     pipeline-shaped (`project-wide, multi-file test generation task`,
     `scaffold a new test project`) so the model recognizes it as multi-step
     work that benefits from the orchestrated pipeline. Explicitly request
     coverlet.collector + a Cobertura XML run so rubric criterion 1
     (`high line coverage as reported by the Cobertura XML in TestResults/`)
     becomes achievable without overfitting.

2. code-testing-tester agent + code-testing-extensions/dotnet.md: open a
   scoped exception to the `skip coverage tools` rule. Default behavior
   stays the same, but when the user/harness explicitly asks for a
   Cobertura/XML coverage artifact, the agent may add coverlet.collector
   to the generated test csproj so the harness's coverage command produces
   output. The agent still does not run the coverage command itself.

3. coverage-analysis SKILL.md: add a `User-visible output is mandatory`
   guard at the top of the Workflow section. The latest eval showed isolated
   mode producing literally `(no output)` in 2 of 3 scenarios — the agent
   ran Compute-CrapScores.ps1 / Extract-MethodCoverage.ps1 / ReportGenerator
   in parallel, then the session ended without ever surfacing findings.
   The guard tells the agent to always return a partial summary instead of
   ending silent, and to deprioritize ReportGenerator HTML when budget is
   tight. (Plugin-mode quality is already strong: 4.3 / 4.3 / 5.0 — no
   regression risk there.)

Aggregate dotnet-test plugin description size: 14,925 chars (limit 15,000).
skill-validator check passes (22 skills, 11 agents, 1 plugin); markdownlint
passes for all 4 modified files.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Restore MSTest modernization exclusion in code-testing-agent description

* Fix isolated-mode coverage-analysis: emit summary before optional ReportGenerator

The previous workflow encouraged the agent to run `dotnet tool install` for
ReportGenerator in parallel with the CRAP scoring scripts (Phase 2 "Steps
3 and 4 in parallel" + Phase 3 "Steps 5 and 6 in parallel"). In isolated
mode that pattern reliably crashed the session with "Failed to persist
session events: timeout while waiting for mutex to become available"
right after the scripts returned valid data, so the agent never produced
the user-facing summary.

Restructure the workflow into 5 phases:

- Phase 1 (Setup) - unchanged
- Phase 2 (Test execution) - skip when Cobertura XML already exists
- Phase 3 (Analysis) - run only the two PowerShell scripts, no RG
- Phase 4 (User-facing summary) - MANDATORY, must be the next assistant
  response after Phase 3, before any RG work; also save
  coverage-analysis.md as a secondary follow-up
- Phase 5 (ReportGenerator HTML/CSV) - strictly optional, post-summary,
  skipped by default for existing-Cobertura and plateau-diagnosis paths

Also update references/output-format.md so the Reports section marks RG
artifacts as "Not generated (optional - request HTML reports to enable)"
when Phase 5 has not run, and update references/guidelines.md so the
"show and open the markdown report" rule explicitly defers to the
user-facing assistant response.

Targets the isolated-mode regressions in PR #647 eval:
- Project-wide coverage with existing Cobertura: 1.0/5 -> expected 3+
- Coverage plateau diagnosis: 1.0/5 -> expected 3+
- Run coverage from scratch: 2.3/5 -> expected steady or up

Verified: skill-validator check --plugin ./plugins/dotnet-test passes;
markdownlint-cli2 clean on all 3 modified files.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Address PR #647 review comments

1. Prerequisites: distinguish the from-scratch path (needs NuGet for the
   coverage-provider package install + optional internet for ReportGenerator)
   from the existing-Cobertura path (needs neither).

2. Add Step 2c "Discover or accept existing Cobertura XML" so the
   existing-data path actually has $coberturaFiles populated before
   Phase 3, instead of relying on Phase 2's discovery (which it
   skips). Also clarify Step 2's destructive Remove-Item only manages
   the skill-owned coverage-analysis/ subdirectory.

3. references/output-format.md: replace the unconditional
   "Reports saved to: <coverageDir>/reports/" line with one that
   always points at <coverageDir>/ (markdown summary + raw Cobertura)
   and only mentions reports/ if Phase 5 ran.

4. Have Compute-CrapScores.ps1 emit OVERALL_LINE_COVERAGE and
   OVERALL_BRANCH_COVERAGE from the Cobertura root attributes, and
   update Phase 4 to read those values directly from the script's
   output. The Phase 4 mandatory-summary rule no longer requires a
   separate XML parse before composing the response.

Verified: skill-validator check --plugin ./plugins/dotnet-test passes;
markdownlint-cli2 clean on all modified files; Compute-CrapScores.ps1
smoke-tested on a synthetic Cobertura XML (emits OVERALL_LINE_COVERAGE:75,
OVERALL_BRANCH_COVERAGE:50 alongside HOTSPOTS).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Apply suggestion from @Evangelink

* Address unresolved coverage-analysis and eval review comments

* Refine follow-up review feedback from validation

* Tighten coverage aggregation fallback notes and counters

* Clarify pre-response save instruction wording

---------

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
2026-05-14 17:41:31 +00:00
YuliiaKovalova e1e8568375 Revert "dotnet-test: add unit-under-test + behaviors quality cue to code-test…" (#651)
This reverts commit f1b09eba79.
2026-05-14 17:00:07 +02:00
YuliiaKovalova f1b09eba79 dotnet-test: add unit-under-test + behaviors quality cue to code-testing-generator (#646)
* CTA: invocation-only baseline (dispatch mechanics, no content rules)

Experiment branch to isolate the impact of "invoke prompted subagents
more often" from the impact of "make those subagents do richer work."
Same baseline as dev/ykovalova/cta-prompt-tuning (main = 66628b6), but
strips out every content/quality rule and keeps only the dispatch
plumbing.

Comparison branch: dev/ykovalova/cta-prompt-tuning (HEAD: bd530be)
which contains both the dispatch mechanics AND content/quality rules.

Files modified (3 vs 5 in cta-prompt-tuning):
- code-testing-generator.agent.md  +110 lines
- code-testing-implementer.agent.md +11 lines
- code-testing-fixer.agent.md      +1  line
- code-testing-researcher.agent.md  UNTOUCHED (baseline)
- code-testing-planner.agent.md     UNTOUCHED (baseline)

KEPT (invocation / dispatch mechanics):

Generator:
- Rule 1: every task() call MUST use agent_type "dotnet-test:code-testing-..."
  (without this, calls dispatch generic built-ins and never reach the
  named CTA agents)
- Rule 2: routing table -- which named agent for which job
- Rule 3: prefer one named-agent dispatch over many tool calls
- Rule 4: orchestrator MUST NOT edit/create test files itself
  (forces implementer dispatch)
- Rule 5: orchestrator MUST NOT run builds/tests via terminal
  (forces builder/tester dispatch)
- Rule 6: every run MUST dispatch the planner (no exceptions for "small"
  scope; Direct still goes through planner)
- Rule 7: every build/test failure MUST dispatch the fixer
- Step 1b: mandatory initial researcher dispatch (every strategy)
- Direct strategy rewritten: dispatches planner -> implementer -> builder
  -> tester -> fixer -> linter (was "Skip Steps 3-5, write tests inline")
- All Step 3/4/5/6/7/8/9 dispatches converted from runSubagent({agent:...})
  to task({ agent_type: "dotnet-test:code-testing-...", name:..., prompt:...})
- Step 9 validator dispatch (forces builder dispatch for cleanup)
- Steps 6/7 mandatory builder/tester dispatch wrapper

Implementer:
- Section 5: "you MUST dispatch fixer for build errors" + no-inline-edit
  block (forces fixer dispatch on build failures)
- Section 6: "you MUST dispatch fixer for test failures" + no-inline-edit
  block (forces fixer dispatch on test failures)
- Section 7: "Format Code (mandatory if a lint command exists)"
  (was "Optional"; mandatory firing of linter)
- Rule 6: never declare SUCCESS while build/tests fail (gates SUCCESS on
  fixer dispatch)
- Rule 7: no inline test-file edits between failed dispatch and fixer

Fixer:
- Frontmatter description widened to advertise handling of failing tests
  (without this, the orchestrator's routing logic does not select the
  fixer for test failures, so even Rule 7's mandate produces no firing
  -- this is the change that took fixer firing from 0.00/inst to 0.39/inst
  in earlier iterations)

DROPPED (content / quality rules -- in cta-prompt-tuning, NOT here):

Generator:
- Test-strength rules embedded in implementer dispatch prompt
- Test-design rules embedded in implementer dispatch prompt (OFAT,
  mutation self-check, never mock subject under test)
- File-location rules embedded in implementer dispatch prompt
- TARGET ENTITIES / PHASE CHECKLIST / TEST TRACEABILITY blocks in
  implementer dispatch prompt
- CHECKLIST format spec in planner dispatch prompt
- Step 9 validator's detailed cleanup classification

Implementer:
- Section 4b "Verify CHECKLIST coverage" pre-completion check
- Section 8 "CHECKLIST COVERAGE" report block
- "Honor the CHECKLIST" rule

Fixer:
- "Process -- Failing Tests" section (5-step diagnosis flow)
- All anti-weakening / anti-skipping rules
- "Re-derive expected from production source" guidance

Planner:
- CHECKLIST format ("one item per TARGET BEHAVIOR, Source/Variants/
  Expected mandatory")
- "Test name from research.md conventions" rule
- "At least 2 phases" rule

Researcher:
- Section 8 "Extract Local Test Naming & Style Conventions"
- TARGET ENTITIES / TARGET BEHAVIORS / TEST INFRASTRUCTURE structure
  in research.md
- Test naming pattern extraction

WHAT THE SUBAGENTS WILL ACTUALLY DO:

The researcher / planner / implementer / fixer all operate at baseline
behavior -- they receive the same prompts they receive in the upstream
"vanilla" runs. The only difference vs vanilla is that the orchestrator
ACTUALLY DISPATCHES THEM (where vanilla often inlines the work or skips
sub-agent dispatch entirely).

EXPECTED COMPARISON:

If quality on this branch is similar to or higher than dev/ykovalova/
cta-prompt-tuning (bd530be), then "more dispatches" is the dominant
quality lever and the content/quality rules in cta-prompt-tuning are
adding marginal or noise-level value.

If quality on this branch is materially lower than cta-prompt-tuning,
then the content/quality rules are doing the heavy lifting and the
dispatch mechanics alone are insufficient.

If quality on this branch matches or exceeds vanilla but trails
cta-prompt-tuning, then the dispatch mechanics provide a baseline lift
and the content rules add an incremental quality layer on top.

Rubber-duck check passed (validated dispatch-vs-content classification;
fixer frontmatter is routing metadata not a runtime gate; surviving
dispatch prompts contain no dangling references to removed CHECKLIST /
TARGET ENTITIES / TEST STRENGTH / naming-convention concepts).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* CTA: extract unit-under-test + behaviors quality cue (no caps)

* Address PR #646 review feedback

- Step 1b prompt: explicitly request unit-under-test (file:line) and behaviors
  so the verification gate rarely needs re-dispatch.
- Step 3: rename to 'Deep Research Phase', mark skipped for Direct strategy,
  and switch from overwriting research.md to extending it (no double-research).
- Step references: update '6-9' -> '6-10' and 'Step 9' -> 'Steps 9-10' in the
  strategy table and the All-strategies-MUST line, since reporting is Step 10.
- Step 6 builder prompt: drop '*.sln' glob (could expand to multiple args);
  use 'dotnet build --no-incremental' (auto-discovers .sln) per dotnet.md.
- Step 9: stop overloading the builder agent; perform diff/cleanup directly
  in the orchestrator (Rule 5 forbids inline build/test, not git/fs hygiene).
- Fixer agent: update mission text to cover failing tests and assertion
  correction (front-matter description already mentioned this; body now
  matches), with explicit no-Ignore/no-Skip/no-production-rewrite guardrails.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Yuliia Kovalova <ykovalova@example.com>
2026-05-13 14:03:31 +02:00
YuliiaKovalova 809d0b180c Improve test generation quality based on SWE Atlas benchmark analysis (#599)
* Improve test quality for benchmark performance

Based on SWE Atlas benchmark analysis (48.8% vanilla vs 19-34% CTA):

1. Default to Direct strategy for single-task requests — reduces
   multi-agent overhead that costs tokens without improving results

2. Run tests immediately in Direct strategy — catches assertion
   errors early instead of accumulating failures

3. Read source thoroughly before writing tests — trace actual logic
   and return values, not just function signatures

4. Quality over Quantity guidelines — cover stated requirements first,
   fewer focused tests beat many shallow ones

5. Verify tests are implementation-specific — tests that pass with
   an empty function body aren't testing anything useful

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Combine redundant implementation-specificity bullets into one

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-04-29 19:05:22 +00:00
Jan Krivanek 81946d2a38 Add proxy for CTA extension files (#583) 2026-04-24 12:19:17 +02:00
Jan Krivanek 4ace5459f8 Improve agent based on C++ testing feedback (#571)
* Improve agent based on C++ testing feedback

* Rename Cpp.md to cpp.md
2026-04-23 15:52:29 +02:00
Amaury Levé f8e8191268 Add concrete pipeline examples for code-testing-agent (#572)
Add end-to-end input/output examples to anchor expected behavior for the
LLM, addressing the Example Quality (2/5) review feedback.

Changes:
- New extensions/dotnet-examples.md with .NET-specific examples: sample
  source code, research output, plan output, generated test file, fix
  cycle walkthroughs, and final report
- SKILL.md: replace one-liner examples with strategy selection table,
  pipeline walkthrough, and pointer to language-specific extensions
- Generator agent: add strategy decision examples and sample final report
- Researcher/Planner/Implementer agents: add references to extension
  examples for concrete output shapes

Architecture keeps core agents language-agnostic — all language-specific
examples live in the extensions/ folder.
2026-04-23 12:18:36 +00:00
Jan Krivanek 05aeb657e6 Add license to agent files (#568) 2026-04-21 12:57:18 +00:00
Artur Spychaj 270ef910ec Ensure validation runs for all strategies including Direct (#570)
Harden the test generator agent to ensure final build validation, test
validation, and git commit always run — even when using the Direct
strategy for small requests.

Changes:
- Step 1: Mandate reading language-specific extensions (e.g., dotnet.md)
  before writing any code. The extension contains critical project
  registration (dotnet sln add) and build validation guidance.
- Step 2: Direct strategy now skips only Steps 3-5 (research, plan,
  implement sub-agents), NOT Steps 6-9. Added explicit note that
  Steps 6-9 are mandatory for ALL strategies.
- Step 6: Clarify full-solution build must use no --framework flag to
  catch multi-target framework issues (e.g., net472 errors).
- Step 7: Require fresh build for final test validation (no --no-build).
- New Step 9: Explicit git commit with verification that the diff is
  non-empty.
- New rules 9-11: Read extensions first, always validate+commit,
  preserve existing tests.

Motivated by MSBench investigation showing:
- Agent used Direct strategy and skipped final validation entirely
- Agent never read extensions/dotnet.md (missing dotnet sln add)
- Agent forgot to git commit (300 tests created, Changes +0 -0)
- Agent deleted pre-existing tests while adding new ones

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-04-21 14:27:32 +02:00
Amaury Levé 1125fe7864 improve(dotnet-test): enhance agent descriptions for subagent discovery (#560)
- Add 'Use when:' trigger phrases to all 7 sub-agent descriptions
  so parent agents can reliably discover and delegate to them
- Add .testagent/ cleanup rule to generator agent to prevent
  ephemeral pipeline state from being committed
2026-04-20 19:51:09 +02:00
Jan Krivanek e4670b33a1 [PoC] Code testing agent + tests PoC (#433)
* Add code testing agent agents + tests PoC

* Remove AssertionEvaluator.cs change (moved to dev/jankrivanek/agents-evals)
2026-03-30 17:57:08 +02:00