mirror of
https://github.com/dotnet/skills.git
synced 2026-09-20 09:49:54 +08:00
fix-dotnet-test-analysis-scripts
25 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
57733bebc8 |
Keep test agent state out of commits (#1108)
* Keep test agent state out of commits Move broad test-generation pipeline state to host scratch storage, worktree-specific Git metadata, or OS temp, and enforce the exclusion in evals. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify absolute test agent state path Use Git's explicit absolute path formatting in both test-generation entry points. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Prune Git metadata from test agent eval guards Avoid scanning nested repositories and align the remaining TESTAGENT_DIR placeholder with the documented format. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify test agent command handoff Require callers to provide exact commands, excerpts, or absolute TESTAGENT_DIR document paths to command-running sub-agents. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Reject all repository-local testagent entries Match .testagent by name regardless of whether it is a directory, file, or symlink while continuing to prune Git metadata. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Verify external test agent artifacts Restore broad-run artifact checks at the Git metadata path and pass the researched lint command and state directory to the linter agent. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Make testagent eval guards constant time Check only the forbidden workspace-root path, including broken symlinks, instead of recursively traversing dependency trees. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Broaden comprehensive test generation Treat explicit requirements as the floor for broad suites and add mutation-relevant equivalence-partition and invariant coverage without test-count padding. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Fix external artifact grader quoting Run state checks directly in the harness shell so TESTAGENT_DIR expands after assignment, with an isolated command probe covering valid and forbidden states. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Run broad skill evals in Git worktrees Initialize the seven broad evaluation roots as Git repositories so TESTAGENT_DIR resolves deterministically and external artifacts remain verifiable. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify non-stageable test agent state Describe the real invariant across the pipeline: state may live under .git metadata but must never be version-controlled workspace content or appear in git status. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Standardize intermediate test state contract Use one TESTAGENT_DIR placeholder, clearer intermediate-state terminology, and detect stageable research, plan, or status files regardless of directory name. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Use one Git root in workspace integrity eval Baseline the fixture from the evaluation root so stageable intermediate-state files remain visible to the directory-independent guard. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Prune Vitest dependencies from state scan Exclude node_modules through per-eval Git metadata so stageable state detection remains fast without modifying fixture content. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Strengthen focused intermediate-state guards Separate shell execution, reject Git-metadata files on focused runs, include ignored state files, and prune node_modules with a pathspec exclusion. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 --------- Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 |
||
|
|
5a06b20cc9 |
Support classic .NET test projects in dotnet-test (#993)
* Support classic .NET test projects Teach dotnet-test skills and agents to preserve non-SDK projects, packages.config dependencies, explicit compile registration, legacy runners, and version-compatible MSTest APIs. Add regression evals for generation, execution, coverage, and authoring. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68 * Address classic test fixture review Tighten the MSTest version grader, make the runner fixture assertion behavioral, and use nameof for the guarded parameter. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68 --------- Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68 |
||
|
|
71414ce000 |
Improve dotnet-test eval coverage and efficiency (#917)
* Improve dotnet-test eval coverage and efficiency Address remaining high-confidence items from #899 by bounding the code-testing pipeline and adding eval coverage for grade-tests and find-untested-sources. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e430fee9-d3df-4ef5-85a4-745ae4b17046 * Fix dotnet-test eval activation and quality Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c5c1a52-4f99-49d6-b503-1bec713a6e98 * Strengthen dotnet-test skill activation Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c5c1a52-4f99-49d6-b503-1bec713a6e98 |
||
|
|
4df4da469a |
Upgrade agentic workflows and fix stale PR cleanup (#916)
* chore: upgrade gh-aw runtime * fix: paginate stale pull request cleanup Upgrade the generated agentic workflow assets and ensure stale PR discovery includes every result page and draft pull requests. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 968a22c2-327f-4d26-8f86-1c59bcd323ea * Improve dotnet-test eval coverage and efficiency Address remaining high-confidence items from #899 by bounding the code-testing pipeline and adding eval coverage for grade-tests and find-untested-sources. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e430fee9-d3df-4ef5-85a4-745ae4b17046 * fix: address agentic workflow review Pin the Copilot setup checkout action and include the cutoff date in stale PR search results. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 968a22c2-327f-4d26-8f86-1c59bcd323ea * Improve test migration skill guidance Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 5e19d263-02a6-45b1-9cfb-424fa4d10863 * fix: add fixture namespace imports Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: a464e6e4-3e45-41fe-b17d-887c8cb8a448 * test: assert GetOrderById grade Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: a464e6e4-3e45-41fe-b17d-887c8cb8a448 --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
77154137e8 |
dotnet-test: make code-testing agent tools Claude Code-compatible + add cross-host portability check (#856)
* dotnet-test: make code-testing agent tools declarations Claude Code-compatible PR #847 added `tools: ["agent", "skill", "read", "search", "edit", "execute"]` to the code-testing-* agents to enable VS Code / Copilot CLI subagent fan-out. Those lowercase aliases map to real tools in VS Code and the Copilot CLI, but Claude Code matches `tools:` against its own vocabulary (Task, Skill, Read, Glob, Grep, Edit, Write, Bash). None of the aliases matched, so when these agents are loaded into Claude Code via --plugin-dir and selected with `claude --agent`, the agent was granted ZERO tools. A tool-less model asked to generate tests emits a textual <tool_call> block and exits after one turn, producing no file changes. Append the Claude Code tool names to each agent's `tools:` list so the same declaration works across all three runtimes (each honors the names it knows and ignores the foreign ones): - Orchestrators (generator, implementer): add Task, Skill, Read, Glob, Grep, Edit, Write, Bash (Task is the Claude Code equivalent of the `agent` fan-out tool). - Workers (researcher, planner, builder, tester, fixer, linter): add Skill, Read, Glob, Grep, Edit, Write, Bash. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> * skill-validator: complete built-in tools + add cross-host tool portability check Two related follow-ups to the agent tools fix: 1. Address the skill-check review feedback. The validator's BuiltInTools set was missing three legitimate host tool spellings that are not case-insensitive matches of existing entries, so they were flagged as non-built-in: - "write" — Claude Code file-creation tool (Copilot CLI / VS Code: "create") - "agent" — Copilot CLI / VS Code subagent fan-out tool (Claude Code: "task") - "execute" — Copilot CLI / VS Code run-command tool (Claude Code: "bash") "agent" and "execute" were already flagged before this branch (introduced by the fan-out PR); adding them to BuiltInTools clears the pre-existing warnings. 2. Add a cross-host tool portability check (CheckAgentToolPortability) so an agent that declares a capability for only one host is flagged. Tool names are matched case-sensitively (hosts resolve tools by exact spelling), so an agent that lists e.g. only "edit" (Copilot / VS Code) without "Edit"/"Write" (Claude Code) is reported as working on one host and silently tool-less on the other. Findings are advisory (do not fail CI) and allowlistable via "agent-tool-portability:AGENT:capability". Wired into the agents loop in CheckCommand and covered by unit tests. Also make the one existing single-host agent (optimizing-dotnet-performance) portable by adding its Claude Code tool spellings, so the new check reports a clean tree. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> |
||
|
|
06dd46e5f1 |
Enable VS Code subagent fan-out for the dotnet-test code-testing agents (#847)
* Add VS Code subagent metadata to dotnet-test code-testing agents
Enable the code-testing-* Research-Plan-Implement pipeline to fan out as
subagents in VS Code while keeping the GitHub Copilot CLI working.
Frontmatter (VS Code "coordinator/worker" pattern; portable tool aliases that
map in both VS Code and the CLI):
- code-testing-generator: tools: [agent, read, search, edit, execute] + an
agents: list (researcher, planner, implementer, builder, tester, fixer,
linter); softened the three runSubagent({ agent, prompt }) blocks to
tool-agnostic delegation prose.
- code-testing-implementer: tools: [agent, read, search, edit, execute] +
agents: (builder, tester, fixer, linter).
- Leaf agents (researcher/planner/builder/tester/fixer/linter):
tools: [read, search, edit, execute].
Why explicit tools (not ["*"]): VS Code has no all-tools wildcard for ools:
and a subagent's ools: overrides its inherited set, so ["*"] matched nothing
and stripped subagents of file tools (they ran without read/edit/search). The
CLI treats ["*"] as all-tools, so this was VS-Code-specific. The portable
aliases agent/read/search/edit/execute map to real tools in both environments;
agents: is ignored by the CLI.
README: document the VS Code chat.subagents.allowInvocationsFromSubagents
setting (off by default) needed for the nested implementer->builder/tester/
fixer/linter layer to fan out on large scopes; the CLI has no such gate.
Validated end to end: VS Code shows researcher->planner->implementer fanning out
with real file I/O (no "subagents lack file tools" warning); CLI fan-out intact
with skills still loading and tests passing. An all-tools baseline used only
tools within this enumerated set, confirming no CLI capability is restricted.
* Include skill tool in code-testing agent tool allowlists
Address PR review: the explicit `tools:` allowlists omitted the `skill`
tool, but every code-testing-* agent's prompt instructs calling skills
(e.g. `code-testing-extensions` for per-language guidance, `test-gap-analysis`,
`assertion-quality`). Because `tools:` is an override, omitting `skill` can
prevent the agents from loading those skills in environments that gate skill
invocation by the allowlist.
Add `skill` to all eight agents:
- orchestrators (generator, implementer): [agent, skill, read, search, edit, execute]
- workers (researcher/planner/builder/tester/fixer/linter): [skill, read, search, edit, execute]
Re-verified in the Copilot CLI: full fan-out (researcher -> planner ->
implementer/tester), the `code-testing-extensions` skill is invoked, and the
generated tests pass.
|
||
|
|
0c98ce6bd9 |
Direct strategy must still run the Step 7 pre-completion gate (#793)
* Direct strategy must still run the Step 7 pre-completion gate The Direct strategy correctly skips the research/plan/implement sub-agents for small single-file tasks, but the wording let agents also skip the Step 7 pre-completion gate (test-gap-analysis + assertion-quality + scenario coverage) — treating a single-file task that enumerates specific behaviors as 'trivially small'. This is the dominant failure mode observed on behavior-enumerating tasks: the agent writes one test file directly and finishes with no assertion-strength or scenario-coverage check, producing weak assertions (mutation survivors) and missing required edge/negative cases. Clarify in both the generator Step 2 strategy table and the code-testing-agent SKILL.md that Direct trades away only the sub-agents, never the gate, and that a request naming a specific symbol or enumerating scenarios is not 'trivially small' and must run the gate. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address review: align Direct gate trigger with Step 7 threshold; clarify gate in SKILL.md - Step 2 Direct cell no longer introduces a separate 'names a specific symbol' gate trigger that contradicted Step 7. It now defers to Step 7's own threshold (>=5 tests, or any enumerated behaviors/scenarios). - SKILL.md now names what/where the gate is: the generator's Step 7 (test-gap-analysis + assertion-quality). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Amaury Levé <evangelink@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
5d717dbdd1 |
Add prompt-scenario coverage check to code-testing-generator gate (#789)
* Add prompt-scenario coverage check to code-testing-generator gate The pre-completion gate already verifies assertion strength (pseudo-mutation and assertion-depth checks), but two recurring failure modes still slip through when the prompt enumerates specific behaviors: - Testing an *adjacent* function/helper instead of the exact feature named in the objective, leaving the requested behavior uncovered. - Covering only a single representative case when the scenario wording implies multiple variations or pins a condition to a specific position or structure. Add a third gate item that maps each enumerated scenario to a dedicated test, requires targeting the exact named function (preferring the canonical existing test file), and requires honoring range/positional qualifiers literally. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address review: genericize example, fix gate-count consistency - Remove benchmark-specific symbol names from the target-the-named-function bullet to avoid overfitting; phrase it generically. - Fix the gate intro that said 'The two skills below' now that there are three numbered items (the third is a prompt self-review, not a skill). - Update Step 8 and Rule 11 so re-running the gate includes the new prompt-scenario coverage check, not just test-gap-analysis + assertion-quality. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Amaury Levé <evangelink@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
f389af86c7 |
code-testing-generator: mandatory pre-completion self-review gate (#768)
* code-testing-generator: mandatory pre-completion self-review gate Replaces the prose `Verify tests are implementation-specific'' bullet in Step 7 of the code-testing-generator agent with a mandatory pre-completion self-review gate that invokes two existing plugin skills: 1. `test-gap-analysis'' (pseudo-mutation check) against the source files tested and the produced test files 2. `assertion-quality'' (trivial/tautological assertion check) against the produced test files Both skills already ship in plugins/dotnet-test; this PR only wires them into the generator's workflow as a mandatory gate before declaring a run complete. The two skills' `When to Use'' sections are extended to list `called by code-testing-generator as a pre-completion self-review step'' as a recognised use case so the model does not refuse the invocation. The skill descriptions are unchanged (frontmatter is already at the 1024-char limit). Rule 11 of the generator agent is updated to list the gate alongside final build, final test, and coverage-gap review as mandatory for ALL strategies including Direct. A matching rubric item is added to the ContosoUniversity scenario in the code-testing-agent eval (yaml and vally) so the LLM judge can verify the gate was actually invoked on the trajectory. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Apply suggestions from code review Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> |
||
|
|
754011b5ee |
Add workspace-integrity guardrail to code-testing agents (#773)
The code-testing-generator/implementer agents could treat an unusual or
scaffolded workspace (e.g. a gutted repo with an injected synthetic module)
as corruption and "repair" it with git checkout/restore/reset/clean or rm,
restoring deleted tracked files and testing the wrong code.
- Replace generator Rule 5 ("Clean git first - stash changes") with an
explicit "Treat the workspace as delivered" rule, and add a "Never mutate
version control" rule. Output must be purely additive test files.
- Add a no-revert/no-clean invariant to the implementer's edit boundaries.
- Add a 'workspace integrity' eval to the code-testing-agent suite
(eval.yaml + eval.vally.yaml). The fixture looks gutted: a metricsd project
whose real core/io modules are committed at HEAD but deleted from the
working tree, leaving only a synthetic 'synthstr' decoy. A git restore would
resurrect the deleted sentinel files; graders fail if they reappear and
require passing pytest tests for the module as delivered.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
||
|
|
49a77a90bf |
code-testing-agent: add polyglot pipeline examples for Python/TypeScript/Go/Java (#708)
* code-testing-agent: add polyglot pipeline examples for Python/TypeScript/Go/Java
The code-testing-agent skill family is polyglot in description but in
practice biased toward .NET because dotnet-examples.md was the only
filled-in pipeline walkthrough. The four sub-agents that participate in
the Research-Plan-Implement pipeline (researcher, planner, implementer,
generator) all pointed to dotnet-examples.md whenever they suggested a
concrete example, which made it harder for the agent to produce
idiomatic non-.NET tests (e.g. in msbench top5-* benchmarks for
Python/Flask, TypeScript/Express).
This change
* adds four new example files mirroring dotnet-examples.md format
(source → research → plan → generated test → fix cycle → final report):
- python-examples.md (pytest, unittest.mock, Mock(spec=...), parametrize)
- typescript-examples.md (Vitest with notes for Jest; it.each,
async tests, fake timers, ESM/CJS fix cycle)
- go-examples.md (standard testing package, table-driven subtests,
hand-written fake repository, injected clock)
- java-examples.md (JUnit 5 + Mockito on Maven, @ParameterizedTest +
@CsvSource, Clock.fixed, Surefire fix cycles)
* updates code-testing-extensions/SKILL.md TOC to list the new files
and clarifies usage instructions to read the matching <language>-
examples.md alongside the base extension
* makes the "Concrete example" pointers in code-testing-generator,
code-testing-implementer, code-testing-planner and
code-testing-researcher agents language-agnostic (list all available
example files instead of hard-coding dotnet-examples.md)
* expands code-testing-researcher project-structure detection list to
cover more Python (tox.ini, noxfile.py, requirements*.txt, uv.lock,
poetry.lock, pdm.lock), JS/TS (.mts/.cts/.jsx, vitest.config.*,
jest.config.*), C++ (CMakeLists.txt, BUILD.bazel, meson.build),
Java/Kotlin (pom.xml, build.gradle[.kts], wrappers), and other
ecosystem files; expands the Identify-Language section accordingly
* extends the "Language-Specific Examples" section in
code-testing-agent/SKILL.md to summarise each example file
Validated with: skill-validator check --plugin ./plugins/dotnet-test
(23 skills, 11 agents — all checks passed) and markdownlint-cli2.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Address review feedback for #708
- go-examples.md: replace the hand-rolled `contains`/`stringIndex`
helpers with `strings.Contains` from the standard library. The hand-rolled
`contains` had subtly wrong semantics — `contains(""abc"", """")` returned
`false` while `strings.Contains` returns `true` — and the helpers are
unnecessary complexity for a code-generation example.
- go-examples.md: in the `go test -run` "wrong selection regex" sample fix
cycle, quote the test name and use `single_item` (matching the underscore
that the surrounding diagnosis text refers to) instead of the unquoted
`single item` which the shell would parse as two separate CLI arguments.
- java-examples.md: the source-tree file list described `Invoice.java` as a
`record` but the `InvoiceService.markAsPaid` example mutates the invoice
via `setStatus(...)` and `setPaidDate(...)` — records are immutable, so
the description was internally inconsistent. Re-describe it as a mutable
POJO with explicit mutators to match the service code.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
||
|
|
2fcfc099b9 | Improve direct scenario (#675) | ||
|
|
dde9f99cf0 |
dotnet-test: drop tools field from agents that declared it (#665)
The four agents (code-testing-generator, test-migration, testability-migration, test-quality-auditor) were the only dotnet-test agents declaring a 'tools:' list. Per Jan's suggestion on #660, drop the field so they inherit the runtime's default tool surface (matching the convention used by every other agent in the repo). Follow-up to #660. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
2a193300ca |
Replace 'terminal' with 'bash' and 'powershell' in agent tools (#660)
The skill-validator only recognizes 'bash' and 'powershell' as built-in shell tools, not 'terminal'. Update the four affected agents in the dotnet-test plugin to use both for cross-platform support. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
3d59e44c7e |
Fix coverage-analysis activation for plateau diagnosis prompts (#647)
* Fix coverage-analysis activation for plateau diagnosis prompts coverage-analysis SKILL.md: - Trim verbose implementation details (provider detection, ReportGenerator) that consumed description budget without aiding skill activation - Add explicit USE FOR keywords: coverage stuck, coverage plateau, can't increase coverage, what's blocking coverage code-testing-agent SKILL.md: - Add 'diagnosing coverage plateaus or CRAP score computation (use coverage-analysis)' to DO NOT USE FOR boundary to prevent test-generation skill from intercepting diagnostic prompts * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Strengthen code-testing-agent activation; harden coverage-analysis isolated mode Address eval regressions reported on PR #647 (run 25813728646): 1. code-testing-agent: `Generate tests for ContosoUniversity ASP.NET Core MVC app` was NOT ACTIVATED in plugin mode (detectedSkills=[], skillEventCount=0, invokedAgents=[]). The model bypassed the skill system entirely. - SKILL.md description: restructure to use the proven `Use when user says ...` pattern with quoted trigger phrases (matching the run-tests skill that consistently activates), make the link to the code-testing-generator sub-agent explicit, and tighten DO NOT USE FOR clauses. - eval prompt (eval.yaml + eval.vally.yaml): make the request pipeline-shaped (`project-wide, multi-file test generation task`, `scaffold a new test project`) so the model recognizes it as multi-step work that benefits from the orchestrated pipeline. Explicitly request coverlet.collector + a Cobertura XML run so rubric criterion 1 (`high line coverage as reported by the Cobertura XML in TestResults/`) becomes achievable without overfitting. 2. code-testing-tester agent + code-testing-extensions/dotnet.md: open a scoped exception to the `skip coverage tools` rule. Default behavior stays the same, but when the user/harness explicitly asks for a Cobertura/XML coverage artifact, the agent may add coverlet.collector to the generated test csproj so the harness's coverage command produces output. The agent still does not run the coverage command itself. 3. coverage-analysis SKILL.md: add a `User-visible output is mandatory` guard at the top of the Workflow section. The latest eval showed isolated mode producing literally `(no output)` in 2 of 3 scenarios — the agent ran Compute-CrapScores.ps1 / Extract-MethodCoverage.ps1 / ReportGenerator in parallel, then the session ended without ever surfacing findings. The guard tells the agent to always return a partial summary instead of ending silent, and to deprioritize ReportGenerator HTML when budget is tight. (Plugin-mode quality is already strong: 4.3 / 4.3 / 5.0 — no regression risk there.) Aggregate dotnet-test plugin description size: 14,925 chars (limit 15,000). skill-validator check passes (22 skills, 11 agents, 1 plugin); markdownlint passes for all 4 modified files. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Restore MSTest modernization exclusion in code-testing-agent description * Fix isolated-mode coverage-analysis: emit summary before optional ReportGenerator The previous workflow encouraged the agent to run `dotnet tool install` for ReportGenerator in parallel with the CRAP scoring scripts (Phase 2 "Steps 3 and 4 in parallel" + Phase 3 "Steps 5 and 6 in parallel"). In isolated mode that pattern reliably crashed the session with "Failed to persist session events: timeout while waiting for mutex to become available" right after the scripts returned valid data, so the agent never produced the user-facing summary. Restructure the workflow into 5 phases: - Phase 1 (Setup) - unchanged - Phase 2 (Test execution) - skip when Cobertura XML already exists - Phase 3 (Analysis) - run only the two PowerShell scripts, no RG - Phase 4 (User-facing summary) - MANDATORY, must be the next assistant response after Phase 3, before any RG work; also save coverage-analysis.md as a secondary follow-up - Phase 5 (ReportGenerator HTML/CSV) - strictly optional, post-summary, skipped by default for existing-Cobertura and plateau-diagnosis paths Also update references/output-format.md so the Reports section marks RG artifacts as "Not generated (optional - request HTML reports to enable)" when Phase 5 has not run, and update references/guidelines.md so the "show and open the markdown report" rule explicitly defers to the user-facing assistant response. Targets the isolated-mode regressions in PR #647 eval: - Project-wide coverage with existing Cobertura: 1.0/5 -> expected 3+ - Coverage plateau diagnosis: 1.0/5 -> expected 3+ - Run coverage from scratch: 2.3/5 -> expected steady or up Verified: skill-validator check --plugin ./plugins/dotnet-test passes; markdownlint-cli2 clean on all 3 modified files. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR #647 review comments 1. Prerequisites: distinguish the from-scratch path (needs NuGet for the coverage-provider package install + optional internet for ReportGenerator) from the existing-Cobertura path (needs neither). 2. Add Step 2c "Discover or accept existing Cobertura XML" so the existing-data path actually has $coberturaFiles populated before Phase 3, instead of relying on Phase 2's discovery (which it skips). Also clarify Step 2's destructive Remove-Item only manages the skill-owned coverage-analysis/ subdirectory. 3. references/output-format.md: replace the unconditional "Reports saved to: <coverageDir>/reports/" line with one that always points at <coverageDir>/ (markdown summary + raw Cobertura) and only mentions reports/ if Phase 5 ran. 4. Have Compute-CrapScores.ps1 emit OVERALL_LINE_COVERAGE and OVERALL_BRANCH_COVERAGE from the Cobertura root attributes, and update Phase 4 to read those values directly from the script's output. The Phase 4 mandatory-summary rule no longer requires a separate XML parse before composing the response. Verified: skill-validator check --plugin ./plugins/dotnet-test passes; markdownlint-cli2 clean on all modified files; Compute-CrapScores.ps1 smoke-tested on a synthetic Cobertura XML (emits OVERALL_LINE_COVERAGE:75, OVERALL_BRANCH_COVERAGE:50 alongside HOTSPOTS). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Apply suggestion from @Evangelink * Address unresolved coverage-analysis and eval review comments * Refine follow-up review feedback from validation * Tighten coverage aggregation fallback notes and counters * Clarify pre-response save instruction wording --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> |
||
|
|
e1e8568375 |
Revert "dotnet-test: add unit-under-test + behaviors quality cue to code-test…" (#651)
This reverts commit
|
||
|
|
f1b09eba79 |
dotnet-test: add unit-under-test + behaviors quality cue to code-testing-generator (#646)
* CTA: invocation-only baseline (dispatch mechanics, no content rules)
Experiment branch to isolate the impact of "invoke prompted subagents
more often" from the impact of "make those subagents do richer work."
Same baseline as dev/ykovalova/cta-prompt-tuning (main =
|
||
|
|
809d0b180c |
Improve test generation quality based on SWE Atlas benchmark analysis (#599)
* Improve test quality for benchmark performance Based on SWE Atlas benchmark analysis (48.8% vanilla vs 19-34% CTA): 1. Default to Direct strategy for single-task requests — reduces multi-agent overhead that costs tokens without improving results 2. Run tests immediately in Direct strategy — catches assertion errors early instead of accumulating failures 3. Read source thoroughly before writing tests — trace actual logic and return values, not just function signatures 4. Quality over Quantity guidelines — cover stated requirements first, fewer focused tests beat many shallow ones 5. Verify tests are implementation-specific — tests that pass with an empty function body aren't testing anything useful Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Combine redundant implementation-specificity bullets into one Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
81946d2a38 | Add proxy for CTA extension files (#583) | ||
|
|
4ace5459f8 |
Improve agent based on C++ testing feedback (#571)
* Improve agent based on C++ testing feedback * Rename Cpp.md to cpp.md |
||
|
|
f8e8191268 |
Add concrete pipeline examples for code-testing-agent (#572)
Add end-to-end input/output examples to anchor expected behavior for the LLM, addressing the Example Quality (2/5) review feedback. Changes: - New extensions/dotnet-examples.md with .NET-specific examples: sample source code, research output, plan output, generated test file, fix cycle walkthroughs, and final report - SKILL.md: replace one-liner examples with strategy selection table, pipeline walkthrough, and pointer to language-specific extensions - Generator agent: add strategy decision examples and sample final report - Researcher/Planner/Implementer agents: add references to extension examples for concrete output shapes Architecture keeps core agents language-agnostic — all language-specific examples live in the extensions/ folder. |
||
|
|
05aeb657e6 | Add license to agent files (#568) | ||
|
|
270ef910ec |
Ensure validation runs for all strategies including Direct (#570)
Harden the test generator agent to ensure final build validation, test validation, and git commit always run — even when using the Direct strategy for small requests. Changes: - Step 1: Mandate reading language-specific extensions (e.g., dotnet.md) before writing any code. The extension contains critical project registration (dotnet sln add) and build validation guidance. - Step 2: Direct strategy now skips only Steps 3-5 (research, plan, implement sub-agents), NOT Steps 6-9. Added explicit note that Steps 6-9 are mandatory for ALL strategies. - Step 6: Clarify full-solution build must use no --framework flag to catch multi-target framework issues (e.g., net472 errors). - Step 7: Require fresh build for final test validation (no --no-build). - New Step 9: Explicit git commit with verification that the diff is non-empty. - New rules 9-11: Read extensions first, always validate+commit, preserve existing tests. Motivated by MSBench investigation showing: - Agent used Direct strategy and skipped final validation entirely - Agent never read extensions/dotnet.md (missing dotnet sln add) - Agent forgot to git commit (300 tests created, Changes +0 -0) - Agent deleted pre-existing tests while adding new ones Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
1125fe7864 |
improve(dotnet-test): enhance agent descriptions for subagent discovery (#560)
- Add 'Use when:' trigger phrases to all 7 sub-agent descriptions so parent agents can reliably discover and delegate to them - Add .testagent/ cleanup rule to generator agent to prevent ephemeral pipeline state from being committed |
||
|
|
e4670b33a1 |
[PoC] Code testing agent + tests PoC (#433)
* Add code testing agent agents + tests PoC * Remove AssertionEvaluator.cs change (moved to dev/jankrivanek/agents-evals) |