mirror of
https://github.com/dotnet/skills.git
synced 2026-09-20 09:49:54 +08:00
fix-dotnet-test-analysis-scripts
27 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9beca0b0be |
Bump vitest (#1145)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.8 to 4.1.11. - [Release notes](https://github.com/vitest-dev/vitest/releases) - [Changelog](https://github.com/vitest-dev/vitest/blob/main/docs/releases.md) - [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.11/packages/vitest) --- updated-dependencies: - dependency-name: vitest dependency-version: 4.1.11 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
57733bebc8 |
Keep test agent state out of commits (#1108)
* Keep test agent state out of commits Move broad test-generation pipeline state to host scratch storage, worktree-specific Git metadata, or OS temp, and enforce the exclusion in evals. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify absolute test agent state path Use Git's explicit absolute path formatting in both test-generation entry points. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Prune Git metadata from test agent eval guards Avoid scanning nested repositories and align the remaining TESTAGENT_DIR placeholder with the documented format. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify test agent command handoff Require callers to provide exact commands, excerpts, or absolute TESTAGENT_DIR document paths to command-running sub-agents. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Reject all repository-local testagent entries Match .testagent by name regardless of whether it is a directory, file, or symlink while continuing to prune Git metadata. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Verify external test agent artifacts Restore broad-run artifact checks at the Git metadata path and pass the researched lint command and state directory to the linter agent. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Make testagent eval guards constant time Check only the forbidden workspace-root path, including broken symlinks, instead of recursively traversing dependency trees. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Broaden comprehensive test generation Treat explicit requirements as the floor for broad suites and add mutation-relevant equivalence-partition and invariant coverage without test-count padding. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Fix external artifact grader quoting Run state checks directly in the harness shell so TESTAGENT_DIR expands after assignment, with an isolated command probe covering valid and forbidden states. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Run broad skill evals in Git worktrees Initialize the seven broad evaluation roots as Git repositories so TESTAGENT_DIR resolves deterministically and external artifacts remain verifiable. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify non-stageable test agent state Describe the real invariant across the pipeline: state may live under .git metadata but must never be version-controlled workspace content or appear in git status. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Standardize intermediate test state contract Use one TESTAGENT_DIR placeholder, clearer intermediate-state terminology, and detect stageable research, plan, or status files regardless of directory name. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Use one Git root in workspace integrity eval Baseline the fixture from the evaluation root so stageable intermediate-state files remain visible to the directory-independent guard. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Prune Vitest dependencies from state scan Exclude node_modules through per-eval Git metadata so stageable state detection remains fast without modifying fixture content. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Strengthen focused intermediate-state guards Separate shell execution, reject Git-metadata files on focused runs, include ignored state files, and prune node_modules with a pathspec exclusion. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 --------- Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 |
||
|
|
94ca0ca748 |
Improve evaluation freshness and scheduled reliability (#1081)
* Improve evaluation freshness and reliability Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> * Address dashboard freshness review Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> * Preserve evidence commit fallback Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 5283cdb3-86e3-41d8-95a9-ebf6b7e0ccbc * Keep watchdog regression assertion current Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 5283cdb3-86e3-41d8-95a9-ebf6b7e0ccbc * Add headroom for MSTest migration evaluation Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 5283cdb3-86e3-41d8-95a9-ebf6b7e0ccbc --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 5283cdb3-86e3-41d8-95a9-ebf6b7e0ccbc |
||
|
|
d3921f7418 |
Strengthen testability skill evaluations (#1057)
* Strengthen testability skill evaluations Raise four dotnet-test evals to eight independent stimuli, add validated fixtures, and resolve code-testing-agent orphan fixtures without speculative routing changes. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> * Relax promo-code eval grader Accept deterministic suffix values beyond one hard-coded literal and match common PascalCase test names. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 * Improve testability skill reliability Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 * Refine testability obstacle graders Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 * Accept qualified Random seams Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 * Improve ambient seam compatibility Replace the C# 12 primary constructor in the copyable Scope sample with a conventional constructor so the guidance works in projects using older language versions. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 * Improve testability skill reliability Refine routing and execution contracts from exact losing transcripts, strengthen behavioral eval checks, and add isolated C# fixtures without increasing repeated runs. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 |
||
|
|
5a06b20cc9 |
Support classic .NET test projects in dotnet-test (#993)
* Support classic .NET test projects Teach dotnet-test skills and agents to preserve non-SDK projects, packages.config dependencies, explicit compile registration, legacy runners, and version-compatible MSTest APIs. Add regression evals for generation, execution, coverage, and authoring. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68 * Address classic test fixture review Tighten the MSTest version grader, make the runner fixture assertion behavioral, and use nameof for the guarded parameter. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68 --------- Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68 |
||
|
|
aa3c4abca5 |
Bump postcss (#988)
Bumps [postcss](https://github.com/postcss/postcss) from 8.5.15 to 8.5.25. - [Release notes](https://github.com/postcss/postcss/releases) - [Changelog](https://github.com/postcss/postcss/blob/main/CHANGELOG.md) - [Commits](https://github.com/postcss/postcss/compare/8.5.15...8.5.25) --- updated-dependencies: - dependency-name: postcss dependency-version: 8.5.25 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
573938df09 |
Close the remaining dotnet-test eval follow-ups (#899) (#971)
* Close the remaining dotnet-test eval follow-ups from #899 Every non-agent dotnet-test eval now clears the 5-trial floor, the two cost P1s are addressed, and the reference-skill coverage gap is recorded as a decision instead of a standing warning. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 024b3241-d7af-418a-b9cd-3bb9ee9bf0ee * Remove the duplicate block that cloned a grade-tests scenario, and gate it The 'production code available' scenario carried a leftover tail from the edit that moved the 'production code unavailable' one. YAML keeps the last duplicate key, so every field of the new scenario was silently overwritten: it shipped as a byte-identical rerun of its predecessor and never loaded the dotnet-production-available fixture it was built around. Parsing the spec and counting stimuli - which is what verified this PR - returns the intended 5 names either way, so only the parser can see it. The eval-quality gate now loads specs with a duplicate-key-strict loader (failing check 9), with a self-test case and the incident recorded. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 024b3241-d7af-418a-b9cd-3bb9ee9bf0ee --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 024b3241-d7af-418a-b9cd-3bb9ee9bf0ee |
||
|
|
030493de5a |
Fix migrate-mstest-v1v2-to-v3 activation and test-skill eval quality (#974)
* Fix migrate-mstest-v1v2-to-v3 skill activation The frontmatter description said DO NOT USE FOR: ... projects already on MSTest v3+, which blocked the skill on every scenario where the packages had already been bumped to 3.x and only the source or settings still needed the v1/v2-to-v3 fixes (Assert object overloads, DataRow strict typing, .testsettings -> .runsettings). It also gated the whole skill behind "the user asks to upgrade MSTest", so a standalone .testsettings conversion never matched. - Rewrite the description around both entry points (pre-upgrade migration and post-upgrade breaking-change fixes) and add the concrete trigger keywords those prompts contain: CS1501/CS1503/CS0121, MSTEST0014, LegacySettings, DeploymentEnabled, per-test TestTimeout, net5.0. Note that the current runner is preserved so "migrate to v3 but keep VSTest" isn't poached by migrate-vstest-to-mtp. - Narrow migrate-mstest-v3-to-v4, which claimed the generic "tests don't compile after upgrading MSTest" phrasing and competed for the same prompts. - Widen the Boundary Gate: a 3.x package version alone no longer ends the migration when v1/v2-era settings or errors remain, so the skill actually performs the requested edits instead of reporting "already migrated". - Add a routing row to the test-migration agent for the same case. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: ad6ff32a-d441-4a7b-b474-2bfaee764740 * Clear the eval-quality gate's test-skill findings The gate reported three classes of debt against the dotnet-test and dotnet-test-migration plugins. All three are addressed here; the four ERRORS it also reports are dotnet-maui allowlist lines and are untouched. Underpowered evals (5). Below five trials the pass gate's sign test cannot reach p <= 0.05 at any effect size, so these five evals could never return a verdict. Each is now at or above the floor and its allowlist line is deleted in the same change, as the ledger's shrink-only rule requires: - coverage-analysis 3 -> 5: adds a refactoring-safety question (the "is this safe to change?" use case named in the skill's Purpose but never exercised) and a branch-vs-line coverage question. Both reuse the existing partial-coverage fixture. - find-untested-sources 4 -> 5: adds a mixed C#/TypeScript repository, which is the only case that exercises the documented engine choice - polyglot tree-sitter rather than the C#-only Roslyn engine. Composed from the two existing fixtures. - generate-testability-wrappers 4 -> 5: adds the ambient-context path (Step 5) for a project with no DI container, where AsyncLocal<T> and scoped disposal are the distinguishing content. - grade-tests 4 -> 5: adds a C# case with the production code present. Every prior C# scenario hides it, so "Unverified" was never tested as a negative, and the D band and the swallowed-exception F were never graded at all. New production-available fixture. - code-testing-agent 3 -> 6 via defaults.runs=2. Scenarios are preferred over runs, but each of these drives a full generate-build-test pipeline (npm ci plus two Vitest runs, pip install plus pytest, a dotnet test build) under a 60m budget, which is the documented case for buying trials with runs. Orphaned fixtures (5). v3-sealed-timeout, mtp-mstest-sdk9, mtp-mstest-sdk10, mtp-mstest-hotreload-installed and vstest-mstest are all superseded first- generation copies: their per-scenario successors differ only in whitespace, a dropped rollForward, or a package version. Both evals are already well above the floor, so wiring them up would add no power. Deleted. Skills with no eval (2 of 4). platform-detection and filter-syntax carry real checkable rules that nothing measured, and several are counterintuitive enough that a baseline is likely to get them wrong - global.json test.runner outranking TestingPlatformDotnetTestSupport on .NET 10+, Microsoft.NET.Test.Sdk not being a VSTest signal, MTP properties living in Directory.Build.props, xUnit v3 dropping VSTest --filter while MSTest on MTP keeps it. Both get a 5-scenario eval with small fixtures and no build step. code-testing-extensions and test-analysis-extensions are left flagged on purpose: their bodies are tables of paths to extension files, so a head-to-head eval would score path recall rather than user value. The content those files hold is already exercised through code-testing-agent's three-language pipeline. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: ad6ff32a-d441-4a7b-b474-2bfaee764740 * Fix the two skill defects behind the v1v2-to-v3 eval losses The first eval run reached 7W/2T/2L, p=0.090, short of the p<=0.05 gate. Both losses trace to skill content that actively misled the agent, and the session transcripts show exactly how. Loss 1 -- 'Migrate MSTest v1 project with assembly reference', skilled scored 0.00 against a 4.17 baseline. The transcript shows the skill loading correctly and the agent then replying, in full: 'To give you specific migration steps, I need to see your project file. Could you share the path to your .csproj?' The project was already in the working directory. Cause: the Inputs table marked 'Project or solution path' as Required=Yes, which reads as a precondition the agent must obtain before doing anything. This is the worst kind of failure for a real user - they describe their project in prose and get a question back instead of an answer. Path is now optional and discovered by globbing, Step 1 leads with locating the project, and a note forbids opening with a request for the path. The same Required=Yes trap was present in migrate-mstest-v3-to-v4 and migrate-vstest-to-mtp, so both are corrected too. Loss 2 -- 'Fix DataRow type mismatch errors', skilled 3.96 against a 5.00 baseline. The skill's breaking-change table said the 16-argument DataRow cap was 'fixed in later v3 versions' and suggested 'refactor test / wrap extra params in array'. On a project already at MSTest 3.8, the agent concluded the valid 17-argument row exceeded the limit and rewrote it - first as new object[] { 17 }, which failed, then second-guessing itself mid-run ('let me check if the latest 3.x actually fixed the 16-arg limit'), finally settling on a (object)17 cast. Churn plus wasted turns on code that was already correct. The vague wording was the problem, so it is replaced with the fact: the cap was introduced in 3.0.1 and removed again in 3.0.3 (microsoft/testfx#1554 and the maintainer's 'please feel free to update to 3.0.3'). On 3.0.3+ a longer row is valid and must be left alone. A general guideline is added alongside it - confirm the diagnostic before editing, because rewriting valid code to dodge a limit the project is not subject to is a defect rather than caution. Both fixes are about what the skill tells a real user, not about the graders; no eval prompt, fixture, or grader is touched. Skill grows ~480 tokens and stays in the 'standard' tier, below the 5,000-token warning threshold. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: ad6ff32a-d441-4a7b-b474-2bfaee764740 * Correct the Assert/DataRow facts and make the fixtures reproduce them The two eval runs on this PR compared byte-identical skill content: run 1 ( |
||
|
|
f2eb897a12 |
Fix dotnet-test findings from the refreshed cross-family eval (#899) (#945)
* Fix dotnet-test findings from the refreshed cross-family eval (#899) Every change below is driven by judge evidence from the losing trials of the refreshed 5-family dotnet-test matrix (runs 30108473397 + recovery runs), not by style preference. Eval measurement fix — the "discovery" P2s were an artifact: - assertion-quality, test-gap-analysis, test-smell-detection, and test-tagging each have a decline stimulus with `constraints.reject_skills: ["*"]`, so the skill cannot activate there by construction. Without `expect_activation: false` the adapter counted those dormant runs as missed activations, which is exactly the 75-88% invocation rates reported in the scorecard. Annotating them (the convention already used by agent.test-quality-auditor) removes the false signal; the non-activations were the only ones observed for these skills. Skill fixes: - test-gap-analysis: baselines won by actually running the suite while the skill reasoned statically and reported survivors that the tests in fact kill. Added Step 4b: confirm every reported survivor by applying it, re-running the covering tests, and reverting; fall back to reasoning only when the suite cannot run, labelled unverified. Calibrated severity down for strong suites. - test-anti-patterns: baselines won on depth, not polish. Added a depth bar — account for every test in scope, give exact expected values in fixes, name the adjacent error-path/boundary gaps, and keep counts consistent. Trimmed three pitfall rows that duplicated the calibration step so the skill stays under the profiler's "comprehensive" threshold. - detect-static-dependencies: losses were all counting accuracy. One authoritative total (no findings parked outside it), classify by the resource touched rather than by the `static` keyword, exclude pure helpers such as Path.Combine from the needs-wrapping total, require file:line, and add the missing randomness/culture/serialization categories. - test-smell-detection: the calibration rule told models to downgrade Sleepy Test for integration tests, which is what lost both losing scenarios. Fixed sleeps now stay High in any category; Mystery Guest and Eager Test still downgrade. - crap-score: losses came from estimating coverage after collection failed. Added the dotnet-coverage/ReportGenerator recovery path and a hard rule never to publish a CRAP score built on assumed coverage. - coverage-analysis: answer the asked question first, reconcile every number against the script output, and list every below-threshold member instead of declaring one method the entire gap. - migrate-static-to-wrapper: migrate exactly what was requested (no adjacent DateTime.Now rewrites, respect intentional-use comments) and never report "build succeeded" when the build or restore failed. - code-testing-agent: quote each requirement verbatim in the evidence table so multi-condition requirements map to a test that covers the whole combination, and cite a clean run rather than a coverage attempt that exited non-zero. Validation: skill-validator check passes (20 skills, 10 agents); markdownlint clean; eval specs parse and the adapter now reports all four decline stimuli as expect-dormant. Refs #899 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1947263a-0ef9-47bd-ac53-5af5afa3ddaa * Strengthen underpowered dotnet-test evals from the PR 945 eval run The PR eval reported 5 of 10 skills as "no credible improvement". Reproducing the gate arithmetic from the artifacts shows the dominant cause is statistical power, not skill quality. The gate is `mean > 0 AND ci_low > 0` with a t-based CI over per-trial scores, which reduces to `sqrt(n) * (mean/sd) > t(n-1)`. The required mean/sd ratio is brutal at small n: n=2 -> 8.98 n=3 -> 2.48 n=4 -> 1.59 n=5 -> 1.24 n=6 -> 1.05 n=8 -> 0.84 Recomputing each reported CI from the per-trial scores reproduces the published numbers exactly, which confirms the mechanism: crap-score [0.4,1.0,0.4] n=3 CI [-0.261, 1.461] migrate-static-to-wrapper [1.0,0.4,0.4] n=3 CI [-0.261, 1.461] test-gap-analysis [0.4,0,0.4,0.4] n=4 CI [-0.018, 0.618] test-anti-patterns [0.4,0,0.4,0,0,0.4] n=6 CI [-0.030, 0.430] code-testing-agent [0,0.4] n=2 CI [-2.341, 2.741] crap-score and migrate-static-to-wrapper won 100% of their trials (3W/0T/0L) and still failed: at n=3 nothing short of three identically-sized wins can clear the gate. That is a property of a thin eval, not of the skill. Scenario counts are raised with discriminating cases, four of them by wiring up fixtures that were already committed but had no stimulus referencing them: - test-gap-analysis 4 -> 6, using the orphaned `report-quality` fixture (trivial auto-properties and an auto-generated .g.cs to skip, private helpers reachable only through the public API, and a deliberately weak Assert.IsTrue that cannot kill arithmetic mutations) and the orphaned `rust-error-propagation` fixture (an untested `?` propagation path and an untested `<=` boundary). - test-anti-patterns 6 -> 8, using the orphaned `pytest-mixed` fixture (which also checks the calibration rule that pytest's bare `assert` must not be flagged) and the orphaned `assertion-problems` fixture (which separates Critical false-confidence assertions from a Low-severity message nit). - crap-score 3 -> 6, with a new `refactor-required` fixture whose numbers are self-consistent: ApplySurcharges has complexity 13 behind a stale "Complexity: 4" comment (CRAP 28.4, needs 77.2% coverage), ClassifyAccount has complexity 17 so coverage alone can never reach CRAP < 15, and RoundToCurrency is 100% covered so its CRAP equals its complexity exactly. - migrate-static-to-wrapper 3 -> 5, adding a DateTimeKind-preservation scenario over the existing fixture and a new `static-helper` fixture where a static class must gain an ambient TimeProvider seam without breaking its callers. - code-testing-agent 2 -> 3, with a compact C# fixture that must extend an existing suite to the untested method only. This eval stays the weakest: each scenario is expensive, so raising `runs` is a better lever than adding more heavyweight scenarios. Verification: - every eval spec parses and all 254 fixture references resolve - the three new fixtures compile; the shipping-quotes fixture restores, builds and its three seed tests pass under `dotnet test` in a clean workspace - skill-validator check passes (20 skills, 10 agents) - markdownlint clean Refs #899 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1947263a-0ef9-47bd-ac53-5af5afa3ddaa * Address review feedback on fixture and counting wording - BillableWeightTests: the ZeroOrNegative test only asserted the zero case, so its name overstated what it covered. Made it data-driven over 0 and -1 so the name matches the assertions. This matters more than usual here: the file is the seed suite for a test-quality eval, and a misleading test name is exactly what these skills are supposed to flag. - detect-static-dependencies: the Step 3 lead-in said to count each "static call pattern", which contradicted the rule immediately below it that instance members reaching the same untestable resource must also be counted. Reworded to "call site" and made the instance-member inclusion explicit. Verified: the fixture restores, builds and now passes 4 tests (was 3); skill-validator check passes; markdownlint clean. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1947263a-0ef9-47bd-ac53-5af5afa3ddaa --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1947263a-0ef9-47bd-ac53-5af5afa3ddaa |
||
|
|
71414ce000 |
Improve dotnet-test eval coverage and efficiency (#917)
* Improve dotnet-test eval coverage and efficiency Address remaining high-confidence items from #899 by bounding the code-testing pipeline and adding eval coverage for grade-tests and find-untested-sources. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e430fee9-d3df-4ef5-85a4-745ae4b17046 * Fix dotnet-test eval activation and quality Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c5c1a52-4f99-49d6-b503-1bec713a6e98 * Strengthen dotnet-test skill activation Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c5c1a52-4f99-49d6-b503-1bec713a6e98 |
||
|
|
78054c1161 |
Migrate LLM evals to the Vally harness (#877)
* Make Vally the sole LLM eval engine, retiring skill-validator evaluate Collapses the parallel skill-validator + Vally eval pipeline into a single Vally-only path while preserving every capability skill-validator provided: PR gate, PR comment, dotnet/skills-data historical push, and the dashboard. The skill-validator `check` linter is retained (skill-check.yml) pending microsoft/vally #463. - evaluation.yml: delete build-validator and the skill-validator `evaluate` job; single `evaluate:` job now uses vally-evaluation.yml. Forward the COPILOT_PAT_0..9 pool onto it (from #911) so the reusable workflow's token gate is satisfied. Rewire comment-on-pr / report-status / publish-* / deploy-dashboard onto `evaluate` and vally-results-* artifacts. - vally-evaluation.yml: pin @microsoft/vally-cli@0.9.0; keep the #911 prepare guard; write the generated per-plugin experiment to $GITHUB_WORKSPACE so relative eval globs resolve. - adapt.mjs: reconcile #887's runCompareWithRetry + conclusive/unmatched hardening with the consolidation helpers (nonActivation, compareByStim, roleToDashboard, etc.). Add consolidate.mjs and gen-experiment.mjs. - build-replay-sessions.ps1: read Vally executor-session-logs/events.jsonl instead of skill-validator session output. - Rename all tests/*/*/eval.vally.yaml to eval.yaml (Vally format is now the only eval format); author Vally evals for the four skills that lacked one. - CONTRIBUTING.md + InvestigatingResults: point contributors at the Vally harness; drop non-public links. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix setup-local-sdk eval for Vally schema The hand-authored setup-local-sdk eval used skill-validator grader/environment schema that Vally rejects at validation (no stimuli executed): - file-contains/file-not-contains graders require 'value:' not 'substring:' (output-contains correctly keeps 'substring:'). - environment.files entries require 'src:'/'dest:' (copy a fixture) rather than inline 'path:'/'content:'. Moved the global.json body into fixtures/global.json. Verified with 'vally lint --eval-spec' (0 errors) across all 97 eval.yaml. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix markdownlint violations in migration docs Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR review feedback on eval migration - Consolidate run-vally-evals.sh into a single neutral entrypoint eng/run-skill-evals.sh (deletes the thin wrapper; adapt.mjs stays in eng/vally-adapter/ since CI invokes it directly) - Rename reusable workflow vally-evaluation.yml -> evaluation-run.yml and update all references/self-trigger regexes - Rename local output dir vally-results -> eval-results - Bake prerequisite preflight (Node 20+, GITHUB_TOKEN/gh) into the runner - Simplify CONTRIBUTING "Running tests locally" prerequisites - Scrub harness name from consumer-facing surfaces Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Reinstate overfitting detection under the Vally harness Overfitting detection previously ran inside skill-validator's `evaluate` command, which the Vally migration replaced, so it silently stopped running. Reinstate it as a standalone step wired into the Vally pipeline: - Add `skill-validator overfitting` command that discovers skills, parses each eval.yaml, runs the existing OverfittingJudge, and emits [{plugin, skill, overfittingResult}] JSON (per-skill failures non-fatal, bounded parallelism). - Add a Vally-format parser bridge (ParseEvalConfigFlexible) so the judge handles the current stimuli/graders eval.yaml format; the legacy scenarios-only parser rejected all 98 evals, so the judge never ran. - adapt.mjs: merge overfitting results onto each verdict via --overfitting <file> (keyed by plugin/skill); output is byte-identical when the flag is absent. - evaluation-run.yml: build skill-validator (shared cache with check), run the judge per leg reading GITHUB_TOKEN, and pass --overfitting to adapt.mjs so the dashboard/skills-data consume it unchanged. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Reconcile #919 Vally eval configs into the sole eval.yaml convention PR #919 landed on main adding fresh Vally specs for three coverage-gap skills (setup-local-sdk, convert-blazor-server-to-webapp, dotnet-webapi) under the interim eval.vally.yaml filename. This branch collapses to a single Vally spec named eval.yaml, so the merge left each skill with both my migrated eval.yaml and #919's newer eval.vally.yaml (a git-clean but semantic duplicate the pipeline would not discover). Adopt #919's reviewed Vally content as the canonical eval.yaml for all three skills and drop the redundant eval.vally.yaml files, preserving the 'sole eval.yaml, zero eval.vally.yaml' invariant. Content is byte-identical to #919's blobs ( |
||
|
|
52ba152ed6 | Update to Vally 0.7.0 (#854) | ||
|
|
33110eeba8 |
code-testing-agent: fix workspace-integrity activation + stabilize Contoso rubric (#806)
* code-testing-agent: fix workspace-integrity activation + stabilize Contoso rubric
Workspace-integrity scenario was NOT ACTIVATED in plugin mode: the prompt
("Generate unit tests for its core module") was terse and small-scoped, so
the runtime did not route to the code-testing-agent skill. Reframe it with
high-level test-generation language ("comprehensive pytest test suite",
"scaffold", "thorough unit tests") that matches the skill description,
while preserving the guardrail anchors: it still points at the on-disk module
without naming it and never implies restoring the gutted tree.
ContosoUniversity rubric item #5 (find-untested-sources) was conditional on
that skill being loaded — only true in plugin mode. In isolated runs the
agent cannot satisfy it, so the judge penalized it asymmetrically, injecting
isolated-vs-plugin variance. Make the conditional deterministic: explicitly
N/A when the skill is not loaded, without lowering the bar when it is.
Applied to both eval.yaml and eval.vally.yaml.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* code-testing-agent: make find-untested-sources rubric decidable from session timeline
Address review feedback: the 'treat as satisfied (N/A) when the skill is
not loaded' clause is not verifiable from the judge's inputs (which show
which tools were called, not which were available). Rewrite the criterion
to be decidable from the session timeline by requiring a source-to-test
pairing map recorded in .testagent/research.md that either cites
find-untested-sources output or documents the equivalent manual approach.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
||
|
|
102663d337 |
Fix code-testing-agent activation for the Flask pytest scenario (#802)
* Fix code-testing-agent activation for the Flask pytest scenario The 'Generate pytest tests for the Flask tasks API' scenario failed to activate code-testing-agent in BOTH isolated and plugin mode: its prompt enumerated, file by file, exactly what to mock/inject/test (TaskService with repo mocked + clock injected, queries.apply_query over fixed lists, both repositories, the blueprint via test_client), acting as an answer key that let the base agent generate tests directly with edit tools instead of routing to the skill's research-plan-implement pipeline. Rewrite the prompt to a realistic, high-level ask (mirroring the ContosoUniversity scenario that does activate): describe the app at a layer level, keep the 'no tests yet', project-wide multi-file framing and the 80% coverage floor, and drop the per-module test checklist. Assertions, rubric and timeout are unchanged. Verified locally that the skill now activates in both isolated and plugin mode. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Use python3 in Flask prompt to match the grader Review feedback: the prompt told the agent to run \python -m ...\ but the grader (and the vally command) invoke \python3\. On Linux runners that may lack a \python\ shim the agent could hit command-not-found. Align the prompt to \python3\ in both eval.yaml and eval.vally.yaml. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
f389af86c7 |
code-testing-generator: mandatory pre-completion self-review gate (#768)
* code-testing-generator: mandatory pre-completion self-review gate Replaces the prose `Verify tests are implementation-specific'' bullet in Step 7 of the code-testing-generator agent with a mandatory pre-completion self-review gate that invokes two existing plugin skills: 1. `test-gap-analysis'' (pseudo-mutation check) against the source files tested and the produced test files 2. `assertion-quality'' (trivial/tautological assertion check) against the produced test files Both skills already ship in plugins/dotnet-test; this PR only wires them into the generator's workflow as a mandatory gate before declaring a run complete. The two skills' `When to Use'' sections are extended to list `called by code-testing-generator as a pre-completion self-review step'' as a recognised use case so the model does not refuse the invocation. The skill descriptions are unchanged (frontmatter is already at the 1024-char limit). Rule 11 of the generator agent is updated to list the gate alongside final build, final test, and coverage-gap review as mandatory for ALL strategies including Direct. A matching rubric item is added to the ContosoUniversity scenario in the code-testing-agent eval (yaml and vally) so the LLM judge can verify the gate was actually invoked on the trajectory. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Apply suggestions from code review Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> |
||
|
|
754011b5ee |
Add workspace-integrity guardrail to code-testing agents (#773)
The code-testing-generator/implementer agents could treat an unusual or
scaffolded workspace (e.g. a gutted repo with an injected synthetic module)
as corruption and "repair" it with git checkout/restore/reset/clean or rm,
restoring deleted tracked files and testing the wrong code.
- Replace generator Rule 5 ("Clean git first - stash changes") with an
explicit "Treat the workspace as delivered" rule, and add a "Never mutate
version control" rule. Output must be purely additive test files.
- Add a no-revert/no-clean invariant to the implementer's edit boundaries.
- Add a 'workspace integrity' eval to the code-testing-agent suite
(eval.yaml + eval.vally.yaml). The fixture looks gutted: a metricsd project
whose real core/io modules are committed at HEAD but deleted from the
working tree, leaving only a synthetic 'synthstr' decoy. A git restore would
resurrect the deleted sentinel files; graders fail if they reappear and
require passing pytest tests for the module as delivered.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
||
|
|
a98fb44e52 |
code-testing-agent: pin down behavior in generated tests (#767)
Adds a new `Write Tests That Pin Down Behavior'' section to `unit-test-generation.prompt.md'' covering five universal test-craft principles: 1. Mutation thinking - each assertion would fail under a plausible bug 2. Property intersections - test combinations, not only coordinate axes 3. Behavior radius - assert on at least one secondary observable 4. Fixture realism - never set the parameter under test to a degenerate value 5. Quick self-review before declaring a test method done Mirrors the same depth requirements in `code-testing-implementer.agent.md'' Step 4 as a cross-language invariant block alongside the existing `Edit boundaries'' rules. Extends the existing `code-testing-agent'' eval rubrics (yaml and vally) with one or two depth-oriented bullets per scenario: * ContosoUniversity: minimal IsNotNull-only assertions + secondary observable check on controller actions * python-flask-tasks: minimal `is not None''-only assertions + at least one combined-property TaskService validation test * typescript-vitest-cart: minimal `toBeDefined''/`toBeTruthy''-only assertions + at least one intersection test (discount + tax + shipping together) Rationale and prior-art references in PR description. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
d9c1e7a801 |
code-testing-agent: mention find-untested-sources for C# discovery (#734)
* code-testing-agent: mention find-untested-sources for C# discovery Adds a conditional pointer (gated on 'when available') to the find-untested-sources skill in two places: - SKILL.md Step 3 (Research Phase): high-level note for C# / .NET multi-file scopes — prefer the helper over manual find/grep/glob walks. - code-testing-researcher.agent.md Section 7 (Discover Preexisting Tests): directive instruction telling the researcher to invoke the helper before manually pairing source <-> test files, and to use its source_to_tests / untested output to fill the research document. Both callouts are phrased as 'when available', so installations without the find-untested-sources skill continue to work via manual discovery. Adds no behavior for non-C# repos. Context: in a 5x136-instance internal experiment on the msbench .NET test bench, adding equivalent pointers to the routed code-testing-agent yielded a 15.67% input-token reduction at neutral pass rate — the model trusted the documented pairing heuristics and skipped its own discovery walk. The helper itself was not invoked in those runs (the Copilot CLI router did not auto-load the sibling skill), so the measured win comes from the doc text causing the model to short-circuit its manual exploration, not from the helper executing. Depends on dotnet/skills#733 (which adds the find-untested-sources skill itself). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * code-testing-agent eval: grade find-untested-sources usage on C# scenario PR #734 adds a doc pointer steering the researcher toward the `find-untested-sources` skill on .NET / C# multi-file tasks. The existing ContosoUniversity scenario is the natural fit: it's the only C# multi-file scenario in this eval, and the new pointer is .NET- scoped. Add one rubric item to that scenario (in both eval.vally.yaml and the legacy eval.yaml) asking the grader to verify the researcher actually leveraged the helper — either by citing its `source_to_tests` / `untested` JSON output in `.testagent/research.md`, or by executing `scripts/Find-UntestedSources.cs` — instead of falling back to manual `find` / `grep` / `glob` walks. Gated on `when available in the workspace` to match the SKILL.md wording, so installs without find-untested-sources are not penalized. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
faca87d669 |
Bump vite, @vitest/coverage-v8 and vitest (#731)
Bumps [vite](https://github.com/vitejs/vite/tree/HEAD/packages/vite) to 8.0.16 and updates ancestor dependencies [vite](https://github.com/vitejs/vite/tree/HEAD/packages/vite), [@vitest/coverage-v8](https://github.com/vitest-dev/vitest/tree/HEAD/packages/coverage-v8) and [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest). These dependencies need to be updated together. Updates `vite` from 5.4.21 to 8.0.16 - [Release notes](https://github.com/vitejs/vite/releases) - [Changelog](https://github.com/vitejs/vite/blob/main/packages/vite/CHANGELOG.md) - [Commits](https://github.com/vitejs/vite/commits/v8.0.16/packages/vite) Updates `@vitest/coverage-v8` from 2.1.9 to 4.1.8 - [Release notes](https://github.com/vitest-dev/vitest/releases) - [Changelog](https://github.com/vitest-dev/vitest/blob/main/docs/releases.md) - [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.8/packages/coverage-v8) Updates `vitest` from 2.1.9 to 4.1.8 - [Release notes](https://github.com/vitest-dev/vitest/releases) - [Changelog](https://github.com/vitest-dev/vitest/blob/main/docs/releases.md) - [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.8/packages/vitest) --- updated-dependencies: - dependency-name: vite dependency-version: 8.0.16 dependency-type: indirect - dependency-name: "@vitest/coverage-v8" dependency-version: 4.1.8 dependency-type: direct:development - dependency-name: vitest dependency-version: 4.1.8 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
812565f8d1 |
Add polyglot evals for code-testing-agent (Python Flask + TypeScript Vitest) (#709)
* Add polyglot evals for code-testing-agent (Python Flask + TypeScript Vitest) Adds two non-.NET evaluation scenarios to the dotnet-test/code-testing-agent eval to measure how well the agent generates idiomatic tests in Python and TypeScript — the languages exercised by the msbench top5-* benchmarks. New fixtures (no pre-existing tests, so passing pytest/vitest proves the agent actually generated them): - fixtures/python-flask-tasks/ — small Flask app with TaskService (injected clock), TaskRepository Protocol + InMemoryTaskRepository, and a /tasks blueprint with HTTP 201/200/400/404/409 surfaces. Validated locally with `python -m pip install -e ".[test]" && python -m pytest`. - fixtures/typescript-vitest-cart/ — small shopping-cart library with an injectable DiscountPolicy seam (No/Percentage policies) and a Cart class with non-trivial merge / clamping semantics. Validated locally with `npm ci && npx vitest run`. package-lock.json committed so `npm ci` is reproducible in CI. New scenarios (added to both eval.yaml and eval.vally.yaml): - "Generate pytest tests for the Flask tasks API (Python polyglot)" — asserts that `python3 -m pip install -e '.[test]' && python3 -m pytest` exits 0 and that at least one `tests/**/test_*.py` file was produced. Rubric rewards using Flask's `test_client()`, mocking `TaskRepository` for service-level tests, asserting HTTP status + JSON body, and covering the empty-title / >200-char / not-found / already-done error paths. - "Generate Vitest tests for the shopping-cart library (TypeScript polyglot)" — asserts that `npm ci && npx vitest run` exits 0 and that at least one `tests/**/*.test.ts` file was produced. Rubric rewards mocking `DiscountPolicy` via `vi.fn()`, covering Cart merge semantics, `updateQuantity` zero-removes, `totals()` discount-clamping, and the `PercentageDiscountPolicy` constructor boundary. Companion to #708 (polyglot examples + sub-agent generification); these scenarios exercise the per-language guidance that PR adds. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address review feedback for #709 - routes.py: guard `create_task` against non-dict JSON payloads (e.g. list) by returning 400 instead of letting `payload.get(...)` raise AttributeError and surface as a 500. This keeps the API behavior stable and makes it easier for generated tests to assert 400 on invalid input shapes. - tsconfig.json: drop `vitest/globals` from `compilerOptions.types`. The fixture's vitest.config.ts sets `globals: false`, so allowing TypeScript to assume globals (describe/it/expect) only enables tests that compile but fail at runtime. Removing it keeps the fixture aligned with the non-global Vitest API the eval is exercising. Both fixtures re-smoke-tested locally: vitest 1/1, pytest 2/2 (including a new test confirming the non-dict body returns 400). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix skill activation for polyglot scenarios in PR #709 Both polyglot scenarios (pytest + Flask, Vitest + TS) scored 5.0/5 quality but reported `⚠️ NOT ACTIVATED` because the SDK's skill router did not load `code-testing-agent` into the agent's context: `activated: false, detectedSkills: []` in skillActivationIsolated / skillActivationPlugin. The .NET ContosoUniversity scenario activates fine because its prompt has stronger `project-wide, multi-file ... high coverage` framing and SKILL.md already covers .NET keywords. Fixes per `eng/skill-validator/src/docs/InvestigatingResults.md` section "Skill not activated": - SKILL.md description: add framework-specific keywords (pytest, Flask/Django, Vitest, Jest, Mocha, JUnit, Node libraries, API, package, project-wide, multi-file) so the router has explicit hooks for polyglot prompts. Trimmed verbose `DO NOT USE FOR` clauses to stay under the 1,024-char skill spec limit (now 1,019 chars). - eval.yaml + eval.vally.yaml: rewrite the two polyglot prompts to mirror the .NET scenario's framing — add `project-wide, multi-file test generation task across the ... layers` and `achieve high coverage`, soften the library-prescriptive bullets (Mock(spec=...), vi.fn()) into capability-level requirements (`mocked or stubbed`, `mock or hand-written stub`). Rubric items unchanged — they remain flexible enough to grade either mock style. Validated locally: `dotnet run --project eng/skill-validator/src -- check --plugin ./plugins/dotnet-test` => ✅ All checks passed (23 skill(s), 11 agent(s), 1 plugin(s)). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR #709 review + sharpen polyglot skill activation **Activation fix for polyglot scenarios (Python + TypeScript)** The previous attempt ( |
||
|
|
04348a46ef |
tests/dotnet-test/code-testing-agent: fix coverage XML lookup path (#657)
The grader expected coverage.cobertura.xml under a workspace-root TestResults/ directory, but `dotnet test --collect:"XPlat Code Coverage"` writes the report under <TestProject>/TestResults/<guid>/coverage.cobertura.xml. With the previous path, `find TestResults …` failed with no-such-directory before even searching, and the glob `TestResults/**/coverage.cobertura.xml` could not match the nested location either, so the assertion failed on otherwise-successful runs. Search recursively from cwd in both eval.vally.yaml (`find . …`) and eval.yaml (`**/TestResults/**/coverage.cobertura.xml`) so the assertion matches the actual output location of XPlat Code Coverage regardless of the test project layout the agent chose. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
3d59e44c7e |
Fix coverage-analysis activation for plateau diagnosis prompts (#647)
* Fix coverage-analysis activation for plateau diagnosis prompts coverage-analysis SKILL.md: - Trim verbose implementation details (provider detection, ReportGenerator) that consumed description budget without aiding skill activation - Add explicit USE FOR keywords: coverage stuck, coverage plateau, can't increase coverage, what's blocking coverage code-testing-agent SKILL.md: - Add 'diagnosing coverage plateaus or CRAP score computation (use coverage-analysis)' to DO NOT USE FOR boundary to prevent test-generation skill from intercepting diagnostic prompts * Potential fix for pull request finding Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> * Strengthen code-testing-agent activation; harden coverage-analysis isolated mode Address eval regressions reported on PR #647 (run 25813728646): 1. code-testing-agent: `Generate tests for ContosoUniversity ASP.NET Core MVC app` was NOT ACTIVATED in plugin mode (detectedSkills=[], skillEventCount=0, invokedAgents=[]). The model bypassed the skill system entirely. - SKILL.md description: restructure to use the proven `Use when user says ...` pattern with quoted trigger phrases (matching the run-tests skill that consistently activates), make the link to the code-testing-generator sub-agent explicit, and tighten DO NOT USE FOR clauses. - eval prompt (eval.yaml + eval.vally.yaml): make the request pipeline-shaped (`project-wide, multi-file test generation task`, `scaffold a new test project`) so the model recognizes it as multi-step work that benefits from the orchestrated pipeline. Explicitly request coverlet.collector + a Cobertura XML run so rubric criterion 1 (`high line coverage as reported by the Cobertura XML in TestResults/`) becomes achievable without overfitting. 2. code-testing-tester agent + code-testing-extensions/dotnet.md: open a scoped exception to the `skip coverage tools` rule. Default behavior stays the same, but when the user/harness explicitly asks for a Cobertura/XML coverage artifact, the agent may add coverlet.collector to the generated test csproj so the harness's coverage command produces output. The agent still does not run the coverage command itself. 3. coverage-analysis SKILL.md: add a `User-visible output is mandatory` guard at the top of the Workflow section. The latest eval showed isolated mode producing literally `(no output)` in 2 of 3 scenarios — the agent ran Compute-CrapScores.ps1 / Extract-MethodCoverage.ps1 / ReportGenerator in parallel, then the session ended without ever surfacing findings. The guard tells the agent to always return a partial summary instead of ending silent, and to deprioritize ReportGenerator HTML when budget is tight. (Plugin-mode quality is already strong: 4.3 / 4.3 / 5.0 — no regression risk there.) Aggregate dotnet-test plugin description size: 14,925 chars (limit 15,000). skill-validator check passes (22 skills, 11 agents, 1 plugin); markdownlint passes for all 4 modified files. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Restore MSTest modernization exclusion in code-testing-agent description * Fix isolated-mode coverage-analysis: emit summary before optional ReportGenerator The previous workflow encouraged the agent to run `dotnet tool install` for ReportGenerator in parallel with the CRAP scoring scripts (Phase 2 "Steps 3 and 4 in parallel" + Phase 3 "Steps 5 and 6 in parallel"). In isolated mode that pattern reliably crashed the session with "Failed to persist session events: timeout while waiting for mutex to become available" right after the scripts returned valid data, so the agent never produced the user-facing summary. Restructure the workflow into 5 phases: - Phase 1 (Setup) - unchanged - Phase 2 (Test execution) - skip when Cobertura XML already exists - Phase 3 (Analysis) - run only the two PowerShell scripts, no RG - Phase 4 (User-facing summary) - MANDATORY, must be the next assistant response after Phase 3, before any RG work; also save coverage-analysis.md as a secondary follow-up - Phase 5 (ReportGenerator HTML/CSV) - strictly optional, post-summary, skipped by default for existing-Cobertura and plateau-diagnosis paths Also update references/output-format.md so the Reports section marks RG artifacts as "Not generated (optional - request HTML reports to enable)" when Phase 5 has not run, and update references/guidelines.md so the "show and open the markdown report" rule explicitly defers to the user-facing assistant response. Targets the isolated-mode regressions in PR #647 eval: - Project-wide coverage with existing Cobertura: 1.0/5 -> expected 3+ - Coverage plateau diagnosis: 1.0/5 -> expected 3+ - Run coverage from scratch: 2.3/5 -> expected steady or up Verified: skill-validator check --plugin ./plugins/dotnet-test passes; markdownlint-cli2 clean on all 3 modified files. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR #647 review comments 1. Prerequisites: distinguish the from-scratch path (needs NuGet for the coverage-provider package install + optional internet for ReportGenerator) from the existing-Cobertura path (needs neither). 2. Add Step 2c "Discover or accept existing Cobertura XML" so the existing-data path actually has $coberturaFiles populated before Phase 3, instead of relying on Phase 2's discovery (which it skips). Also clarify Step 2's destructive Remove-Item only manages the skill-owned coverage-analysis/ subdirectory. 3. references/output-format.md: replace the unconditional "Reports saved to: <coverageDir>/reports/" line with one that always points at <coverageDir>/ (markdown summary + raw Cobertura) and only mentions reports/ if Phase 5 ran. 4. Have Compute-CrapScores.ps1 emit OVERALL_LINE_COVERAGE and OVERALL_BRANCH_COVERAGE from the Cobertura root attributes, and update Phase 4 to read those values directly from the script's output. The Phase 4 mandatory-summary rule no longer requires a separate XML parse before composing the response. Verified: skill-validator check --plugin ./plugins/dotnet-test passes; markdownlint-cli2 clean on all modified files; Compute-CrapScores.ps1 smoke-tested on a synthetic Cobertura XML (emits OVERALL_LINE_COVERAGE:75, OVERALL_BRANCH_COVERAGE:50 alongside HOTSPOTS). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Apply suggestion from @Evangelink * Address unresolved coverage-analysis and eval review comments * Refine follow-up review feedback from validation * Tighten coverage aggregation fallback notes and counters * Clarify pre-response save instruction wording --------- Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> |
||
|
|
0f31228b58 |
Fix eval timeouts and skill activation for dotnet-test evals (#641)
* Fix evaluation timeouts and improve prompts for dotnet-test skills
- writing-mstest-tests: increase global vally timeout from 3m to 5m to match
the 300s eval.yaml timeout for the complex 'Modernize legacy test patterns' stimulus
- code-testing-agent: increase timeout from 40m/2400s to 60m/3600s
- mtp-hot-reload: increase timeout from 6m to 10m and remove overly strict
expect_tools constraint on launchSettings stimulus
- migrate-mstest-v1v2-to-v3: rephrase prompt for assertion overload errors to
mention CS1501 and cover AreNotEqual/AreSame alongside AreEqual
* Increase eval timeouts for mtp-hot-reload and migrate-mstest-v1v2-to-v3 scenarios
* Strengthen activation prompts and add missing vally scenarios
- migrate-mstest-v1v2-to-v3: Add explicit 'MSTest v2 to v3' migration
language in .testsettings and DataRow prompts (both eval.yaml and
eval.vally.yaml)
- writing-mstest-tests: Add 'MSTest' keyword and skill-specific
terminology to all prompts that lacked activation signal
- writing-mstest-tests: Add 5 missing vally scenarios (string
assertions, comparison assertions, collection/null/reference
assertions, conditional execution, parallelization)
- Add service-registry fixture files for the collection assertions
scenario
* Improve skill activation via SKILL.md descriptions and prompt alignment
Three-pronged fix for systematic NOT ACTIVATED failures:
1. SKILL.md descriptions: Restructure writing-mstest-tests, migrate-
mstest-v1v2-to-v3, and code-testing-agent descriptions to use the
proven 'Use when user says' pattern with quoted trigger phrases
(matching the working run-tests pattern)
2. Vally prompts: Each prompt now contains at least one exact quoted
trigger phrase from its SKILL.md description to ensure strong
description-to-prompt matching
3. Specific fixes:
- writing-mstest-tests: All 15 prompts now include MSTest-specific
trigger phrases (Assert.ThrowsExactly, sealed test class, DataRow,
DynamicData, Assert.HasCount, Assert.IsInstanceOfType, etc.)
- migrate-mstest-v1v2-to-v3: Add '.NET 5 dropped' and 'framework
compatibility issues' to trigger phrases; strengthen dropped-TFM
prompt
- code-testing-agent: Add 'generate comprehensive unit tests' and
'achieve high code coverage' to trigger phrases
* Revert SKILL.md description changes that caused activation regression
The description restructuring in
|
||
|
|
9433a31e67 | Initial port of evals to Vally (#615) | ||
|
|
17df9a10eb |
Fix code-testing-agent eval timeout by removing unused NuGet packages (#551)
The ContosoUniversity test fixture had 17 unused NuGet packages (bootstrap, jQuery, Modernizr, Antlr4, WebGrease, Microsoft.Build.Tasks.Core, etc.) that bloated NuGet restore and build time. The multi-agent test generation pipeline performs multiple build/test cycles, causing cumulative delays that exceeded the 2400s scenario timeout on CI. Changes: - Remove 17 unused PackageReference entries from ContosoUniversity.csproj - Remove stale ItemGroup entries (Global.asax, Scripts, compile removes) - Remove NU1701 warning suppression (no longer needed) - Increase command_timeout from 180s to 300s for dotnet test assertion (accounts for first-time restore + build + test + coverage collection) |
||
|
|
400661df09 |
run-tests: add critical rules, detection guidance, and troubleshooting (#514)
* run-tests: add critical rules, detection guidance, and troubleshooting - Add 'Critical Rules' table to prevent cross-platform VSTest/MTP mistakes - Expand Step 1 detection with explicit file-by-file lookup table - Add 'dotnet --version' as first detection action - Add Troubleshooting section with 7 common error patterns - Expand Common Pitfalls from 3 to 7 entries - Strengthen negative guidance for SDK version-specific syntax * Improve description * Improve evaluation * Add expected and rejected tools * Change * Improve do not use sections * Improve prompt * Increase eval timeouts for scenarios hitting the limit - run-tests: trx reporting MTP SDK 9 (360s -> 480s) - run-tests: blame-hang MTP SDK 10 (300s -> 420s) - run-tests: combine filter criteria VSTest (180s -> 300s) - mtp-hot-reload: hot reload SDK 9 (360s -> 480s) - mtp-hot-reload: hot reload filter (180s -> 300s) - code-testing-agent: ContosoUniversity (1800s -> 2400s) |
||
|
|
e4670b33a1 |
[PoC] Code testing agent + tests PoC (#433)
* Add code testing agent agents + tests PoC * Remove AssertionEvaluator.cs change (moved to dev/jankrivanek/agents-evals) |