mirror of
https://github.com/dotnet/skills.git
synced 2026-09-20 09:49:54 +08:00
fix-dotnet-test-analysis-scripts
10 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9beca0b0be |
Bump vitest (#1145)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.8 to 4.1.11. - [Release notes](https://github.com/vitest-dev/vitest/releases) - [Changelog](https://github.com/vitest-dev/vitest/blob/main/docs/releases.md) - [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.11/packages/vitest) --- updated-dependencies: - dependency-name: vitest dependency-version: 4.1.11 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
d3921f7418 |
Strengthen testability skill evaluations (#1057)
* Strengthen testability skill evaluations Raise four dotnet-test evals to eight independent stimuli, add validated fixtures, and resolve code-testing-agent orphan fixtures without speculative routing changes. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> * Relax promo-code eval grader Accept deterministic suffix values beyond one hard-coded literal and match common PascalCase test names. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 * Improve testability skill reliability Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 * Refine testability obstacle graders Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 * Accept qualified Random seams Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 * Improve ambient seam compatibility Replace the C# 12 primary constructor in the copyable Scope sample with a conventional constructor so the guidance works in projects using older language versions. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 * Improve testability skill reliability Refine routing and execution contracts from exact losing transcripts, strengthen behavioral eval checks, and add isolated C# fixtures without increasing repeated runs. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 84f7c88c-c8e7-4d8f-96c9-421de725ab31 |
||
|
|
5a06b20cc9 |
Support classic .NET test projects in dotnet-test (#993)
* Support classic .NET test projects Teach dotnet-test skills and agents to preserve non-SDK projects, packages.config dependencies, explicit compile registration, legacy runners, and version-compatible MSTest APIs. Add regression evals for generation, execution, coverage, and authoring. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68 * Address classic test fixture review Tighten the MSTest version grader, make the runner fixture assertion behavioral, and use nameof for the guarded parameter. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68 --------- Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68 |
||
|
|
aa3c4abca5 |
Bump postcss (#988)
Bumps [postcss](https://github.com/postcss/postcss) from 8.5.15 to 8.5.25. - [Release notes](https://github.com/postcss/postcss/releases) - [Changelog](https://github.com/postcss/postcss/blob/main/CHANGELOG.md) - [Commits](https://github.com/postcss/postcss/compare/8.5.15...8.5.25) --- updated-dependencies: - dependency-name: postcss dependency-version: 8.5.25 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
573938df09 |
Close the remaining dotnet-test eval follow-ups (#899) (#971)
* Close the remaining dotnet-test eval follow-ups from #899 Every non-agent dotnet-test eval now clears the 5-trial floor, the two cost P1s are addressed, and the reference-skill coverage gap is recorded as a decision instead of a standing warning. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 024b3241-d7af-418a-b9cd-3bb9ee9bf0ee * Remove the duplicate block that cloned a grade-tests scenario, and gate it The 'production code available' scenario carried a leftover tail from the edit that moved the 'production code unavailable' one. YAML keeps the last duplicate key, so every field of the new scenario was silently overwritten: it shipped as a byte-identical rerun of its predecessor and never loaded the dotnet-production-available fixture it was built around. Parsing the spec and counting stimuli - which is what verified this PR - returns the intended 5 names either way, so only the parser can see it. The eval-quality gate now loads specs with a duplicate-key-strict loader (failing check 9), with a self-test case and the incident recorded. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 024b3241-d7af-418a-b9cd-3bb9ee9bf0ee --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 024b3241-d7af-418a-b9cd-3bb9ee9bf0ee |
||
|
|
f2eb897a12 |
Fix dotnet-test findings from the refreshed cross-family eval (#899) (#945)
* Fix dotnet-test findings from the refreshed cross-family eval (#899) Every change below is driven by judge evidence from the losing trials of the refreshed 5-family dotnet-test matrix (runs 30108473397 + recovery runs), not by style preference. Eval measurement fix — the "discovery" P2s were an artifact: - assertion-quality, test-gap-analysis, test-smell-detection, and test-tagging each have a decline stimulus with `constraints.reject_skills: ["*"]`, so the skill cannot activate there by construction. Without `expect_activation: false` the adapter counted those dormant runs as missed activations, which is exactly the 75-88% invocation rates reported in the scorecard. Annotating them (the convention already used by agent.test-quality-auditor) removes the false signal; the non-activations were the only ones observed for these skills. Skill fixes: - test-gap-analysis: baselines won by actually running the suite while the skill reasoned statically and reported survivors that the tests in fact kill. Added Step 4b: confirm every reported survivor by applying it, re-running the covering tests, and reverting; fall back to reasoning only when the suite cannot run, labelled unverified. Calibrated severity down for strong suites. - test-anti-patterns: baselines won on depth, not polish. Added a depth bar — account for every test in scope, give exact expected values in fixes, name the adjacent error-path/boundary gaps, and keep counts consistent. Trimmed three pitfall rows that duplicated the calibration step so the skill stays under the profiler's "comprehensive" threshold. - detect-static-dependencies: losses were all counting accuracy. One authoritative total (no findings parked outside it), classify by the resource touched rather than by the `static` keyword, exclude pure helpers such as Path.Combine from the needs-wrapping total, require file:line, and add the missing randomness/culture/serialization categories. - test-smell-detection: the calibration rule told models to downgrade Sleepy Test for integration tests, which is what lost both losing scenarios. Fixed sleeps now stay High in any category; Mystery Guest and Eager Test still downgrade. - crap-score: losses came from estimating coverage after collection failed. Added the dotnet-coverage/ReportGenerator recovery path and a hard rule never to publish a CRAP score built on assumed coverage. - coverage-analysis: answer the asked question first, reconcile every number against the script output, and list every below-threshold member instead of declaring one method the entire gap. - migrate-static-to-wrapper: migrate exactly what was requested (no adjacent DateTime.Now rewrites, respect intentional-use comments) and never report "build succeeded" when the build or restore failed. - code-testing-agent: quote each requirement verbatim in the evidence table so multi-condition requirements map to a test that covers the whole combination, and cite a clean run rather than a coverage attempt that exited non-zero. Validation: skill-validator check passes (20 skills, 10 agents); markdownlint clean; eval specs parse and the adapter now reports all four decline stimuli as expect-dormant. Refs #899 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1947263a-0ef9-47bd-ac53-5af5afa3ddaa * Strengthen underpowered dotnet-test evals from the PR 945 eval run The PR eval reported 5 of 10 skills as "no credible improvement". Reproducing the gate arithmetic from the artifacts shows the dominant cause is statistical power, not skill quality. The gate is `mean > 0 AND ci_low > 0` with a t-based CI over per-trial scores, which reduces to `sqrt(n) * (mean/sd) > t(n-1)`. The required mean/sd ratio is brutal at small n: n=2 -> 8.98 n=3 -> 2.48 n=4 -> 1.59 n=5 -> 1.24 n=6 -> 1.05 n=8 -> 0.84 Recomputing each reported CI from the per-trial scores reproduces the published numbers exactly, which confirms the mechanism: crap-score [0.4,1.0,0.4] n=3 CI [-0.261, 1.461] migrate-static-to-wrapper [1.0,0.4,0.4] n=3 CI [-0.261, 1.461] test-gap-analysis [0.4,0,0.4,0.4] n=4 CI [-0.018, 0.618] test-anti-patterns [0.4,0,0.4,0,0,0.4] n=6 CI [-0.030, 0.430] code-testing-agent [0,0.4] n=2 CI [-2.341, 2.741] crap-score and migrate-static-to-wrapper won 100% of their trials (3W/0T/0L) and still failed: at n=3 nothing short of three identically-sized wins can clear the gate. That is a property of a thin eval, not of the skill. Scenario counts are raised with discriminating cases, four of them by wiring up fixtures that were already committed but had no stimulus referencing them: - test-gap-analysis 4 -> 6, using the orphaned `report-quality` fixture (trivial auto-properties and an auto-generated .g.cs to skip, private helpers reachable only through the public API, and a deliberately weak Assert.IsTrue that cannot kill arithmetic mutations) and the orphaned `rust-error-propagation` fixture (an untested `?` propagation path and an untested `<=` boundary). - test-anti-patterns 6 -> 8, using the orphaned `pytest-mixed` fixture (which also checks the calibration rule that pytest's bare `assert` must not be flagged) and the orphaned `assertion-problems` fixture (which separates Critical false-confidence assertions from a Low-severity message nit). - crap-score 3 -> 6, with a new `refactor-required` fixture whose numbers are self-consistent: ApplySurcharges has complexity 13 behind a stale "Complexity: 4" comment (CRAP 28.4, needs 77.2% coverage), ClassifyAccount has complexity 17 so coverage alone can never reach CRAP < 15, and RoundToCurrency is 100% covered so its CRAP equals its complexity exactly. - migrate-static-to-wrapper 3 -> 5, adding a DateTimeKind-preservation scenario over the existing fixture and a new `static-helper` fixture where a static class must gain an ambient TimeProvider seam without breaking its callers. - code-testing-agent 2 -> 3, with a compact C# fixture that must extend an existing suite to the untested method only. This eval stays the weakest: each scenario is expensive, so raising `runs` is a better lever than adding more heavyweight scenarios. Verification: - every eval spec parses and all 254 fixture references resolve - the three new fixtures compile; the shipping-quotes fixture restores, builds and its three seed tests pass under `dotnet test` in a clean workspace - skill-validator check passes (20 skills, 10 agents) - markdownlint clean Refs #899 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1947263a-0ef9-47bd-ac53-5af5afa3ddaa * Address review feedback on fixture and counting wording - BillableWeightTests: the ZeroOrNegative test only asserted the zero case, so its name overstated what it covered. Made it data-driven over 0 and -1 so the name matches the assertions. This matters more than usual here: the file is the seed suite for a test-quality eval, and a misleading test name is exactly what these skills are supposed to flag. - detect-static-dependencies: the Step 3 lead-in said to count each "static call pattern", which contradicted the rule immediately below it that instance members reaching the same untestable resource must also be counted. Reworded to "call site" and made the instance-member inclusion explicit. Verified: the fixture restores, builds and now passes 4 tests (was 3); skill-validator check passes; markdownlint clean. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1947263a-0ef9-47bd-ac53-5af5afa3ddaa --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1947263a-0ef9-47bd-ac53-5af5afa3ddaa |
||
|
|
71414ce000 |
Improve dotnet-test eval coverage and efficiency (#917)
* Improve dotnet-test eval coverage and efficiency Address remaining high-confidence items from #899 by bounding the code-testing pipeline and adding eval coverage for grade-tests and find-untested-sources. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e430fee9-d3df-4ef5-85a4-745ae4b17046 * Fix dotnet-test eval activation and quality Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c5c1a52-4f99-49d6-b503-1bec713a6e98 * Strengthen dotnet-test skill activation Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 9c5c1a52-4f99-49d6-b503-1bec713a6e98 |
||
|
|
754011b5ee |
Add workspace-integrity guardrail to code-testing agents (#773)
The code-testing-generator/implementer agents could treat an unusual or
scaffolded workspace (e.g. a gutted repo with an injected synthetic module)
as corruption and "repair" it with git checkout/restore/reset/clean or rm,
restoring deleted tracked files and testing the wrong code.
- Replace generator Rule 5 ("Clean git first - stash changes") with an
explicit "Treat the workspace as delivered" rule, and add a "Never mutate
version control" rule. Output must be purely additive test files.
- Add a no-revert/no-clean invariant to the implementer's edit boundaries.
- Add a 'workspace integrity' eval to the code-testing-agent suite
(eval.yaml + eval.vally.yaml). The fixture looks gutted: a metricsd project
whose real core/io modules are committed at HEAD but deleted from the
working tree, leaving only a synthetic 'synthstr' decoy. A git restore would
resurrect the deleted sentinel files; graders fail if they reappear and
require passing pytest tests for the module as delivered.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
||
|
|
faca87d669 |
Bump vite, @vitest/coverage-v8 and vitest (#731)
Bumps [vite](https://github.com/vitejs/vite/tree/HEAD/packages/vite) to 8.0.16 and updates ancestor dependencies [vite](https://github.com/vitejs/vite/tree/HEAD/packages/vite), [@vitest/coverage-v8](https://github.com/vitest-dev/vitest/tree/HEAD/packages/coverage-v8) and [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest). These dependencies need to be updated together. Updates `vite` from 5.4.21 to 8.0.16 - [Release notes](https://github.com/vitejs/vite/releases) - [Changelog](https://github.com/vitejs/vite/blob/main/packages/vite/CHANGELOG.md) - [Commits](https://github.com/vitejs/vite/commits/v8.0.16/packages/vite) Updates `@vitest/coverage-v8` from 2.1.9 to 4.1.8 - [Release notes](https://github.com/vitest-dev/vitest/releases) - [Changelog](https://github.com/vitest-dev/vitest/blob/main/docs/releases.md) - [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.8/packages/coverage-v8) Updates `vitest` from 2.1.9 to 4.1.8 - [Release notes](https://github.com/vitest-dev/vitest/releases) - [Changelog](https://github.com/vitest-dev/vitest/blob/main/docs/releases.md) - [Commits](https://github.com/vitest-dev/vitest/commits/v4.1.8/packages/vitest) --- updated-dependencies: - dependency-name: vite dependency-version: 8.0.16 dependency-type: indirect - dependency-name: "@vitest/coverage-v8" dependency-version: 4.1.8 dependency-type: direct:development - dependency-name: vitest dependency-version: 4.1.8 dependency-type: direct:development ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
812565f8d1 |
Add polyglot evals for code-testing-agent (Python Flask + TypeScript Vitest) (#709)
* Add polyglot evals for code-testing-agent (Python Flask + TypeScript Vitest) Adds two non-.NET evaluation scenarios to the dotnet-test/code-testing-agent eval to measure how well the agent generates idiomatic tests in Python and TypeScript — the languages exercised by the msbench top5-* benchmarks. New fixtures (no pre-existing tests, so passing pytest/vitest proves the agent actually generated them): - fixtures/python-flask-tasks/ — small Flask app with TaskService (injected clock), TaskRepository Protocol + InMemoryTaskRepository, and a /tasks blueprint with HTTP 201/200/400/404/409 surfaces. Validated locally with `python -m pip install -e ".[test]" && python -m pytest`. - fixtures/typescript-vitest-cart/ — small shopping-cart library with an injectable DiscountPolicy seam (No/Percentage policies) and a Cart class with non-trivial merge / clamping semantics. Validated locally with `npm ci && npx vitest run`. package-lock.json committed so `npm ci` is reproducible in CI. New scenarios (added to both eval.yaml and eval.vally.yaml): - "Generate pytest tests for the Flask tasks API (Python polyglot)" — asserts that `python3 -m pip install -e '.[test]' && python3 -m pytest` exits 0 and that at least one `tests/**/test_*.py` file was produced. Rubric rewards using Flask's `test_client()`, mocking `TaskRepository` for service-level tests, asserting HTTP status + JSON body, and covering the empty-title / >200-char / not-found / already-done error paths. - "Generate Vitest tests for the shopping-cart library (TypeScript polyglot)" — asserts that `npm ci && npx vitest run` exits 0 and that at least one `tests/**/*.test.ts` file was produced. Rubric rewards mocking `DiscountPolicy` via `vi.fn()`, covering Cart merge semantics, `updateQuantity` zero-removes, `totals()` discount-clamping, and the `PercentageDiscountPolicy` constructor boundary. Companion to #708 (polyglot examples + sub-agent generification); these scenarios exercise the per-language guidance that PR adds. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address review feedback for #709 - routes.py: guard `create_task` against non-dict JSON payloads (e.g. list) by returning 400 instead of letting `payload.get(...)` raise AttributeError and surface as a 500. This keeps the API behavior stable and makes it easier for generated tests to assert 400 on invalid input shapes. - tsconfig.json: drop `vitest/globals` from `compilerOptions.types`. The fixture's vitest.config.ts sets `globals: false`, so allowing TypeScript to assume globals (describe/it/expect) only enables tests that compile but fail at runtime. Removing it keeps the fixture aligned with the non-global Vitest API the eval is exercising. Both fixtures re-smoke-tested locally: vitest 1/1, pytest 2/2 (including a new test confirming the non-dict body returns 400). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Fix skill activation for polyglot scenarios in PR #709 Both polyglot scenarios (pytest + Flask, Vitest + TS) scored 5.0/5 quality but reported `⚠️ NOT ACTIVATED` because the SDK's skill router did not load `code-testing-agent` into the agent's context: `activated: false, detectedSkills: []` in skillActivationIsolated / skillActivationPlugin. The .NET ContosoUniversity scenario activates fine because its prompt has stronger `project-wide, multi-file ... high coverage` framing and SKILL.md already covers .NET keywords. Fixes per `eng/skill-validator/src/docs/InvestigatingResults.md` section "Skill not activated": - SKILL.md description: add framework-specific keywords (pytest, Flask/Django, Vitest, Jest, Mocha, JUnit, Node libraries, API, package, project-wide, multi-file) so the router has explicit hooks for polyglot prompts. Trimmed verbose `DO NOT USE FOR` clauses to stay under the 1,024-char skill spec limit (now 1,019 chars). - eval.yaml + eval.vally.yaml: rewrite the two polyglot prompts to mirror the .NET scenario's framing — add `project-wide, multi-file test generation task across the ... layers` and `achieve high coverage`, soften the library-prescriptive bullets (Mock(spec=...), vi.fn()) into capability-level requirements (`mocked or stubbed`, `mock or hand-written stub`). Rubric items unchanged — they remain flexible enough to grade either mock style. Validated locally: `dotnet run --project eng/skill-validator/src -- check --plugin ./plugins/dotnet-test` => ✅ All checks passed (23 skill(s), 11 agent(s), 1 plugin(s)). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> * Address PR #709 review + sharpen polyglot skill activation **Activation fix for polyglot scenarios (Python + TypeScript)** The previous attempt ( |