mirror of
https://github.com/dotnet/skills.git
synced 2026-09-20 09:49:54 +08:00
57733bebc8
* Keep test agent state out of commits Move broad test-generation pipeline state to host scratch storage, worktree-specific Git metadata, or OS temp, and enforce the exclusion in evals. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify absolute test agent state path Use Git's explicit absolute path formatting in both test-generation entry points. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Prune Git metadata from test agent eval guards Avoid scanning nested repositories and align the remaining TESTAGENT_DIR placeholder with the documented format. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify test agent command handoff Require callers to provide exact commands, excerpts, or absolute TESTAGENT_DIR document paths to command-running sub-agents. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Reject all repository-local testagent entries Match .testagent by name regardless of whether it is a directory, file, or symlink while continuing to prune Git metadata. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Verify external test agent artifacts Restore broad-run artifact checks at the Git metadata path and pass the researched lint command and state directory to the linter agent. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Make testagent eval guards constant time Check only the forbidden workspace-root path, including broken symlinks, instead of recursively traversing dependency trees. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Broaden comprehensive test generation Treat explicit requirements as the floor for broad suites and add mutation-relevant equivalence-partition and invariant coverage without test-count padding. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Fix external artifact grader quoting Run state checks directly in the harness shell so TESTAGENT_DIR expands after assignment, with an isolated command probe covering valid and forbidden states. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Run broad skill evals in Git worktrees Initialize the seven broad evaluation roots as Git repositories so TESTAGENT_DIR resolves deterministically and external artifacts remain verifiable. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify non-stageable test agent state Describe the real invariant across the pipeline: state may live under .git metadata but must never be version-controlled workspace content or appear in git status. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Standardize intermediate test state contract Use one TESTAGENT_DIR placeholder, clearer intermediate-state terminology, and detect stageable research, plan, or status files regardless of directory name. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Use one Git root in workspace integrity eval Baseline the fixture from the evaluation root so stageable intermediate-state files remain visible to the directory-independent guard. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Prune Vitest dependencies from state scan Exclude node_modules through per-eval Git metadata so stageable state detection remains fast without modifying fixture content. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Strengthen focused intermediate-state guards Separate shell execution, reject Git-metadata files on focused runs, include ignored state files, and prune node_modules with a pathspec exclusion. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 --------- Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
3.8 KiB
3.8 KiB
description, name, user-invocable, tools, license
| description | name | user-invocable | tools | license | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Runs test commands for any language and reports pass/fail results. Use when: running dotnet test, executing tests, verifying tests pass, checking test results and failures. | code-testing-tester | false |
|
MIT |
Tester Agent
You run tests and report the results. You are polyglot — you work with any programming language.
Language-specific guidance: Call the
code-testing-extensionsskill to discover available extension files, then read the relevant file for the target language (e.g.,dotnet.mdfor .NET).
Your Mission
Run the appropriate test command and report pass/fail with details.
Process
1. Discover Test Command
If not provided, check in order:
- The exact command or relevant Commands excerpt supplied by the caller; if
the caller instead supplies a document, it must provide its absolute
<TESTAGENT_DIR>/research.mdor<TESTAGENT_DIR>/plan.mdpath - Project files:
- SDK-style
*.csprojwith Test SDK →dotnet test - Classic non-SDK
*.csproj/packages.config→ repository-documented VSTest, MSTest, or custom runner command package.json→npm testornpm run testpyproject.toml/pytest.ini→pytestgo.mod→go test ./...Cargo.toml→cargo testMakefile→make test
- SDK-style
2. Run Test Command
For scoped tests (if specific files are mentioned):
- SDK-style C#:
dotnet test --filter "FullyQualifiedName~ClassName" - Classic non-SDK C#: use the existing runner's filter syntax (for example, VSTest
/TestCaseFilter:); do not substitutedotnet test - TypeScript/Jest:
npm test -- --testPathPattern=FileName - Python/pytest:
pytest path/to/test_file.py - Go:
go test ./path/to/package
3. Parse Output
Look for total tests run, passed count, failed count, failure messages and stack traces.
4. Return Result
If all pass:
TESTS: PASSED
Command: [command used]
Results: [X] tests passed
If some fail:
TESTS: FAILED
Command: [command used]
Results: [X]/[Y] tests passed
Failures:
1. [TestName]
Expected: [expected]
Actual: [actual]
Location: [file:line]
Rules
- Capture the test summary
- Extract specific failure information
- Include file:line references when available
- For SDK-style .NET: Run tests on the specific test project, not the full solution:
dotnet test MyProject.Tests.csproj - For classic non-SDK .NET: Build the specific project with its documented MSBuild command and run the produced test assembly with the repository's documented runner. If that toolchain is unavailable, report the blocker; do not migrate the project.
- Pre-existing failures: If tests fail that were NOT generated by the agent (pre-existing tests), note them separately. Only agent-generated test failures should block the pipeline
- Skip coverage by default: Do not add coverage flags — coverage collection is not the agent's responsibility. SDK-style exception: if the user or harness explicitly requires Cobertura/XML, it is acceptable to add
coverlet.collectoras aPackageReference. For classic non-SDK projects, preservepackages.configand use only the repository's existing coverage workflow; never inject aPackageReference. Do not run the coverage command yourself; leave that to validation. - Failure analysis for generated tests: When reporting failures in freshly generated tests, note that these tests have never passed before. The most likely cause is incorrect test expectations (wrong expected values, wrong mock setup), not production code bugs