Commit Graph

19 Commits

Author SHA1 Message Date
Amaury Levé 57733bebc8 Keep test agent state out of commits (#1108)
* Keep test agent state out of commits

Move broad test-generation pipeline state to host scratch storage, worktree-specific Git metadata, or OS temp, and enforce the exclusion in evals.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify absolute test agent state path

Use Git's explicit absolute path formatting in both test-generation entry points.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Prune Git metadata from test agent eval guards

Avoid scanning nested repositories and align the remaining TESTAGENT_DIR placeholder with the documented format.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify test agent command handoff

Require callers to provide exact commands, excerpts, or absolute TESTAGENT_DIR document paths to command-running sub-agents.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Reject all repository-local testagent entries

Match .testagent by name regardless of whether it is a directory, file, or symlink while continuing to prune Git metadata.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Verify external test agent artifacts

Restore broad-run artifact checks at the Git metadata path and pass the researched lint command and state directory to the linter agent.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Make testagent eval guards constant time

Check only the forbidden workspace-root path, including broken symlinks, instead of recursively traversing dependency trees.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Broaden comprehensive test generation

Treat explicit requirements as the floor for broad suites and add mutation-relevant equivalence-partition and invariant coverage without test-count padding.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Fix external artifact grader quoting

Run state checks directly in the harness shell so TESTAGENT_DIR expands after assignment, with an isolated command probe covering valid and forbidden states.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Run broad skill evals in Git worktrees

Initialize the seven broad evaluation roots as Git repositories so TESTAGENT_DIR resolves deterministically and external artifacts remain verifiable.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify non-stageable test agent state

Describe the real invariant across the pipeline: state may live under .git metadata but must never be version-controlled workspace content or appear in git status.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Standardize intermediate test state contract

Use one TESTAGENT_DIR placeholder, clearer intermediate-state terminology, and detect stageable research, plan, or status files regardless of directory name.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Use one Git root in workspace integrity eval

Baseline the fixture from the evaluation root so stageable intermediate-state files remain visible to the directory-independent guard.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Prune Vitest dependencies from state scan

Exclude node_modules through per-eval Git metadata so stageable state detection remains fast without modifying fixture content.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Strengthen focused intermediate-state guards

Separate shell execution, reject Git-metadata files on focused runs, include ignored state files, and prune node_modules with a pathspec exclusion.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

---------

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
2026-09-03 16:36:32 -07:00
Amaury Levé 5a06b20cc9 Support classic .NET test projects in dotnet-test (#993)
* Support classic .NET test projects

Teach dotnet-test skills and agents to preserve non-SDK projects, packages.config dependencies, explicit compile registration, legacy runners, and version-compatible MSTest APIs. Add regression evals for generation, execution, coverage, and authoring.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68

* Address classic test fixture review

Tighten the MSTest version grader, make the runner fixture assertion behavioral, and use nameof for the guarded parameter.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68

---------

Copilot-Session: fdfec89f-b610-479c-a6c7-c2936b300e68
2026-08-12 08:38:51 -07:00
Amaury Levé 4df4da469a Upgrade agentic workflows and fix stale PR cleanup (#916)
* chore: upgrade gh-aw runtime

* fix: paginate stale pull request cleanup

Upgrade the generated agentic workflow assets and ensure stale PR discovery includes every result page and draft pull requests.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 968a22c2-327f-4d26-8f86-1c59bcd323ea

* Improve dotnet-test eval coverage and efficiency

Address remaining high-confidence items from #899 by bounding the code-testing pipeline and adding eval coverage for grade-tests and find-untested-sources.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: e430fee9-d3df-4ef5-85a4-745ae4b17046

* fix: address agentic workflow review

Pin the Copilot setup checkout action and include the cutoff date in stale PR search results.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 968a22c2-327f-4d26-8f86-1c59bcd323ea

* Improve test migration skill guidance

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 5e19d263-02a6-45b1-9cfb-424fa4d10863

* fix: add fixture namespace imports

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: a464e6e4-3e45-41fe-b17d-887c8cb8a448

* test: assert GetOrderById grade

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: a464e6e4-3e45-41fe-b17d-887c8cb8a448

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-07-20 13:20:14 +00:00
Amaury Levé 77154137e8 dotnet-test: make code-testing agent tools Claude Code-compatible + add cross-host portability check (#856)
* dotnet-test: make code-testing agent tools declarations Claude Code-compatible

PR #847 added `tools: ["agent", "skill", "read", "search", "edit", "execute"]`
to the code-testing-* agents to enable VS Code / Copilot CLI subagent fan-out.
Those lowercase aliases map to real tools in VS Code and the Copilot CLI, but
Claude Code matches `tools:` against its own vocabulary (Task, Skill, Read,
Glob, Grep, Edit, Write, Bash). None of the aliases matched, so when these
agents are loaded into Claude Code via --plugin-dir and selected with
`claude --agent`, the agent was granted ZERO tools. A tool-less model asked
to generate tests emits a textual <tool_call> block and exits after one turn,
producing no file changes.

Append the Claude Code tool names to each agent's `tools:` list so the same
declaration works across all three runtimes (each honors the names it knows and
ignores the foreign ones):

- Orchestrators (generator, implementer): add Task, Skill, Read, Glob, Grep,
  Edit, Write, Bash (Task is the Claude Code equivalent of the `agent`
  fan-out tool).
- Workers (researcher, planner, builder, tester, fixer, linter): add Skill,
  Read, Glob, Grep, Edit, Write, Bash.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* skill-validator: complete built-in tools + add cross-host tool portability check

Two related follow-ups to the agent tools fix:

1. Address the skill-check review feedback. The validator's BuiltInTools set was
   missing three legitimate host tool spellings that are not case-insensitive
   matches of existing entries, so they were flagged as non-built-in:
   - "write"   — Claude Code file-creation tool (Copilot CLI / VS Code: "create")
   - "agent"   — Copilot CLI / VS Code subagent fan-out tool (Claude Code: "task")
   - "execute" — Copilot CLI / VS Code run-command tool (Claude Code: "bash")
   "agent" and "execute" were already flagged before this branch (introduced by
   the fan-out PR); adding them to BuiltInTools clears the pre-existing warnings.

2. Add a cross-host tool portability check (CheckAgentToolPortability) so an
   agent that declares a capability for only one host is flagged. Tool names are
   matched case-sensitively (hosts resolve tools by exact spelling), so an agent
   that lists e.g. only "edit" (Copilot / VS Code) without "Edit"/"Write"
   (Claude Code) is reported as working on one host and silently tool-less on the
   other. Findings are advisory (do not fail CI) and allowlistable via
   "agent-tool-portability:AGENT:capability". Wired into the agents loop in
   CheckCommand and covered by unit tests.

Also make the one existing single-host agent (optimizing-dotnet-performance)
portable by adding its Claude Code tool spellings, so the new check reports a
clean tree.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
2026-07-07 09:00:22 +02:00
Petr Pokorny 06dd46e5f1 Enable VS Code subagent fan-out for the dotnet-test code-testing agents (#847)
* Add VS Code subagent metadata to dotnet-test code-testing agents

Enable the code-testing-* Research-Plan-Implement pipeline to fan out as
subagents in VS Code while keeping the GitHub Copilot CLI working.

Frontmatter (VS Code "coordinator/worker" pattern; portable tool aliases that
map in both VS Code and the CLI):
- code-testing-generator: tools: [agent, read, search, edit, execute] + an
  agents: list (researcher, planner, implementer, builder, tester, fixer,
  linter); softened the three runSubagent({ agent, prompt }) blocks to
  tool-agnostic delegation prose.
- code-testing-implementer: tools: [agent, read, search, edit, execute] +
  agents: (builder, tester, fixer, linter).
- Leaf agents (researcher/planner/builder/tester/fixer/linter):
  tools: [read, search, edit, execute].

Why explicit tools (not ["*"]): VS Code has no all-tools wildcard for 	ools:
and a subagent's 	ools: overrides its inherited set, so ["*"] matched nothing
and stripped subagents of file tools (they ran without read/edit/search). The
CLI treats ["*"] as all-tools, so this was VS-Code-specific. The portable
aliases agent/read/search/edit/execute map to real tools in both environments;
agents: is ignored by the CLI.

README: document the VS Code chat.subagents.allowInvocationsFromSubagents
setting (off by default) needed for the nested implementer->builder/tester/
fixer/linter layer to fan out on large scopes; the CLI has no such gate.

Validated end to end: VS Code shows researcher->planner->implementer fanning out
with real file I/O (no "subagents lack file tools" warning); CLI fan-out intact
with skills still loading and tests passing. An all-tools baseline used only
tools within this enumerated set, confirming no CLI capability is restricted.

* Include skill tool in code-testing agent tool allowlists

Address PR review: the explicit `tools:` allowlists omitted the `skill`
tool, but every code-testing-* agent's prompt instructs calling skills
(e.g. `code-testing-extensions` for per-language guidance, `test-gap-analysis`,
`assertion-quality`). Because `tools:` is an override, omitting `skill` can
prevent the agents from loading those skills in environments that gate skill
invocation by the allowlist.

Add `skill` to all eight agents:
- orchestrators (generator, implementer): [agent, skill, read, search, edit, execute]
- workers (researcher/planner/builder/tester/fixer/linter): [skill, read, search, edit, execute]

Re-verified in the Copilot CLI: full fan-out (researcher -> planner ->
implementer/tester), the `code-testing-extensions` skill is invoked, and the
generated tests pass.
2026-06-30 19:42:39 +02:00
Amaury Levé 754011b5ee Add workspace-integrity guardrail to code-testing agents (#773)
The code-testing-generator/implementer agents could treat an unusual or
scaffolded workspace (e.g. a gutted repo with an injected synthetic module)
as corruption and "repair" it with git checkout/restore/reset/clean or rm,
restoring deleted tracked files and testing the wrong code.

- Replace generator Rule 5 ("Clean git first - stash changes") with an
  explicit "Treat the workspace as delivered" rule, and add a "Never mutate
  version control" rule. Output must be purely additive test files.
- Add a no-revert/no-clean invariant to the implementer's edit boundaries.
- Add a 'workspace integrity' eval to the code-testing-agent suite
  (eval.yaml + eval.vally.yaml). The fixture looks gutted: a metricsd project
  whose real core/io modules are committed at HEAD but deleted from the
  working tree, leaving only a synthetic 'synthstr' decoy. A git restore would
  resurrect the deleted sentinel files; graders fail if they reappear and
  require passing pytest tests for the module as delivered.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-16 13:49:34 +00:00
Amaury Levé a98fb44e52 code-testing-agent: pin down behavior in generated tests (#767)
Adds a new `Write Tests That Pin Down Behavior'' section to
`unit-test-generation.prompt.md'' covering five universal
test-craft principles:

1. Mutation thinking - each assertion would fail under a plausible bug
2. Property intersections - test combinations, not only coordinate axes
3. Behavior radius - assert on at least one secondary observable
4. Fixture realism - never set the parameter under test to a degenerate value
5. Quick self-review before declaring a test method done

Mirrors the same depth requirements in
`code-testing-implementer.agent.md'' Step 4 as a cross-language
invariant block alongside the existing `Edit boundaries'' rules.

Extends the existing `code-testing-agent'' eval rubrics (yaml and
vally) with one or two depth-oriented bullets per scenario:

* ContosoUniversity: minimal IsNotNull-only assertions + secondary
  observable check on controller actions
* python-flask-tasks: minimal `is not None''-only assertions + at
  least one combined-property TaskService validation test
* typescript-vitest-cart: minimal `toBeDefined''/`toBeTruthy''-only
  assertions + at least one intersection test (discount + tax +
  shipping together)

Rationale and prior-art references in PR description.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-16 11:14:34 +00:00
Amaury Levé 0a387fcdfe code-testing-implementer: add mandatory harness-discovery check (#757)
* code-testing-implementer: add mandatory harness-discovery check

Generic CI/benchmark verifiers run the framework's default discovery from the repo root. Tests that pass via a scoped command (e.g. 'dotnet test MyProject.Tests.csproj', 'bundle exec rspec subgem/spec', 'Invoke-Pester -Path ./tools') but are invisible to the harness count as 0 generated tests.

Repeatedly observed in msbench cta-nightly runs:

- new C# test project never 'dotnet sln add'ed (ocelot)

- Pester tests placed under custom directories invisible to default Invoke-Pester (homebrew, ruby/ruby)

- RSpec specs placed in a sub-gem's spec/ dir invisible from repo root (fastlane)

Changes:

- code-testing-implementer: capture baseline test count in Step 2, new Step 7 'Verify Harness Discovery (MANDATORY)' with concrete failure examples, HARNESS_DISCOVERY line in report template, reminder after Step 3 to revisit registration if Step 4 creates a new project.

- code-testing-researcher: instruct to record BOTH scoped test command and harness-equivalent discovery command in .testagent/research.md.

- dotnet.md: strengthen 'Registering a new test project' heading to MANDATORY when dotnet new was used; add 'Harness Discovery Check' section with 'dotnet test <solution> --list-tests' from repo root.

- powershell.md: add 'Harness Discovery Check' section with default-config Invoke-Pester from repo root.

- ruby.md: add gem-monorepo trap guidance (fastlane, ruby/ruby) in Test Placement Contract; add 'Harness Discovery Check' section with 'bundle exec rspec --dry-run' from repo root.

* Address review feedback on harness-discovery check

- dotnet.md: replace stale `Step 8 cleanup` reference with the actual Step 3 (`Register Test Project with Build System`).
- dotnet.md: demote `Harness Discovery Check` to `###` so it nests under the parent `## .csproj / .sln Handling` section alongside `### Registering...`.
- dotnet.md: replace `\s\{4\}` (non-POSIX) in the grep regex with 4 literal spaces so the count works under BRE.
- ruby.md: group the rake/rails fallback with `{ ...; } | wc -l` so the pipe applies to both branches of `||`.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-12 12:39:29 +00:00
Amaury Levé 49a77a90bf code-testing-agent: add polyglot pipeline examples for Python/TypeScript/Go/Java (#708)
* code-testing-agent: add polyglot pipeline examples for Python/TypeScript/Go/Java

The code-testing-agent skill family is polyglot in description but in
practice biased toward .NET because dotnet-examples.md was the only
filled-in pipeline walkthrough. The four sub-agents that participate in
the Research-Plan-Implement pipeline (researcher, planner, implementer,
generator) all pointed to dotnet-examples.md whenever they suggested a
concrete example, which made it harder for the agent to produce
idiomatic non-.NET tests (e.g. in msbench top5-* benchmarks for
Python/Flask, TypeScript/Express).

This change

* adds four new example files mirroring dotnet-examples.md format
  (source → research → plan → generated test → fix cycle → final report):
  - python-examples.md (pytest, unittest.mock, Mock(spec=...), parametrize)
  - typescript-examples.md (Vitest with notes for Jest; it.each,
    async tests, fake timers, ESM/CJS fix cycle)
  - go-examples.md (standard testing package, table-driven subtests,
    hand-written fake repository, injected clock)
  - java-examples.md (JUnit 5 + Mockito on Maven, @ParameterizedTest +
    @CsvSource, Clock.fixed, Surefire fix cycles)
* updates code-testing-extensions/SKILL.md TOC to list the new files
  and clarifies usage instructions to read the matching <language>-
  examples.md alongside the base extension
* makes the "Concrete example" pointers in code-testing-generator,
  code-testing-implementer, code-testing-planner and
  code-testing-researcher agents language-agnostic (list all available
  example files instead of hard-coding dotnet-examples.md)
* expands code-testing-researcher project-structure detection list to
  cover more Python (tox.ini, noxfile.py, requirements*.txt, uv.lock,
  poetry.lock, pdm.lock), JS/TS (.mts/.cts/.jsx, vitest.config.*,
  jest.config.*), C++ (CMakeLists.txt, BUILD.bazel, meson.build),
  Java/Kotlin (pom.xml, build.gradle[.kts], wrappers), and other
  ecosystem files; expands the Identify-Language section accordingly
* extends the "Language-Specific Examples" section in
  code-testing-agent/SKILL.md to summarise each example file

Validated with: skill-validator check --plugin ./plugins/dotnet-test
(23 skills, 11 agents — all checks passed) and markdownlint-cli2.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Address review feedback for #708

- go-examples.md: replace the hand-rolled `contains`/`stringIndex`
  helpers with `strings.Contains` from the standard library. The hand-rolled
  `contains` had subtly wrong semantics — `contains(""abc"", """")` returned
  `false` while `strings.Contains` returns `true` — and the helpers are
  unnecessary complexity for a code-generation example.
- go-examples.md: in the `go test -run` "wrong selection regex" sample fix
  cycle, quote the test name and use `single_item` (matching the underscore
  that the surrounding diagnosis text refers to) instead of the unquoted
  `single item` which the shell would parse as two separate CLI arguments.
- java-examples.md: the source-tree file list described `Invoice.java` as a
  `record` but the `InvoiceService.markAsPaid` example mutates the invoice
  via `setStatus(...)` and `setPaidDate(...)` — records are immutable, so
  the description was internally inconsistent. Re-describe it as a mutable
  POJO with explicit mutators to match the service code.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-02 13:42:46 +00:00
Amaury Levé ac653c62e4 dotnet-test: add cross-language edit-boundary rules to code-testing-implementer (#689)
The code-testing-* agent family generates tests for 12 languages (not just
.NET). Until now, no instruction told the implementer to keep its changes
additive: it would routinely modify existing test files non-additively
(deleting/reformatting lines) or edit non-test production code to make
something easier to test. Both behaviors cause the change to be rejected
by:

- the msbench 'simple' test verifier (any 'dels > 0' on a test file, or
  any change to a non-test file, sets reward=0)
- most real-world test-quality gates, code-review policies, and CI

Add explicit cross-language rules in the implementer agent prompt:

1. Existing test files are append-only (no reformat/reorder/remove).
2. Do not modify non-test source files; surface untestable seams as
   follow-ups for the testability-migration agent.
3. Prefer new test files over edits to existing ones when equivalent.
4. Build-system manifests may be edited only for project/dependency
   registration.

Rule 6 in the Rules section now points at the Step 4 detail so the
implementer reads it on every phase.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-05-27 09:42:22 +02:00
YuliiaKovalova e1e8568375 Revert "dotnet-test: add unit-under-test + behaviors quality cue to code-test…" (#651)
This reverts commit f1b09eba79.
2026-05-14 17:00:07 +02:00
YuliiaKovalova f1b09eba79 dotnet-test: add unit-under-test + behaviors quality cue to code-testing-generator (#646)
* CTA: invocation-only baseline (dispatch mechanics, no content rules)

Experiment branch to isolate the impact of "invoke prompted subagents
more often" from the impact of "make those subagents do richer work."
Same baseline as dev/ykovalova/cta-prompt-tuning (main = 66628b6), but
strips out every content/quality rule and keeps only the dispatch
plumbing.

Comparison branch: dev/ykovalova/cta-prompt-tuning (HEAD: bd530be)
which contains both the dispatch mechanics AND content/quality rules.

Files modified (3 vs 5 in cta-prompt-tuning):
- code-testing-generator.agent.md  +110 lines
- code-testing-implementer.agent.md +11 lines
- code-testing-fixer.agent.md      +1  line
- code-testing-researcher.agent.md  UNTOUCHED (baseline)
- code-testing-planner.agent.md     UNTOUCHED (baseline)

KEPT (invocation / dispatch mechanics):

Generator:
- Rule 1: every task() call MUST use agent_type "dotnet-test:code-testing-..."
  (without this, calls dispatch generic built-ins and never reach the
  named CTA agents)
- Rule 2: routing table -- which named agent for which job
- Rule 3: prefer one named-agent dispatch over many tool calls
- Rule 4: orchestrator MUST NOT edit/create test files itself
  (forces implementer dispatch)
- Rule 5: orchestrator MUST NOT run builds/tests via terminal
  (forces builder/tester dispatch)
- Rule 6: every run MUST dispatch the planner (no exceptions for "small"
  scope; Direct still goes through planner)
- Rule 7: every build/test failure MUST dispatch the fixer
- Step 1b: mandatory initial researcher dispatch (every strategy)
- Direct strategy rewritten: dispatches planner -> implementer -> builder
  -> tester -> fixer -> linter (was "Skip Steps 3-5, write tests inline")
- All Step 3/4/5/6/7/8/9 dispatches converted from runSubagent({agent:...})
  to task({ agent_type: "dotnet-test:code-testing-...", name:..., prompt:...})
- Step 9 validator dispatch (forces builder dispatch for cleanup)
- Steps 6/7 mandatory builder/tester dispatch wrapper

Implementer:
- Section 5: "you MUST dispatch fixer for build errors" + no-inline-edit
  block (forces fixer dispatch on build failures)
- Section 6: "you MUST dispatch fixer for test failures" + no-inline-edit
  block (forces fixer dispatch on test failures)
- Section 7: "Format Code (mandatory if a lint command exists)"
  (was "Optional"; mandatory firing of linter)
- Rule 6: never declare SUCCESS while build/tests fail (gates SUCCESS on
  fixer dispatch)
- Rule 7: no inline test-file edits between failed dispatch and fixer

Fixer:
- Frontmatter description widened to advertise handling of failing tests
  (without this, the orchestrator's routing logic does not select the
  fixer for test failures, so even Rule 7's mandate produces no firing
  -- this is the change that took fixer firing from 0.00/inst to 0.39/inst
  in earlier iterations)

DROPPED (content / quality rules -- in cta-prompt-tuning, NOT here):

Generator:
- Test-strength rules embedded in implementer dispatch prompt
- Test-design rules embedded in implementer dispatch prompt (OFAT,
  mutation self-check, never mock subject under test)
- File-location rules embedded in implementer dispatch prompt
- TARGET ENTITIES / PHASE CHECKLIST / TEST TRACEABILITY blocks in
  implementer dispatch prompt
- CHECKLIST format spec in planner dispatch prompt
- Step 9 validator's detailed cleanup classification

Implementer:
- Section 4b "Verify CHECKLIST coverage" pre-completion check
- Section 8 "CHECKLIST COVERAGE" report block
- "Honor the CHECKLIST" rule

Fixer:
- "Process -- Failing Tests" section (5-step diagnosis flow)
- All anti-weakening / anti-skipping rules
- "Re-derive expected from production source" guidance

Planner:
- CHECKLIST format ("one item per TARGET BEHAVIOR, Source/Variants/
  Expected mandatory")
- "Test name from research.md conventions" rule
- "At least 2 phases" rule

Researcher:
- Section 8 "Extract Local Test Naming & Style Conventions"
- TARGET ENTITIES / TARGET BEHAVIORS / TEST INFRASTRUCTURE structure
  in research.md
- Test naming pattern extraction

WHAT THE SUBAGENTS WILL ACTUALLY DO:

The researcher / planner / implementer / fixer all operate at baseline
behavior -- they receive the same prompts they receive in the upstream
"vanilla" runs. The only difference vs vanilla is that the orchestrator
ACTUALLY DISPATCHES THEM (where vanilla often inlines the work or skips
sub-agent dispatch entirely).

EXPECTED COMPARISON:

If quality on this branch is similar to or higher than dev/ykovalova/
cta-prompt-tuning (bd530be), then "more dispatches" is the dominant
quality lever and the content/quality rules in cta-prompt-tuning are
adding marginal or noise-level value.

If quality on this branch is materially lower than cta-prompt-tuning,
then the content/quality rules are doing the heavy lifting and the
dispatch mechanics alone are insufficient.

If quality on this branch matches or exceeds vanilla but trails
cta-prompt-tuning, then the dispatch mechanics provide a baseline lift
and the content rules add an incremental quality layer on top.

Rubber-duck check passed (validated dispatch-vs-content classification;
fixer frontmatter is routing metadata not a runtime gate; surviving
dispatch prompts contain no dangling references to removed CHECKLIST /
TARGET ENTITIES / TEST STRENGTH / naming-convention concepts).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* CTA: extract unit-under-test + behaviors quality cue (no caps)

* Address PR #646 review feedback

- Step 1b prompt: explicitly request unit-under-test (file:line) and behaviors
  so the verification gate rarely needs re-dispatch.
- Step 3: rename to 'Deep Research Phase', mark skipped for Direct strategy,
  and switch from overwriting research.md to extending it (no double-research).
- Step references: update '6-9' -> '6-10' and 'Step 9' -> 'Steps 9-10' in the
  strategy table and the All-strategies-MUST line, since reporting is Step 10.
- Step 6 builder prompt: drop '*.sln' glob (could expand to multiple args);
  use 'dotnet build --no-incremental' (auto-discovers .sln) per dotnet.md.
- Step 9: stop overloading the builder agent; perform diff/cleanup directly
  in the orchestrator (Rule 5 forbids inline build/test, not git/fs hygiene).
- Fixer agent: update mission text to cover failing tests and assertion
  correction (front-matter description already mentioned this; body now
  matches), with explicit no-Ignore/no-Skip/no-production-rewrite guardrails.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Yuliia Kovalova <ykovalova@example.com>
2026-05-13 14:03:31 +02:00
YuliiaKovalova 809d0b180c Improve test generation quality based on SWE Atlas benchmark analysis (#599)
* Improve test quality for benchmark performance

Based on SWE Atlas benchmark analysis (48.8% vanilla vs 19-34% CTA):

1. Default to Direct strategy for single-task requests — reduces
   multi-agent overhead that costs tokens without improving results

2. Run tests immediately in Direct strategy — catches assertion
   errors early instead of accumulating failures

3. Read source thoroughly before writing tests — trace actual logic
   and return values, not just function signatures

4. Quality over Quantity guidelines — cover stated requirements first,
   fewer focused tests beat many shallow ones

5. Verify tests are implementation-specific — tests that pass with
   an empty function body aren't testing anything useful

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Combine redundant implementation-specificity bullets into one

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-04-29 19:05:22 +00:00
Jan Krivanek 81946d2a38 Add proxy for CTA extension files (#583) 2026-04-24 12:19:17 +02:00
Amaury Levé f8e8191268 Add concrete pipeline examples for code-testing-agent (#572)
Add end-to-end input/output examples to anchor expected behavior for the
LLM, addressing the Example Quality (2/5) review feedback.

Changes:
- New extensions/dotnet-examples.md with .NET-specific examples: sample
  source code, research output, plan output, generated test file, fix
  cycle walkthroughs, and final report
- SKILL.md: replace one-liner examples with strategy selection table,
  pipeline walkthrough, and pointer to language-specific extensions
- Generator agent: add strategy decision examples and sample final report
- Researcher/Planner/Implementer agents: add references to extension
  examples for concrete output shapes

Architecture keeps core agents language-agnostic — all language-specific
examples live in the extensions/ folder.
2026-04-23 12:18:36 +00:00
Jan Krivanek 05aeb657e6 Add license to agent files (#568) 2026-04-21 12:57:18 +00:00
Amaury Levé 1125fe7864 improve(dotnet-test): enhance agent descriptions for subagent discovery (#560)
- Add 'Use when:' trigger phrases to all 7 sub-agent descriptions
  so parent agents can reliably discover and delegate to them
- Add .testagent/ cleanup rule to generator agent to prevent
  ephemeral pipeline state from being committed
2026-04-20 19:51:09 +02:00
Artur Spychaj 3dbe832796 Add 'dotnet sln add' step to code-testing-implementer (#522)
* Add 'dotnet sln add' step to code-testing-implementer

When a new test project is created, register it with the solution
file using 'dotnet sln add' so that 'dotnet test <solution>'
discovers the tests.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Address review: improve dotnet sln add guidance

- Use solution identified in research/plan rather than searching
- Account for .slnf solution filter scenarios
- Prefer 'dotnet test --solution' over 'dotnet test <solution>'

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Move dotnet sln add guidance to dotnet.md extension

The implementer is polyglot, so 'dotnet sln add' doesn't belong
there. Replace with a generic 'register with build system' step
that defers to extensions/, and add the .NET-specific detail
(sln/slnx/slnf handling) to extensions/dotnet.md.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Address review: fix solution registration guidance in dotnet.md

- Include .slnf in 'don't substitute' warning alongside .sln/.slnx
- Qualify '--solution' flag as SDK 10+/MTP-only; fall back to
  positional form for older SDKs

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-04-15 14:38:03 +02:00
Jan Krivanek e4670b33a1 [PoC] Code testing agent + tests PoC (#433)
* Add code testing agent agents + tests PoC

* Remove AssertionEvaluator.cs change (moved to dev/jankrivanek/agents-evals)
2026-03-30 17:57:08 +02:00