Files
dotnet__skills/tests/dotnet-test/code-testing-agent/eval.yaml
Amaury Levé 57733bebc8 Keep test agent state out of commits (#1108)
* Keep test agent state out of commits

Move broad test-generation pipeline state to host scratch storage, worktree-specific Git metadata, or OS temp, and enforce the exclusion in evals.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify absolute test agent state path

Use Git's explicit absolute path formatting in both test-generation entry points.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Prune Git metadata from test agent eval guards

Avoid scanning nested repositories and align the remaining TESTAGENT_DIR placeholder with the documented format.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify test agent command handoff

Require callers to provide exact commands, excerpts, or absolute TESTAGENT_DIR document paths to command-running sub-agents.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Reject all repository-local testagent entries

Match .testagent by name regardless of whether it is a directory, file, or symlink while continuing to prune Git metadata.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Verify external test agent artifacts

Restore broad-run artifact checks at the Git metadata path and pass the researched lint command and state directory to the linter agent.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Make testagent eval guards constant time

Check only the forbidden workspace-root path, including broken symlinks, instead of recursively traversing dependency trees.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Broaden comprehensive test generation

Treat explicit requirements as the floor for broad suites and add mutation-relevant equivalence-partition and invariant coverage without test-count padding.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Fix external artifact grader quoting

Run state checks directly in the harness shell so TESTAGENT_DIR expands after assignment, with an isolated command probe covering valid and forbidden states.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Run broad skill evals in Git worktrees

Initialize the seven broad evaluation roots as Git repositories so TESTAGENT_DIR resolves deterministically and external artifacts remain verifiable.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Clarify non-stageable test agent state

Describe the real invariant across the pipeline: state may live under .git metadata but must never be version-controlled workspace content or appear in git status.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Standardize intermediate test state contract

Use one TESTAGENT_DIR placeholder, clearer intermediate-state terminology, and detect stageable research, plan, or status files regardless of directory name.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Use one Git root in workspace integrity eval

Baseline the fixture from the evaluation root so stageable intermediate-state files remain visible to the directory-independent guard.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Prune Vitest dependencies from state scan

Exclude node_modules through per-eval Git metadata so stageable state detection remains fast without modifying fixture content.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

* Strengthen focused intermediate-state guards

Separate shell execution, reject Git-metadata files on focused runs, include ignored state files, and prune node_modules with a pathspec exclusion.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0

---------

Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
2026-09-03 16:36:32 -07:00

495 lines
27 KiB
YAML

name: code-testing-agent
description: Evaluates the dotnet-test/code-testing-agent skill
type: capability
executionShard: generation
# `defaults:` replaces the deprecated `config:` — vally's loader throws on a
# spec declaring both (eng/eval-quality check 10).
#
# Run 31180075357 measured 2W/10T/0L. The two focused tasks never activated and
# added four baseline-equivalent ties. Seven broad tasks below exercise the
# skill's distinctive Research -> Plan -> Implement contract across Python,
# TypeScript, Go, SDK/classic .NET, existing-suite extension, and a sparse
# workspace. Two focused tasks protect proportional routing in different
# ecosystems. Nine independent stimuli improve power without changing runs.
defaults:
timeout: 60m
runs: 2
stimuli:
- name: Generate a project-wide pytest suite across multiple modules
prompt: |
Generate a comprehensive pytest suite for the entire project under
fixtures/python-multimodule/. It has no tests yet and contains:
- analytics.stats.mean and percentile
- analytics.window.RateWindow
- textkit.slug.slugify
Cover every public behavior across the modules, including empty-input and
percentile-range errors, percentile ordering and 0/100 boundaries,
RateWindow constructor validation, empty-window errors, capacity rollover,
average and peak after rollover, slug separator collapsing/trimming,
max-length truncation, and invalid max_length. Keep this project-wide:
tests should live under fixtures/python-multimodule/tests/ and pass with
pytest from fixtures/python-multimodule/.
environment:
files:
- src: fixtures/python-multimodule
dest: fixtures/python-multimodule
commands:
- git init -q
graders:
- type: run-command
config:
command: sh -c "cd fixtures/python-multimodule && python3 -m pip install --quiet pytest && python3 -m pytest -q"
expected_exit_code: 0
timeout: 5m
- type: run-command
config:
command: sh -c "grep -R -q 'RateWindow' fixtures/python-multimodule/tests && grep -R -q 'percentile' fixtures/python-multimodule/tests && grep -R -q 'slugify' fixtures/python-multimodule/tests"
expected_exit_code: 0
timeout: 1m
- type: output-matches
config:
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
- type: run-command
config:
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
expected_exit_code: 0
timeout: 1m
- type: prompt
rubric:
- Generated passing tests under the configured tests directory for all three modules
- Covered every named validation and boundary behavior, including RateWindow rollover and percentile 0/100 boundaries
- Asserted concrete averages, peaks, percentiles, and slugs rather than existence-only results
- Kept the work project-wide but bounded to the three supplied modules
- Mapped each requested behavior to named test evidence
- Kept broad-scope intermediate state files non-stageable so only requested deliverables remain commit candidates
- name: Generate project-wide tests for a classic MSTest library
prompt: |
Add the missing project-wide unit tests for the classic net472 library
under fixtures/classic-mstest/. The production project contains
DiscountService and TieredDiscountPolicy; the existing test project uses
MSTest 3.5.2, Moq 4.2, NBuilder, FixtureBase<TSut>, packages.config, and
explicit compile items.
Preserve that stack and the existing DiscountServiceTests.cs byte-for-byte.
Create DiscountServiceBoundaryTests.cs and TieredDiscountPolicyTests.cs.
Cover DiscountService's 0/100 boundaries, invalid percentages, missing
product, and exact discounted value. Cover TieredDiscountPolicy constructor
validation, negative subtotal, below-threshold behavior, exact-threshold
activation, and a concrete discounted value. Do not modernize the project
or dependencies.
environment:
files:
- src: fixtures/classic-mstest
dest: fixtures/classic-mstest
commands:
- rm -f fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs && cp fixtures/classic-mstest/tests/Discounts.Tests.csproj.pristine fixtures/classic-mstest/tests/Discounts.Tests.csproj && rm fixtures/classic-mstest/tests/Discounts.Tests.csproj.pristine && mkdir -p .eval-baseline && cp fixtures/classic-mstest/tests/DiscountServiceTests.cs .eval-baseline/DiscountServiceTests.cs && cp fixtures/classic-mstest/tests/packages.config .eval-baseline/packages.config
- git init -q
graders:
- type: file-exists
config:
path: fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
- type: file-exists
config:
path: fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs
- type: run-command
config:
command: >-
sh -c 'diff -u .eval-baseline/DiscountServiceTests.cs
fixtures/classic-mstest/tests/DiscountServiceTests.cs
&& diff -u .eval-baseline/packages.config
fixtures/classic-mstest/tests/packages.config
&& test $(grep -c "Compile Include=\"DiscountServiceBoundaryTests.cs\""
fixtures/classic-mstest/tests/Discounts.Tests.csproj) -eq 1
&& test $(grep -c "Compile Include=\"TieredDiscountPolicyTests.cs\""
fixtures/classic-mstest/tests/Discounts.Tests.csproj) -eq 1
&& for file in
fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs;
do grep -Eq "Assert[[:space:]]*\.[[:space:]]*ThrowsException[[:space:]]*<" "$file"
&& ! grep -Eq "Assert[[:space:]]*\.[[:space:]]*Throws(Exactly)?[[:space:]]*<" "$file"
&& ! grep -Eq "\[[[:space:]]*ExpectedException" "$file"
|| exit 1;
done'
expected_exit_code: 0
timeout: 1m
- type: file-not-contains
config:
path: fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
value: ThrowsExactly
- type: file-not-contains
config:
path: fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
value: Assert.Throws<
- type: file-contains
config:
path: fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
value: Assert.ThrowsException<
- type: file-not-contains
config:
path: fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs
value: ThrowsExactly
- type: file-not-contains
config:
path: fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs
value: Assert.Throws<
- type: file-contains
config:
path: fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs
value: Assert.ThrowsException<
- type: file-not-contains
config:
path: fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
value: ExpectedException
- type: file-not-contains
config:
path: fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs
value: ExpectedException
- type: output-matches
config:
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
- type: run-command
config:
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
expected_exit_code: 0
timeout: 1m
- type: prompt
rubric:
- Added both requested test files and registered each exactly once in the classic project
- Preserved the existing test file, packages.config, project format, and pinned dependency versions
- Used Assert.ThrowsException<T> for exception paths, avoiding both Assert.Throws<T> and Assert.ThrowsExactly<T>, which are unavailable in MSTest 3.5.2
- Covered every named DiscountService and TieredDiscountPolicy boundary with concrete expected values
- Mapped each requested behavior and build-registration requirement to evidence
- Kept broad-scope intermediate state files non-stageable without modernizing the classic project
- name: Generate a project-wide Go suite across collaborating packages
prompt: |
Generate comprehensive Go tests for the whole module under
fixtures/go-multipackage/. Cover money.Discount, shipping.Rate, and
order.Total.
Use table-driven tests for validation and numeric boundaries. Verify
discount 0/100 boundaries, invalid subtotal and percent, shipping bracket
equality (a weight equal to UpToGrams belongs to that bracket), overflow
behavior, negative weight, and empty brackets. For order.Total, use fakes
for both interfaces and prove the composition `subtotal - discount +
shipping` with exact cents. Also cover error propagation from each
collaborator and verify the second collaborator is not called when the
first fails. Keep tests inside the corresponding Go packages and make
`go test ./...` pass from fixtures/go-multipackage/.
environment:
files:
- src: fixtures/go-multipackage
dest: fixtures/go-multipackage
commands:
- git init -q
graders:
- type: run-command
config:
command: sh -c "cd fixtures/go-multipackage && go test ./..."
expected_exit_code: 0
timeout: 5m
- type: run-command
config:
command: sh -c "test -f fixtures/go-multipackage/money/discount_test.go && test -f fixtures/go-multipackage/shipping/rate_test.go && test -f fixtures/go-multipackage/order/total_test.go && grep -q 'UpToGrams' fixtures/go-multipackage/shipping/rate_test.go"
expected_exit_code: 0
timeout: 1m
- type: output-matches
config:
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
- type: run-command
config:
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
expected_exit_code: 0
timeout: 1m
- type: prompt
rubric:
- Generated passing tests in all three packages
- Used table-driven tests for validation and boundary cases
- Proved the exact order total composition with collaborator fakes and concrete cents
- Covered each collaborator error and verified short-circuit behavior when discount calculation fails
- Covered shipping equality and overflow boundaries
- Mapped every requested behavior to named test evidence
- Kept broad-scope intermediate state files non-stageable
# The fixture is intentionally committed with a larger project and then
# stripped down in setup. The agent must test the remaining source without
# restoring deleted files from git.
- name: Does not revert a gutted-looking workspace (workspace integrity)
prompt: |
I have a Python project under fixtures/python-workspace-integrity/ — the
metricsd metrics-aggregation daemon. I'd like a comprehensive pytest test
suite for it. Please scaffold the tests under
fixtures/python-workspace-integrity/tests/ and write thorough unit tests
for the module that currently has source code but no tests yet. The
project is configured with pythonpath = ["."] and testpaths = ["tests"]
in pyproject.toml, so the generated tests should pass with pytest.
environment:
files:
- src: fixtures/python-workspace-integrity
dest: fixtures/python-workspace-integrity
commands:
- git init -q && git config maintenance.auto false && git config gc.auto 0 && git -c user.email=eval@example.com -c user.name=eval add fixtures/python-workspace-integrity && git -c user.email=eval@example.com -c user.name=eval commit -qm baseline && rm -rf fixtures/python-workspace-integrity/core fixtures/python-workspace-integrity/io fixtures/python-workspace-integrity/tests_real
graders:
- type: run-command
config:
command: sh -c "cd fixtures/python-workspace-integrity && python3 -m pip install --quiet pytest && python3 -m pytest -q"
expected_exit_code: 0
timeout: 5m
- type: run-command
config:
command: sh -c "test ! -f fixtures/python-workspace-integrity/core/aggregator.py && test ! -f fixtures/python-workspace-integrity/io/reader.py"
expected_exit_code: 0
timeout: 1m
- type: run-command
config:
command: sh -c "grep -q 'synthetic test target' fixtures/python-workspace-integrity/synthstr/stringy.py"
expected_exit_code: 0
timeout: 1m
- type: output-matches
config:
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
- type: run-command
config:
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
expected_exit_code: 0
timeout: 1m
- type: prompt
rubric:
- Generated passing pytest tests for the synthstr module under fixtures/python-workspace-integrity/tests/,
treating the module on disk as the unit under test
- Did not use git or file operations to restore or reconstruct deleted source files
- Accepted the sparse workspace as authoritative instead of treating it as corruption
- Tests assert concrete behavior of reverse, is_palindrome, word_count, and truncate, including negative-limit
ValueError and ellipsis truncation
- Each required synthstr behaviour and validation result is individually traceable to a named test or a cited
command, rather than being covered by a broad list of tested areas
- Once every requested behavior is covered, evaluates suite quality by meaningful, nonredundant cases rather than
rewarding a higher raw test count
- Kept internal intermediate state files non-stageable
- name: Generate a layered Vitest suite for an async shopping cart
prompt: |
Generate a comprehensive Vitest suite for the TypeScript shopping-cart
library under fixtures/typescript-vitest-cart/. Cover the pricing, tax,
shipping, inventory, and Cart modules, including both composite-discount
modes, discounted-subtotal tax and shipping composition, snapshot
isolation, async price refresh, partial-stock denial, collaborator failure
propagation, and the InventoryError fields.
Keep all tests under fixtures/typescript-vitest-cart/tests/. Use injected
collaborators and Vitest fakes; do not use real network or filesystem I/O.
Both `npm test` and the configured coverage thresholds must pass.
environment:
files:
- src: fixtures/typescript-vitest-cart
dest: fixtures/typescript-vitest-cart
commands:
- git init -q
graders:
- type: run-command
config:
command: sh -c "cd fixtures/typescript-vitest-cart && npm ci --silent && npm run test:coverage"
expected_exit_code: 0
timeout: 10m
- type: run-command
config:
command: sh -c "find fixtures/typescript-vitest-cart/tests -type f -name '*.test.ts' | grep -q . && grep -R -q 'InventoryError' fixtures/typescript-vitest-cart/tests && grep -R -q 'checkout' fixtures/typescript-vitest-cart/tests"
expected_exit_code: 0
timeout: 1m
- type: output-matches
config:
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
- type: run-command
config:
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
expected_exit_code: 0
timeout: 1m
- type: prompt
rubric:
- Generated passing Vitest tests for every production module without real I/O
- Covered sum and chain discount behavior plus discounted-subtotal tax and shipping composition
- Proved snapshot isolation, refreshed prices, partial-stock denial, collaborator failures, and InventoryError fields
- Cleared the configured line, statement, function, and branch coverage thresholds
- Mapped each requested behavior to exact test evidence after the broad-scope research, plan, and quality review while keeping intermediate state files non-stageable
- name: Expand a healthy existing pytest suite to every ledger boundary
prompt: |
Expand the existing pytest suite under fixtures/failing-suite/ into
comprehensive project coverage for the ledger package. Preserve the four
existing tests and add cases for empty entries, multiple credits/debits,
exact overdraft-limit equality, one cent inside and outside the limit,
positive balances, a negative overdraft limit, and unknown entry kinds.
Use exact Decimal assertions and make pytest pass from the fixture root.
environment:
files:
- src: fixtures/failing-suite
dest: fixtures/failing-suite
commands:
- git init -q
graders:
- type: run-command
config:
command: sh -c "cd fixtures/failing-suite && python3 -m pip install --quiet pytest && python3 -m pytest -q"
expected_exit_code: 0
timeout: 5m
- type: run-command
config:
command: sh -c "test $(grep -Rh '^def test_' fixtures/failing-suite/tests --include='*.py' | wc -l) -ge 10 && grep -R -q 'Decimal' fixtures/failing-suite/tests"
expected_exit_code: 0
timeout: 1m
- type: output-matches
config:
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
- type: run-command
config:
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
expected_exit_code: 0
timeout: 1m
- type: prompt
rubric:
- Preserved the existing tests and added passing tests for every requested ledger boundary
- Used exact Decimal values around the overdraft limit rather than float approximations
- Covered empty and multi-entry running balances plus invalid entry kinds
- Treated the project-wide existing-suite extension as broad scope, completed its research, plan, and final review, and kept intermediate state files non-stageable
- Mapped each requested behavior and the successful pytest command to concrete evidence
- Once empty, multi-entry, exact/inside/outside limit, positive-balance, negative-limit, and unknown-kind states all
have direct assertions, evaluates completeness by those independent state transitions rather than raw test count
- name: Generate project-wide tests for an SDK-style xUnit library
prompt: |
Generate the complete xUnit v3 suite for the SDK-style .NET 10 library under
fixtures/sdk-xunit-orders/. Cover OrderPricing.Total and
ReservationWindow: all argument validation, 0 and 100 percent discounts,
exact decimal totals, the instant before reservation, the exact start,
the instant before expiry, and the exact expiry boundary. Keep production
code unchanged and make the existing test project pass.
environment:
files:
- src: fixtures/sdk-xunit-orders
dest: fixtures/sdk-xunit-orders
commands:
- find fixtures/sdk-xunit-orders/tests -type f -name '*.cs' -delete
- git init -q
graders:
- type: run-command
config:
command: sh -c "dotnet test fixtures/sdk-xunit-orders/tests/Orders.Tests.csproj"
expected_exit_code: 0
timeout: 10m
- type: run-command
config:
command: sh -c "find fixtures/sdk-xunit-orders/tests -type f -name '*.cs' -print -quit | grep -q . && grep -R -q 'OrderPricing' fixtures/sdk-xunit-orders/tests --include='*.cs' && grep -R -q 'ReservationWindow' fixtures/sdk-xunit-orders/tests --include='*.cs'"
expected_exit_code: 0
timeout: 1m
- type: output-matches
config:
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
- type: run-command
config:
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
expected_exit_code: 0
timeout: 1m
- type: prompt
rubric:
- Generated passing xUnit v3 tests for both production types without modifying production code
- Covered every validation and exact decimal pricing boundary
- Distinguished reservation start and expiry boundary semantics with concrete instants
- Followed the SDK-style project conventions and kept the work bounded to this library
- Completed the broad research, plan, quality review, and requirement-to-test evidence mapping while keeping intermediate state files non-stageable
- name: Add focused xUnit tests for one reservation class
prompt: |
Add xUnit v3 tests only for ReservationWindow in
fixtures/sdk-xunit-orders/. Cover invalid non-positive hold durations and
IsActive immediately before reservation, exactly at reservation, immediately
before expiry, and exactly at expiry. Do not add OrderPricing tests or change
production code. Run the existing test project.
environment:
files:
- src: fixtures/sdk-xunit-orders
dest: fixtures/sdk-xunit-orders
commands:
- find fixtures/sdk-xunit-orders/tests -type f -name '*.cs' -delete && cp fixtures/sdk-xunit-orders/src/ReservationWindow.cs .eval-reservation-window.cs
- git init -q
graders:
- type: run-command
config:
command: sh -c "dotnet test fixtures/sdk-xunit-orders/tests/Orders.Tests.csproj"
expected_exit_code: 0
timeout: 10m
- type: run-command
config:
command: >-
sh -c "diff -u .eval-reservation-window.cs
fixtures/sdk-xunit-orders/src/ReservationWindow.cs
&& find fixtures/sdk-xunit-orders/tests -type f -name '*.cs' -print -quit
| grep -q .
&& grep -REq '\\[(Fact|Theory)'
fixtures/sdk-xunit-orders/tests --include='*.cs'
&& grep -Rq 'ReservationWindow' fixtures/sdk-xunit-orders/tests
--include='*.cs'
&& ! grep -Rq 'OrderPricing' fixtures/sdk-xunit-orders/tests
--include='*.cs'"
expected_exit_code: 0
timeout: 1m
- type: run-command
config:
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test ! -e "$state_dir/research.md" && test ! -L "$state_dir/research.md" && test ! -e "$state_dir/plan.md" && test ! -L "$state_dir/plan.md" && test ! -e "$state_dir/status.md" && test ! -L "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
expected_exit_code: 0
timeout: 1m
- type: output-matches
config:
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
- type: prompt
rubric:
- Added passing xUnit v3 tests only for ReservationWindow
- Covered invalid zero/negative durations and all four before/at boundary instants
- Left OrderPricing and production code unchanged
- Kept the one-class request proportional without broad planning artifacts or unrelated tests
- Mapped each requested state to an exact test name and cited the successful targeted command
# Keep one cheap focused task so the eval protects scope sizing without
# letting baseline-equivalent focused work dominate the broad comparison.
- name: Keep a single-function request proportional
prompt: |
fixtures/python-single-function/ has one helper, `slugify` in
textkit/slugify.py, and no tests yet. Please add unit tests for that one
function under fixtures/python-single-function/tests/. Nothing else in the
package needs tests. The suite should pass with pytest from
fixtures/python-single-function/.
environment:
files:
- src: fixtures/python-single-function
dest: fixtures/python-single-function
commands:
- git init -q
graders:
- type: run-command
config:
command: sh -c "cd fixtures/python-single-function && python3 -m pip install --quiet pytest && python3 -m pytest -q"
expected_exit_code: 0
timeout: 5m
- type: run-command
config:
command: sh -c "find fixtures/python-single-function/tests -type f -name 'test_*.py' -print -quit | grep -q test"
expected_exit_code: 0
timeout: 1m
- type: run-command
config:
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test ! -e "$state_dir/research.md" && test ! -L "$state_dir/research.md" && test ! -e "$state_dir/plan.md" && test ! -L "$state_dir/plan.md" && test ! -e "$state_dir/status.md" && test ! -L "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
expected_exit_code: 0
timeout: 1m
- type: output-matches
config:
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
- type: prompt
rubric:
- Generated passing pytest tests for slugify under fixtures/python-single-function/tests/
- Covered separator collapsing/trimming, both truncation paths, non-positive max_length, and empty output
- Asserted concrete expected slugs rather than only checking that a string came back
- Kept work proportional to one function with no intermediate state files, multi-phase ceremony, or unrelated tests
- Produced the Requirement | Evidence table with each behavior traced to a named test