mirror of
https://github.com/dotnet/skills.git
synced 2026-09-20 09:49:54 +08:00
57733bebc8
* Keep test agent state out of commits Move broad test-generation pipeline state to host scratch storage, worktree-specific Git metadata, or OS temp, and enforce the exclusion in evals. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify absolute test agent state path Use Git's explicit absolute path formatting in both test-generation entry points. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Prune Git metadata from test agent eval guards Avoid scanning nested repositories and align the remaining TESTAGENT_DIR placeholder with the documented format. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify test agent command handoff Require callers to provide exact commands, excerpts, or absolute TESTAGENT_DIR document paths to command-running sub-agents. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Reject all repository-local testagent entries Match .testagent by name regardless of whether it is a directory, file, or symlink while continuing to prune Git metadata. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Verify external test agent artifacts Restore broad-run artifact checks at the Git metadata path and pass the researched lint command and state directory to the linter agent. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Make testagent eval guards constant time Check only the forbidden workspace-root path, including broken symlinks, instead of recursively traversing dependency trees. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Broaden comprehensive test generation Treat explicit requirements as the floor for broad suites and add mutation-relevant equivalence-partition and invariant coverage without test-count padding. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Fix external artifact grader quoting Run state checks directly in the harness shell so TESTAGENT_DIR expands after assignment, with an isolated command probe covering valid and forbidden states. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Run broad skill evals in Git worktrees Initialize the seven broad evaluation roots as Git repositories so TESTAGENT_DIR resolves deterministically and external artifacts remain verifiable. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Clarify non-stageable test agent state Describe the real invariant across the pipeline: state may live under .git metadata but must never be version-controlled workspace content or appear in git status. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Standardize intermediate test state contract Use one TESTAGENT_DIR placeholder, clearer intermediate-state terminology, and detect stageable research, plan, or status files regardless of directory name. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Use one Git root in workspace integrity eval Baseline the fixture from the evaluation root so stageable intermediate-state files remain visible to the directory-independent guard. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Prune Vitest dependencies from state scan Exclude node_modules through per-eval Git metadata so stageable state detection remains fast without modifying fixture content. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 * Strengthen focused intermediate-state guards Separate shell execution, reject Git-metadata files on focused runs, include ignored state files, and prune node_modules with a pathspec exclusion. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0 --------- Copilot-Session: 35c50c03-2dda-4919-981e-fd5f6b7938f0
495 lines
27 KiB
YAML
495 lines
27 KiB
YAML
name: code-testing-agent
|
|
description: Evaluates the dotnet-test/code-testing-agent skill
|
|
type: capability
|
|
executionShard: generation
|
|
# `defaults:` replaces the deprecated `config:` — vally's loader throws on a
|
|
# spec declaring both (eng/eval-quality check 10).
|
|
#
|
|
# Run 31180075357 measured 2W/10T/0L. The two focused tasks never activated and
|
|
# added four baseline-equivalent ties. Seven broad tasks below exercise the
|
|
# skill's distinctive Research -> Plan -> Implement contract across Python,
|
|
# TypeScript, Go, SDK/classic .NET, existing-suite extension, and a sparse
|
|
# workspace. Two focused tasks protect proportional routing in different
|
|
# ecosystems. Nine independent stimuli improve power without changing runs.
|
|
defaults:
|
|
timeout: 60m
|
|
runs: 2
|
|
stimuli:
|
|
- name: Generate a project-wide pytest suite across multiple modules
|
|
prompt: |
|
|
Generate a comprehensive pytest suite for the entire project under
|
|
fixtures/python-multimodule/. It has no tests yet and contains:
|
|
|
|
- analytics.stats.mean and percentile
|
|
- analytics.window.RateWindow
|
|
- textkit.slug.slugify
|
|
|
|
Cover every public behavior across the modules, including empty-input and
|
|
percentile-range errors, percentile ordering and 0/100 boundaries,
|
|
RateWindow constructor validation, empty-window errors, capacity rollover,
|
|
average and peak after rollover, slug separator collapsing/trimming,
|
|
max-length truncation, and invalid max_length. Keep this project-wide:
|
|
tests should live under fixtures/python-multimodule/tests/ and pass with
|
|
pytest from fixtures/python-multimodule/.
|
|
environment:
|
|
files:
|
|
- src: fixtures/python-multimodule
|
|
dest: fixtures/python-multimodule
|
|
commands:
|
|
- git init -q
|
|
graders:
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "cd fixtures/python-multimodule && python3 -m pip install --quiet pytest && python3 -m pytest -q"
|
|
expected_exit_code: 0
|
|
timeout: 5m
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "grep -R -q 'RateWindow' fixtures/python-multimodule/tests && grep -R -q 'percentile' fixtures/python-multimodule/tests && grep -R -q 'slugify' fixtures/python-multimodule/tests"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: output-matches
|
|
config:
|
|
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
|
|
- type: run-command
|
|
config:
|
|
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: prompt
|
|
rubric:
|
|
- Generated passing tests under the configured tests directory for all three modules
|
|
- Covered every named validation and boundary behavior, including RateWindow rollover and percentile 0/100 boundaries
|
|
- Asserted concrete averages, peaks, percentiles, and slugs rather than existence-only results
|
|
- Kept the work project-wide but bounded to the three supplied modules
|
|
- Mapped each requested behavior to named test evidence
|
|
- Kept broad-scope intermediate state files non-stageable so only requested deliverables remain commit candidates
|
|
|
|
- name: Generate project-wide tests for a classic MSTest library
|
|
prompt: |
|
|
Add the missing project-wide unit tests for the classic net472 library
|
|
under fixtures/classic-mstest/. The production project contains
|
|
DiscountService and TieredDiscountPolicy; the existing test project uses
|
|
MSTest 3.5.2, Moq 4.2, NBuilder, FixtureBase<TSut>, packages.config, and
|
|
explicit compile items.
|
|
|
|
Preserve that stack and the existing DiscountServiceTests.cs byte-for-byte.
|
|
Create DiscountServiceBoundaryTests.cs and TieredDiscountPolicyTests.cs.
|
|
Cover DiscountService's 0/100 boundaries, invalid percentages, missing
|
|
product, and exact discounted value. Cover TieredDiscountPolicy constructor
|
|
validation, negative subtotal, below-threshold behavior, exact-threshold
|
|
activation, and a concrete discounted value. Do not modernize the project
|
|
or dependencies.
|
|
environment:
|
|
files:
|
|
- src: fixtures/classic-mstest
|
|
dest: fixtures/classic-mstest
|
|
commands:
|
|
- rm -f fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs && cp fixtures/classic-mstest/tests/Discounts.Tests.csproj.pristine fixtures/classic-mstest/tests/Discounts.Tests.csproj && rm fixtures/classic-mstest/tests/Discounts.Tests.csproj.pristine && mkdir -p .eval-baseline && cp fixtures/classic-mstest/tests/DiscountServiceTests.cs .eval-baseline/DiscountServiceTests.cs && cp fixtures/classic-mstest/tests/packages.config .eval-baseline/packages.config
|
|
- git init -q
|
|
graders:
|
|
- type: file-exists
|
|
config:
|
|
path: fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
|
|
- type: file-exists
|
|
config:
|
|
path: fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs
|
|
- type: run-command
|
|
config:
|
|
command: >-
|
|
sh -c 'diff -u .eval-baseline/DiscountServiceTests.cs
|
|
fixtures/classic-mstest/tests/DiscountServiceTests.cs
|
|
&& diff -u .eval-baseline/packages.config
|
|
fixtures/classic-mstest/tests/packages.config
|
|
&& test $(grep -c "Compile Include=\"DiscountServiceBoundaryTests.cs\""
|
|
fixtures/classic-mstest/tests/Discounts.Tests.csproj) -eq 1
|
|
&& test $(grep -c "Compile Include=\"TieredDiscountPolicyTests.cs\""
|
|
fixtures/classic-mstest/tests/Discounts.Tests.csproj) -eq 1
|
|
&& for file in
|
|
fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
|
|
fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs;
|
|
do grep -Eq "Assert[[:space:]]*\.[[:space:]]*ThrowsException[[:space:]]*<" "$file"
|
|
&& ! grep -Eq "Assert[[:space:]]*\.[[:space:]]*Throws(Exactly)?[[:space:]]*<" "$file"
|
|
&& ! grep -Eq "\[[[:space:]]*ExpectedException" "$file"
|
|
|| exit 1;
|
|
done'
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: file-not-contains
|
|
config:
|
|
path: fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
|
|
value: ThrowsExactly
|
|
- type: file-not-contains
|
|
config:
|
|
path: fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
|
|
value: Assert.Throws<
|
|
- type: file-contains
|
|
config:
|
|
path: fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
|
|
value: Assert.ThrowsException<
|
|
- type: file-not-contains
|
|
config:
|
|
path: fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs
|
|
value: ThrowsExactly
|
|
- type: file-not-contains
|
|
config:
|
|
path: fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs
|
|
value: Assert.Throws<
|
|
- type: file-contains
|
|
config:
|
|
path: fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs
|
|
value: Assert.ThrowsException<
|
|
- type: file-not-contains
|
|
config:
|
|
path: fixtures/classic-mstest/tests/DiscountServiceBoundaryTests.cs
|
|
value: ExpectedException
|
|
- type: file-not-contains
|
|
config:
|
|
path: fixtures/classic-mstest/tests/TieredDiscountPolicyTests.cs
|
|
value: ExpectedException
|
|
- type: output-matches
|
|
config:
|
|
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
|
|
- type: run-command
|
|
config:
|
|
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: prompt
|
|
rubric:
|
|
- Added both requested test files and registered each exactly once in the classic project
|
|
- Preserved the existing test file, packages.config, project format, and pinned dependency versions
|
|
- Used Assert.ThrowsException<T> for exception paths, avoiding both Assert.Throws<T> and Assert.ThrowsExactly<T>, which are unavailable in MSTest 3.5.2
|
|
- Covered every named DiscountService and TieredDiscountPolicy boundary with concrete expected values
|
|
- Mapped each requested behavior and build-registration requirement to evidence
|
|
- Kept broad-scope intermediate state files non-stageable without modernizing the classic project
|
|
|
|
- name: Generate a project-wide Go suite across collaborating packages
|
|
prompt: |
|
|
Generate comprehensive Go tests for the whole module under
|
|
fixtures/go-multipackage/. Cover money.Discount, shipping.Rate, and
|
|
order.Total.
|
|
|
|
Use table-driven tests for validation and numeric boundaries. Verify
|
|
discount 0/100 boundaries, invalid subtotal and percent, shipping bracket
|
|
equality (a weight equal to UpToGrams belongs to that bracket), overflow
|
|
behavior, negative weight, and empty brackets. For order.Total, use fakes
|
|
for both interfaces and prove the composition `subtotal - discount +
|
|
shipping` with exact cents. Also cover error propagation from each
|
|
collaborator and verify the second collaborator is not called when the
|
|
first fails. Keep tests inside the corresponding Go packages and make
|
|
`go test ./...` pass from fixtures/go-multipackage/.
|
|
environment:
|
|
files:
|
|
- src: fixtures/go-multipackage
|
|
dest: fixtures/go-multipackage
|
|
commands:
|
|
- git init -q
|
|
graders:
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "cd fixtures/go-multipackage && go test ./..."
|
|
expected_exit_code: 0
|
|
timeout: 5m
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "test -f fixtures/go-multipackage/money/discount_test.go && test -f fixtures/go-multipackage/shipping/rate_test.go && test -f fixtures/go-multipackage/order/total_test.go && grep -q 'UpToGrams' fixtures/go-multipackage/shipping/rate_test.go"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: output-matches
|
|
config:
|
|
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
|
|
- type: run-command
|
|
config:
|
|
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: prompt
|
|
rubric:
|
|
- Generated passing tests in all three packages
|
|
- Used table-driven tests for validation and boundary cases
|
|
- Proved the exact order total composition with collaborator fakes and concrete cents
|
|
- Covered each collaborator error and verified short-circuit behavior when discount calculation fails
|
|
- Covered shipping equality and overflow boundaries
|
|
- Mapped every requested behavior to named test evidence
|
|
- Kept broad-scope intermediate state files non-stageable
|
|
|
|
# The fixture is intentionally committed with a larger project and then
|
|
# stripped down in setup. The agent must test the remaining source without
|
|
# restoring deleted files from git.
|
|
- name: Does not revert a gutted-looking workspace (workspace integrity)
|
|
prompt: |
|
|
I have a Python project under fixtures/python-workspace-integrity/ — the
|
|
metricsd metrics-aggregation daemon. I'd like a comprehensive pytest test
|
|
suite for it. Please scaffold the tests under
|
|
fixtures/python-workspace-integrity/tests/ and write thorough unit tests
|
|
for the module that currently has source code but no tests yet. The
|
|
project is configured with pythonpath = ["."] and testpaths = ["tests"]
|
|
in pyproject.toml, so the generated tests should pass with pytest.
|
|
environment:
|
|
files:
|
|
- src: fixtures/python-workspace-integrity
|
|
dest: fixtures/python-workspace-integrity
|
|
commands:
|
|
- git init -q && git config maintenance.auto false && git config gc.auto 0 && git -c user.email=eval@example.com -c user.name=eval add fixtures/python-workspace-integrity && git -c user.email=eval@example.com -c user.name=eval commit -qm baseline && rm -rf fixtures/python-workspace-integrity/core fixtures/python-workspace-integrity/io fixtures/python-workspace-integrity/tests_real
|
|
graders:
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "cd fixtures/python-workspace-integrity && python3 -m pip install --quiet pytest && python3 -m pytest -q"
|
|
expected_exit_code: 0
|
|
timeout: 5m
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "test ! -f fixtures/python-workspace-integrity/core/aggregator.py && test ! -f fixtures/python-workspace-integrity/io/reader.py"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "grep -q 'synthetic test target' fixtures/python-workspace-integrity/synthstr/stringy.py"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: output-matches
|
|
config:
|
|
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
|
|
- type: run-command
|
|
config:
|
|
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: prompt
|
|
rubric:
|
|
- Generated passing pytest tests for the synthstr module under fixtures/python-workspace-integrity/tests/,
|
|
treating the module on disk as the unit under test
|
|
- Did not use git or file operations to restore or reconstruct deleted source files
|
|
- Accepted the sparse workspace as authoritative instead of treating it as corruption
|
|
- Tests assert concrete behavior of reverse, is_palindrome, word_count, and truncate, including negative-limit
|
|
ValueError and ellipsis truncation
|
|
- Each required synthstr behaviour and validation result is individually traceable to a named test or a cited
|
|
command, rather than being covered by a broad list of tested areas
|
|
- Once every requested behavior is covered, evaluates suite quality by meaningful, nonredundant cases rather than
|
|
rewarding a higher raw test count
|
|
- Kept internal intermediate state files non-stageable
|
|
|
|
- name: Generate a layered Vitest suite for an async shopping cart
|
|
prompt: |
|
|
Generate a comprehensive Vitest suite for the TypeScript shopping-cart
|
|
library under fixtures/typescript-vitest-cart/. Cover the pricing, tax,
|
|
shipping, inventory, and Cart modules, including both composite-discount
|
|
modes, discounted-subtotal tax and shipping composition, snapshot
|
|
isolation, async price refresh, partial-stock denial, collaborator failure
|
|
propagation, and the InventoryError fields.
|
|
|
|
Keep all tests under fixtures/typescript-vitest-cart/tests/. Use injected
|
|
collaborators and Vitest fakes; do not use real network or filesystem I/O.
|
|
Both `npm test` and the configured coverage thresholds must pass.
|
|
environment:
|
|
files:
|
|
- src: fixtures/typescript-vitest-cart
|
|
dest: fixtures/typescript-vitest-cart
|
|
commands:
|
|
- git init -q
|
|
graders:
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "cd fixtures/typescript-vitest-cart && npm ci --silent && npm run test:coverage"
|
|
expected_exit_code: 0
|
|
timeout: 10m
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "find fixtures/typescript-vitest-cart/tests -type f -name '*.test.ts' | grep -q . && grep -R -q 'InventoryError' fixtures/typescript-vitest-cart/tests && grep -R -q 'checkout' fixtures/typescript-vitest-cart/tests"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: output-matches
|
|
config:
|
|
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
|
|
- type: run-command
|
|
config:
|
|
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: prompt
|
|
rubric:
|
|
- Generated passing Vitest tests for every production module without real I/O
|
|
- Covered sum and chain discount behavior plus discounted-subtotal tax and shipping composition
|
|
- Proved snapshot isolation, refreshed prices, partial-stock denial, collaborator failures, and InventoryError fields
|
|
- Cleared the configured line, statement, function, and branch coverage thresholds
|
|
- Mapped each requested behavior to exact test evidence after the broad-scope research, plan, and quality review while keeping intermediate state files non-stageable
|
|
|
|
- name: Expand a healthy existing pytest suite to every ledger boundary
|
|
prompt: |
|
|
Expand the existing pytest suite under fixtures/failing-suite/ into
|
|
comprehensive project coverage for the ledger package. Preserve the four
|
|
existing tests and add cases for empty entries, multiple credits/debits,
|
|
exact overdraft-limit equality, one cent inside and outside the limit,
|
|
positive balances, a negative overdraft limit, and unknown entry kinds.
|
|
Use exact Decimal assertions and make pytest pass from the fixture root.
|
|
environment:
|
|
files:
|
|
- src: fixtures/failing-suite
|
|
dest: fixtures/failing-suite
|
|
commands:
|
|
- git init -q
|
|
graders:
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "cd fixtures/failing-suite && python3 -m pip install --quiet pytest && python3 -m pytest -q"
|
|
expected_exit_code: 0
|
|
timeout: 5m
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "test $(grep -Rh '^def test_' fixtures/failing-suite/tests --include='*.py' | wc -l) -ge 10 && grep -R -q 'Decimal' fixtures/failing-suite/tests"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: output-matches
|
|
config:
|
|
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
|
|
- type: run-command
|
|
config:
|
|
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: prompt
|
|
rubric:
|
|
- Preserved the existing tests and added passing tests for every requested ledger boundary
|
|
- Used exact Decimal values around the overdraft limit rather than float approximations
|
|
- Covered empty and multi-entry running balances plus invalid entry kinds
|
|
- Treated the project-wide existing-suite extension as broad scope, completed its research, plan, and final review, and kept intermediate state files non-stageable
|
|
- Mapped each requested behavior and the successful pytest command to concrete evidence
|
|
- Once empty, multi-entry, exact/inside/outside limit, positive-balance, negative-limit, and unknown-kind states all
|
|
have direct assertions, evaluates completeness by those independent state transitions rather than raw test count
|
|
|
|
- name: Generate project-wide tests for an SDK-style xUnit library
|
|
prompt: |
|
|
Generate the complete xUnit v3 suite for the SDK-style .NET 10 library under
|
|
fixtures/sdk-xunit-orders/. Cover OrderPricing.Total and
|
|
ReservationWindow: all argument validation, 0 and 100 percent discounts,
|
|
exact decimal totals, the instant before reservation, the exact start,
|
|
the instant before expiry, and the exact expiry boundary. Keep production
|
|
code unchanged and make the existing test project pass.
|
|
environment:
|
|
files:
|
|
- src: fixtures/sdk-xunit-orders
|
|
dest: fixtures/sdk-xunit-orders
|
|
commands:
|
|
- find fixtures/sdk-xunit-orders/tests -type f -name '*.cs' -delete
|
|
- git init -q
|
|
graders:
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "dotnet test fixtures/sdk-xunit-orders/tests/Orders.Tests.csproj"
|
|
expected_exit_code: 0
|
|
timeout: 10m
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "find fixtures/sdk-xunit-orders/tests -type f -name '*.cs' -print -quit | grep -q . && grep -R -q 'OrderPricing' fixtures/sdk-xunit-orders/tests --include='*.cs' && grep -R -q 'ReservationWindow' fixtures/sdk-xunit-orders/tests --include='*.cs'"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: output-matches
|
|
config:
|
|
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
|
|
- type: run-command
|
|
config:
|
|
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test -f "$state_dir/research.md" && test -f "$state_dir/plan.md" && test -f "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: prompt
|
|
rubric:
|
|
- Generated passing xUnit v3 tests for both production types without modifying production code
|
|
- Covered every validation and exact decimal pricing boundary
|
|
- Distinguished reservation start and expiry boundary semantics with concrete instants
|
|
- Followed the SDK-style project conventions and kept the work bounded to this library
|
|
- Completed the broad research, plan, quality review, and requirement-to-test evidence mapping while keeping intermediate state files non-stageable
|
|
|
|
- name: Add focused xUnit tests for one reservation class
|
|
prompt: |
|
|
Add xUnit v3 tests only for ReservationWindow in
|
|
fixtures/sdk-xunit-orders/. Cover invalid non-positive hold durations and
|
|
IsActive immediately before reservation, exactly at reservation, immediately
|
|
before expiry, and exactly at expiry. Do not add OrderPricing tests or change
|
|
production code. Run the existing test project.
|
|
environment:
|
|
files:
|
|
- src: fixtures/sdk-xunit-orders
|
|
dest: fixtures/sdk-xunit-orders
|
|
commands:
|
|
- find fixtures/sdk-xunit-orders/tests -type f -name '*.cs' -delete && cp fixtures/sdk-xunit-orders/src/ReservationWindow.cs .eval-reservation-window.cs
|
|
- git init -q
|
|
graders:
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "dotnet test fixtures/sdk-xunit-orders/tests/Orders.Tests.csproj"
|
|
expected_exit_code: 0
|
|
timeout: 10m
|
|
- type: run-command
|
|
config:
|
|
command: >-
|
|
sh -c "diff -u .eval-reservation-window.cs
|
|
fixtures/sdk-xunit-orders/src/ReservationWindow.cs
|
|
&& find fixtures/sdk-xunit-orders/tests -type f -name '*.cs' -print -quit
|
|
| grep -q .
|
|
&& grep -REq '\\[(Fact|Theory)'
|
|
fixtures/sdk-xunit-orders/tests --include='*.cs'
|
|
&& grep -Rq 'ReservationWindow' fixtures/sdk-xunit-orders/tests
|
|
--include='*.cs'
|
|
&& ! grep -Rq 'OrderPricing' fixtures/sdk-xunit-orders/tests
|
|
--include='*.cs'"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: run-command
|
|
config:
|
|
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test ! -e "$state_dir/research.md" && test ! -L "$state_dir/research.md" && test ! -e "$state_dir/plan.md" && test ! -L "$state_dir/plan.md" && test ! -e "$state_dir/status.md" && test ! -L "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: output-matches
|
|
config:
|
|
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
|
|
- type: prompt
|
|
rubric:
|
|
- Added passing xUnit v3 tests only for ReservationWindow
|
|
- Covered invalid zero/negative durations and all four before/at boundary instants
|
|
- Left OrderPricing and production code unchanged
|
|
- Kept the one-class request proportional without broad planning artifacts or unrelated tests
|
|
- Mapped each requested state to an exact test name and cited the successful targeted command
|
|
|
|
# Keep one cheap focused task so the eval protects scope sizing without
|
|
# letting baseline-equivalent focused work dominate the broad comparison.
|
|
- name: Keep a single-function request proportional
|
|
prompt: |
|
|
fixtures/python-single-function/ has one helper, `slugify` in
|
|
textkit/slugify.py, and no tests yet. Please add unit tests for that one
|
|
function under fixtures/python-single-function/tests/. Nothing else in the
|
|
package needs tests. The suite should pass with pytest from
|
|
fixtures/python-single-function/.
|
|
environment:
|
|
files:
|
|
- src: fixtures/python-single-function
|
|
dest: fixtures/python-single-function
|
|
commands:
|
|
- git init -q
|
|
graders:
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "cd fixtures/python-single-function && python3 -m pip install --quiet pytest && python3 -m pytest -q"
|
|
expected_exit_code: 0
|
|
timeout: 5m
|
|
- type: run-command
|
|
config:
|
|
command: sh -c "find fixtures/python-single-function/tests -type f -name 'test_*.py' -print -quit | grep -q test"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: run-command
|
|
config:
|
|
command: state_dir="$(git rev-parse --path-format=absolute --git-path testagent)" && test ! -e "$state_dir/research.md" && test ! -L "$state_dir/research.md" && test ! -e "$state_dir/plan.md" && test ! -L "$state_dir/plan.md" && test ! -e "$state_dir/status.md" && test ! -L "$state_dir/status.md" && test -z "$(git ls-files --cached --others -- ':(glob)**/research.md' ':(glob)**/plan.md' ':(glob)**/status.md' ':(exclude,glob)**/node_modules/**')"
|
|
expected_exit_code: 0
|
|
timeout: 1m
|
|
- type: output-matches
|
|
config:
|
|
pattern: \|\s*Requirement\s*\|\s*Evidence\s*\|
|
|
- type: prompt
|
|
rubric:
|
|
- Generated passing pytest tests for slugify under fixtures/python-single-function/tests/
|
|
- Covered separator collapsing/trimming, both truncation paths, non-positive max_length, and empty output
|
|
- Asserted concrete expected slugs rather than only checking that a string came back
|
|
- Kept work proportional to one function with no intermediate state files, multi-phase ceremony, or unrelated tests
|
|
- Produced the Requirement | Evidence table with each behavior traced to a named test
|