mirror of
https://github.com/max-sixty/worktrunk.git
synced 2026-09-14 20:00:38 +08:00
codex/remove-codex-cloud-specific-tests
24 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1b278042de |
chore(ci): weekly renovation 2026-08-16 (#3826)
## Summary Weekly CI renovation check found the following updates: - `worktrunk`: 0.72.0 → 0.74.0 (MSRV 1.96, compatible with our 1.96.0) — `ci.yaml` ×2, `nightly.yaml` - `nushell`: 0.114.1 → 0.115.0 — `nightly.yaml`, `benchmarks.yaml`, `coverage.yaml`, `actions/test-setup`, and `scripts/codex-cloud/Taskfile.yaml` - `pre-commit`: 4.6.1 → 4.6.2 — `scripts/codex-cloud/Taskfile.yaml` - `PowerShell`: 7.6.4 → 7.6.5 — `scripts/codex-cloud/Taskfile.yaml` and the root `Taskfile.yaml`'s `setup-web` task The Codex Cloud archive checksums were recomputed from the new upstream tarballs, and the resulting `Taskfile.yaml` digest (`f14dbc89…`) is copied into both README launcher commands. The `setup-web` PowerShell pin came in as a follow-up commit: the initial sweep only grepped `.rs`/`.md`/`.toml` for stale versions, so the root `Taskfile.yaml`'s `PWSH_VERSION="7.6.4"` was missed. Nothing tests the two PowerShell pins against each other, so that one drifts silently — worth a note for future renovation runs. The `powershell_7.6.5-1.deb_amd64.deb` asset the `setup-web` branch downloads is present in the v7.6.5 release. ## Already up to date - Rust stable is 1.97.1, so MSRV and toolchain stay at 1.96 (latest stable − 1) — `Cargo.toml`, `tests/helpers/wt-perf/Cargo.toml`, `rust-toolchain.toml` need no change, and `flake.lock` is untouched. - `cargo-insta` 1.48.0, `cargo-nextest` 0.9.143, `cargo-llvm-cov` 0.8.7, `cargo-msrv` 0.19.3, `cargo-affected` 0.4.0, `cargo-udeps` 0.1.61, `lychee` 0.24.2 - Task 3.52.0 (mise, Codex Cloud) - Runner images: ubuntu-24.04, macos-15, windows-2022 ## Held back: zola 0.22.1 → 0.23.3 Not bumped. Zola 0.23.0 shipped [Tera2 + refactoring](https://github.com/getzola/zola/pull/3105), which is a templating-engine swap rather than a routine release. Building `docs/` with the 0.23.3 binary fails at the first line of `templates/base.html`: ``` ERROR error: Unknown tag --> base.html:1:4 | 1 | {% import "macros.html" as macros %} | ^^^^^^ ``` `templates/base.html` and `templates/macros.html` are the two files that use the `import`/`macro` pair, so the migration looks small, but it is template work with its own review rather than a pin bump — kept out of this PR so the rest can land. Raised separately. <details><summary>Verification</summary> - Every version above was read from the upstream source of truth: `crates.io` for the cargo tools, `nushell/nushell` and `PowerShell/PowerShell` releases, PyPI for pre-commit, and `static.rust-lang.org/dist/channel-rust-stable.toml` for Rust stable (1.97.1). - Checksums were computed from the downloaded archives and the extracted binaries were run (`nu --version` → `0.115.0`); the archive layouts (`nu-<ver>-x86_64-unknown-linux-gnu/nu`, top-level `pwsh`) are unchanged, so the `install_binary` paths still resolve. - All six edited YAML files parse. - The nushell bump was exercised against the shell-integration suite: `cargo test --features shell-integration-tests --test integration -- nushell` with 0.115.0 on `PATH`. 13 of 14 pass; `test_nushell_install_target_is_a_vendor_autoload_dir` fails — but it fails identically on the currently-pinned 0.114.1, and passes on *both* versions when run alone. It is a pre-existing shared-state race in the sandbox, not a regression from this bump: the test asserts against the real user `$nu.vendor-autoload-dirs` entry rather than one under its temp `HOME` (nu resolves the home dir from the passwd database, so the test's `HOME` override does not move it), and a sibling uninstall test in the same filter removes `wt.nu` from that shared directory. Noted rather than fixed here — it is unrelated to the pins. - The zola failure above was reproduced with the official 0.23.3 `x86_64-unknown-linux-gnu` release binary against this repo's `docs/`. </details> --------- Co-authored-by: worktrunk-bot <254187624+worktrunk-bot@users.noreply.github.com> |
||
|
|
18b408ce98 |
Add shared Codex Cloud environment setup (#3810)
## Summary - add a repository-owned Codex Cloud Taskfile, exposed through root setup and maintenance tasks - share setup and maintenance preparation in one task instead of two scripts - document concise, checksum-gated environment commands - preserve the proven UID 1000, `tini`, pinned-tool, and retry behavior - keep tool pins, archive checksums, Task version, and launcher digests synchronized by test and maintenance guidance ## Why Worktrunk's full suite needs dependencies and process/permission semantics beyond the stock universal image. The working configuration previously lived only in one saved environment, where other contributors could neither review nor reuse it. The dedicated Taskfile sits beside the project Taskfile without making unrelated task edits invalidate the Cloud environment hash. Root wrapper tasks launch it as a new Task process so its repository-relative paths retain their own Taskfile context. ## Security Codex checks out the task branch before setup or maintenance. Each saved launcher verifies the dedicated Taskfile's fixed SHA-256 digest before executing it as root. `MISE_NO_CONFIG=1` also prevents branch-controlled mise configuration from running before the verified Taskfile. Approved changes require updating the digest in environment settings, which invalidates the cache. Repository-sensitive Rustup, pre-commit, and Cargo work runs as the image's UID 1000 `ubuntu` user; root is limited to the verified system and ownership preparation. ## Validation - root and direct Taskfile discovery expose `setup-codex` and `maintain-codex` - YAML parsing and extracted Bash syntax pass - warning-level ShellCheck passes - applicable pre-commit hooks pass - positive and negative checksum checks pass - the launcher-sync integration test passes - independent adversarial, abstraction-level, and current-head code reviews are clean - exact Taskfile validation passed in Codex Cloud task `task_e_6a7e53a87e908325bf50ea3413ed521c` - setup and maintenance launchers passed - root and direct Task discovery passed - Cargo identity probe returned UID 1000 - `cargo run -- hook pre-merge --yes` exited 0 - 4,601/4,601 tests passed; all gate components passed - HEAD remained `288953dcb18466f64b8e5355192ae297f4862240` and the checkout remained clean - final-head Cloud task `task_e_6a8089197ecc8325afd723d41beb5c50` reached `READY` with no diff after about 38 minutes on `dab4a5dcd5f343c3e5ea0a1183e9fcc5d8271a08`; its transcript was unavailable, so no finer-grained result is claimed - all 18 applicable final-head checks pass on Linux, macOS, and Windows, including coverage, `codecov/patch`, and current-head tend review > _This was written by Codex on behalf of @max-sixty_ |
||
|
|
4d7f8c944e |
fix(env): drop Intel macOS from the flake, and install pwsh and jq for web setup (#3776)
Three follow-ups from #3768, plus a bug the verification turned up. **The flake stops declaring outputs for `x86_64-darwin`.** nixpkgs drops Intel macOS in 26.11: evaluating anything for that system against `nixos-unstable` (`26.11pre-git`) throws. The rev `flake.lock` pins is 26.05-era, so it still evaluates today, carrying nixpkgs' own warning that "26.05 will be the last release to support x86_64-darwin". Naming three systems rather than `eachDefaultSystem` drops Nix support for Intel Macs now, ahead of that bump. Release binaries are untouched: `dist-workspace.toml` still ships `x86_64-apple-darwin` and nightly still tests it on `macos-15-intel`. **`git` leaves the devShell's `packages`.** It arrives with the `checks`, which crane folds in via `inputsFrom`, the same mechanism that already supplied `python3`, `procps` and `lsof`. **`task setup-web` installs `pwsh` and `jq`.** Without them Claude Code web can't run `--features shell-integration-tests`, which is what the pre-merge gate runs. PowerShell comes from the release `.deb` rather than the tarball, because pwsh aborts at startup without libicu and only the `.deb` declares that dependency for apt to resolve. The verification loop runs each tool instead of looking for it on PATH, since the tarball install left a `pwsh` that was on PATH and still aborted. **A `set -e` abort found while testing that.** The `sources.list.d` cleanup was an `&&` chain, and under `set -e` a chain ending false takes the whole task down. This one ends false on an unmatched glob and on a `.list` file with no `[` line, so setup was dying before it installed anything on a stock Debian box as well as an empty one. It's an `if` now. ## Verification No `nix` on the machine this was written on, so the flake was checked in a `nixos/nix` container and the Taskfile block in an amd64 Debian one. <details> <summary>flake: three systems evaluate, x86_64-darwin is gone, git survives its deletion</summary> ``` == devShell evaluates per system == x86_64-linux OK g172vwl0g339zsxx9l6mz5pca6w9jbcx-nix-shell.drv aarch64-linux OK pgvq71zs48bx3naddncms954jyqpl0bl-nix-shell.drv aarch64-darwin OK d17q1772si0x0hj1lgpnin8wiq4zlpr2-nix-shell.drv x86_64-darwin FAIL: flake does not provide attribute 'devShells.x86_64-darwin.default' == systems the flake declares == ["aarch64-darwin","aarch64-linux","x86_64-linux"] == tools in the x86_64-linux devShell == git: present jq: present nushell: present powershell: present python3: present procps: present lsof: present fish: present zsh: present bash: present gh: present pre-commit: present == nixfmt --check flake.nix == clean (exit 0) ``` The x86_64-darwin claim, checked against nixpkgs directly rather than inferred: ``` == nixos-unstable lib.version == "26.11pre-git" == x86_64-darwin eval on nixos-unstable == error, pointing at release-notes#x86_64-darwin-26.11 == x86_64-darwin eval on the pinned rev (flake.lock) == evaluation warning: Nixpkgs 26.05 will be the last release to support x86_64-darwin "hello-2.12.3" ``` Not verified: nothing was built, only evaluated. The nightly `nix-flake` job runs `nix flake check` on PRs touching `flake.nix`, which covers that on x86_64-linux. </details> <details> <summary>setup-web: the block run under Task's own interpreter, in an amd64 Debian container</summary> The edited block was extracted into a minimal Taskfile and run by `task` itself, so mvdan/sh parses it rather than bash. `curl` and nushell are container prereqs, not part of what's under test. ``` === running the extracted block under Task === Installing shell-integration test dependencies... pwsh installed bash available zsh available fish available nu available pwsh available jq available task exit: 0 === does the installed pwsh actually run? === 7.6.4 jq-1.6 /usr/bin/pwsh === rerun is idempotent === Installing shell-integration test dependencies... bash available zsh available fish available nu available pwsh available jq available ``` Two earlier runs are why the shape changed. The first died at the `sources.list.d` glob. The second installed PowerShell from the release tarball: every tool reported "available" and `pwsh` then aborted with `Couldn't find a valid ICU package installed on the system`, which is what moved the install to the `.deb` and the check from `command -v` to `--version`. </details> `cargo run -- hook pre-merge --yes` passes: 4574 tests, 1 skipped. ## Notes `task setup-web` still requires nushell to be present rather than installing it, unchanged here. > _This was written by Claude Code on behalf of max-sixty_ --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dd453304ed |
test: fold the mock stub into the wt binary (#3712)
`cargo test --test integration` neither built nor rebuilt `mock-stub`: a
target filter deselects the dummy test that pulled it in, so a fresh
tree panicked ("mock-stub binary not found") and a warm one could run a
stale stub. The chain that compensated — the separate helper package,
its dummy `builds.rs`, the `default-members` entry, nextest's
experimental `build-bins` setup script, `workspace_bin()` — existed only
because cargo-dist ships every `[[bin]]`, and carried a TODO to collapse
once that changed. dist 0.30.2 does support per-binary exclusion now
(`[dist.binaries]`, since 0.29.0; verified with `dist plan`), but the
TODO's plan has a hole it predates: `cargo install worktrunk` installs
every feature-satisfied `[[bin]]`, and dist config doesn't govern
crates.io installs.
So the mock commands are now the `wt` binary itself. `mock_commands`
links `wt` under the mock's name (`gh`, `glab`, …), and `main()`
dispatches to the ported playback (`testing::mock_stub`) when
`WORKTRUNK_TEST_MOCK_CONFIG_DIR` is set, argv[0] is a foreign name, and
the config dir holds `<argv0>.json` for it. The existence check keeps wt
under a foreign argv[0] *without* a config being wt — the
argv0-validation security test symlinks it as `wt;touch` and must reach
wt's own rejection. The shipped binary already compiles the whole
`testing` module unconditionally, so this adds no new category of test
code to it. Windows links `wt.exe` with `hard_link` (copy fallback for
cross-drive dirs); the debug binary is ~67 MB, so per-mock copies stay
the fallback.
Every runner is now safe by construction — cargo rebuilds a package's
own binaries whenever its integration tests build, so there is no
separate artifact to go missing or stale. Deleted: the helper package,
the setup script plus `experimental = ["setup-scripts"]`, the
`default-members` trick, `workspace_bin()`, and `wt_bin()`'s dead
compile-time branch (unit-test targets get neither the runtime variable
nor the `option_env!` value, so the runtime resolution is the one
mechanism).
Validated locally: the pre-merge gate's full suite passes (4542/4542;
its doc step also caught unescaped `argv[0]` intra-doc links in the new
comments, fixed and `cargo doc -Dwarnings` re-verified). With all stub
artifacts purged from `target/`, plain `cargo test --test integration`
on mock-dependent tests builds them and passes — the command that used
to hit the trap. Two integration tests pin the dispatch's argv[0] edges
(an empty argv[0], and a non-UTF8 one), alongside the existing
`wt;touch` carve-out test.
The first commit is the investigation that preceded the fix: it verified
the `wt` binary itself was never subject to the staleness the mock-stub
was, and documented that in `tests/CLAUDE.md`; the fix then narrows that
paragraph further, since the gap it scoped no longer exists.
A two-reviewer subagent round (one prosecuting the diff against the
failure modes documented in the repo's own mock history — the #401/#407
Windows era, #547, #654, #127, #2544, #2730, #2744 — the other
adversarial) then hardened the dispatch. The reserved-name guard is
case-insensitive, matching the config probe, which goes through a
filesystem that equates `WT.json` with `wt.json` on macOS and Windows;
`command_name()` reads `args_os` — `env::args()` panics on a non-Unicode
argument, and this runs inside `main()` on every invocation (caught by
the tend review) — and returns `None` for a degenerate argv[0] instead
of panicking; the `.exe` suffix is stripped explicitly rather than via
`file_stem`, so a dotted mock name (`python3.11`) resolves identically
on every platform; and `copy_mock_binary` is now private —
`MockConfig::write` writes `<name>.json` before linking and is the only
way to create a mock, so a link cannot exist without its config, and the
dispatch's missing-config fall-through can only mean "wt under a foreign
name" (the argv0-validation tests' `wt;touch`), never a half-configured
mock that silently runs real wt with the mocked tool's arguments. The
`Option<&str>` mock helpers whose `None` arm produced exactly such
configless links lost the arm (every caller passed `Some`), and 25
redundant standalone link calls went with it.
The review also surfaced the one remaining spawn-a-stale-binary path
outside the suite: `wt-perf timeline` resolved a sibling `wt` by path,
checked only existence, and told the user to build it manually — so
`cargo run -p wt-perf -- timeline` after a `src/` edit silently measured
stale code. It now builds `wt` first and takes the artifact path from
cargo's `--message-format=json` report rather than deriving a sibling
location, so target-dir and profile overrides can't divert the build
away from where it's resolved; a release wt-perf builds and measures a
release wt. The build runs before the timeline's wall-clock measurement
starts, cargo's progress streams on stderr, and stdout keeps the
`--chrome` JSON contract.
> _This was written by Claude Code on behalf of max-sixty_
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
4554f50ce6 |
perf(tests): stop leaking a temp dir per test, and measure where the suite's CPU goes (#3604)
Measured where the test suite's CPU actually goes, fixed what was doing
real extra work, and added `task profile-tests` so the measurement is
repeatable.
## Baseline
`cargo nextest run --features shell-integration-tests` on an 18-core
M-series machine: 4,570 tests, ~95s wall, **988 CPU-seconds (332 user +
655 sys)**.
Two thirds is kernel time, so the cost is process creation and
filesystem churn rather than computation. The integration binary is 86%
of summed self-time: 2,184 tests at ~0.6s each, each spawning `wt` (a 65
MB debug binary, ~11ms CPU per spawn against 2.6ms for a trivial
process) and `git` against a fresh fixture copy. It sits in a broad
middle, not a few outliers: 73% of self-time is in tests taking 0.25 to
2.0s.
## One leaked temp directory per test
`isolated_test_cwd()` held a `TempDir` in a `LazyLock`. Statics don't
run destructors at process exit and nextest runs one process per test,
so every test leaked an empty directory into the system temp root.
Measured with `TMPDIR` pointed at a fresh directory: **704 per
integration-suite run**. This machine had accumulated **454,907 entries,
353,268 of them empty strays older than a day**.
Stale entries are cheap to ignore but expensive to enumerate, and
`git::recover::recover_from_path` reads every ancestor directory of a
deleted CWD:
| temp root | `test_recover_from_path_returns_none_for_unrelated_path` |
|---|---|
| 454k entries | 14.2s (34.4s on a quieter run) |
| empty | 0.27s |
One fixed directory replaces it. Leaks per run: 704 to 0, verified
across full suite runs, and the directory is still empty after ~9,000
test executions.
I also checked whether a crowded temp root slows ordinary temp
operations. It does not: create, populate and delete of a fixture-sized
tree ran at 21ms/iter in a 454k-entry parent against 40 to 60ms in an
empty one. The leak's cost is concentrated entirely in code that
enumerates.
## Fixture temp dirs no longer sit in the shared temp root
The fixtures created their temp directories directly in the system temp
dir, among however many entries the machine had put there.
`test_temp_root()` (`$TMPDIR/wt`) roots them one level in, and
`test_tempdir()` replaces `TempDir::new()` across every `TestRepo`
constructor, the mock-command helper, the `temp_home` fixture and the
recovery tests. That test again: **14.2s to 0.06s**.
Two constraints worth knowing, both found by trying the more aggressive
version first:
- **A cache-dir root fails.** `~/Library/Caches` is under `/Users/`,
which a conditional `includeIf "gitdir:/Users/"` matches, so 16 picker
tests fail their commits under `commit.gpgsign` — they drive git through
`Repository::run_command`, the production API with no isolated config.
`step_promote` above was not a one-off: the suite was hermetic against
host git config only because macOS puts `$TMPDIR` outside `/Users/`.
That is the hole the next section closes — at the layer that covers
in-process git, not by choosing where temp files live.
- **The root's name is load-bearing.** A unix socket path cannot exceed
`sun_path`, 104 bytes on macOS. The canonicalized per-user `$TMPDIR` is
56, and `test_copy_ignored_skips_non_regular_files` binds a listener 89
bytes in. `worktrunk-tests` (16 bytes with its slash) overflowed by two;
`wt` costs 3 of the 14 spare.
The ~240 tests that call `tempfile` directly still use the system temp
dir. They are transient and a clean run leaks nothing, so converting
them is tidiness rather than a fix.
## A test that passed by accident of TMPDIR location
`step_promote::test_promote_bare_repo_with_worktrees` drove git through
bare `Cmd::new("git")` instead of `configure_git_env`, so the host's
config applied. A conditional `includeIf "gitdir:/Users/"` enabling
`commit.gpgsign` fails its commit, but only when `TMPDIR` sits inside
the matched tree. macOS puts `TMPDIR` under `/var/folders`, so it
passed; pointing `TMPDIR` anywhere under the home directory failed it.
## The suite was not actually isolated from the developer's git config
Chasing the temp-root change turned up a real hole. `TestRepo` exposes
the production `Repository` type, and `Repository::run_command` builds a
plain `Cmd::new("git")` with no `GIT_CONFIG_GLOBAL`, so it inherits the
test process's own environment. Separately, a bare `wt_command()` had no
`GIT_CONFIG_GLOBAL` at all and fell through to `~/.gitconfig`. 280
`Repository::at/current/discover` constructions across 31 test files and
156 direct `run_command*` calls in test code sat on that path.
Signing was only the symptom that surfaced. Confirmed leaking from a
real developer config: `commit.gpgsign`, `core.fsmonitor` (spawns a
daemon per fixture repo), `worktree.guessremote` (changes `git worktree
add`, directly under test), `help.autocorrect = prompt` (a mistyped git
command blocks). Structurally: `core.hooksPath`, `credential.helper`,
`filter.*` clean/smudge and `diff.external` all execute arbitrary
programs; `url.*.insteadOf` rewrites remotes; `merge.conflictstyle`,
`diff.context`, `rebase.autostash`, `fetch.prune`, `push.default` all
change what the code under test observes. That set is unbounded, which
is why hardening each fixture's local config was rejected — a denylist
can't cover it, and local config cannot unset an inherited `[include]`
or `credential.helper` at all.
The floor is one constant, `shell_exec::HERMETIC_TEST_GIT_ENV` — the
deny pair pointing `GIT_CONFIG_GLOBAL` and `GIT_CONFIG_SYSTEM` at a path
that does not exist, plus the settings the suite needs through
`GIT_CONFIG_COUNT` / `GIT_CONFIG_KEY_n` / `GIT_CONFIG_VALUE_n` — applied
to every child at its spawn site. There is no git-config file anywhere
in the repo. For the spawn sites the harness owns (`git_test_env`,
`configure_git_cmd`, `isolate_subprocess_env` for `wt` children,
`pty_env_vars` for `env_clear`ed PTY children) that's ordinary per-child
env. For the git that *production* code spawns while a test drives it
in-process, the test never holds the command — so the harness flips an
atomic latch (`shell_exec::enable_hermetic_test_env`, called from the
fixture constructors), and `Cmd`, the choke point every production spawn
passes through, applies the floor to each child while the latch is set.
Setting env on a child is safe; it was setting the test process's *own*
env that wasn't (`set_var` races the other test threads), and the latch
dissolves the need for it — no `unsafe`, no pre-`main` constructor, no
cargo `[env]`. Because the latch lives in the binary, every runner
agrees by construction: `cargo test`, nextest, `cargo llvm-cov`, `cargo
bench`, an IDE, a debugger, a directly executed
`target/debug/deps/integration-*`. `.config/nextest.toml` was tested and
rejected, and is now ruled out standing: nextest 0.9.132 has no `[env]`
key, and a `$NEXTEST_ENV` setup script would miss the three non-nextest
runners CI uses (`cargo llvm-cov`, the Nix `cargo test` derivation,
`cargo bench`). A runner-specific knob doesn't fail loudly when another
runner misses it — it yields a different result, usually in the coverage
job whose numbers gate a merge. `tests/CLAUDE.md` → One Result Per Test,
Whatever Runs It records the rule; `.config/nextest.toml` points at it
from the place someone would be tempted.
Acceptance test, since the suite passed before only by accident of
`$TMPDIR` sitting outside `/Users/`: with `test_temp_root()` temporarily
pointed under `$HOME` so a conditional `includeIf "gitdir:/Users/"`
fires, the picker tests go from **20 of 38 failing to 38 of 38
passing**.
**The cost:** the latch is a test-serving switch compiled into
`shell_exec` — one static, one relaxed load per spawn, marked
`TODO(hermetic-env)` with the structural alternative (threading an
explicit env value through `Repository`). `cargo run -- <cmd>` is
untouched: nothing in production latches it, so a developer's own
invocations keep their aliases, credential helper, and identity. The one
production git spawn that bypasses `Cmd` (the fsmonitor daemon launch)
re-applies the floor by hand. (Two earlier shapes were tried and
replaced: cargo's `[env]`, which taxed every `cargo run` and vanished
whenever a test binary ran outside cargo, and a pre-`main` constructor
crate, which every test target had to link and which put the floor on
developers' `cargo bench`-adjacent runs too.)
## In-process tests were reading the developer's worktrunk config too
`config_path()`'s third priority is the real
`~/.config/worktrunk/config.toml`, and a lib-crate test cannot set the
second for itself — `set_var` is `unsafe` and this crate forbids
`unsafe`. So a test reaching priority 3 got the developer's own config.
A `panic!` build of the guard named the live callers immediately:
`git::repository::tests::prewarm_*` read it on every run, and
`set_skip_shell_integration_prompt` /
`set_skip_commit_generation_prompt` reach the same resolver to
**write**.
Priority 3 is now absent under `#[cfg(test)]`, which is compiled out of
the real binary — so unlike the git floor above, this one costs `cargo
run` nothing. It returns `None` rather than panicking like
`approvals_path()`: that guards a mutation target where silent absence
would let a test believe it saved something, whereas this is a lookup
whose absent state is already handled — `require_config_path()` turns it
into an error, so a write still fails loudly while a best-effort read
preloads nothing.
The guard covers lib-crate tests only; `src/commands/` and `src/output/`
link the lib in non-test mode. Nothing there exercises the fall-through
today (968 bin-crate tests create no config under a scratch `$HOME`), so
that is a requirement on new tests, recorded in `tests/CLAUDE.md`.
`system_config_path()` stays unguarded on purpose — machine-wide file,
and `config::deprecation`'s `PendingDefault` rules need the lookup.
## Merging main's parallel fix
#3620 attacked the same in-process hole from the other side, writing
`LOCAL_TEST_CONFIG` into each fixture's own `.git/config`. The two
compose rather than compete and both are kept: the floor denies the
host's config to every git, and the local config supplies what a
hermetic in-process git still needs — an identity, which the floor
deliberately withholds because it has per-command homes already
(`git_test_env`, `LOCAL_TEST_CONFIG`) and `useConfigOnly` fails loudly
if a path misses both. main's structure (`TestConfigPaths`,
`TestRepo::bare`) is kept as-is.
Two docstrings were true on each side and false together:
`test_gitconfig_path` restated the gitconfig inline, and
`LOCAL_TEST_CONFIG` said in-process git reads the developer's
`~/.gitconfig`, which is what the floor prevents. `test_gitconfig_path`
also held its `TempDir` in a `LazyLock`, the leak this PR removes. Both
are moot now: the function is gone with the files.
## Why there is no gitconfig file
The isolation went through two file-based shapes before this one, and
neither earned its keep. Denial never needed a file, because
`GIT_CONFIG_GLOBAL` and `GIT_CONFIG_SYSTEM` *are* the denial; the file
existed only to *set* things, and `GIT_CONFIG_COUNT` is git's
environment spelling of `-c`. Deleting both files also deletes the
`[include]` that kept them from drifting, the per-process write that
broke `test (windows)` on a shared path, and the config-path argument
threaded through 22 call sites.
The floor that remains is the deny pair plus two settings.
`user.useConfigOnly` is a backstop: denial alone leaves git *guessing*
an identity from the OS username and hostname rather than failing, which
is the one way a hermetic suite could still author a commit as the
developer. Nothing exercises it, and that is the reason to keep it.
`rerere.enabled = false` is *set* rather than left unset, so the suite's
rerere state cannot depend on what a fixture happens to carry.
`commit.gpgsign`, `advice.mergeConflict` and `advice.resolveConflict`
are gone: denial leaves git on its own default for the first, and the
snapshot layer strips the gutter-prefixed `hint:` lines the other two
quieted (they vary across git versions), so nothing depends on
suppressing them at the source. The long-dead
`tests/fixtures/template-repo/` fixture went with them.
Once the latch made denial universal, the redundant copies went too:
`git_test_env` no longer restates the deny pair per command (the floor
is denial's only writer, so its value is uniform across every
transport), `LOCAL_TEST_CONFIG` dropped its `commit.gpgsign` (denial
guarantees the default), the platform-dependent `NULL_DEVICE` constant
is deleted, and the `.env.GIT_CONFIG_GLOBAL` snapshot redaction is gone
— the recorded value is one cross-platform constant, so there is nothing
volatile to redact.
I removed `rerere.enabled` first, on a local measurement that was wrong,
and CI failed on all three platforms. The standard fixture is built once
into `target/debug/wt-test-fixtures/` and copied per test, and that
cached copy held an `rr-cache` directory left by a rebase run while the
floor still enabled rerere. Git turns rerere on by itself whenever
`rr-cache` exists, so every local test kept the behavior the change had
just removed, while CI built the fixture fresh and lost it.
`tests/CLAUDE.md` now records the trap: clear the fixture cache before
trusting a local measurement of a git-config change.
**Why the floor can't live in a fixture:** every other test variable is
set on a *child* — `git_test_env` on a git command,
`configure_cli_command` on a `wt` subprocess. In-process git is not a
child the test configures: `Repository::run_command` is production code
building a plain `Cmd::new("git")`, and the test never holds that
command, so there is no place to set env on it — while setting the test
process's own environment is the one thing a test can't do safely
(`set_var` races the other test threads). The identity did move to the
fixtures and the per-command env this way; the denial reaches
production's children through the latch at `Cmd`, the choke point they
all pass through.
Two things fall out of `-c` semantics, both pinned:
- **It outranks a repository's own config**, where a global file would
yield to it. So `init.defaultBranch` cannot live in the floor:
`default_branch.rs` sets that key in a repo to prove `wt` reads it, and
an entry would silently win. The three harness `git init` calls that
relied on the floor now name their branch, as the other two already did.
- **A PTY child is `env_clear`ed**, so it inherits nothing and used to
get the floor through the file. It now gets the family from
`configure_pty_command`, the choke point every PTY spawn routes through,
and again from `pty_env_vars`, whose vector declares a PTY `wt` child's
complete environment. Each copy is pinned by its own test, because the
settings only quiet advice and refuse a guessed identity, so no PTY
assertion would catch their loss.
Four tests wrote their own gitconfig to get `init.defaultBranch` plus an
identity; the harness supplies both, so those writes are gone too. Net
45 lines lighter, and no snapshot changed.
## Measurement
`task profile-tests` builds first, then runs the suite under bash's
`time` keyword (task's own interpreter, mvdan/sh, parses `time` but
hardcodes `user`/`sys` to zero): CPU totals on the console, every
per-test duration in the default profile's `junit.xml`. It began three
sizes larger — a scratch-`TMPDIR` leak check that dragged a `sun_path`
byte budget into the Taskfile, a `/usr/bin/time` dependency GNU-less
Linux lacks, and a `perf` nextest profile whose console slow-listing
restated what junit already carries — and each piece fell to the same
question, whether the measuring goal needed it. Method and how to read
the numbers: `tests/CLAUDE.md` under Profiling the Suite.
## Found, measured, not changed
**The gate keeps a duplicate cargo artifact set.** `RUSTFLAGS='-D
warnings'` on the insta step is part of cargo's fingerprint, so it forks
all 343 crates into a second artifact set: 261 CPU-seconds to prime plus
a duplicate of `target/debug/deps` (`target/debug` here is 35 GB, in a
52 GB `target/`). I removed it and then put it back: the clippy step
that would cover it runs on ubuntu only, while the cross-platform matrix
runs `wt hook pre-merge --yes insta`, so this RUSTFLAGS is the only
thing denying warnings on macOS and Windows. `[lints.rust] warnings =
"deny"` would keep that coverage everywhere without forking the graph
(verified locally: fails on a planted unused variable, recompiles only
worktrunk and wt-perf, feature-check commands still pass), but it also
makes plain local builds fail on warnings and newly exposes `cargo msrv
verify` and the minimal-versions job. That is a workflow call. The
duplication is a one-time cost per artifact set rather than per-edit, so
it is second-order next to the ~600 CPU-seconds each suite run costs.
**`recover_from_path` enumerates every ancestor up to `/`.** At each
ancestor it reads the directory and stats `.git` in every child.
Bounding the child scan to the first *existing* ancestor would preserve
both documented layouts (sibling and nested) and both of that module's
regression tests, but it would break a custom `worktree-path` layout
where the repo is a child of a higher ancestor, so it needs a decision
about which layouts recovery must support.
**Nothing in the gate is the biggest cost; concurrency is.** The five
`[[pre-merge]]` keys are one table, so they run concurrently, and each
is a cargo command that wants the whole machine. They serialize on
cargo's build-directory lock (`Blocking waiting for file lock on build
directory` appears in every run) while test execution overlaps another
step's build. The `lockfile` comment says it "must be first", which
concurrent execution does not provide. Separately, agent worktrees run
whole gates at once: during this work a second worktree ran its own `wt
hook pre-merge` alongside mine, load average hit **154 on 18 cores**,
and the same suite took 147.7s instead of 92.9s.
## Tried and rejected
`[profile.dev] debug = "line-tables-only"` first measured 14% less CPU,
but per-spawn CPU was unchanged, which did not fit the proposed
mechanism. Re-measuring both configurations on a quiet machine gave
602.6s (baseline) against 606.5s (line-tables). The original delta was
contention from a sibling worktree running its own suite.
macOS Gatekeeper (`syspolicyd`) looked like a candidate at 230% CPU, but
300 spawns of the freshly built `wt` cost it 0s, and 0.2s after a
relink, against 0.4s during a 10s idle baseline.
## Verification
`cargo run -- hook pre-merge --yes` green, 4,611 tests passed, run
against a freshly built fixture cache after the switch to the latch. The
latch was verified to be the only source of the variables: the invoking
shell carries no `GIT_CONFIG_*`, and the meta-tests
(`in_process_git_reads_only_the_hermetic_config` asserting the *origin*
of every resolved setting, `pty_env_vars_carry_the_git_config_floor`,
and the `isolate_subprocess_env` scrub test asserting the floor is
re-set after the scrub) pin each transport. Across earlier runs one hit
a single intermittent PTY failure in `shell_wrapper` (exit 127) that
passes 3/3 in isolation and whose code path never touches `wt_command()`
or the shared cwd; a later run was green under heavier load than the one
that failed.
## Review round
A full review of the branch (three finder lenses, findings adversarially
verified) landed one more commit:
- **The wrapper-suite PTY children never got the floor.**
`configure_pty_command` env-clears, skips the `Cmd` latch, and sets the
real `HOME`, and the shell-wrapper call sites layer only fixture paths
and an identity on top. Every git those ~89 tests ran therefore read the
developer's real `~/.gitconfig` (and lost the gpgsign shield when
`LOCAL_TEST_CONFIG` dropped it). The floor now rides that choke point,
pinned by `configure_pty_command_carries_the_git_config_floor`; every
raw `CommandBuilder` site was checked to route through it.
- **`task profile-tests` could never report CPU.** mvdan/sh's `time`
hardcodes `user`/`sys` to `0m0.000s`, so the numbers the docs said to
track were unproducible; now `bash -c 'time "$@"' bash ...`. Measured
post-fix: user 5m47s / sys 12m8s, a 67.7% kernel share, confirming the
two-thirds claim above; the integration mean measured ~1s and the docs
were corrected from ~0.6s.
- Smaller: two raw `Command::new("git")` asserts in `remove.rs` tests
now go through `configure_git_cmd`; the hermetic meta-test keys on
`--show-scope` scopes rather than git's origin-path spelling; the dead
`template-repo` fixture is deleted; three test `git init`s name `-b
main`; doc corrections (sun_path arithmetic, recover-walk attribution,
dead `[TEST_GIT_CONFIG]` remnants, volatile counts).
Final gate on the finished tree: green, 4,612 tests. Deferred with
rationale: ~985 snapshots carry stale env-block metadata (insta never
compares it; it churns on future re-records), `$TMPDIR/wt` is not
per-user on Linux (a root run poisons it for later users), and
`spawn_detached_exec_*` / `step tether` do not hand-apply the floor (no
in-process test reaches them today).
> _This was written by Claude Code on behalf of Maximilian_
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
1b35950a9c |
test: simplify the integration suite (#3657)
This reduces duplicated and false-confidence integration coverage while preserving the suite's semantic and user-facing contracts. ## What changed - Replaces three overlapping list-layout suites with two representative CLI integrations, leaving exhaustive geometry at the direct layout layer. - Groups Git error render variants into labeled family snapshots and removes command-by-shell wrapper cross-products while retaining shell-specific conformance and regression cases. - Updates test-authoring guidance around boundary choice, minimal contrasts, and PTY use, and runs local and CI coverage through Nextest isolation. ## Results - Test catalog: 4,642 to 4,562 - Snapshots: 1,193 to 1,131 - Warm all-feature runtime: 84.70s to 78.17-80.35s - Full coverage: 97.32% of lines ## Testing - `cargo run -- hook pre-merge --yes` - `cargo llvm-cov nextest --features shell-integration-tests --summary-only` > _This was written by Claude Code on behalf of max_. |
||
|
|
9cb128328f |
refactor(tests): deny network transports for in-process git, consolidate fixtures (#3620)
An in-process `Repository::run_command()` spawns git with the test process's own environment, so none of the harness env isolation reaches it — no `GIT_CONFIG_GLOBAL`, and crucially no `GIT_ALLOW_PROTOCOL`, meaning a unit test driving the library directly could make a real network connect (reproduced: `ls-remote` against a fixture-style `https://` URL attempted the wire). The repo's own config is the one layer such a spawn still reads, so every fixture repo now carries `LOCAL_TEST_CONFIG` (identity, `commit.gpgsign`, `protocol.allow = never` with a `file` exception), written by a single `TestRepo::assemble` that every constructor routes through. The consolidation on top of that fix: - **`git_test_env()`** is the one home for the nine git-isolation settings; `configure_git_cmd` (Command), `configure_git_env` (`Cmd`), and `pty_env_vars` (PTY) all consume it. Composes with #3618's PTY layering — the env doc in `tests/CLAUDE.md` now reads four layers. - **The test gitconfig is one process-static file** (`test_gitconfig_path()`); the per-fixture copies (including two hand-rolled, drifted ones in `bare_repository.rs`) are gone. - **`BareRepoTest` and `NestedBareRepoTest` compose over `TestRepo`** (new `TestRepo::bare_at()`), so they inherit `assemble`'s guarantees structurally. Their public APIs are unchanged, deliberately — ~220 call-site references and the bare-repo snapshots are written against them. - **The committed `tests/fixtures/standard/` is deleted** (83 files), along with its Taskfile generator and the `_git` rename/sed dance. `build_standard_fixture` constructs the template once per `target/<profile>` (private-dir build + atomic rename, so parallel nextest processes race benignly) and `copy_standard_fixture` keeps only copy + per-copy gitdir rewrite. Two review-worthy subtleties in the fixture rebuild: - The old fixture's commits were stamped `-08:00` (the generator machine's timezone) and **~170 snapshots embed the resulting commit SHAs**, so the builder pins that timestamp as an explicit constant and `test_standard_fixture_template_reproduces_pinned_shas` asserts the four exact SHAs — the template reproduces the committed fixture bit-for-bit. - The old fixture's index files carried an untracked-cache extension from the generator machine's config, which made `git status` print `warning: untracked cache is disabled…` into exactly two `for_each` snapshots. The clean-room template drops it; those two snapshots are regenerated here (that's the whole content diff, plus documented-cosmetic env-block lines). Tested by the full pre-merge gate including `shell-integration-tests` (4583 tests) plus a post-merge-with-main revalidation. > _This was written by Claude Code on behalf of max_ --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4cf3006041 | fix(taskfile): drop codex from bench list without leaving an empty element (#3611) | ||
|
|
c532d2dabd | docs(llm-commits): bump Codex commit-generation model to gpt-5.6-luna (#3430) | ||
|
|
16deb5e6c4 | fix: use --safe-mode for Claude commit.generation to preserve apiKeyHelper auth (#3170) | ||
|
|
958b9c188a |
refactor(picker): migrate to skim 4.8 (ratatui), drop vendored skim-tuikit (#3137)
Migrates the `wt switch` picker from skim 0.20.5 (tuikit backend) to
skim 4.8.0 (ratatui/crossterm), removes the `vendor/skim-tuikit/` patch
tree we carried against the old line, and makes the full test suite
green under the new backend.
## Why the vendor tree goes away
We vendored skim-tuikit for two patches: `alt-<digit>` key parsing and
an `Output::flush` `write_all` fix for dropped bytes under PTY pressure.
Both are moot in 4.8. Key parsing is native — `binds.rs` routes every
single character (digits included) through `KeyCode::Char` rather than a
hardcoded `alt-a..alt-z` arm list. The flush fix is subsumed by stdlib
`BufWriter::flush` in the ratatui backend, which loops over partial
writes correctly. So the migration carries zero vendor patches against
skim, and the build machinery that kept the vendored source alive
(Taskfile `vendor-diff`, flake `vendorSrc`/`extraDummyScript`) is
deleted too.
## The picker rewrite
skim 4.x is a different API. `WorktreeSkimItem::display()` now returns
`ratatui::text::Line` instead of an `AnsiString`. alt-r removal rides
skim's `reload(remove {})` token (the selected row's `output()`,
expanded into the reload command) instead of the old
`as_any().downcast_ref` path (which never worked across compilation
units) plus a signal file. Preview-tab switching and the progressive
list redraw are driven by `Skim::event_sender()` +
`Event::Render`/`Event::RunPreview`.
## The regression that motivated the bulk of this
skim 4.x renders on demand. There is no 100ms tuikit heartbeat
re-rendering the frame, so the progressive `wt switch` list rendered
blank while collection ran in the background — the original report. The
fix pokes `Event::Render` from the list-collect callbacks (throttled to
16ms). The same on-demand model surfaced four more regressions, all
fixed here and verified head-to-head against the 0.20 picker in a
terminal: preview-tab lag, current-row highlight, Shift-Tab back-cycle
(crossterm reports Shift-Tab under three distinct `KeyEvent` shapes, so
all three are bound), and alt-r removal (4.x `execute-silent` is
fire-and-forget and raced its reader).
## Help text and the test harness
Two things the integration suite caught that the unit tests did not (the
prior pass ran `--lib --bins` only):
- **Styled help.** skim 0.20 transitively pulled in
`clap/unstable-markdown`, which renders worktrunk's `///` doc-comments
as styled help. skim 4 with `default-features = false` drops it, so help
reverted to raw text (`[experimental]` → `\[experimental\]`). Restored
by depending on `unstable-markdown` explicitly, keeping help output
identical and the `test_help` / `test_step_alias` /
`test_docs_are_in_sync` snapshots passing unchanged.
- **PTY harness.** skim 4.x queries cursor position (`ESC[6n`) at
startup in partial-height mode and blocks in `select()` for the reply;
`portable_pty` is a bare PTY and never answered, so every
`switch_picker` test failed init with "Cursor position detection timed
out." The Unix harness now answers the query, mirroring the existing
ConPTY responder. skim 4.x also draws the list/preview separator one
column left, so the snapshot panel-split columns shifted to match. The
regenerated picker snapshots show two cosmetic changes: the match
counter no longer overlaps the preview tab header, and the HEAD column
shows the full short-SHA instead of a truncated one.
## Minimal-versions CI
The skim 4.8 dep tree pins tighter floors than our manifests declared,
so the nightly `minimal-versions` job needed updating. It now runs `-Z
direct-minimal-versions` — minimize only our own direct deps and let
transitive crates resolve normally — and the manifests raise each
under-specified floor to the minimum the workspace builds against
(largely mirroring skim 4.8's own requirements). Full `-Z
minimal-versions` would instead drag in skim's transitive TUI/image
stack (ansi-to-tui, ratatui's `instability` macro, color-eyre,
ratatui-image → image/avif), whose crates under-declare their floors and
don't compile at the picked versions; direct minimization confines the
check to floors we own, so no transitive pins are needed. Normal
resolution (the committed `Cargo.lock`) is unchanged. Full floor list
and the `signal-hook` libc-pin detail are in the `ci(min-versions)`
commit message.
## Reviewing
Start at `src/commands/picker/mod.rs` (the `run_skim` entry point, the
action keybinds, and `parse_reload_remove_token`), then
`progressive_handler.rs` (the render pokes) and `items.rs` (the
`Line`-based `display()`). Test-harness changes are in
`tests/common/pty.rs` (the `ESC[6n` responder) and
`tests/integration_tests/switch_picker.rs` (the panel-split columns).
The rest is the dependency swap, the vendor and build-machinery
deletions, and regenerated snapshots.
All 3980 tests pass (`cargo run -- hook pre-merge --yes`), including the
`shell-integration-tests` PTY suite that exercises the picker end-to-end
across the list, previews, scroll, create/remove, and accept flows. No
automated test drives a real terminal; the interactive surface was also
checked by hand against the 0.20 picker.
> _This was written by Claude Code on behalf of max_
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
7d16e63fad |
refactor(commit-generation): drop the CLAUDECODE= nesting workaround (#2979)
worktrunk stripped `CLAUDECODE` before spawning the
`[commit.generation]` LLM command, and the recommended Claude Code
command carried a leading `CLAUDECODE=`, both to get past Claude Code's
nested-session check that rejected `claude -p` launched from inside
another Claude Code session.
That check is gone. It's absent from the `claude` binary across every
current build (2.1.157 through 2.1.162), and `CLAUDECODE=1 claude -p …`
runs cleanly (exit 0, empty stderr). So the workaround is no longer
needed.
This drops it: `src/llm.rs` no longer calls `.env_remove("CLAUDECODE")`,
and the recommended command loses the `CLAUDECODE=` prefix in the source
of truth (`dev/config.example.toml`), the `wt config --help` text, and
the Taskfile bench command. Docs, the skill reference, and the config
help snapshots are regenerated to match.
Caveat: there's no published guarantee the nested-session check stays
gone, and a Claude Code old enough to still have it (older than roughly
late May 2026) would again block nested commit generation. Since the
check is absent from the implementation across all current builds, that
risk is low.
Tests: 1236 lib + 693 integration pass; fmt and clippy clean.
> _This was written by Claude Code on behalf of max_
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
d65d5028fa | docs(commit-generation): bump Codex model to gpt-5.4-mini (#2949) | ||
|
|
33bab089bf |
docs(taskfile): fix ONGs typo in build-social-cards desc (#2947)
|
||
|
|
eac94579d3 |
fix(picker): vendor skim-tuikit with write_all in Output::flush (#2226)
## Summary Fixes intermittent partial-first-render in `wt switch`'s picker — under tmux PTY pressure, rows past the first ~1024 bytes silently vanished. Root cause is in `skim-tuikit`: `Output::flush` calls `stdout.write(&self.buffer)`, which `io::Write::write` is allowed to satisfy with a short write. The unwritten remainder is then discarded by the immediate `self.buffer.clear()`. Instrumenting the flush made it obvious: ``` flush: buf=8464 wrote=1024 dropped=7440 flush: buf=3382 wrote=1024 dropped=2358 flush: buf=6096 wrote=1024 dropped=5072 ``` Fix is a one-line `write` → `write_all`. Carried as `[patch.crates-io]` pointing at `vendor/skim-tuikit/`, which is byte-identical to `skim-tuikit` 0.6.6 from crates.io except that single line and an added `LICENSE` file (the crates.io tarball omits it; MIT requires the notice). Upstream PR: [skim-rs/skim#1056](https://github.com/skim-rs/skim/pull/1056). Drop the vendor once released. ## Notes - `task vendor-diff` prints the exact patch we carry — makes it easy to audit. - `.pre-commit-config.yaml` and `.typos.toml` extended to skip `vendor/` so upstream source isn't linted against our style. - `skim-rs/tuikit` (the older pre-rename repo) has the same bug but is archived; upstream fix targets `skim-rs/skim @ feat/dynamic-header` where the current `skim-tuikit` 0.6.x crate actually lives. ## Test plan - [x] `pre-commit run --all-files` passes - [x] Reproduced original partial render in tmux (intermittent, ~50% of runs with 14 items at cold start) - [x] Post-fix: 15/15 reliable full renders in the same repro - [ ] CI 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
4ef64c47ad |
feat(opencode): add OpenCode integration (activity tracking, plugin, config) (#1807)
This PR adds OpenCode integration: activity tracking markers in `wt list`, plugin installation via CLI, `wt config show` diagnostics, and LLM commit generation detection. Continues the work started in #1295 and #1533 which added OpenCode to the docs and example config. ## What's included - **Activity tracking plugin** (`dev/opencode-plugin.ts`): maps `session.status`/`session.idle`/`session.deleted` to branch markers, same pattern as Claude Code's plugin - **`wt config plugins opencode install/uninstall`**: installs the plugin to `~/.config/opencode/plugins/worktrunk.ts` — source embedded via `include_str!()`, no npm needed. Sits under `wt config plugins` alongside Claude Code's `wt config plugins claude`. - **`wt config show` OPENCODE section**: shows plugin install status with actionable hints (only when `opencode` is on PATH) - **`LlmTool::OpenCode` variant**: detected via PATH for commit generation auto-config ## Docs approach Kept deliberately low-profile — no dedicated docs page, no README mention. OpenCode is discoverable via `wt config show` and a mention in tips-patterns. If it becomes popular, docs prominence can increase. > _This was written by Claude Code on behalf of @max-sixty_ --------- Co-authored-by: Maximilian Roos <m@maxroos.com> Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
5a2895cd03 |
Add LLM tool commands to example config as single source of truth (#1531)
The Claude/Codex recommended commands were duplicated across 4 locations (Rust code, llm-commits.md, config.md, Taskfile) with no automated sync — they were already drifting (missing `CLAUDECODE=` prefix, missing `-c system_prompt=''`). Now the double-commented entries in `config.example.toml` are the single source of truth. `recommended_config()` parses them at runtime via `include_str!` + `LazyLock`, and a new sync test verifies `llm-commits.md` matches. Also fixes the existing drift in all copies. > _This was written by Claude Code on behalf of @max-sixty_ --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
9b77ab2d93 |
fix: add --no-session-persistence to Claude commit generation (#1454)
Without this flag, `claude -p` saves the commit generation conversation to session history. When the user later runs `claude --continue`, it resumes that ephemeral commit session instead of their previous working session. > _This was written by Claude Code on behalf of @max-sixty_ --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
c7128ce11b | Add nushell support (#964) | ||
|
|
e154b9ee50 |
Add LLM setup prompt for first-time commit configuration (#867)
* Add LLM setup prompt for first-time commit configuration Implement one-time interactive prompt when users attempt commit/merge/squash without LLM configuration. Detects available tools (claude, codex) and offers auto-configuration with preview on `?`. Adds `skip-commit-generation-prompt` flag to suppress re-prompting, `set_commit_generation_command()` config method, and reusable `prompt_yes_no_preview()` utility. Updates commit message templates to clarify format requirements, fixes Claude Code command syntax (spaces to equals in flags), and adds visual styling to section headers in documentation. * Update snapshots for template format changes Co-Authored-By: Claude <noreply@anthropic.com> * Simplify commit template format guidance Remove redundant trivial-changes line from template format section. Fix duplicate "# Other" header from merge. Document skip-commit-generation-prompt in first-run prompts section alongside skip-shell-integration-prompt. Co-Authored-By: Claude <noreply@anthropic.com> * Add unit tests for command detection functions Cover command_exists() and detect_llm_tool() with basic unit tests to improve test coverage on the commit generation module. Co-Authored-By: Claude <noreply@anthropic.com> * Add PTY tests for commit generation prompt Test the interactive prompt flow for LLM commit configuration: - No tool found → sets skip flag - User declines → sets skip flag - User accepts → saves config - User requests preview → shows preview Uses fake claude script to test the prompt path that requires a detected LLM tool. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
cbe65a65ea |
docs: add Codex config and optimize Claude Code for LLM commits (#837)
* added verbose flag Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com> * docs: add Codex config and optimize Claude Code for LLM commits Add OpenAI Codex CLI as an option for LLM commit message generation. Configure with gpt-5.1-codex-mini model, low reasoning effort, and read-only sandbox. Uses jq to parse JSON output. Update Claude Code config with optimization flags that disable tools, skills, settings, and system prompt for faster text-only output. Add bench-llm-commits task to Taskfile for comparing tool performance. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude Haiku 4.5 <noreply@anthropic.com> |
||
|
|
cd8ac3e060 |
perf(tests): add pre-built fixture to speed up Windows CI (#634)
* perf(tests): add pre-built fixture to speed up Windows CI Instead of creating repos from scratch for each test, copy a pre-built fixture with worktrees and remote already configured. This should significantly reduce test time on Windows where process spawning is slow. - Add tests/fixtures/standard/ with repo, 3 worktrees, and bare remote - Update TestRepo to copy fixture instead of running git init/commit - Add fixture cleanup to doc-generating tests for clean output - Remove git sample hooks and boilerplate from fixture Co-Authored-By: Claude <noreply@anthropic.com> * fix(tests): correctly clean fixture state for isolated tests - Change `remove_fixture_worktrees()` to take `&mut self` and clear the worktrees map so `add_worktree()` can recreate branches - Add remote removal to select and approval tests to prevent `origin/main` from appearing in snapshots - Update select snapshots for fixture-based commit hashes Co-Authored-By: Claude <noreply@anthropic.com> * fix(tests): use relative paths in fixture for portability The fixture gitdir files contained absolute paths to the local development machine, causing CI failures. Updated to use relative paths which work after the _git → .git rename during fixture copy. Co-Authored-By: Claude <noreply@anthropic.com> * fix(tests): remove invalid remote from bare repo fixture The bare repository in the fixture had a remote configured with an absolute path to the local machine, which wouldn't exist on CI. Bare repositories don't need remotes configured. Co-Authored-By: Claude <noreply@anthropic.com> * fix(tests): remove origin in approval tests for consistent project_id Tests manually create approvals using a computed project_id based on repo path. With the fixture's origin remote, worktrunk computes a different project_id from the remote URL. Removing origin makes both use the same fallback (directory path). Co-Authored-By: Claude <noreply@anthropic.com> * fix(tests): improve fixture copy robustness for CI - Add platform-specific copy: `cp -r` on Unix, `robocopy` on Windows - Add verification that essential paths exist after copy - Add verification that origin.git is a valid git repository after rename - Include stderr in error messages for better debugging Co-Authored-By: Claude <noreply@anthropic.com> * fix(tests): add .gitkeep to preserve empty git directories in fixture Git doesn't track empty directories, so objects/info, objects/pack, and refs/tags were missing on CI after checkout. These directories are required for a valid git repository. Co-Authored-By: Claude <noreply@anthropic.com> * fix(tests): ensure LF line endings in fixture git files Add .gitattributes to force LF line endings for git internal files (packed-refs, HEAD, config, etc.) to prevent CRLF corruption on Windows. Co-Authored-By: Claude <noreply@anthropic.com> * fix(tests): handle file copies separately from directories on Windows robocopy only works with directories, not files. Use Rust's std::fs::copy for individual files like .gitattributes, and robocopy only for directories. Co-Authored-By: Claude <noreply@anthropic.com> * fix(tests): enable rerere in test git config for consistent output Git's rerere feature produces "Recorded preimage" messages during conflicts. Enable this in test config so output is consistent between local machines and CI. Co-Authored-By: Claude <noreply@anthropic.com> * fix(tests): restore accidentally deleted hook_clear tests The tests test_hook_clear_no_approvals and test_hook_clear_with_approvals were accidentally removed during the fixture refactoring. This restores them with the necessary remote removal so project_identifier matches "repo" (directory name) instead of the fixture's remote URL. Also adds remote removal to test_hook_show_approval_status for consistency. Co-Authored-By: Claude <noreply@anthropic.com> * chore: remove unreferenced ping_pong snapshots These snapshots used feature_a but the tests now use feature, leaving these orphaned. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
90564bcc94 |
Add asset deletion safeguard and refactor demo build setup
Require manual publish when asset files are deleted to prevent accidental removals. Refactor demo build to only copy starship config for docs target and simplify setup_demo_output usage. |
||
|
|
3669b74f4b | refactor: rename Taskfile.yml to Taskfile.yaml |