mirror of
https://github.com/max-sixty/worktrunk.git
synced 2026-09-14 20:00:38 +08:00
codex/remove-codex-cloud-specific-tests
9 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1b278042de |
chore(ci): weekly renovation 2026-08-16 (#3826)
## Summary Weekly CI renovation check found the following updates: - `worktrunk`: 0.72.0 → 0.74.0 (MSRV 1.96, compatible with our 1.96.0) — `ci.yaml` ×2, `nightly.yaml` - `nushell`: 0.114.1 → 0.115.0 — `nightly.yaml`, `benchmarks.yaml`, `coverage.yaml`, `actions/test-setup`, and `scripts/codex-cloud/Taskfile.yaml` - `pre-commit`: 4.6.1 → 4.6.2 — `scripts/codex-cloud/Taskfile.yaml` - `PowerShell`: 7.6.4 → 7.6.5 — `scripts/codex-cloud/Taskfile.yaml` and the root `Taskfile.yaml`'s `setup-web` task The Codex Cloud archive checksums were recomputed from the new upstream tarballs, and the resulting `Taskfile.yaml` digest (`f14dbc89…`) is copied into both README launcher commands. The `setup-web` PowerShell pin came in as a follow-up commit: the initial sweep only grepped `.rs`/`.md`/`.toml` for stale versions, so the root `Taskfile.yaml`'s `PWSH_VERSION="7.6.4"` was missed. Nothing tests the two PowerShell pins against each other, so that one drifts silently — worth a note for future renovation runs. The `powershell_7.6.5-1.deb_amd64.deb` asset the `setup-web` branch downloads is present in the v7.6.5 release. ## Already up to date - Rust stable is 1.97.1, so MSRV and toolchain stay at 1.96 (latest stable − 1) — `Cargo.toml`, `tests/helpers/wt-perf/Cargo.toml`, `rust-toolchain.toml` need no change, and `flake.lock` is untouched. - `cargo-insta` 1.48.0, `cargo-nextest` 0.9.143, `cargo-llvm-cov` 0.8.7, `cargo-msrv` 0.19.3, `cargo-affected` 0.4.0, `cargo-udeps` 0.1.61, `lychee` 0.24.2 - Task 3.52.0 (mise, Codex Cloud) - Runner images: ubuntu-24.04, macos-15, windows-2022 ## Held back: zola 0.22.1 → 0.23.3 Not bumped. Zola 0.23.0 shipped [Tera2 + refactoring](https://github.com/getzola/zola/pull/3105), which is a templating-engine swap rather than a routine release. Building `docs/` with the 0.23.3 binary fails at the first line of `templates/base.html`: ``` ERROR error: Unknown tag --> base.html:1:4 | 1 | {% import "macros.html" as macros %} | ^^^^^^ ``` `templates/base.html` and `templates/macros.html` are the two files that use the `import`/`macro` pair, so the migration looks small, but it is template work with its own review rather than a pin bump — kept out of this PR so the rest can land. Raised separately. <details><summary>Verification</summary> - Every version above was read from the upstream source of truth: `crates.io` for the cargo tools, `nushell/nushell` and `PowerShell/PowerShell` releases, PyPI for pre-commit, and `static.rust-lang.org/dist/channel-rust-stable.toml` for Rust stable (1.97.1). - Checksums were computed from the downloaded archives and the extracted binaries were run (`nu --version` → `0.115.0`); the archive layouts (`nu-<ver>-x86_64-unknown-linux-gnu/nu`, top-level `pwsh`) are unchanged, so the `install_binary` paths still resolve. - All six edited YAML files parse. - The nushell bump was exercised against the shell-integration suite: `cargo test --features shell-integration-tests --test integration -- nushell` with 0.115.0 on `PATH`. 13 of 14 pass; `test_nushell_install_target_is_a_vendor_autoload_dir` fails — but it fails identically on the currently-pinned 0.114.1, and passes on *both* versions when run alone. It is a pre-existing shared-state race in the sandbox, not a regression from this bump: the test asserts against the real user `$nu.vendor-autoload-dirs` entry rather than one under its temp `HOME` (nu resolves the home dir from the passwd database, so the test's `HOME` override does not move it), and a sibling uninstall test in the same filter removes `wt.nu` from that shared directory. Noted rather than fixed here — it is unrelated to the pins. - The zola failure above was reproduced with the official 0.23.3 `x86_64-unknown-linux-gnu` release binary against this repo's `docs/`. </details> --------- Co-authored-by: worktrunk-bot <254187624+worktrunk-bot@users.noreply.github.com> |
||
|
|
00ab0ffa26 |
Canonicalize benchmark fixtures and variants (#3761)
Benchmark fixtures still encoded the benchmark that first needed each repository state, which left overlapping recipes and variants after the earlier harness consolidation. This change reduces the fixture catalog to two provenance-based bases: `Generated` builds an ordinary Git repository locally, while `Imported` copies the pinned `rust-lang/rust` corpus. Worktree, branch, and remote-ref populations remain parameters on `Generated`; prune candidates and backdrop are overlays that work with either base. The generated base deliberately combines heterogeneous worktree states, history-spread branches, and optional remote refs so ordinary list, completion, picker, first-output, alias, remove, and prune benchmarks can share it. Imported history-spread branches and clean base-tip worktrees carry their own commits, preserving the base populations without making them incidental prune candidates when overlays advance the default branch. The benchmark matrix now keeps single-factor contrasts: list scaling uses the 1- and 8-worktree endpoints; alias dispatch has a startup floor, two population endpoints, and one warm/cold variable-resolution pair; completion keeps one full-surface case; remove and prune vary cache or hook state only where the command exercises it. Historical recipes, redundant cache rows, and intermediate scaling points are removed. Manual setup paths live under `target/`, and the benchmark guide documents the resulting fixture and cache model. Tests: `cargo run -- hook pre-merge --yes` after merging current `main` (4,571 tests); targeted Criterion test-mode runs; `cargo test -p wt-perf`; benchmark check, clippy, formatting, and diff checks. > _This was written by Codex on behalf of max-sixty_ |
||
|
|
59c7320a78 |
ci: read TEND_BOT_TOKEN from an environment in every job (#3748)
Closes worktrunk's half of the `tend` environment migration (tend's `TODO.md`, "Finish moving the operational secrets into the `tend` environment", item 2). The environment is a secret scope, not a deploy target: its deployment branch policy is what stops a workflow pushed to a feature branch from reading the bot's PAT. That gate closes only when the repo-level copy of the secret is gone, since a job naming an environment still reads repo-level secrets. worktrunk kept one because these hand-maintained workflows read `TEND_BOT_TOKEN` outside tend's generated set. ## What changed | Job | Environment | Why that one | |---|---|---| | `benchmarks.yaml` `append-gist` (new) | `tend` | `schedule`-gated, so it runs on `main` | | `benchmarks.yaml` `create-issue-on-benchmark-failure` | `tend` | already `schedule`-gated | | `nightly.yaml` `create-issue-on-nightly-failure` | `tend` | already `schedule`-gated | | `release.yaml` `publish-winget` | `release` | runs on a `v*` tag push | | `release.yaml` `publish-homebrew` | `release` | runs on a `v*` tag push | Every job reading `TEND_BOT_TOKEN` now names an environment, so deleting the repo-level copy breaks nothing. ### The gist append moved into its own job The `benchmarks` job has no `if` gate, so putting `tend` on it would refuse a `workflow_dispatch` against a non-`main` ref — on-demand runs against a chosen branch are what that trigger is documented for. A job GitHub skips never requests its environment, so moving the append into a `schedule`-gated `append-gist` job keeps the policy off the dispatch path entirely. It reads `target/criterion` back from the artifact the `benchmarks` job already uploads. `create-issue-on-benchmark-failure` now `needs` both jobs, so a failed append still files an issue — previously it failed the `benchmarks` job directly. ### Why the release jobs get `release`, not `tend` A tag push is not bot-steerable and tag creation and update are already restricted to admins by the "Tag operations" ruleset, so a tag policy is a real boundary. The tag entry cannot go on `tend`: `tend check`'s `check_environment` pins that policy to exactly the protected branches and its `--fix` deletes anything else. `release` already exists with a `v*` tag policy and already holds `AUR_SSH_PRIVATE_KEY`, so no new environment and no new credential — `TEND_BOT_TOKEN` is seeded into it as a second copy. ### `deployment: false` Jobs naming `tend` use the mapping form. GitHub files a deployment record for every job that names an environment, against whatever ref the run belongs to; under `pull_request_target` that is the PR's own head, which is why PR timelines grew a "worktrunk-bot deployed to tend" line on every push. `deployment: false` drops the record and keeps the gate. The release jobs keep their records, which land in no PR timeline. ## Follow-up: one step, after merge `TEND_BOT_TOKEN` is **already seeded into the `release` environment** (read from the local `worktrunk-bot` gh config dir at `~/.config/gh-bots/worktrunk-bot`, so no new credential was minted and nothing was pasted). Verified: the token resolves to `worktrunk-bot` and has push on both `max-sixty/winget-pkgs` and `max-sixty/homebrew-worktrunk`. That leaves one step, and it must come **after** this PR merges: ``` gh secret delete TEND_BOT_TOKEN --repo max-sixty/worktrunk ``` Not before. On `main` today the gist append, both `create-issue-on-*-failure` jobs, and the two publish jobs still read the token with no environment named, so deleting the repo-level copy first would break the next benchmarks cron (03:47 UTC daily). Merging this PR is what makes the deletion safe. ## What this does not fix - `repo-secret-allowlist` still fails on `CLAUDE_CODE_OAUTH_TOKEN`, also at repo level. It cannot be read back either; separate item. - `environment-deployments` still fails on the generated `tend-*.yaml`, which pin tend 0.1.13 and carry the bare `environment: tend`. Those are regenerated by the published tend, and a regen also carries an unrelated `gh api --paginate` fix in `tend-mention.yaml`, so it belongs in its own PR. ## Verification - The `append-gist` script was run end-to-end against a real `benchmark-results-*` artifact from run 30976221483. It emits 40 rows whose `bench` names match the live gist's existing rows exactly, confirming the artifact round-trip preserves paths relative to `target/criterion`. - That a skipped `if` short-circuits the environment gate is confirmed by run 31066000517: `publish-cargo`, which names `environment: release`, completed as *skipped* on a `pull_request` from a non-tag ref rather than failing on the policy. - `actionlint` reports the same eight pre-existing shellcheck notes as `main`; no new findings. `pre-commit` passes. > _This was written by Claude Code on behalf of max-sixty_ |
||
|
|
970976bd32 |
Consolidate benchmark recipes and cases (#3721)
Benchmark fixtures had accumulated around individual call sites, leaving the same repository shapes and command modes expressed several ways. This change makes repository state the organizing concept: benchmark groups select semantic `FixtureRecipe`s, share table-driven cases, and retain separate fixtures only when a controlled contrast, destructive precondition, or disproportionate setup cost requires one. The real-repository list benchmarks now share one pinned `rust-lang/rust` fixture with eight worktrees and fifty branches spread across history. The list matrix keeps default, branch, warm, and cold coverage without maintaining several “real” repository handles. Remove and prune cases share the same case machinery, while the destructive large-repository prune state remains separate. The scheduled workflow now converts Criterion estimates directly with `jq`, removing the one-off Python converter and its tests. The benchmark guide records the canonical-fixture principle and the remaining recipe-to-group mapping. Tests: `cargo run -- hook pre-merge --yes` (4,551 tests); `cargo bench --bench list large_repository -- --test`; `cargo test -p wt-perf`; benchmark check, clippy, formatting, and diff checks. > _This was written by Codex on behalf of max-sixty_. |
||
|
|
a7c4a7a120 |
refactor(wt-perf): place fixtures under target/wt-perf, not the cache dir (#3547)
## What Move every `wt-perf` on-disk fixture out of the per-user cache dir and under the cargo target dir at `<target>/wt-perf/`, resolved by a renamed `wt_perf_fixture_dir()`: - `setup <config>` fixtures → `<target>/wt-perf/<config>` (was `~/.cache/wt-perf/<config>`). - The rust-lang/rust clone and `prune-real` fixtures → `<target>/wt-perf/bench-repos/` (was `~/.cache/wt-perf/bench-repos/`). - `wt_perf_cache_dir()` → `wt_perf_fixture_dir()`. The target dir is derived from the **running executable's own path** (it lives inside whichever dir cargo built into — `<target>/debug/wt-perf`, `<target>/release/deps/<bench>`), so it honors `CARGO_TARGET_DIR`, a config-file `build.target-dir`, and cargo-llvm-cov's `target/llvm-cov-target/` — none of which a bare `CARGO_TARGET_DIR` env read covers. Falls back to `<workspace>/target` if the binary isn't under a recognizable profile dir. The `etcetera` dependency is dropped. - The `WT_PERF_CACHE_DIR` env override (and the empty-value guard it required) is removed; the benchmarks workflow now caches the deterministic `target/wt-perf/bench-repos` path directly. - In-process throwaway fixtures (`create_repo`) are unchanged — they keep using `tempfile::TempDir`. This reverses the location decision from #3542, which had moved these into `~/.cache/wt-perf`. ## Why `target/` is the conventional home for build/test-generated artifacts — gitignored, reaped by `cargo clean`, and already where criterion writes (`target/criterion`). Keeping every wt-perf fixture there gives one predictable location under the repo that `cargo clean` fully resets. Deriving the target dir from the running binary (rather than assuming `<workspace>/target`) keeps that property intact even when the target dir is relocated. ## Tradeoff (accepted) `target/` is **not** shared across git worktrees (worktrees don't share it) and is wiped by `cargo clean`. So the ~15 GiB `prune-real` rust clone re-clones per worktree and after every `cargo clean` — the cost #3542 avoided by using the cache dir. This is deliberate and documented on `wt_perf_fixture_dir`; it's cheap for the synthetic `setup` fixtures, which rebuild in seconds. ## Testing - `cargo build -p wt-perf --all-targets`, `cargo test -p wt-perf` — clean. - New unit test `target_dir_from_exe_finds_cargo_target` covers the resolution logic (relocated `CARGO_TARGET_DIR`, bench binary under `release/deps/`, cargo-llvm-cov's nested target, closest-profile-wins, and the outside-any-target fallback). - `pre-commit run --all-files` (fmt, clippy, yaml, typos, cargo-lock, custom hooks) — clean. - Smoke-tested end-to-end: `setup branches-2` lands at `<worktree>/target/wt-perf/branches-2` and writes nothing to `~/.cache/`; a copy of the binary run from a fake `<dir>/debug/wt-perf` correctly resolves fixtures to `<dir>/wt-perf/`, proving a relocated target dir is honored. > _This was written by Claude Code on behalf of Maximilian_ --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4121dc979c |
refactor(wt-perf): cache fixtures in the platform cache dir, not /tmp or target/ (#3542)
## What Move `wt-perf`'s benchmark/debug fixture repos out of `std::env::temp_dir()` and `target/bench-repos/` into the per-user cache dir, resolved by a new `wt_perf_cache_dir()`: - `$WT_PERF_CACHE_DIR` if set, else `<cache>/wt-perf` via `etcetera::choose_base_strategy().cache_dir()` — `~/.cache/wt-perf` (or `$XDG_CACHE_HOME/wt-perf`) on Linux **and** macOS. This matches worktrunk's own base-dir convention (`src/config/user/path.rs` resolves user config through the same XDG-on-macOS strategy). - `setup <config>` fixtures → `<cache>/<config>` (was `temp_dir()/wt-perf-<config>`). `--path` still overrides. - Real cloned fixtures (`prune-real`, the rust-lang/rust clone) → `<cache>/bench-repos/` (was `target/bench-repos/`). - In-process throwaway fixtures (`create_repo`) are unchanged — they keep using `tempfile::TempDir`. ## Why Both old homes were wrong for a *reused* fixture (these are given stable names, persist between runs, and are referenced from docs and `-C <path>` — that's a cache, not a temp file): - **A fixed name in a shared `/tmp` (Linux)** collides across users: `/tmp` is sticky (`1777`), so a second user's `setup` hits `remove_dir_all().unwrap()` on a dir they don't own → `EPERM` → panic. It's also a predictable-path/TOCTOU hazard, and `systemd-tmpfiles` reaps `/tmp` per-file at 10 days by `max(atime,mtime,ctime)` — on a `noatime` mount an actively-*read* fixture still loses cold `.git/objects/pack/*.pack` mid-use → silent git corruption. - **On macOS**, `temp_dir()` is already `/var/folders/.../T` (per-user, `0700`), so the docs' `/tmp/wt-perf-*` paths were simply wrong there. - **`target/`** is per-worktree (git worktrees don't share it) and wiped by `cargo clean`, so the ~15 GiB rust clone was re-cloned per worktree — worst for the most expensive fixture, in a worktree-heavy workflow. The cache dir is per-user (no collision, no TOCTOU), stable and discoverable (docs/`-C` references still work), shared across worktrees (clone once per machine), and survives `cargo clean` — matching sccache, cargo, rustup, Go, Bazel, and Hugging Face, which all place large reusable caches there rather than in `/tmp`. ## Migration notes - First bench/setup run after this rebuilds the cache under the new location (the old `target/bench-repos/` is no longer read). Existing fixtures can be dropped with `rm -rf target/bench-repos`. - `.github/workflows/benchmarks.yaml` sets `WT_PERF_CACHE_DIR` and caches `$WT_PERF_CACHE_DIR/bench-repos`, so the cached path and the path wt-perf writes stay pinned together (no drift if a runner sets `$XDG_CACHE_HOME`); key unchanged. - Docs (`benches/CLAUDE.md`, `src/commands/CLAUDE.md`) updated to the new paths, noting that `wt-perf setup` prints the exact path. ## Testing - `cargo build --workspace --all-targets`, `cargo clippy --workspace --all-targets --features shell-integration-tests -- -D warnings` — clean. - `cargo test -p wt-perf` — passes. - Smoke-tested default resolution (`~/.cache/wt-perf/…` on macOS), `$WT_PERF_CACHE_DIR` override, and the `prune-real --path` rejection (now names the resolved cache path). > _This was written by Claude Code on behalf of Maximilian_ --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d7e3aac9a0 | ci: bump pinned tool versions (cargo-nextest, worktrunk, nushell) (#3429) | ||
|
|
0beeb3b738 | chore(ci): weekly renovation 2026-07-05 (#3368) | ||
|
|
820666741f |
ci: run benchmarks on their own daily schedule, split from nightly (#3337)
The criterion suite ran as a job inside `nightly.yaml`: ~80 min, checking performance rather than correctness, yet feeding nightly's failure aggregator and dominating its wall-clock. This moves it into a standalone `benchmarks.yaml` on its own daily schedule, decoupled from nightly's correctness checks. `benchmarks.yaml` runs on its own cron (3:47 UTC, offset from nightly's 5:37 so the two long runs don't contend) plus `workflow_dispatch`. The env block is carried over so the shared `test` cache key still matches (per `.github/CLAUDE.md`); the gist time-series append stays gated to `schedule` events; and a dedicated `create-issue-on-benchmark-failure` job keeps a broken bench harness visible, since benchmarks no longer ride nightly's aggregator. One behavior change: benchmarks no longer run on PRs at all. Previously the job ran on a PR carrying the `nightly` label or touching `Cargo.*`. `workflow_dispatch` against a branch covers on-demand bench runs. `nightly.yaml` loses the `benchmarks` job, its header line, and its `needs` entry in the failure aggregator. Docs updated: the `CLAUDE.md` benchmarks note (benchmarks no longer gate merges since they're off PRs) and the `running-tend` skill's `setup-nu` call-site list. This branch also merges current `main`, which includes #3324 (drop `cache-apt-pkgs-action` for plain `apt-get`); `benchmarks.yaml` installs zsh/fish the same way rather than reintroducing the dropped action. > _This was written by Claude Code on behalf of max_ Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |