mirror of
https://github.com/max-sixty/worktrunk.git
synced 2026-09-14 20:00:38 +08:00
9c37aae5b9
Five tend jobs on 2026-08-18/19 hung in `tend-setup`'s `sudo apt-get update` and were killed at GitHub's 360-minute job cap, each burning six hours without ever booting the agent. The two calls are now bounded with `timeout 300`, so an unreachable mirror fails the job in five minutes instead of six hours. Verification is the next tend run: a healthy `apt-get update` on these runners takes ~15 s, so a green setup step confirms the bound is headroom rather than a new failure mode. The cost isn't only compute. `tend-review` [32095910705](https://github.com/max-sixty/worktrunk/actions/runs/32095910705) was the only review #3843 was ever going to get — `pull_request_target` doesn't re-fire without a push, and that dependabot PR merged at 15:52Z with `reviews: []` while its review job sat in apt. That is the "missed review, not late review" case [max-sixty/tend#962](https://github.com/max-sixty/tend/issues/962) was closed pending; I've posted the evidence there, since the job-level `timeout-minutes` question it raises belongs upstream and is a sizing call, not a bug fix. <details><summary>Evidence: the five wedged runs</summary> All five died in the same step (`__self.__run_2`, `Install shells (zsh, fish)`), with the same `duration_ms` — ~21.59 M, i.e. the 360-minute cap — and the same terminal `##[error]The operation was canceled.` GitHub reports a job-timeout kill as `cancelled`, which is why none of them show as `failure`. | run | workflow | started | killed | step duration | what it stranded | |---|---|---|---|---|---| | [32095910705](https://github.com/max-sixty/worktrunk/actions/runs/32095910705) | tend-review | 08-18 03:33Z | 09:33Z | 21,591,590 ms | #3843 merged unreviewed | | [32223334370](https://github.com/max-sixty/worktrunk/actions/runs/32223334370) | tend-nightly | 08-19 06:25Z | 12:25Z | 21,587,692 ms | one nightly sweep (next tick recovered) | | [32226926018](https://github.com/max-sixty/worktrunk/actions/runs/32226926018) | tend-mention | 08-19 07:14Z | 13:15Z | 21,608,063 ms | `max-sixty`'s "can we simplify?" on #3846 | | [32230273316](https://github.com/max-sixty/worktrunk/actions/runs/32230273316) | tend-review-runs | 08-19 07:57Z | 13:57Z | 21,596,592 ms | the 08-19 sweep (this run's window widened to 48 h to cover it) | | [32232437597](https://github.com/max-sixty/worktrunk/actions/runs/32232437597) | tend-notifications | 08-19 08:24Z | 14:24Z | 21,593,512 ms | one poll (next tick recovered) | The last log line before each six-hour silence is mid-`apt-get update`, in the mirror-fallback loop — repeated `Ign:` against `azure.archive.ubuntu.com`, then `Get:` against `archive.ubuntu.com`, then nothing: ``` 2026-08-19T07:59:14.1259813Z Ign:14 http://azure.archive.ubuntu.com/ubuntu noble-updates/main amd64 Packages 2026-08-19T07:59:14.1262122Z Ign:15 http://azure.archive.ubuntu.com/ubuntu noble-updates/main Translation-en [ ~6 hours of nothing ] 2026-08-19T13:57:55.8550938Z ##[error]The operation was canceled. ``` Because the kill lands before the tend action's failure handler, none of the five uploaded a session-log artifact or filed a `tend-outage` row — the outage issue was empty all window, and only the run census caught them. The stranded mention did recover, but slowly and by accident: `tend-notifications` [32258453408](https://github.com/max-sixty/worktrunk/actions/runs/32258453408) picked up the unanswered comment at 13:30Z and replied at 13:36Z, 6 h 22 m after it was posted, at $3.00. The safety net worked; it isn't a substitute for the job not wedging. </details> <details><summary>Why `timeout(1)` and not `timeout-minutes`</summary> `timeout-minutes` isn't among the keys accepted on a step inside a composite action — [the metadata-syntax reference](https://docs.github.com/en/actions/reference/workflows-and-actions/metadata-syntax) lists `run`, `shell`, `if`, `name`, `id`, `env`, `working-directory`, `uses`, `with`, `continue-on-error` and nothing else — so the bound has to live in the `run:` body. `sudo timeout …` rather than `timeout sudo …` puts the timer inside the privileged process, so it can signal apt directly. `-k 30` is what makes the bound a bound: plain `timeout` sends SIGTERM and then gives up, so an apt — or one of the acquire-method children it spawns — that doesn't die on that signal would run the step back out to the 360-minute cap with a bound in place that reads as if it applied. Either exit fails the step: 124 when SIGTERM lands it, 137 after the SIGKILL escalation. Scoped to `tend-setup` because that's where the five occurrences are and where no human is watching a wedged job. The same unbounded `sudo apt-get update` appears in `.github/actions/test-setup/action.yaml`, `nightly.yaml` (×2), `coverage.yaml`, and `benchmarks.yaml` — identical exposure, zero observed occurrences, and a hang there is loud because it blocks a PR someone is waiting on. Happy to widen if you'd rather have it uniform; it seemed like your call rather than mine. </details> --------- Co-authored-by: worktrunk-bot <254187624+worktrunk-bot@users.noreply.github.com>