## Summary
Two confirmed bugs in the showcase harness + operator dashboard that
together produced a false-green D3 on the dashboard while the e2e probe
pipeline was actually dead.
### Bug 1 — Browser-pool death spiral
(`showcase/harness/src/probes/helpers/browser-pool.ts`)
When a chromium process crashed, `recycleSlot` relaunched it; if that
launch threw, the slot was permanently removed via `slots.splice`. A
single transient launch failure (OOM spike, fd exhaustion) thus
monotonically shrank the pool to empty with no self-heal, after which
every e2e probe failed with `launcher-error` / `BrowserPool acquire
timeout`.
Fix:
- The recycle retries the relaunch with bounded backoff (3 attempts,
doubling backoff).
- On exhaustion the slot is **kept** (capacity preserved) and parked as
`relaunchPending`; the next `acquire()` lazily re-attempts the launch.
- `acquire()` re-initializes the pool as a backstop when every slot has
been lost (`slots.length === 0`).
- Existing acquire/release/recycle semantics and logging style
preserved.
### Bug 2 — Stale-green e2e rows read as healthy
(`shell-dashboard/src/lib/cell-model.ts`,
`src/components/depth-utils.ts`)
When the e2e driver stops writing `e2e:<slug>/<feature>` rows, the last
row freezes. A green row then reads as a healthy D3 forever, so the
depth ladder shows a false-green D3 that masks a dead probe pipeline
instead of surfacing it.
Fix:
- `buildCellModel` and `deriveDepth` treat a green e2e row whose
`observed_at` is older than `E2E_STALE_AFTER_MS` (6h, matching the
original stale-window model) as degraded/amber, so it no longer credits
D3.
- Only green is downgraded — a stale red/degraded row already signals a
problem and is left as-is.
- An unparseable timestamp is treated as not-stale, so staleness is
never inferred from bad data.
## Test plan
- [x] New unit test: pool recovers after a transient relaunch failure
(does NOT shrink to 0) and re-inits when emptied — confirmed red→green.
- [x] New unit tests: a stale green `e2e:` row downgrades to amber (not
green) and does not credit D3; a fresh green row stays green; a stale
red row stays red — confirmed red→green.
- [x] Existing fixtures with hardcoded old `observed_at` updated to
default to a recent timestamp so legitimate green rows aren't treated as
stale.
- [x] `showcase/harness` full suite: 1646 tests pass.
- [x] `shell-dashboard` full suite: 630 pass (1 env-gated build-spike
skipped).
- [x] `packages/**` test graph: 17 projects pass.
- [x] typecheck clean (both packages); oxlint 0 errors; oxfmt clean.
Note: there is a separate operational lever (reduce `BROWSER_POOL_SIZE`
/ e2e concurrency, raise harness memory) handled outside this PR; this
is the code fix only.
Re-validate each pending slot inside relaunchPendingSlots before claiming it so
two concurrent invocations can't double-launch a slot (leak + double-publish).
Contain recycle throws with a logged catch and gate recycleSlot entry on shutdown.
resolveD5 skips missing sub-rows and folds only present rows, so it does not
match isD5Green's every-key-present requirement; clarify staleness IS mirrored.
The deferred-relaunch background timer produced a new race in three review
rounds. Remove it entirely and rely on lazy, on-demand recovery: a failed
recycle parks the slot relaunchPending (never evicts), and the next acquire
re-attempts the launch via relaunchPendingSlots, serving any queued waiter
through handOff. The no-eviction + lazy-recovery design already prevents the
original drain-to-0 bug without a timer. Also clamp stats().inUse to >= 0.
A stale-green D5 sub-row was masked when a fresh-green sibling won the
all-green tie, since green is the lowest rank and staleness was checked
only on the post-fold winner. Downgrade each green-but-stale sub-row to
degraded before folding so any stale-green forces amber, matching
depth-utils isD5Green per-key semantics regardless of map order.
Re-arm the deferred relaunch when a waiter parks after a recycle has
already failed with no waiter queued, so it is not stranded to timeout.
Make scheduleDeferredRelaunch idempotent via deferredRelaunchActive so
the acquire-kick and recycle-kick cannot spawn two timer loops. Guard
recycleSlot entry and the relaunchPending filter with a single
isSlotBusy predicate covering both recyclingSlots and relaunchingSlots,
so a late disconnect for a pending slot's old browser cannot double
launch and leak a process. Strengthen the double-publish test to fail
loud and pin exact-once, and add red-green tests for both fixes.
Consolidate isE2eGreenAndFresh into isGreenAndFresh; both had identical
staleness semantics. Single helper now serves D3/D5/D6, removing a drift
hazard. No behavior change.
Reconcile the stale stats()/reinit()/acquire() comments: all-pending recovery
goes through relaunchPendingSlots, not the size-0 reinit backstop (which is the
truly-empty never-launched case). Rename the misleading reinit test to assert
the actual relaunchPendingSlots mechanism. Chain track()'s cleanup so a future
recovery throw cannot float as an unhandled rejection, and log a -1 sentinel
slotIndex for the never-added reinit close. Fix the self-contradicting
drainMicrotasks JSDoc.
Clarify the wrapper's role (it forces noStore:false because unstable_noStore is
unavailable in middleware/Edge). Pure rename across shell, shell-docs, and
shell-dashboard: definitions, middleware call sites, and tests. No behavior
change.
Memoize getToken so the Railway config is not re-read and the deprecation
warning is not re-emitted on every GraphQL request; failures are never cached.
Resolve the GHCR username with || so an empty GHCR_USERNAME falls through to
GITHUB_ACTOR, and use a truthy cache-hit guard so an empty token is never
memoized.
loadAllFixtures silently skipped files lacking a fixtures array and let
JSON.parse throw without the path; throw a clear error naming the file in both
cases and attach the caught error as cause. Export the helper and add red-green
tests covering both throw branches.
Add a regression spec that fails if a fleet-scoped check reads the narrowed
single-service snapshot (the historic spurious-WARN/REFUSE bug), an executable
lint banning direct @*_snapshot reads outside the sanctioned accessors, and
harden PopenSpy with UnexpectedPopen + honest exit-code stamping (raise only on
spawn failure, not on a configured non-zero exit).
Reject ASCII control chars, a colon (port suffix), and any char outside the
DNS-label charset; switch the Host brand from a string-literal __brand to a
non-exported unique symbol so a stray `as Host` cast from outside the module
is a type error. asHost stays the sole runtime constructor. Adds positive +
negative test coverage incl. interior-tab rejection.
Extend the e2e staleness downgrade to resolveD5/resolveD6 in cell-model and to
depth-utils so a frozen-green pipeline no longer credits D5/D6 as healthy. Also
correct misleading comments/test-doc nits flagged in review.
relaunchPendingSlots only pushed recovered slots to `available` and never resolved
queued waiters, so the deferred relaunch scheduled by a failed recycle (scheduled
precisely because waiters were queued) left the waiter stranded for its full 30s
acquire timeout while a live browser sat idle. It also had no re-entry guard, so an
acquire()-driven call racing the deferred timer could both launch a browser for the
same pending slot — leaking one process and/or publishing the slot to `available`
twice (handing the same browser to two probes).
- Extract a single `handOff` helper (waiter-first, then `available` with an includes
guard) and use it in reinit, the recycle success path, and relaunchPendingSlots so
all three share one correct invariant.
- Add a `relaunchingSlots` re-entry guard: a slot is marked before its launch await
and skipped by concurrent invocations, cleared in finally.
- Replace the one-shot deferred relaunch with a self-rescheduling timer (bounded
backoff) that retries while pending slots and waiters coexist, so a waiter is no
longer stranded when a single deferred attempt also fails. Stops on shutdown.
- Track reinit/relaunch launch promises so shutdown() drains them, closing the leak
where a launch resolved after shutdown awaited the in-flight set.
- Route failure and close logging through the injected logger (with slot index)
instead of console.error / swallowed catches, so failures reach the harness pipeline.
- stats().size now counts only live (non-pending) capacity, making the empty-pool
backstop observable (size 0 when every slot is pending).
Tests: fix the vacuous reinit test (failAtCalls now exhausts all retries for every
slot and asserts size 0 before recovery), and add deferred-waiter recovery and
no-double-publish tests. Full harness suite green (1648 tests), typecheck green.
When the e2e driver stops writing e2e:<slug>/<feature> status rows, the last
row freezes. A green row then reads as a healthy D3 forever, so the depth
ladder shows a false-green D3 that masks a dead probe pipeline (a wedged
browser pool, an outage) instead of surfacing it.
buildCellModel (cell-model.ts) and deriveDepth (depth-utils.ts) now treat a
green e2e row whose observed_at is older than E2E_STALE_AFTER_MS (6h, matching
the original stale-window model) as degraded/amber, so it no longer credits
D3. Only green is downgraded — a stale red/degraded row already signals a
problem and is left as-is. An unparseable timestamp is treated as not-stale so
staleness is never inferred from bad data.
Existing test fixtures that hardcoded old observed_at timestamps now default
to a recent value so green rows are not treated as stale; the new staleness
tests pass explicit timestamps.
When a chromium process crashes, recycleSlot relaunches it; if that launch
threw, the slot was permanently removed via slots.splice. A single transient
launch failure (OOM spike, fd exhaustion) thus monotonically shrank the pool
to empty with no self-heal, after which every e2e probe failed with
launcher-error / BrowserPool acquire timeout.
The recycle now retries the relaunch with bounded backoff; on exhaustion it
keeps the slot (capacity preserved) and parks it as relaunchPending for a lazy
retry on the next acquire. acquire() also re-initializes the pool as a backstop
when every slot has been lost (slots.length === 0), so the pool always
self-heals rather than wedging.
The eval-webhook server only exposed /health, but the verify-deploy
probe (showcase/scripts/verify-deploy.drivers.webhooks.ts) and every
other API-shaped showcase backend (agent, eval, pocketbase) standardize
on /api/health. Add the /api/health route so the service matches the
SSOT convention the probe expects; keep /health for back-compat with
any pre-existing callers.
## Summary
Hardening pass over the showcase deploy pipeline and its operator
tooling. Five focused areas:
- **`bin/railway` rollback shell-injection fix** —
`RollbackCommitCommand` shelled out via backtick subshells interpolating
the `--sha`/`--env` flags. Replaced with `IO.popen` argument arrays (no
shell), with `--sha` validated against an anchored hex regex and `--env`
validated against the known env set before any git invocation. Both `git
ls-tree` and `git show` now gate on `$?.exitstatus` before parsing their
output (a failed `git show` previously merged stderr into stdout and fed
it to `YAML.safe_load`). New spec covers injection rejection, git-show
failure, and empty-snapshot paths.
- **Promote snapshot encapsulation** — replaced the bare `@full_*`/`@*`
snapshot ivars in `PromoteCommand` with named `fleet_*`/`target_*`
accessors so fleet-scoped vs per-service snapshot selection is
name-enforced rather than comment-enforced.
- **Branded `Host` type** — `ProbeTarget.host` is now a branded `Host`
produced only by `asHost` (rejects scheme, path, empty, whitespace,
userinfo, query, fragment), branded at the verify-pipeline ingress;
`ProbeTarget` fields are `readonly`.
- **`deploy-to-railway.ts` SSOT migration** — uses the shared
`resolveRailwayToken` and sources project/env IDs from the railway-envs
SSOT; reads the GHCR username from env (fail-loud) instead of a
hardcoded handle.
- **Workflow safety + observability** — collapsed the promote workflow
to an input-agnostic concurrency group so promotes can't race the same
Railway service; added `#oss-alerts` failure notifications to the build
and validate workflows (which previously had silent red paths).
- **Harness probe test fixes** — repointed the aimock fixture-coverage
probe to the per-integration D4/D6/shared layout and aligned
d5-multimodal assertions to the current auto-send sentinels.
## Test plan
- [ ] CI: Validate Showcase, build-checks, python tests green
- [ ] CI: static/quality (oxfmt + lint + typecheck) green
- [ ] Local: Ruby promote suite (`ruby spec/all_tests.rb`) — 103/342
green
- [ ] Local: showcase-scripts vitest green; harness probe tests green
- [ ] Local: oxfmt --check clean; actionlint clean
Repoint the aimock fixture-coverage probe to the per-integration d4/d6/shared layout; align
d5-multimodal assertions to the current auto-send button sentinels.
Replace the divergent getToken with resolveRailwayToken; source project/env IDs from the
railway-envs SSOT; read the GHCR username from env (trimmed, fail-loud) instead of a hardcoded
handle.
Introduce a branded Host produced by asHost (rejects scheme/path/empty/whitespace/userinfo/query/
fragment); brand at the resolveProbeTargets ingress; make ProbeTarget fields readonly; add asHost
+ override-seam tests.
Replace RollbackCommitCommand backtick subshells with IO.popen arg-arrays; validate --sha (hex)
and --env (known set) before any git call; gate both git ls-tree and git show on $?.exitstatus
before parsing; add injection + git-show-failure + empty-snapshot spec. Also encapsulate promote
snapshots behind fleet_*/target_* accessors so fleet-vs-target selection is name-enforced.
Addresses PR #5087 review comment from @tylerslaton:
> When we open the cookbook section, the sidebar should update to include
> only the recipes. Similar to how Reference works today.
Mirror the dedicated-route approach Reference uses (app/reference/page.tsx
+ app/reference/[...slug]/page.tsx), reusing the existing MDX flow via
DocsPageView's pre-built navTree prop:
- app/cookbook/page.tsx — landing route. Builds a navTree scoped to the
cookbook subdir (buildNavTree(CONTENT_DIR/cookbook, 'cookbook')) and
passes it to DocsPageView so the sidebar shows only cookbook entries.
- app/cookbook/[...slug]/page.tsx — catch-all for /cookbook/<recipe>,
using the same scoped navTree.
- Remove the '---Cookbook---' divider and '...cookbook' spread from
showcase/shell-docs/src/content/docs/meta.json so cookbook no longer
appears in the Documentation sidebar (only via the navbar tab).
Verified locally:
- /cookbook and /cookbook/daytona serve 200 with sidebar scoped to
Overview + Daytona only (active highlight tracks the current page).
- / has exactly one /cookbook anchor — the navbar tab — and no cookbook
entries in the Documentation sidebar.
- /built-in-agent sidebar has zero /cookbook entries.
OSS-222
the integration was added to the railway-services SSOT and smoke.yml nameExcludes; this
adds it to the sibling probe configs the probe-config-parity invariant requires to stay
in lockstep.
deploy.yml moved SSOT consumption into resolve-verify-matrix.ts; the test now asserts
deploy.yml invokes that script AND the script reads railway-envs.generated.json + filters
probe.staging===true, rather than scanning the YAML for literals that moved one hop down.
runner-stage ENV NEXT_PUBLIC_COMMIT_SHA/BRANCH expanded empty because Docker ARGs are
per-stage; re-declaring them in the runner stage (mirroring shell-docs) restores build-arg
values at runtime. Verified via local buildx.
CLI accepts optional positional SERVICE + --digest REF (the showcase_promote.yml per-service
loop contract); validates against the SSOT; narrows the promotion target while fleet-scoped
preflight (service-set parity, expected prod domains) keeps reading full snapshots so a
healthy fleet isn't spuriously refused; adds red-green specs.
Non-functional cleanup pass on the showcase deploy-pipeline integration
branch. All changes are scoped to comment rot, log severity for already-
demoted runtime-config fields, length-aware env-name coalescing (a
deliberately-empty primary no longer masks a populated alternate), and
test-quality tightening. No production behavior change beyond the
specific items below.
Changes by area:
- shell/shell-dashboard/shell-docs runtime-config.ts: factor the
`process.env[primary] ?? process.env[alt]` chain into a shared
length-aware `readEnvPair` helper. The prior `??` form treated
`PRIMARY=""` as set, masking a populated alternate; the helper now
treats empty-string as unset and falls through to the alternate.
- shell-docs runtime-config.ts: demote the two recoverable URL fields
(`intelligenceSignupUrl`, `posthogHost`) from console.info to
console.warn. The `FATAL-CONFIG:` Sentry-alert prefix is preserved
only on the true sentinels; the demoted fields now clear prod log-
aggregation thresholds without raising ops alerts.
- All three shells' runtime-config.ts: prefix log lines with the shell
name (e.g. `[shell-docs runtime-config]`) so the shared log stream
identifies which shell emitted the line.
- shell-docs runtime-config-serialize.ts: rewrite the U+2028 / U+2029
RegExp arguments using six-character ASCII backslash-u escape
sequences (was: literal codepoints in the string arg). The literal
codepoints are line terminators that a formatter or editor could
silently strip, breaking the security-critical XSS escape. The
ASCII form is robust to any such pass.
- shell-docs use-google-analytics.test.ts: de-tautologize the hook-
order test. It now asserts `usePathname(` and `useEffect(` both
exist in the source, so deleting all hooks would fail the test
rather than trivially satisfying the early-return path.
- shell-dashboard baseline-types.test.ts: update the partner-count
expectation from 25 to 26 -- the 26th entry (Cloudflare) is a
legitimate integration that landed independently; the test was
stale and had nothing to do with this branch.
- scripts/resolve-verify-matrix.ts: drop the `FIX 7 --` plan-
internal prefix from a comment; keep the explanation.
- shell-docs/.env.example: correct the `NEXT_PUBLIC_SHELL_URL`
fallback claim (sentinel, not canonical prod host) and document
the remaining 7 consumed env vars with their FATAL/warn/silent
semantics so the example matches runtime-config.ts.
Skipped:
- C-SENTINEL-DEDUP (`http://ops.invalid` shared constant across
shell-dashboard's next.config.ts and runtime-config.ts): both
files are at different module levels (root vs src/lib) and the
string appears once in each; extracting to a shared module would
widen the diff into a refactor for marginal benefit. Skipped per
the spec's "if it widens diff awkwardly, skip" guidance.
- C-SSRTEST: already exhaustively covered. Each of the three shells
has an SSR placeholder test that exercises every URL field via
`new URL()` parseability and (for shell-docs) the analytics-key
empty-string semantics. Treated as a no-op.
Validation: shell + shell-dashboard + shell-docs runtime-config /
serialize / GA tests green; bin/showcase Ruby suite green (87 runs);
showcase/scripts resolve-verify-matrix + aggregate-build-results +
lint-rule-no-public-env green (79 runs).
Six fixes addressing CR findings on the Option-B runtime URL-injection migration:
1. SSR_PLACEHOLDER must be parseable URL sentinels — `new URL("")` throws on
SSR causing 500s for any consumer that constructs URLs from runtime-config
fields. Use `.invalid`-TLD sentinels (RFC 2606) for URL fields; analytics
keys stay empty string. Add `suppressHydrationWarning` on consumers that
render the placeholder server-side and the real value post-hydration
(integration-grid, page-actions popover).
2. Hook-order: move `usePathname()`/`useEffect` ABOVE the early-return in
use-google-analytics. Gate the effect bodies on `GA_ID` instead so React
sees a stable hook order across renders.
3. `readUrl`/`readKey` accept either bare or `NEXT_PUBLIC_*`-prefixed env
names via a fallback chain — covers both server-only and inlined-public
variable conventions without forcing a rename across deploy targets.
4. Extract `serializeRuntimeConfig` to `lib/runtime-config-serialize.ts` so
the OWASP-escape behavior (XSS via </script>, U+2028/U+2029 line-terminator
injection) can be unit-tested without importing the layout into vitest.
5. Reclassify `intelligenceSignupUrl`/`posthogHost` from FATAL-CONFIG to
info-level in shell-docs — these are optional integrations, not hard
wiring failures, so absence should not poison the error stream.
6. Comment-rot cleanup: drop "Option B", B12, "the bug we are fixing", fix
"four substrings"→"three substrings" miscounts, and refresh shell-docs
.env.example to describe the runtime-injection contract instead of a
stale next.config throw claim.
V1: shell + shell-docs `next build` succeeds (no Edge-runtime crash on
`unstable_noStore`).
V2: `OPS_BASE_URL=` shell-dashboard `next build` no longer throws —
`next.config.ts` is now a phase-aware function that emits a sentinel
destination at build time and throws only at start (PHASE_PRODUCTION_BUILD
from next/constants).
Tests: shell-docs 72/72, shell 12/12, shell-dashboard runtime-config 16/16
(pre-existing baseline-partner-count failure unchanged).
Closing hardening pass on the showcase deploy-gate's verify-matrix
resolver. The 7-agent review confirmed the gate is correct; this
commit fixes the residual rough edges.
- showcase_deploy.yml: correct the false §3 ok-non-empty comment.
The empty-intersection case can coexist with redeploy_red=false
(every redeploy succeeded, just none probe-eligible) — that's a
correctly-green run, not a red one.
- showcase_deploy.yml: tighten the summary.json shape guard to catch
PARTIAL drift (TOTAL>0 && WITH_STATUS<TOTAL). The previous all-or-
nothing TOTAL>0 && WITH_STATUS==0 check silently dropped drifted
rows on a mixed summary. Validated locally on mixed/normal/empty/
total-drift jq samples.
- resolve-verify-matrix.ts: add asSupportedEventName narrowing helper
+ use it in the CLI. Replaces the unchecked `as` cast — type system
and runtime now tell one story. Resolver's internal eventName
guard becomes defense-in-depth for direct (test) callers.
- resolve-verify-matrix.ts: make the workflow_run boundary total —
summaryPresent MUST be exactly "true"/"false". Any other value
(including "" from a step-id-rename wiring break) throws now
instead of silently emitting has_services=false.
- resolve-verify-matrix.ts: drop the try/catch around
fileURLToPath(import.meta.url) in `invokedDirectly`. The catch
used to swallow ESM-interop failures and silently no-op the CLI
(exit 0, no GITHUB_OUTPUT write → verify skipped = false-green).
- resolve-verify-matrix.ts: reword parseSsotServices JSDoc to
distinguish schema-drift from truncation (the two are different
failure modes, not one conflated story).
- showcase_build.yml: comment addendum on the redeploy-summary
upload — swapping the guard to `if: always()` would red the
legitimate services=='' path (no summary written), trading the
already-closed false-green for a false-red on every non-buildable
push.
- resolve-verify-matrix.cli.test.ts: switch to spawnSync so stderr
is captured on both zero and non-zero exit (execFileSync only
exposes stderr on throw). Hard-code two stable probe-eligible
names ("aimock", "harness") for the sorted-CSV test rather than
picking probe[0]/probe[1] off the live SSOT — the prior test was
tautological (already-sorted in, sorted out) and would silently
pass if the resolver did nothing.
- resolve-verify-matrix.cli.test.ts: add CLI coverage for the
dropped-token ::warning:: path (FIX 3 — the entire drift-detection
contract had zero CLI coverage), the unexpected-EVENT_NAME error
(FIX 5), and the workflow_run-summary_present total boundary
(FIX 7, both "" and "True" inputs).
- resolve-verify-matrix.test.ts: add unit coverage for the new
workflow_run summaryPresent boundary (empty + "True" + the
workflow_dispatch ignores-summaryPresent regression).
Red-green: 6 tests RED before code changes (FIX 3 warning, FIX 5
unknown EVENT_NAME, FIX 7 unit + CLI ×2 for "" and "True"); 79
tests GREEN after.
Validation: 4 vitest files / 79 tests passing; 87/87 ruby specs
passing; actionlint findings unchanged vs integration baseline
(8 → 8, identical diff); yaml.safe_load OK on both workflows.
A 7-agent review of the verify-matrix resolver and its surrounding workflow plumbing found three
boundary surfaces that could silently produce a GREEN deploy on a broken release, plus an
untested CLI contract that CI compares against the literal strings 'true' / 'false'.
FIX 1 — Validate the SSOT shape in loadSsotServices(). The prior `JSON.parse(...) as
{services: SsotService[]}` was an unchecked cast: a truncated/drifted SSOT (emitter crashed
mid-write, or schema renamed) parses fine but silently shrinks/empties the probe-eligible set
→ some redeployed services go unverified, or verify is skipped on a real redeploy. Extract a
pure exported parseSsotServices(raw, path) that requires the shape we depend on (non-empty
services array; each entry has a non-empty string name, an optional string|null dispatchName,
and a probe object with a boolean staging). Throw `::error::SSOT <path> malformed: <detail>`
on any violation. Also re-check existsSync(SSOT_JSON) after the regenerate-if-missing
execFileSync — a regen that exits 0 without writing must not proceed to a useless JSON.parse
crash. Drop the defensive `probe?.staging` once shape is guaranteed.
FIX 2 — Validate summary.json shape in the redeploy-gate bash. The bullseye false-green
surface: if redeploy-env.ts's schema ever drifts (e.g. `status` → `state`, `ok` → `success`),
every `jq select(.status==...)` yields empty → redeploy_red=false AND ok_services="" →
resolver skips verify → GREEN CI on a real unverified redeploy. Add a TOTAL vs WITH_STATUS
shape guard right after loading the summary: if TOTAL > 0 && WITH_STATUS == 0, emit
::error::summary.json has $TOTAL entries but none with status ok|error (schema drift?) and
exit 1. The legitimate empty-array path (TOTAL=0) is preserved.
FIX 3 — Fail loud on unknown eventName in resolveVerifyMatrix. The prior code fell through to
the workflow_run intersection branch for ANY unrecognized eventName (typo, unexpected
trigger), silently emitting has_services=false → indistinguishable from a legit "summary
absent" skip. Add an explicit guard so only workflow_run / workflow_dispatch are accepted;
anything else throws ::error::resolve-verify-matrix: unexpected eventName '<value>'. Tighten
the eventName parameter type to the literal union.
FIX 4 — Trim ok tokens + warn on dropped tokens in okCsvToCanonicalNames. Split, then
.map(t => t.trim()).filter(Boolean) so "a, b" (spaces) matches. Collect tokens that match NO
SSOT service (by name or dispatchName) and have the CLI wrapper emit ::warning::ok_services
tokens dropped (no SSOT match): <list> on stderr when non-empty — surfaces SSOT/build drift.
The pure function stays IO-free; logging lives in the wrapper.
FIX 5 — CLI wrapper integration test. New resolve-verify-matrix.cli.test.ts spawns
`npx tsx showcase/scripts/resolve-verify-matrix.ts` with a temp $GITHUB_OUTPUT file across
four scenarios and asserts the temp file contents EXACTLY (the workflow YAML compares
has_services against the literal strings 'true'/'false', so the byte-for-byte format is part
of the contract). Uses the real railway-envs.generated.json so the loader exercise is real.
FIX 6 — Cleanup. Remove the dead `env: DISPATCH_SERVICE: ...` block on the redeploy-gate
step (the next step redeclares it — leftover from the extraction). Soften the §3
decision-table all-errors bullet to match resolve-verify-matrix.ts's careful wording, and
append that when the success-set is empty (or the intersection collapses to empty), verify
is skipped and the gate reds independently. Append to showcase_build.yml's "Upload redeploy
summary" path-(A) comment that `if-no-files-found: error` still reds path (A) even if a
future change adds `if: always()`.
Tests: red→green for FIX 1/3/4/5 verified locally. Resolve-verify-matrix vitest count:
12 → 28. Full requested suite (resolve-verify-matrix + cli + aggregate-build-results +
lint-rule-no-public-env): 72 passed. showcase/bin ruby specs: 87 runs / 0 failures / 0
errors / 0 skips. actionlint baseline preserved (8 findings, identical to integration tip).
Extract the inline bash+jq decision logic from showcase_deploy.yml's
resolve-matrix job into showcase/scripts/resolve-verify-matrix.ts, a
pure function with a vitest suite. The bash had produced two confirmed
bugs across prior CR rounds, so making it testable is the lasting fix.
Issue A (the bug this PR fixes): when summary_present=true but
ok_services is empty (every service errored on redeploy), the old bash
skipped the intersection and fell through to the full probe-eligible
fleet, gratuitously probing every service against stale :latest. The
resolver now returns has_services=false in that case — enforce-redeploy
-gate independently reds the workflow on redeploy_red=true, so this
case is already loud; there is nothing left to verify.
Parity preserved for unchanged cases:
- workflow_dispatch + 'all'/empty → full probe-eligible set
- workflow_dispatch + specific svc → that one (unknown → error exit)
- workflow_run + summary_present=false → has_services=false
- workflow_run + present + ok non-empty → intersection with probe-
eligible (SSOT key OR dispatchName aliases both resolve)
Also clarified the Upload-redeploy-summary comment in showcase_build.yml
to document both red paths (hard crash → redeploy step exits non-zero;
exit-0-but-no-file → if-no-files-found:error reds the step) so no
false-green path is possible.
Tests: 12-case vitest suite covers each decision-table row plus the
Issue A fix (written red-first; failed against a naive full-fleet
fallback, passed once the early return was added). CLI parity verified
against the real generated SSOT for the three representative env-var
combinations (workflow_run + present + ok=[a,c]; workflow_run + present
+ ok empty; workflow_dispatch + 'all').
D1 — showcase_deploy.yml false-red fix
======================================
The build workflow legitimately uploads no `redeploy-summary` artifact when it ran
(push touched `showcase/**` so `paths:` matched) but `detect-changes` found no
buildable service, so `redeploy-staging` was skipped. The build still concludes
`success`, so `showcase_deploy.yml` fires on `workflow_run` and `resolve-matrix`
runs. `actions/download-artifact@v4` with `name:` HARD-FAILS on a missing
artifact, so the unguarded download was failing the job, and a downstream guard
that trips `enforce-redeploy-gate` on `resolve-matrix.result == 'failure'` was
flipping the workflow RED — a false-red on a routine showcase-docs/script change.
Add an artifact-existence pre-check using `actions/github-script` (pinned by SHA,
matching the existing repo convention) that lists the artifacts for
`workflow_run.id` via `actions: read` (already granted to `resolve-matrix`) and
sets `summary_present=true|false`. Gate the existing download step on
`summary_present == 'true'`. Keep NO `continue-on-error`, so the C1 property
holds: when the artifact exists but the download genuinely fails, the job still
fails loud and `enforce-redeploy-gate` correctly reds the workflow. When the
artifact is legitimately absent, the bash gate's existing `[ ! -f "$SUMMARY" ]`
branch no-ops (`redeploy_red=false`, `ok_services=""`) — nothing was
redeployed, so there is nothing to gate.
Updated the step comment block to enumerate the three distinct cases now
handled: workflow_dispatch (no download); workflow_run + artifact absent
(graceful skip); workflow_run + artifact present (download with fail-loud).
L1-L5 — env lint rule hardening
===============================
- L1: route the destructuring (VariableDeclarator/ObjectPattern) branch through
the shared `staticKeyName()` helper so the computed-string-key form
`const { ["NEXT_PUBLIC_X"]: y } = process.env` and the no-expression
template-literal form `const { [\`NEXT_PUBLIC_X\`]: y } = process.env` are
caught with the same parity as the bracket-member read.
- L2: unwrap a wrapping `ChainExpression` at the top of `isProcessEnv()` so
`process.env?.X` is matched robustly across parser flavors; corrected the
helper's doc comment to describe the actual semantics.
- L3: export `BANNED_KEYS` from the rule module and have the table-driven test
dynamically import the rule's own Set instead of hand-mirroring it — the
test set now cannot drift from the rule.
- L4: added override-scoping fixtures for `showcase/shell/src/**` and
`showcase/shell-dojo/src/**`; the `.oxlintrc.json` override list already
includes these, but the test now exercises them so an accidental drop is
caught.
- L5: expanded the file-header "Out of scope" doc list to include bulk-iteration
reads (`Object.keys/values/entries(process.env)`, for-in, spread
`{...process.env}`), rest-pattern destructuring, compound-assignment LHS, and
update operators. Documentation-only — the deliberate non-coverage is now
auditable.
Validation
==========
- RED→GREEN confirmed for L1 (two new destructuring computed-key tests) and L3
(dynamic `await import(...)` of BANNED_KEYS failed pre-fix with
"Rule module did not export a non-empty BANNED_KEYS Set", green after export).
- vitest: 38 passed (was 34 baseline + 4 new); aggregate-build-results 6 passed.
- Ruby promote suite: 87 runs, 251 assertions, 0 failures (unchanged).
- python3 yaml.safe_load: showcase_deploy.yml + showcase_build.yml +
showcase_promote.yml all parse OK.
- actionlint: zero NEW findings on the changed file. The pre-existing
showcase_build.yml SC2086/SC2129/runner-label findings are identical on the
integration baseline (unchanged by this commit).
Seven-agent CR surfaced correctness defects in the build/deploy/promote
pipeline and in the no-public-env-shell-read oxlint rule. This commit
closes the false-green paths and broadens lint coverage.
Workflow fixes:
- showcase_deploy.yml: drop `continue-on-error: true` on the redeploy-summary
artifact download. The dispatch path is already guarded by the `if:
workflow_run` clause, so the bash "no summary" branch handles legitimate
manual dispatches. A genuine workflow_run download failure must now fail
loud instead of silently widening verify to the full service set against
stale `:latest`.
- showcase_build.yml: redeploy-staging now intersects the build matrix with
the aggregator success set (`needs.aggregate-build-results.outputs.results`,
status == "success") before producing the redeploy CSV. Failed/skipped
slots no longer get redeployed (which would just re-pull stale `:latest`
and look healthy).
- showcase_build.yml: `notify-all-builds-failed` now additionally requires
`needs.build.result == 'failure'` so it doesn't Slack-spam when the build
job was SKIPPED (verify-image-refs upstream failure).
- showcase_build.yml: `notify` now lists [build, aggregate-build-results,
redeploy-staging] in `needs:` so aggregator/redeploy failures still emit
a Slack signal. `if: failure()` still skips when none of the needs failed.
- showcase_build.yml: `set -euo pipefail` on the Prepare build args step
so a transient $GITHUB_OUTPUT write failure can't ship images without
COMMIT_SHA/BRANCH baked in.
- showcase_deploy.yml: `enforce-redeploy-gate` now also trips on a
resolve-matrix failure (`needs.resolve-matrix.result == 'failure'`) so
an upstream crash that leaves `redeploy_red` empty can't bypass the gate.
- Doc-comment accuracy: drop stale `(PR #5093)` reference; correct the
env-IDs source-of-truth comment; document the optional `skip_build` field
in ALL_SERVICES; clarify that health_path is informational and verify
uses per-service drivers; add the missing `resolve-targets` step 0 to the
promote workflow's "Order:" header.
Aggregator fix (RED-GREEN):
- aggregate-build-results.ts: throw on zero slot dirs. The job is gated
upstream on has_changes == 'true', so zero slot dirs is a broken artifact
download, not a legitimate empty build set. Silently emitting
any_success=false + results=[] is indistinguishable from "all builds
failed" and lets the deploy workflow fall back to probing the full
service set against stale `:latest`. Refuse the ambiguity.
- aggregate-build-results.test.ts: existing empty-INPUT_DIR test was
updated to assert the throw (was: return []).
Oxlint rule (RED-GREEN):
- no-public-env-shell-read.mjs: handle destructuring reads
(const { NEXT_PUBLIC_X } = process.env and aliased form), template-literal
computed keys (process.env[\`NEXT_PUBLIC_X\`]), and explicitly skip
assignment-LHS / `delete` targets (writes are not reads). Optional
chaining already worked through the existing MemberExpression path.
Aliasing (`const e = process.env; e.X`) is intentionally documented as
out of scope (needs scope tracking). Description sharpened to say the
rule guards a specific banned-key set, not all NEXT_PUBLIC_* reads.
- .oxlintrc.json: tighten the off-override glob from
`showcase/**/*runtime-config*` to
`showcase/**/lib/runtime-config*.{ts,tsx}` so it only silences the
intended implementation files, not arbitrary paths containing that
substring.
- lint-rule-no-public-env.test.ts: rewritten as table-driven coverage of
every BANNED_KEYS entry (dotted + bracket-string forms), every ALLOWED
key (asserting non-firing), all new variants from the rule expansion,
the assignment/delete non-fire cases, and override scoping
(runtime-config exempt; packages exempt; shell-tree non-runtime-config
flagged).
Validation:
- actionlint on all three workflows: 8 pre-existing findings (depot label,
pre-existing SC2086 infos in untouched steps); my edits add zero.
- python3 yaml.safe_load: all three workflows OK.
- vitest aggregate-build-results.test.ts: 6/6 pass (incl. new throw test).
- vitest lint-rule-no-public-env.test.ts: 34/34 pass.
- vitest full showcase/scripts suite: 1654/1654 pass across 46 files.
- ruby showcase/bin/spec/all_tests.rb: 87 runs, 0 failures.
- Intersection jq proof (matrix a,b,c × success a,c) → "a,c"; all-failed
→ ""; skipped status excluded.
Nine correctness fixes to bin/railway PromoteCommand, each red-green tested.
- P2 in-flight race-check now compares deployed digest against the digest
captured in @promote_refs (P1-resolved), not svc["digest"] which is nil
for tag-form staging — the check was dead code. Also: parse JSON-string
Deployment.meta; sort fetch_latest_staging_deployments by createdAt desc.
- @promote_refs is RESET (not memoized) at the top of check_p1_ghcr_digests,
so a reused command instance cannot carry stale A-era refs into a B-era
promote. execute_promotion hard-guards against a nil @promote_refs.
- execute_promotion pre-validates that every prod-matched service has a
digest-shaped @promote_refs entry BEFORE pinning anything, eliminating
the partial-promotion-on-missing-ref hazard.
- execute_promotion rescue broadens to MutationError + GraphQL::Error +
StandardError so a transient mid-loop failure still surfaces the
PARTIAL-PROMOTION recovery report; dedup the duplicate warn line and
note that source.image may already be partially advanced on Railway.
- check_p1_ghcr_digests emits REFUSE: P1 ... "no image" for an imageless
staging service (instead of a silent skip that surfaced later as a
misleading "internal error").
- check_p1_ghcr_digests per-service rescue broadens to StandardError so a
non-GHCR error (e.g. ArgumentError, network) does not bypass the rescue
and crash the loop, discarding earlier services' findings.
- pin_and_verify raises ArgumentError immediately if called with a
tag-form image (instead of 30s of futile retries + misleading error).
- pin_and_verify timestamp gate is non-vacuous: a non-nil observed
updatedAt is ALWAYS required, even when pre_update_ts is nil
(which previously collapsed the gate to digest-equality alone).
- run_staging_probe rescues Errno::ENOENT / StandardError around the
IO.popen launch so a missing npx produces a clean ok:false summary
instead of a raw stack trace bubbling out of P3.
Spec hygiene: drop the unused FakeGQL class in test_promote_execute.rb
(it referenced an uninitialized @after_image); give the unresolvable-tag
fixture a placeholder digest so it never builds a malformed "...@" ref;
test_promote_p2.rb tests now capture both streams and assert against the
combined output, matching the convention used elsewhere in the suite.
PromoteCommand had a TOCTOU window: resolved_prod_image(svc) was called
twice for every staging service — once in check_p1_ghcr_digests (where
the resolved digest was manifest_exists-verified), and again in
execute_promotion (whose result is what actually got pinned). Because
staging is a mutable :latest tag, a concurrent push between P1 and
execute could make the two resolutions return different digests, and
prod would be pinned to a digest P1 never verified. It also doubled
the GHCR round-trip per service.
Resolve+verify each staging service's digest exactly once during P1,
store the result on @promote_refs (service_name => digest-pinned ref),
and reuse that exact ref in execute_promotion. If a service has no
entry (P1 didn't run or didn't pass), refuse rather than silently fall
back to a tag.
Also:
- check_p1_ghcr_digests had a method-level rescue Railway::GHCR::Error
that replaced the entire findings array with one entry — so a GHCR
error on service N discarded findings already accumulated for
services 1..N-1. Move the rescue inside the per-service iteration
so each error becomes its own REFUSE finding and the loop continues.
- self.pin_and_verify asserted serviceInstanceUpdate == true but
discarded the serviceInstanceRedeploy result. A failed redeploy
could pass verification because the update mutation had already
advanced source.image+updatedAt. Require truthy redeploy result;
raise MutationError otherwise, symmetric with the update check.
- check_p2_staging_deployments already guarded meta.is_a?(Hash) so it
doesn't crash on a String meta, but the silent skip of the in-flight
race-check was invisible. Add a WARN finding so the skip is visible.
SUCCESS status remains the real gate (still REFUSE).
- execute_promotion now tracks already-pinned services and, on a
mid-loop MutationError, emits a loud PARTIAL PROMOTION report
naming both the already-pinned services and the failing one with
a pointer at bin/railway rollback-commit. Auto-rollback is left as
a follow-up — the goal here is just to make the mixed-state loud
and actionable rather than a quiet exit 1.
The showcase deploy model is STAGING = mutable :latest tag, PROD = immutable
@sha256: digest (P6 enforces both shapes). SnapshotCommand#build_snapshot
stored the raw serviceInstance.source.image, so for staging svc["image"] was
the :latest TAG. execute_promotion was pinning THAT mutable tag to prod via
serviceInstanceUpdate, defeating the immutable-prod invariant before
pin_and_verify raised on the nil expected_digest.
Fix: add PromoteCommand#resolved_prod_image — returns the staging svc as
@sha256:-pinned (pass-through if already pinned; resolves the tag via the
shared GHCR client otherwise; returns nil if the tag cannot be resolved).
execute_promotion now refuses (P0) rather than pin a mutable tag, and
check_p1_ghcr_digests verifies the resolved digest (it previously SKIPPED
tag-form images entirely, so :latest was never P1-checked).
Also:
- P2 race-check guards latest["meta"] when Railway returns a JSON String
(deserialized as Ruby String, not Hash) — .dig used to crash with
NoMethodError. SUCCESS status remains the real gate.
- Remove dead --include-startcommand flag (never read; doubly inert because
P6 REFUSEs on any startCommand divergence).
- Spec hygiene: P3 skip-test raises if probe runs under --no-require-staging-
green; P6 warn-proceed stubs execute_promotion to isolate the gate and
asserts rc==0; test_ghcr_token teardown unconditionally deletes
GITHUB_TOKEN/GHCR_TOKEN/RAILWAY_TOKEN before restoring priors.
70 runs, 204 assertions, 0 failures (up from 66/188 baseline).
A RAILWAY_TOKEN secret with trailing whitespace/newline (common from op
read, heredoc, shell export) was returned verbatim and produced invalid
Authorization: Bearer headers and silent Railway 401s. Trim the env-var
lane and treat whitespace-only as UNSET so the config-file fallback runs.
Also adds a missing should-have-thrown guard in the NO_HOME test and
removes a stale gateIgnore clause from findUntrackedServices docstring.
Distinguish ENOENT (treat as drift) from other read errors in
emit-railway-envs-json.ts --check; non-ENOENT errors now exit 2 with the
real error on stderr instead of being silently coerced into a misleading
'stale' message or an overwrite on a false drift signal.
Add a --out=<path> override so tests can write to a temp directory and
never mutate the tracked railway-envs.generated.json artifact. Rewrite
the emit-railway-envs-json test to use mkdtempSync + --out, switch the
staleness assertion to spawnSync so it asserts on exit code 1 plus the
stale-diagnostic substring, and add coverage for the new EISDIR
fail-loud path. After this change git status is clean post-test.
A token read from ~/.railway/config.json with surrounding whitespace or a
trailing newline passed nonEmpty() but was returned verbatim, so an
'Authorization: Bearer <token>' header could carry CR/LF or stray spaces
(Node HTTP rejects invalid header chars; Railway 401s otherwise). Trim
each return path so the canonical token is always emitted; whitespace-only
values still fall through. Also clarify the JSDoc that this resolver does
not consult process.env.RAILWAY_TOKEN.