The old "Showcase: Build & Deploy" workflow used a single concurrency group
that cancelled in-flight builds on every push to main. When multiple PRs
merged in quick succession, most service builds got cancelled and never
deployed.
Split into two workflows:
1. showcase_build.yml ("Showcase: Build & Push") - triggered on push to main,
builds Docker images and pushes to GHCR. Has NO concurrency group so every
run completes. Railway auto-update picks up the new :latest tag.
2. showcase_deploy.yml ("Showcase: Verify Deploy") - triggered by workflow_run
from the build workflow. Polls Railway to verify each service picked up the
new image and is healthy. Uses cancel-in-progress since verification is
idempotent. Posts results to showcase-harness via webhook.
Also updates showcase_capture-previews.yml to trigger from the renamed build
workflow.
- deploy workflow: add shared/scripts/manifest paths to shell-dashboard
and shell-docs filters (previously triggered implicitly by committed
JSON diffs in those directories)
- capture-previews: add generate-registry step before capture; use
git add -f for the gitignored registry.json
- e2e smoke test: document generator dependency in import comment
The capture job can take ~30 minutes, during which other commits
routinely land on main. Without a rebase-and-retry loop, the
devops-bot push loses the race and fails with "fetch first".
Observed in runs 24799601181 and 24809331709 on 2026-04-22,
both failing at the same step with identical remote-rejected-push
output. Bounded to 5 attempts so a persistent failure still
surfaces rather than looping forever.
The workflow runs exclusively on main-branch pushes, completed
"Showcase: Build & Deploy" runs, and manual dispatch — all production
events where silent failures (e.g. the GH013 PROTECT_OUR_MAIN
regression that motivated PR #4159) must surface in #oss-alerts
rather than getting buried in the Actions tab.
Mirrors the failure-alert pattern from showcase_validate.yml:
- Hoist SLACK_WEBHOOK_OSS_ALERTS into a job-level env var so
step-level `if:` expressions can reference it (secrets.* is
not a valid named-value inside `if:`).
- Best-effort "Extract failure details" step pulls the failed
step name and first meaningful error line from the jobs API +
`gh run view --log-failed`, truncated to 300 chars.
- `slackapi/slack-github-action@v2.1.0` with toJSON(format(...))
wrapping to safely JSON-encode any dynamic values.
- Fallback `::warning::` log when the webhook secret is unset so
the gap is still visible in the workflow output.
Gated on `failure() && env.SLACK_WEBHOOK != ''`. No `github.event_name`
filter needed — this workflow has no pull_request trigger, so every
failure is an actionable production event.