mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
f1b2595dcd
## Summary An **on-demand** tool to answer "is prod caught up with staging right now?". The showcase deploy model is **staging = mutable `:latest`** (continuously rebuilt) and **prod = immutable `@sha256:`** (advances only on an explicit promote), so a prod column can sit **behind** a green staging. **There is no scheduled drift alert — by design.** Prod lagging staging is **often intentional**: changes are batched and promoted deliberately, so a recurring "N columns stale" alert would be pure noise. This tool is therefore manual-only: a maintainer runs it when they want to check, and it tells them the current state. - **`bin/railway reconcile-prod`** — for every prod-eligible (`probe.prod == true`) service, compares the **prod serving digest** (the `@sha256:` from `SnapshotCommand.build_snapshot(PRODUCTION_ENV_ID)`) against the **staging running digest** (reuses `PromoteCommand#staging_running_digest`, the same source the promote pin uses). Classifies each: - `green` — prod == staging (in sync) - `stale` — prod != staging **and** staging is resolvable (prod is behind a green staging) - `gray` — staging running digest not resolvable, or the service has no prod snapshot entry yet — informational, **not** stale - Prints a readable per-service table + summary; **exits nonzero iff any service is stale**; `--json` for machine output. **Read-only — no promotes/mutations.** - **`showcase/scripts/reconcile-prod-gate.sh`** — wrapper mirroring `lint-prod-gate.sh`: surfaces the table into `$GITHUB_STEP_SUMMARY`, optionally captures `--json` to `RECONCILE_JSON`, and propagates the exit-code verdict (never swallows a non-zero). - **`.github/workflows/showcase_reconcile.yml`** — **`workflow_dispatch` only** (no cron). Regenerates the SSOT JSON (`EMIT_SKIP_OXFMT=1`, same as the promote workflow's resolve/promote jobs), runs the gate with the Railway/GHCR auth env, renders the reconcile table to the **GH step summary**, and uploads the `--json` as a `reconcile-json` artifact. **No Slack.** The run exits nonzero on a stale column so a manual run visibly flags drift. `timeout-minutes: 10`. - **Tests** — Ruby minitest (`test_reconcile_prod.rb`: classification + exit-code + `--json` shape + dispatcher registration) and a bats gate test (`reconcile-prod-gate.bats`). Wired the gate script into the `showcase_validate.yml` shellcheck list. ### What changed from the original scheduled-alert design The first cut of this PR shipped a daily cron + auto-post to #oss-alerts on any stale column. Per owner feedback, that was reshaped to on-demand only: the `schedule:` trigger and the Slack-on-stale step were removed (intentional/deliberate staleness is not a bug, so an unsolicited recurring alert is noise). The CLI command, the gate wrapper, and all tests are unchanged. ## Gates - `ruby showcase/bin/spec/test_reconcile_prod.rb` → **9 runs, 20 assertions, 0 failures** - `bats showcase/scripts/__tests__/reconcile-prod-gate.bats` → **6 ok** - `shellcheck -s bash showcase/scripts/reconcile-prod-gate.sh` → **clean** - `actionlint .github/workflows/showcase_reconcile.yml` → **clean** (the pre-existing `depot-ubuntu-24.04-4` custom-runner-label warning is on `showcase_validate.yml`, predates this PR — my only change there is one line in the shellcheck list) ## Test plan - [ ] CI green (Ruby suite, bats suite, actionlint/shellcheck, commitlint) - [ ] Optional: read-only `workflow_dispatch` run of `showcase_reconcile.yml` to confirm it runs against live prod/staging (safe — no mutations)