with_fast_sleeper mutated the process-global RETRY_DELAY_SEC via
remove_const/const_set. The clean seam (pin_and_verify(sleeper:)) is not
reachable from cmd.run without changing bin/railway production logic, so keep
the swap but make it bulletproof against run-order state leakage: capture the
original before mutating, track whether the swap happened so a mid-setup
failure never leaves the const perturbed, restore in ensure even on raise,
and silence the "already initialized constant" warning locally.
Also reword the header/test comments so they no longer hardcode the literal
"5 public hosts" count, referring to the EXPECTED_DOMAINS[PRODUCTION_ENV_ID]
set instead so the prose can't drift from the SSOT the fixture derives from.
install_fleet_fixture derived one prod service per SSOT public host and then
unconditionally appended make_prod_service(target). When the target already
owns a public prod host (e.g. "docs" owns docs.copilotkit.ai) this listed the
same prod service twice — one domain-bearing, one with custom_domains:[] —
a malformed fleet shape that contradicts the helper's "no public domain of
its own" contract and was only masked by find_service first-match + .uniq.
Guard the append so the bare target is added only when it is NOT already a
derived domain owner. Add a red-green test asserting the derived prod
snapshot contains no duplicate service names or service_ids and that the
target still appears exactly once.
PinCommand#run resolves the service id via RollbackCommand#resolve_service_id
BEFORE the --dry-run early-return, which issues a real GraphQL call and
die!s (exit) on a tokenless CI runner — aborting the whole minitest
process before the summary. Stub resolve_service_id at the class level so
the colon-split assertions run hermetically. Verified green under an
unset RAILWAY_TOKEN + isolated HOME.
install_fleet_fixture hardcoded the 5 public prod hosts and the tests
hardcoded service names, so any change to the SSOT
(railway-envs.generated.json) would break these tests with confusing
phantom-domain WARN / parity die! failures — a brittle gate around the
promote logic rather than the logic itself.
Derive the fixture's domain-bearing prod services from the same constant
the prod code reads (Railway::EXPECTED_DOMAINS[PRODUCTION_ENV_ID]),
mapping each public host back to its owning SSOT service, and select the
target/sibling from Railway::STAGING_SERVICES instead of bare strings.
Documents the derive-from-SSOT invariant. Behavior is identical for the
current SSOT — all four existing tests stay green.
PinCommand#run and PromoteCommand#image_shape stripped the tag with a
first-colon split(":", 2), which cuts at the registry PORT colon and
corrupts a host:PORT/org/img:tag ref (e.g. localhost:5000/img:latest →
base "localhost"). Switch both call sites to the existing last-colon
String#rsplit_colon helper (already used by GHCR#parse_image_ref) so the
tag-stripping is consistent and port-safe. Canonical ghcr.io/...:tag refs
(no port) are unaffected — this is a latent correctness fix.
Adds red-green unit coverage proving a host:PORT/img:tag ref now parses
and pins correctly, and that an empty-tag port ref is no longer
misclassified as :tag.
The `bin/railway promote` help banner advertised preflight checks and
effects that were never implemented in any commit (verified via full
git history on showcase/bin/railway — all phrases trace to the original
56e85dfd80 add, never as working logic):
- MOVES "autoUpdate=disabled flag": execute_promotion only pins the prod
image digest + redeploys (serviceInstanceUpdate sets source.image only);
auto_updates_disabled is a vestigial snapshot field hardcoded to nil and
never mutated.
- VERIFY-REFUSE "PB superuser auth" / "PB collection parity": no such
checks exist; POCKETBASE_SUPERUSER_* are only key-presence entries in
CRITICAL_ENV_KEYS, and PocketBase auth/collection logic lives entirely in
the harness, never reached by promote.
- VERIFY-REFUSE "cross-env URL leak scan": never implemented; snapshots
capture env-key NAMES only (values are never compared), making such a
scan impossible by construction.
- WARN "sealed-var heuristics": isSealed is fetched in the env-vars query
but never read; no heuristic consumes it.
Rewrite MOVES/VERIFY-REFUSE/WARN/IGNORE to list only the real preflight
(P1 GHCR digest, P2 staging deployment + race, P3 staging live-green, P6
startCommand/healthcheckPath/image-shape parity, service-set parity,
critical env-key parity; P6 region/replicas/restartPolicy/env-key-set and
expected-prod-domains WARNs) and the real effect (pin prod image to the
staging digest + redeploy). Doc-text-only; no logic changed.
Renumber test_snapshot_ivar_lint.rb ALLOWED_LINES by +3 to track the
line shift from the (longer) banner — the lint is line-number-pinned by
design and instructs hand-renumbering on any shift above its region.
check_service_set_parity scoped only the staging-only arm to the
single-service target; the prod-only arm was still computed over the
full fleet, so any prod-only service (e.g. a deprecated harness-legacy)
REFUSEd every unrelated single-service promote — the exact mirror of
the bug #5324 fixed. Scope both arms to the target for single-service
promotes; full-fleet promotes (target nil) keep both arms at full
strictness.
Tests: add prod-only tolerance red-green test, strengthen the
target-absent test to assert target-scoping (unrelated staging-only
sibling ignored), rewrite the stale snapshot-narrowing comments to
describe the real fleet_*/& [target] contract, drop the dead
FLEET_PUBLIC_PROD_HOSTS constant, and renumber the ivar-lint allowlist
for the one-line shift in bin/railway.
A single-service `promote <svc>` ran check_service_set_parity over the
FULL staging vs FULL prod fleet and REFUSEd whenever staging carried any
service prod lacks. The live staging fleet legitimately contains 13
staging-only services — harness-workers (SSOT-modeled) and 12 starter-*
demos — so an otherwise-clean single-service promote (e.g. docs) is
blocked with `REFUSE: services in staging not in prod`.
Scope the "staging not in prod" REFUSE to the promote TARGET when a
single-service promote is in effect (intersect the staging-only set with
[target]). The target-absent-from-prod footgun still REFUSEs (target is
in the intersection), the "prod not in staging" arm is unchanged, and
full-fleet promotes (no --service) retain full strictness.
Complements #5322.
build_snapshot enumerates every project service and queries each one's
serviceInstance. The only guard was `next if inst.nil?` — it handled a
NULL result but not a THROWN `GraphQL: ServiceInstance not found` error
(a half-deleted service that still appears in the project service list
but has no instance in the env). That error bubbled to Railway.run's
top-level `rescue GraphQL::Error` and aborted the ENTIRE promote with an
opaque exit 2 before any preflight/divergence logic ran (run 27144525566
killed the docs promote this way).
Scope the rescue narrowly to ONLY the per-service "ServiceInstance not
found" message — log+skip that one service exactly like the nil case —
so every other GraphQL failure (auth, rate-limit, schema drift) still
propagates fail-loud. Adds red-green coverage: a single thrown not-found
is skipped (healthy services still snapshot), while an unrelated GraphQL
error still raises.
Make the showcase dev tool faithful to staging by construction. Two changes:
1. `showcase test --d5/--d6` now drives the fleet CONTROL-PLANE (producer ->
probe_jobs queue -> worker -> result-aggregator) instead of the legacy
in-process runLevel() driver. The new cli/control-plane-run.ts replicates
the deep/full producer tick exactly as runControlPlane wires it
(createE2eDeepServiceEnumerator / createServiceEnumerator over
createJobProducer + createFleetQueueClient), enqueues one operator-triggered
tick, and polls local PocketBase for the run's terminal cells. The running
worker fleet claims + runs the driver + the aggregator writes the d5/d6
status cells, so the dev tool exercises the IDENTICAL wiring + concurrency
as staging. The old in-process path stays available behind `--direct`.
2. `showcase up --dev` adds a docker-compose.dev.yml overlay that bind-mounts
each integration's source and overrides the run command with a stack-aware
hot-reload entrypoint (shared/dev/dev-entrypoint.sh: uvicorn --reload for
FastAPI agents, langgraph dev for graphs, next dev for the frontend). Edit a
source file and the component reloads in place with no image rebuild. The
built-image mode remains the faithful/staging-equivalent default.
find_previous_deployment returned the second-newest SUCCESS regardless of the
head deploy's status. When the head deploy FAILED/CRASHED (precisely when
rollback is invoked) the newest SUCCESS is the last-known-good target, but the
old code skipped it and rolled back one good deploy too far. Now select the
newest SUCCESS strictly older than the current head (drop head, first SUCCESS
in the remainder), and fail loud when the fetched window is saturated and no
target is found — directing the operator to pass --to explicitly rather than
silently returning nil.
Also guard the one unguarded accessor in EnvDiffCommand#diff_services with
`|| []` so a snapshot lacking "services" hits the documented exit-2 path
instead of a raw NoMethodError.
Renumber the ivar-lint allowlist to track the added lines in bin/railway.
Two pre-existing bin/railway bugs surfaced by code review:
1. RollbackCommand#find_previous_deployment selected "the second SUCCESS in
reverse-chronological order" but never sorted the deployments — it trusted
Railway's arbitrary GraphQL connection order, so the wrong deploy could be
chosen as the rollback target. Now sort by createdAt descending before
selecting, mirroring the sibling fetch_latest_staging_deployments which
documents and guards against this exact hazard.
2. EnvDiffCommand advertised behavior it never implemented:
- The banner claimed it compares "custom domains", but run() only diffed
digest/startCommand/env_keys. custom_domains was captured in the snapshot
yet never compared. FIX: extract a pure diff_services helper and add a
custom-domain set comparison (both directions).
- --ignore-env-scoped was parsed into options[:ignore_env_scoped] but never
read, and the env_scoped? helper (+ ENV_SCOPED_URL_MARKERS) it relied on
was defined but never called. The flag had nothing meaningful to act on:
env-diff only compares env-var KEY names (snapshots never carry values),
and custom domains are compared as exact sets (silently dropping
"env-scoped" domains would defeat the comparison). DECISION: remove the
dead flag, the unused env_scoped? helper, and the now-orphan
ENV_SCOPED_URL_MARKERS constant rather than leave false advertising.
Tests: TDD red->green for both. test_rollback_sort feeds deployments in
non-chronological edge order (with a newer FAILED deploy as a decoy) and
asserts the true second-newest SUCCESS is returned. test_env_diff exercises
diff_services directly for identical and differing custom_domains. The
snapshot-ivar lint allowlist is renumbered (numbers only) to track the line
shifts from removing the dead code and adding the rollback comment.
bearer_for ended with a blanket `rescue StandardError => nil` plus
`return nil if status >= 400`. Now that bearer_for ALWAYS performs the
/token exchange (even when a token is present), that swallow masked real
failures: a non-2xx /token response, a malformed JSON body, or a transport
error all collapsed to nil, after which manifest_exists issued the manifest
HEAD anonymously (no Authorization). That silently violated manifest_exists's
documented raise-on-transport/5xx contract and conflated "no token supplied"
with "supplied token failed to exchange".
bearer_for now:
- raises GHCR::Error on a >=400 /token response WHEN a token was supplied
(token absent still returns nil — the legitimate anonymous-public fallback);
- raises GHCR::Error on JSON::ParserError for an unparseable 200 body;
- no longer swallows StandardError, so transport exceptions
(Errno::ECONNREFUSED, Net::ReadTimeout, ...) propagate.
Call-site enumeration (bearer_for is called only by manifest_exists and
resolve_digest):
- manifest_exists: documents "Raises GHCR::Error on 5xx or transport failure",
so a raised exchange error is consistent with — and strengthens — its own
contract. Assumption holds.
- resolve_digest: already raises GHCR::Error on its own manifest HEAD >=400 and
does not rescue bearer_for, so a raised exchange error propagates exactly as
its other failures do. Assumption holds.
Neither caller is broken by bearer_for raising; both already propagate
GHCR::Error to their callers.
Tests: add 3 red-green cases in test_ghcr_bearer.rb (token-present 401 raises;
token-present malformed body raises; token-present exchange failure issues NO
anonymous manifest HEAD). The existing manifest_exists/digest fakes were only
green because the old swallow masked the fake's "no fake response" error — they
never modeled the mandatory /token exchange; added a successful /token fake to
each so they exercise the real path. Renumbered the snapshot-ivar-lint
allowlist (+13 lines) to track the bin/railway line drift.
bearer_for returned a raw GitHub/Actions token, which GHCR's OCI manifest
endpoint rejects with HTTP 403. Always perform the /token exchange instead:
authenticate with Basic base64("x-access-token:<token>") when a token is
present, fall back to the anonymous exchange for public packages, and use
the minted token as the manifest-read bearer. This fixes the docs promote
that has deterministically 403'd in CI.
Route the exchange through the injectable @http seam (matching resolve_digest
/ manifest_exists) so it is testable. Renumber the snapshot-ivar lint
allowlist for the +21 line shift.
Add a regression spec that fails if a fleet-scoped check reads the narrowed
single-service snapshot (the historic spurious-WARN/REFUSE bug), an executable
lint banning direct @*_snapshot reads outside the sanctioned accessors, and
harden PopenSpy with UnexpectedPopen + honest exit-code stamping (raise only on
spawn failure, not on a configured non-zero exit).
Replace RollbackCommitCommand backtick subshells with IO.popen arg-arrays; validate --sha (hex)
and --env (known set) before any git call; gate both git ls-tree and git show on $?.exitstatus
before parsing; add injection + git-show-failure + empty-snapshot spec. Also encapsulate promote
snapshots behind fleet_*/target_* accessors so fleet-vs-target selection is name-enforced.
CLI accepts optional positional SERVICE + --digest REF (the showcase_promote.yml per-service
loop contract); validates against the SSOT; narrows the promotion target while fleet-scoped
preflight (service-set parity, expected prod domains) keeps reading full snapshots so a
healthy fleet isn't spuriously refused; adds red-green specs.
Nine correctness fixes to bin/railway PromoteCommand, each red-green tested.
- P2 in-flight race-check now compares deployed digest against the digest
captured in @promote_refs (P1-resolved), not svc["digest"] which is nil
for tag-form staging — the check was dead code. Also: parse JSON-string
Deployment.meta; sort fetch_latest_staging_deployments by createdAt desc.
- @promote_refs is RESET (not memoized) at the top of check_p1_ghcr_digests,
so a reused command instance cannot carry stale A-era refs into a B-era
promote. execute_promotion hard-guards against a nil @promote_refs.
- execute_promotion pre-validates that every prod-matched service has a
digest-shaped @promote_refs entry BEFORE pinning anything, eliminating
the partial-promotion-on-missing-ref hazard.
- execute_promotion rescue broadens to MutationError + GraphQL::Error +
StandardError so a transient mid-loop failure still surfaces the
PARTIAL-PROMOTION recovery report; dedup the duplicate warn line and
note that source.image may already be partially advanced on Railway.
- check_p1_ghcr_digests emits REFUSE: P1 ... "no image" for an imageless
staging service (instead of a silent skip that surfaced later as a
misleading "internal error").
- check_p1_ghcr_digests per-service rescue broadens to StandardError so a
non-GHCR error (e.g. ArgumentError, network) does not bypass the rescue
and crash the loop, discarding earlier services' findings.
- pin_and_verify raises ArgumentError immediately if called with a
tag-form image (instead of 30s of futile retries + misleading error).
- pin_and_verify timestamp gate is non-vacuous: a non-nil observed
updatedAt is ALWAYS required, even when pre_update_ts is nil
(which previously collapsed the gate to digest-equality alone).
- run_staging_probe rescues Errno::ENOENT / StandardError around the
IO.popen launch so a missing npx produces a clean ok:false summary
instead of a raw stack trace bubbling out of P3.
Spec hygiene: drop the unused FakeGQL class in test_promote_execute.rb
(it referenced an uninitialized @after_image); give the unresolvable-tag
fixture a placeholder digest so it never builds a malformed "...@" ref;
test_promote_p2.rb tests now capture both streams and assert against the
combined output, matching the convention used elsewhere in the suite.
PromoteCommand had a TOCTOU window: resolved_prod_image(svc) was called
twice for every staging service — once in check_p1_ghcr_digests (where
the resolved digest was manifest_exists-verified), and again in
execute_promotion (whose result is what actually got pinned). Because
staging is a mutable :latest tag, a concurrent push between P1 and
execute could make the two resolutions return different digests, and
prod would be pinned to a digest P1 never verified. It also doubled
the GHCR round-trip per service.
Resolve+verify each staging service's digest exactly once during P1,
store the result on @promote_refs (service_name => digest-pinned ref),
and reuse that exact ref in execute_promotion. If a service has no
entry (P1 didn't run or didn't pass), refuse rather than silently fall
back to a tag.
Also:
- check_p1_ghcr_digests had a method-level rescue Railway::GHCR::Error
that replaced the entire findings array with one entry — so a GHCR
error on service N discarded findings already accumulated for
services 1..N-1. Move the rescue inside the per-service iteration
so each error becomes its own REFUSE finding and the loop continues.
- self.pin_and_verify asserted serviceInstanceUpdate == true but
discarded the serviceInstanceRedeploy result. A failed redeploy
could pass verification because the update mutation had already
advanced source.image+updatedAt. Require truthy redeploy result;
raise MutationError otherwise, symmetric with the update check.
- check_p2_staging_deployments already guarded meta.is_a?(Hash) so it
doesn't crash on a String meta, but the silent skip of the in-flight
race-check was invisible. Add a WARN finding so the skip is visible.
SUCCESS status remains the real gate (still REFUSE).
- execute_promotion now tracks already-pinned services and, on a
mid-loop MutationError, emits a loud PARTIAL PROMOTION report
naming both the already-pinned services and the failing one with
a pointer at bin/railway rollback-commit. Auto-rollback is left as
a follow-up — the goal here is just to make the mixed-state loud
and actionable rather than a quiet exit 1.
The showcase deploy model is STAGING = mutable :latest tag, PROD = immutable
@sha256: digest (P6 enforces both shapes). SnapshotCommand#build_snapshot
stored the raw serviceInstance.source.image, so for staging svc["image"] was
the :latest TAG. execute_promotion was pinning THAT mutable tag to prod via
serviceInstanceUpdate, defeating the immutable-prod invariant before
pin_and_verify raised on the nil expected_digest.
Fix: add PromoteCommand#resolved_prod_image — returns the staging svc as
@sha256:-pinned (pass-through if already pinned; resolves the tag via the
shared GHCR client otherwise; returns nil if the tag cannot be resolved).
execute_promotion now refuses (P0) rather than pin a mutable tag, and
check_p1_ghcr_digests verifies the resolved digest (it previously SKIPPED
tag-form images entirely, so :latest was never P1-checked).
Also:
- P2 race-check guards latest["meta"] when Railway returns a JSON String
(deserialized as Ruby String, not Hash) — .dig used to crash with
NoMethodError. SUCCESS status remains the real gate.
- Remove dead --include-startcommand flag (never read; doubly inert because
P6 REFUSEs on any startCommand divergence).
- Spec hygiene: P3 skip-test raises if probe runs under --no-require-staging-
green; P6 warn-proceed stubs execute_promotion to isolate the gate and
asserts rc==0; test_ghcr_token teardown unconditionally deletes
GITHUB_TOKEN/GHCR_TOKEN/RAILWAY_TOKEN before restoring priors.
70 runs, 204 assertions, 0 failures (up from 66/188 baseline).
P3 is the live re-probe gate spec §7.2 requires: CI history is not
authoritative for staging-green because showcase_deploy.yml uses
cancel-in-progress (the most recent CI run may have aborted before
the probe ran). Promote shells out to Workstream A's parameterized
verify-deploy.ts entrypoint at promote time and refuses on a red
result. Default-on; can be disabled with --no-require-staging-green
(prints "P3 SKIPPED" so the bypass is visible in logs).
run_staging_probe builds a clean child env (RAILWAY_TOKEN, GHCR_TOKEN,
GITHUB_TOKEN, PATH, HOME) and IO.popens `npx --yes tsx
showcase/scripts/verify-deploy.ts --env staging --services <csv>`.
Exit 0 = green, non-zero = red; the last 10 lines of stdout become
the human summary in the REFUSE message.
3 new P3 tests (red probe REFUSE, skip when flag off, green probe
pass). Also stubs run_staging_probe in the P1 and P2 test fakes so
those tests don't shell out to tsx (full suite stays sub-10ms).
Adds the P6 staging/prod parity matrix to promote:
- REFUSE on startCommand, healthcheckPath, or image-shape divergence
(staging is expected :tag/mutable, prod is expected :digest/pinned).
- WARN on region, replicas, restartPolicy, or env-var KEY-set divergence.
WARN findings refuse the promote unless --confirm-divergence is set,
at which point they print "[--confirm-divergence set] proceeding past
N WARN finding(s)" and continue.
- Env var VALUES are never compared (staging/prod hold different
secrets/URLs by design); the NOTE wired in commit #2's
run_with_preflight_only is preserved.
Snapshot schema bumped to v2:
- SERVICE_INSTANCE_QUERY adds healthcheckPath, region, numReplicas,
restartPolicyType.
- SnapshotCommand#build_snapshot maps them onto healthcheck_path,
region, replicas, restart_policy in the snapshot hash.
- SnapshotIO.SCHEMA_VERSION = 2 with SUPPORTED_VERSIONS = [1, 2] so
rollback-commit can still replay v1 snapshots from historical SHAs.
PromoteCommand.image_shape classifies a ref as :digest / :tag /
:missing / :other.
New tests: 6 P6 cases (startCommand REFUSE, healthcheckPath REFUSE,
image-shape REFUSE, WARN-without-confirm refusal, WARN-with-confirm
proceed, every-run NOTE). Two new snapshot tests: v2 captures new
fields end-to-end, and SnapshotIO.read accepts both v1 and v2.
serviceInstanceUpdate returns a Boolean scalar — the spec requires we
confirm that boolean is true AND re-query serviceInstance to confirm
BOTH source.image advanced to the new digest AND updatedAt strictly
advanced past the pre-mutation value. Image-equality alone is
insufficient: a no-op re-pin to the current value would otherwise
appear green.
Implementation:
- PromoteCommand::SERVICE_INSTANCE_RECHECK_QUERY: minimal query adding
updatedAt (kept separate from snapshot's SERVICE_INSTANCE_QUERY to
avoid disturbing snapshot behavior).
- PromoteCommand.pin_and_verify: pre-query updatedAt, run
serviceInstanceUpdate, assert boolean true, run serviceInstanceRedeploy,
then re-query up to RETRY_COUNT=3 times with RETRY_DELAY_SEC=10s
apart. Each retry must observe BOTH gates green (image match AND
updatedAt > pre_update_ts). Otherwise raises PromoteCommand::MutationError.
- execute_promotion now calls pin_and_verify (instead of
RestoreCommand.pin_and_redeploy) so promote inherits the verification.
MutationError is caught and converted to exit 1.
5 new P5 tests: boolean=false refusal, happy-path success,
image-advanced-but-ts-stale refusal, all-retries-stale refusal, and
late-third-retry success.
Implements check_p2_staging_deployments via
RollbackCommand::DEPLOYMENTS_QUERY (input:{serviceId,environmentId})
against STAGING_ENV_ID. For each staging service we are promoting,
requires that the most recent staging deployment is SUCCESS and its
deployed image digest matches the digest we are about to promote.
A mismatch indicates an in-flight build that landed mid-promote;
refusing prevents racing a newer digest into prod.
Refuse messages:
- "no staging deployments found" when edges is empty.
- "latest staging deployment status is <STATUS>, not SUCCESS".
- "in-flight race - latest staging deployment is <NEW> but snapshot
has <OLD>. Re-snapshot and retry."
Also adjusts test_promote_p1 fakes: preflight checks accumulate
findings before any short-circuit, so P2 still issues its gql query
even on a P1 REFUSE. Replaces the raise-on-call FakeGQL with a
benign empty-deployments FakeGQLEmpty so P1 cases stay focused.
3 new P2 tests (FAILED status, digest race, clean SUCCESS).
Full suite: 50 runs, 141 assertions, 0 failures.
Splits the previously-monolithic PromoteCommand#run into:
- capture_snapshots: pulls staging + prod snapshots (test-injectable).
- run_with_preflight_only: runs the P1..P6 preconditions, prints the
mandatory "env var VALUES are not compared" NOTE, gates on
REFUSE/WARN findings, then defers to execute_promotion.
- execute_promotion: the actual pin+redeploy loop.
Adds the P1 GHCR digest existence gate: every staging-side image
(@sha256:...) is HEAD-checked via GHCR.manifest_exists before any
serviceInstanceUpdate is issued. Tri-state result drives explicit
REFUSE messages (:missing -> garbage-collected hint; :auth_failed ->
GHCR_TOKEN / GITHUB_TOKEN hint). P2/P3/P6 land as []-returning stubs
for later phases (D.2 / D.3 / D.5).
Also adds parser flags --confirm-divergence,
--require-staging-green / --no-require-staging-green and
default_options that flips require_staging_green default-on per
spec §7.2 P3.
3 new P1 tests (REFUSE on :missing, REFUSE on :auth_failed with the
expected hint, clean pass on :exists). Existing suite stays green.
Adds two foundational helpers for the promote-hardening P1 check:
- Railway::Auth.ghcr_token: resolves a GHCR bearer separately from the
Railway API token. Prefers GHCR_TOKEN, falls back to GITHUB_TOKEN
(CI workflow token with packages:read). Returns nil if neither is set
so callers can refuse rather than silently fall through to anonymous.
- Railway::GHCR#manifest_exists: digest-existence HEAD against
ghcr.io/v2/<repo>/manifests/<sha256:...>. Returns tri-state
(:exists/:missing/:auth_failed); raises on 5xx; raises ArgumentError
if caller passes an unpinned tag (programmer error — P1 verifies the
concrete bytes about to ship).
GHCR.new default token source switched from ENV["GHCR_TOKEN"] direct to
Railway::Auth.ghcr_token, which adds the GITHUB_TOKEN fallback. Audited
all GHCR.new callers: BaseCommand#ghcr uses the default (intended new
behavior); all existing spec callers pass explicit token: kwarg so they
are unaffected.
Tests: 4 ghcr_token cases + 5 manifest_exists cases.
bin/railway now derives EXPECTED_DOMAINS from
showcase/scripts/railway-envs.generated.json instead of maintaining a
parallel Ruby hash. The TS railway-envs.ts is canonical; CI guards drift
via `emit-railway-envs-json.ts --check`. Adds a Minitest parity test that
boots Ruby and asserts its derived EXPECTED_DOMAINS matches the SSOT
JSON (public hosts only, env-id keys match SSOT envIds). Refs spec §3a.
[BLITZ:A6]
The tool was written from a spec but never exercised against live Railway,
so several queries reference fields that do not exist in the public
schema. Live runs (including the CI lint-prod step) failed with errors
like `Cannot query field "domains" on type "Project"`. Unit tests passed
because they only covered parsing and IO, not the GraphQL shape.
Verified the live schema via introspection (2026-05) and corrected
every mismatch:
- Project has no `domains` field. The previous `customDomains: domains
{ customDomains { ... } }` block on the Project selection is gone.
Custom domains now come from `serviceInstance.domains.customDomains`
for the env we are inspecting.
- Service has no `serviceInstances` field. We can no longer enumerate
per-env instances by nesting under Service. The new flow is:
1. SERVICES_LIST_QUERY -> list services in the project
2. SERVICE_INSTANCE_QUERY -> per (service, env), fetch source,
startCommand, latestDeployment, domains
3. ENVIRONMENT_VARIABLES_QUERY -> all variables in the env, then
group keys by Variable.serviceId for per-service env_keys
- `serviceInstanceDeployV2` does not accept an `image` argument; its
signature is (commitSha, environmentId, serviceId). To pin a service
to a specific image we now use
serviceInstanceUpdate(input: { source: { image } })
followed by serviceInstanceRedeploy. RestoreCommand exposes a
pin_and_redeploy class method that PromoteCommand and PinCommand
reuse.
- `deploymentRollback` returns scalar Boolean, so the previous
`deploymentRollback(id: $id) { id }` was invalid GraphQL — selection
sets are not allowed on scalars. Dropped the selection set.
- Auth.token now reads `user.accessToken` from ~/.railway/config.json
first. The legacy `user.token` field is a short CLI session token
(4 chars on a fresh login) that does not authenticate against the
public GraphQL API and was producing silent "Not Authorized" errors.
Added spec/test_snapshot_graphql.rb with a FakeGQL that returns realistic
shapes, plus regression guards that fail the build if anyone reintroduces
`Project.domains`, `Service.serviceInstances`, `serviceInstanceDeployV2`
with an image arg, or a selection set on `deploymentRollback`.
Smoke-tested live (read-only) against the showcase project:
- lint-prod: "OK: all production services digest-pinned." (27 services)
- snapshot --env staging: full YAML with images, startCommands, env_keys
- env-diff staging production: 32 drift findings (expected, staging is
not digest-pinned)
- resolve-digest ghcr.io/copilotkit/showcase-aimock:latest: digest
returned successfully
No mutating subcommands (restore, rollback, promote, pin) were run live.
Make the lint-prod audit result legible without having to click into the
workflow logs. Two surfaces, both rendered from the same JSON payload:
1. `$GITHUB_STEP_SUMMARY` — structured markdown block at the top of every
workflow run page. Shows on every event (push, pull_request,
workflow_dispatch).
2. Sticky PR comment — one comment per PR, keyed by the HTML marker
`<!-- lint-prod-sticky-comment -->`. Re-runs update the same comment via
`gh api -X PATCH` instead of creating duplicates. Plain `gh` CLI only,
no third-party action.
Both surfaces show: one-line status, a table of the unpinned services only
(not all 27), and a Pacific-time run timestamp with the finding count.
To support the renderer, add `--format json` to `lint-prod`:
{services:[{name,source,status}], findings:N, timestamp:"ISO8601"}
The workflow consumes this shape and also writes `findings` to
`$GITHUB_OUTPUT` so downstream jobs (future Slack alert) can compare runs.
Idempotent: re-running the workflow finds the existing comment by marker
and PATCHes it — never duplicates.
Add --exit-zero flag to bin/railway lint-prod that makes the command exit 0
even when digest-pinning findings exist. Findings still print to stdout so
they remain visible in the CI step log.
Wire the showcase_lint_prod.yml workflow to pass --exit-zero so the job
cannot block PRs while we soak the check against real production state.
Once we have confidence the findings are clean, remove --exit-zero from
the workflow step to flip lint-prod to enforcing.
Updates README to document the advisory-mode behavior and the path to
flipping the check to enforcing.
PR #4449 removed the eval tier system while unaware #4448 had just
merged. Restoring orchestrator, matrix, scope, config, GHA workflow,
and CLI wiring with all 8 CR fixes intact.