Bug #2: promote used serviceInstanceRedeploy, which replays the EXISTING
deployment and never pulls the newly-pinned digest, so prod could keep
serving stale. Switch to serviceInstanceDeployV2 to spawn a NEW deployment
that pulls the pinned digest, then verify_serving_digest! fail-loud asserts
the new deployment reaches SUCCESS and its meta.imageDigest == the pinned
digest. Update the promote mock-GraphQL fixtures across the spec suite to
return serviceInstanceDeployV2 + meta.imageDigest accordingly.
build_snapshot enumerates every project service and queries each one's
serviceInstance. The only guard was `next if inst.nil?` — it handled a
NULL result but not a THROWN `GraphQL: ServiceInstance not found` error
(a half-deleted service that still appears in the project service list
but has no instance in the env). That error bubbled to Railway.run's
top-level `rescue GraphQL::Error` and aborted the ENTIRE promote with an
opaque exit 2 before any preflight/divergence logic ran (run 27144525566
killed the docs promote this way).
Scope the rescue narrowly to ONLY the per-service "ServiceInstance not
found" message — log+skip that one service exactly like the nil case —
so every other GraphQL failure (auth, rate-limit, schema drift) still
propagates fail-loud. Adds red-green coverage: a single thrown not-found
is skipped (healthy services still snapshot), while an unrelated GraphQL
error still raises.
Adds the P6 staging/prod parity matrix to promote:
- REFUSE on startCommand, healthcheckPath, or image-shape divergence
(staging is expected :tag/mutable, prod is expected :digest/pinned).
- WARN on region, replicas, restartPolicy, or env-var KEY-set divergence.
WARN findings refuse the promote unless --confirm-divergence is set,
at which point they print "[--confirm-divergence set] proceeding past
N WARN finding(s)" and continue.
- Env var VALUES are never compared (staging/prod hold different
secrets/URLs by design); the NOTE wired in commit #2's
run_with_preflight_only is preserved.
Snapshot schema bumped to v2:
- SERVICE_INSTANCE_QUERY adds healthcheckPath, region, numReplicas,
restartPolicyType.
- SnapshotCommand#build_snapshot maps them onto healthcheck_path,
region, replicas, restart_policy in the snapshot hash.
- SnapshotIO.SCHEMA_VERSION = 2 with SUPPORTED_VERSIONS = [1, 2] so
rollback-commit can still replay v1 snapshots from historical SHAs.
PromoteCommand.image_shape classifies a ref as :digest / :tag /
:missing / :other.
New tests: 6 P6 cases (startCommand REFUSE, healthcheckPath REFUSE,
image-shape REFUSE, WARN-without-confirm refusal, WARN-with-confirm
proceed, every-run NOTE). Two new snapshot tests: v2 captures new
fields end-to-end, and SnapshotIO.read accepts both v1 and v2.
The tool was written from a spec but never exercised against live Railway,
so several queries reference fields that do not exist in the public
schema. Live runs (including the CI lint-prod step) failed with errors
like `Cannot query field "domains" on type "Project"`. Unit tests passed
because they only covered parsing and IO, not the GraphQL shape.
Verified the live schema via introspection (2026-05) and corrected
every mismatch:
- Project has no `domains` field. The previous `customDomains: domains
{ customDomains { ... } }` block on the Project selection is gone.
Custom domains now come from `serviceInstance.domains.customDomains`
for the env we are inspecting.
- Service has no `serviceInstances` field. We can no longer enumerate
per-env instances by nesting under Service. The new flow is:
1. SERVICES_LIST_QUERY -> list services in the project
2. SERVICE_INSTANCE_QUERY -> per (service, env), fetch source,
startCommand, latestDeployment, domains
3. ENVIRONMENT_VARIABLES_QUERY -> all variables in the env, then
group keys by Variable.serviceId for per-service env_keys
- `serviceInstanceDeployV2` does not accept an `image` argument; its
signature is (commitSha, environmentId, serviceId). To pin a service
to a specific image we now use
serviceInstanceUpdate(input: { source: { image } })
followed by serviceInstanceRedeploy. RestoreCommand exposes a
pin_and_redeploy class method that PromoteCommand and PinCommand
reuse.
- `deploymentRollback` returns scalar Boolean, so the previous
`deploymentRollback(id: $id) { id }` was invalid GraphQL — selection
sets are not allowed on scalars. Dropped the selection set.
- Auth.token now reads `user.accessToken` from ~/.railway/config.json
first. The legacy `user.token` field is a short CLI session token
(4 chars on a fresh login) that does not authenticate against the
public GraphQL API and was producing silent "Not Authorized" errors.
Added spec/test_snapshot_graphql.rb with a FakeGQL that returns realistic
shapes, plus regression guards that fail the build if anyone reintroduces
`Project.domains`, `Service.serviceInstances`, `serviceInstanceDeployV2`
with an image arg, or a selection set on `deploymentRollback`.
Smoke-tested live (read-only) against the showcase project:
- lint-prod: "OK: all production services digest-pinned." (27 services)
- snapshot --env staging: full YAML with images, startCommands, env_keys
- env-diff staging production: 32 drift findings (expected, staging is
not digest-pinned)
- resolve-digest ghcr.io/copilotkit/showcase-aimock:latest: digest
returned successfully
No mutating subcommands (restore, rollback, promote, pin) were run live.