Commit Graph

5 Commits

Author SHA1 Message Date
Jordan Ritter a5186c50c2 feat(showcase): ruby promote preflight (service-ref/replicate/resource) + lint-prod starter coverage 2026-06-19 12:23:23 -07:00
Jordan Ritter 6865c3d1b8 fix(showcase): activate prod pin via serviceInstanceDeployV2 + verify running==pinned
Bug #2: promote used serviceInstanceRedeploy, which replays the EXISTING
deployment and never pulls the newly-pinned digest, so prod could keep
serving stale. Switch to serviceInstanceDeployV2 to spawn a NEW deployment
that pulls the pinned digest, then verify_serving_digest! fail-loud asserts
the new deployment reaches SUCCESS and its meta.imageDigest == the pinned
digest. Update the promote mock-GraphQL fixtures across the spec suite to
return serviceInstanceDeployV2 + meta.imageDigest accordingly.
2026-06-18 16:20:43 -07:00
Jordan Ritter b43d64ea49 fix(showcase): tolerate per-service "ServiceInstance not found" in promote snapshot
build_snapshot enumerates every project service and queries each one's
serviceInstance. The only guard was `next if inst.nil?` — it handled a
NULL result but not a THROWN `GraphQL: ServiceInstance not found` error
(a half-deleted service that still appears in the project service list
but has no instance in the env). That error bubbled to Railway.run's
top-level `rescue GraphQL::Error` and aborted the ENTIRE promote with an
opaque exit 2 before any preflight/divergence logic ran (run 27144525566
killed the docs promote this way).

Scope the rescue narrowly to ONLY the per-service "ServiceInstance not
found" message — log+skip that one service exactly like the nil case —
so every other GraphQL failure (auth, rate-limit, schema drift) still
propagates fail-loud. Adds red-green coverage: a single thrown not-found
is skipped (healthy services still snapshot), while an unrelated GraphQL
error still raises.
2026-06-08 10:45:07 -07:00
Jordan Ritter 8114eda52c feat(showcase): extend snapshot schema (v2) with healthcheck/region/replicas/restartPolicy + add P6 parity matrix
Adds the P6 staging/prod parity matrix to promote:
- REFUSE on startCommand, healthcheckPath, or image-shape divergence
  (staging is expected :tag/mutable, prod is expected :digest/pinned).
- WARN on region, replicas, restartPolicy, or env-var KEY-set divergence.
  WARN findings refuse the promote unless --confirm-divergence is set,
  at which point they print "[--confirm-divergence set] proceeding past
  N WARN finding(s)" and continue.
- Env var VALUES are never compared (staging/prod hold different
  secrets/URLs by design); the NOTE wired in commit #2's
  run_with_preflight_only is preserved.

Snapshot schema bumped to v2:
- SERVICE_INSTANCE_QUERY adds healthcheckPath, region, numReplicas,
  restartPolicyType.
- SnapshotCommand#build_snapshot maps them onto healthcheck_path,
  region, replicas, restart_policy in the snapshot hash.
- SnapshotIO.SCHEMA_VERSION = 2 with SUPPORTED_VERSIONS = [1, 2] so
  rollback-commit can still replay v1 snapshots from historical SHAs.

PromoteCommand.image_shape classifies a ref as :digest / :tag /
:missing / :other.

New tests: 6 P6 cases (startCommand REFUSE, healthcheckPath REFUSE,
image-shape REFUSE, WARN-without-confirm refusal, WARN-with-confirm
proceed, every-run NOTE). Two new snapshot tests: v2 captures new
fields end-to-end, and SnapshotIO.read accepts both v1 and v2.
2026-05-29 11:45:11 -07:00
Jordan Ritter 9def341da4 fix(showcase): align bin/railway GraphQL with Railway public schema
The tool was written from a spec but never exercised against live Railway,
so several queries reference fields that do not exist in the public
schema. Live runs (including the CI lint-prod step) failed with errors
like `Cannot query field "domains" on type "Project"`. Unit tests passed
because they only covered parsing and IO, not the GraphQL shape.

Verified the live schema via introspection (2026-05) and corrected
every mismatch:

- Project has no `domains` field. The previous `customDomains: domains
  { customDomains { ... } }` block on the Project selection is gone.
  Custom domains now come from `serviceInstance.domains.customDomains`
  for the env we are inspecting.
- Service has no `serviceInstances` field. We can no longer enumerate
  per-env instances by nesting under Service. The new flow is:
    1. SERVICES_LIST_QUERY -> list services in the project
    2. SERVICE_INSTANCE_QUERY -> per (service, env), fetch source,
       startCommand, latestDeployment, domains
    3. ENVIRONMENT_VARIABLES_QUERY -> all variables in the env, then
       group keys by Variable.serviceId for per-service env_keys
- `serviceInstanceDeployV2` does not accept an `image` argument; its
  signature is (commitSha, environmentId, serviceId). To pin a service
  to a specific image we now use
    serviceInstanceUpdate(input: { source: { image } })
  followed by serviceInstanceRedeploy. RestoreCommand exposes a
  pin_and_redeploy class method that PromoteCommand and PinCommand
  reuse.
- `deploymentRollback` returns scalar Boolean, so the previous
  `deploymentRollback(id: $id) { id }` was invalid GraphQL — selection
  sets are not allowed on scalars. Dropped the selection set.
- Auth.token now reads `user.accessToken` from ~/.railway/config.json
  first. The legacy `user.token` field is a short CLI session token
  (4 chars on a fresh login) that does not authenticate against the
  public GraphQL API and was producing silent "Not Authorized" errors.

Added spec/test_snapshot_graphql.rb with a FakeGQL that returns realistic
shapes, plus regression guards that fail the build if anyone reintroduces
`Project.domains`, `Service.serviceInstances`, `serviceInstanceDeployV2`
with an image arg, or a selection set on `deploymentRollback`.

Smoke-tested live (read-only) against the showcase project:
- lint-prod: "OK: all production services digest-pinned." (27 services)
- snapshot --env staging: full YAML with images, startCommands, env_keys
- env-diff staging production: 32 drift findings (expected, staging is
  not digest-pinned)
- resolve-digest ghcr.io/copilotkit/showcase-aimock:latest: digest
  returned successfully

No mutating subcommands (restore, rollback, promote, pin) were run live.
2026-05-28 11:05:25 -07:00