- Reject empty/whitespace-only service in parseBuildOutputs,
mergeBuildResultFiles, and buildResultArtifactName (was only
buildResultArtifactName, and only for length===0).
- mergeBuildResultFiles now throws on duplicate service names across
slots so an upstream dispatch-name collision surfaces instead of
letting a failure+success pair spuriously look like a success in
successSet (fail-loud discipline).
- Derive BuildOutcome union and VALID_STATUSES set from a single
BUILD_OUTCOMES as-const tuple plus a compile-time exhaustiveness
assignment so they cannot drift.
- Extract validateServiceBuildResult shared validator used by both
parseBuildOutputs (context: 'entry[i]') and mergeBuildResultFiles
(context: 'slot[i]') — one set of rules, one error format, both
now include the offending index.
- Trim shouldRedeployStaging JSDoc to what/why only (dropped the
external-caller enumeration) and note in the module header that
the single-result.json-per-slot invariant is enforced workflow-side,
not by the parser. shouldRedeployStaging([]) === false behavior
unchanged.
Tests: 16 -> 23 passing, all new red-then-green; tsc 0.
Treat whitespace-only token strings as empty so the resolver falls through
to the next candidate instead of returning a bearer that fails at Railway
with a confusing 401.
Guard against non-object JSON input (null/undefined/string/number/array)
defensively before property access — the config originates from
JSON.parse of ~/.railway/config.json (untrusted).
Replace non-null assertions on re-accessed expressions with
locally-captured values so the nonEmpty narrowing actually applies; no
behavior change for the happy path.
Comment cleanups: replace stale line-number reference with a symbolic
one, reword the Ruby-parity note to acknowledge the per-project token
fallback that is intentionally not honored, drop the rotting
"(43+ chars)" parenthetical, and soften the deprecation timeline to
"a future release."
A.7 rewrote showcase_deploy.yml to be SSOT-driven (no hardcoded options
list, no inline ALL_SERVICES). Webhooks remains canonically wired via
the SSOT (probe.staging:true, dispatchName:"webhooks"); test now pins
that contract instead of the now-obsolete literal regexes.
P3 is the live re-probe gate spec §7.2 requires: CI history is not
authoritative for staging-green because showcase_deploy.yml uses
cancel-in-progress (the most recent CI run may have aborted before
the probe ran). Promote shells out to Workstream A's parameterized
verify-deploy.ts entrypoint at promote time and refuses on a red
result. Default-on; can be disabled with --no-require-staging-green
(prints "P3 SKIPPED" so the bypass is visible in logs).
run_staging_probe builds a clean child env (RAILWAY_TOKEN, GHCR_TOKEN,
GITHUB_TOKEN, PATH, HOME) and IO.popens `npx --yes tsx
showcase/scripts/verify-deploy.ts --env staging --services <csv>`.
Exit 0 = green, non-zero = red; the last 10 lines of stdout become
the human summary in the REFUSE message.
3 new P3 tests (red probe REFUSE, skip when flag off, green probe
pass). Also stubs run_staging_probe in the P1 and P2 test fakes so
those tests don't shell out to tsx (full suite stays sub-10ms).
Adds the P6 staging/prod parity matrix to promote:
- REFUSE on startCommand, healthcheckPath, or image-shape divergence
(staging is expected :tag/mutable, prod is expected :digest/pinned).
- WARN on region, replicas, restartPolicy, or env-var KEY-set divergence.
WARN findings refuse the promote unless --confirm-divergence is set,
at which point they print "[--confirm-divergence set] proceeding past
N WARN finding(s)" and continue.
- Env var VALUES are never compared (staging/prod hold different
secrets/URLs by design); the NOTE wired in commit #2's
run_with_preflight_only is preserved.
Snapshot schema bumped to v2:
- SERVICE_INSTANCE_QUERY adds healthcheckPath, region, numReplicas,
restartPolicyType.
- SnapshotCommand#build_snapshot maps them onto healthcheck_path,
region, replicas, restart_policy in the snapshot hash.
- SnapshotIO.SCHEMA_VERSION = 2 with SUPPORTED_VERSIONS = [1, 2] so
rollback-commit can still replay v1 snapshots from historical SHAs.
PromoteCommand.image_shape classifies a ref as :digest / :tag /
:missing / :other.
New tests: 6 P6 cases (startCommand REFUSE, healthcheckPath REFUSE,
image-shape REFUSE, WARN-without-confirm refusal, WARN-with-confirm
proceed, every-run NOTE). Two new snapshot tests: v2 captures new
fields end-to-end, and SnapshotIO.read accepts both v1 and v2.
serviceInstanceUpdate returns a Boolean scalar — the spec requires we
confirm that boolean is true AND re-query serviceInstance to confirm
BOTH source.image advanced to the new digest AND updatedAt strictly
advanced past the pre-mutation value. Image-equality alone is
insufficient: a no-op re-pin to the current value would otherwise
appear green.
Implementation:
- PromoteCommand::SERVICE_INSTANCE_RECHECK_QUERY: minimal query adding
updatedAt (kept separate from snapshot's SERVICE_INSTANCE_QUERY to
avoid disturbing snapshot behavior).
- PromoteCommand.pin_and_verify: pre-query updatedAt, run
serviceInstanceUpdate, assert boolean true, run serviceInstanceRedeploy,
then re-query up to RETRY_COUNT=3 times with RETRY_DELAY_SEC=10s
apart. Each retry must observe BOTH gates green (image match AND
updatedAt > pre_update_ts). Otherwise raises PromoteCommand::MutationError.
- execute_promotion now calls pin_and_verify (instead of
RestoreCommand.pin_and_redeploy) so promote inherits the verification.
MutationError is caught and converted to exit 1.
5 new P5 tests: boolean=false refusal, happy-path success,
image-advanced-but-ts-stale refusal, all-retries-stale refusal, and
late-third-retry success.
Implements check_p2_staging_deployments via
RollbackCommand::DEPLOYMENTS_QUERY (input:{serviceId,environmentId})
against STAGING_ENV_ID. For each staging service we are promoting,
requires that the most recent staging deployment is SUCCESS and its
deployed image digest matches the digest we are about to promote.
A mismatch indicates an in-flight build that landed mid-promote;
refusing prevents racing a newer digest into prod.
Refuse messages:
- "no staging deployments found" when edges is empty.
- "latest staging deployment status is <STATUS>, not SUCCESS".
- "in-flight race - latest staging deployment is <NEW> but snapshot
has <OLD>. Re-snapshot and retry."
Also adjusts test_promote_p1 fakes: preflight checks accumulate
findings before any short-circuit, so P2 still issues its gql query
even on a P1 REFUSE. Replaces the raise-on-call FakeGQL with a
benign empty-deployments FakeGQLEmpty so P1 cases stay focused.
3 new P2 tests (FAILED status, digest race, clean SUCCESS).
Full suite: 50 runs, 141 assertions, 0 failures.
Splits the previously-monolithic PromoteCommand#run into:
- capture_snapshots: pulls staging + prod snapshots (test-injectable).
- run_with_preflight_only: runs the P1..P6 preconditions, prints the
mandatory "env var VALUES are not compared" NOTE, gates on
REFUSE/WARN findings, then defers to execute_promotion.
- execute_promotion: the actual pin+redeploy loop.
Adds the P1 GHCR digest existence gate: every staging-side image
(@sha256:...) is HEAD-checked via GHCR.manifest_exists before any
serviceInstanceUpdate is issued. Tri-state result drives explicit
REFUSE messages (:missing -> garbage-collected hint; :auth_failed ->
GHCR_TOKEN / GITHUB_TOKEN hint). P2/P3/P6 land as []-returning stubs
for later phases (D.2 / D.3 / D.5).
Also adds parser flags --confirm-divergence,
--require-staging-green / --no-require-staging-green and
default_options that flips require_staging_green default-on per
spec §7.2 P3.
3 new P1 tests (REFUSE on :missing, REFUSE on :auth_failed with the
expected hint, clean pass on :exists). Existing suite stays green.
Adds two foundational helpers for the promote-hardening P1 check:
- Railway::Auth.ghcr_token: resolves a GHCR bearer separately from the
Railway API token. Prefers GHCR_TOKEN, falls back to GITHUB_TOKEN
(CI workflow token with packages:read). Returns nil if neither is set
so callers can refuse rather than silently fall through to anonymous.
- Railway::GHCR#manifest_exists: digest-existence HEAD against
ghcr.io/v2/<repo>/manifests/<sha256:...>. Returns tri-state
(:exists/:missing/:auth_failed); raises on 5xx; raises ArgumentError
if caller passes an unpinned tag (programmer error — P1 verifies the
concrete bytes about to ship).
GHCR.new default token source switched from ENV["GHCR_TOKEN"] direct to
Railway::Auth.ghcr_token, which adds the GITHUB_TOKEN fallback. Audited
all GHCR.new callers: BaseCommand#ghcr uses the default (intended new
behavior); all existing spec callers pass explicit token: kwarg so they
are unaffected.
Tests: 4 ghcr_token cases + 5 manifest_exists cases.
New workflow_dispatch-only workflow that promotes a staging-tested digest
to prod. Four jobs in strict order:
1. verify-staging-precondition (live re-probe, refuse on red)
2. promote (bin/railway promote; spec §7 preconditions P1..P6)
3. verify-prod (verify-deploy.ts --env prod)
4. notify (Slack #oss-alerts on red; never #engr)
Driven entirely off the railway-envs.ts SSOT. Refs spec §3.
showcase_deploy.yml is now SSOT-driven: env id and service set both come
from railway-envs.generated.json (no inline matrix, no hardcoded
RAILWAY_ENV_ID). Probes staging on workflow_run (was incorrectly probing
prod). Reads redeploy-env's per-service JSON summary uploaded as an
artifact and fails the workflow on any staging status:error while
verify still runs against the success-set. Preserves PR #5093's
redeploy-env.ts exit semantics unchanged. Refs spec §3.
bin/railway now derives EXPECTED_DOMAINS from
showcase/scripts/railway-envs.generated.json instead of maintaining a
parallel Ruby hash. The TS railway-envs.ts is canonical; CI guards drift
via `emit-railway-envs-json.ts --check`. Adds a Minitest parity test that
boots Ruby and asserts its derived EXPECTED_DOMAINS matches the SSOT
JSON (public hosts only, env-id keys match SSOT envIds). Refs spec §3a.
[BLITZ:A6]
New TS probe driven off railway-envs SSOT. Accepts --env staging|prod and
optional --services CSV; iterates SERVICES where probe[env]===true and
dispatches to per-driver feature-level verifiers. Refuses to start when a
probe-required service is missing a domain for the requested env (no
silent skip). HTTP 200 is necessary but not sufficient.
Adds verify-deploy.drivers.ts dispatch with exhaustive never check on
ProbeDriver, plus one stub module per ProbeDriver literal (shell, docs,
dashboard, dojo, harness, eval, aimock, pocketbase, webhooks, agent).
Stubs fail loud with an explicit "not yet implemented" error so any
accidental real-network invocation surfaces; per-driver feature-level
impls land as subsequent micro-tasks. Refs spec section 3 / section 3.5.
[BLITZ:A5]
Adds optional REDEPLOY_SUMMARY_JSON path; when set, redeploy-env writes
a structured per-service record array {service,status,error?}. PR #5093's
exit-code contract is preserved (staging=0, prod=1-on-failure).
Consumed by showcase_deploy.yml to fail the workflow on staging per-service
errors without changing the script's exit semantics. Refs spec §3.
getRuntimeConfig() in each shell's runtime-config.client.ts threw when
typeof window === 'undefined'. But Next.js App Router executes 'use
client' component bodies on the SERVER during initial SSR, so any client
component that called getRuntimeConfig() in its render body 500'd the
page. shell-dashboard already had the fix.
Mirror shell-dashboard's pattern: return a typed SSR_PLACEHOLDER (empty
strings for URL/key fields; {} for shell-dojo whose RuntimeConfig is
empty) when window is undefined. Keep the loud throw when window IS
present but window.__SHOWCASE_CONFIG__ is missing — that's a genuine
wiring bug and should not be masked.
Updated shell-docs and shell client tests: replace 'throws on server'
case with 'returns SSR sentinel placeholder' assertion matching each
shell's RuntimeConfig shape. shell-dojo has no client test so verified
via tsc only.
Close-out proof for Option B: load each public shell (shell, shell-docs,
shell-dashboard) on staging and prod and assert that
(a) the inlined `window.__SHOWCASE_CONFIG__` matches the env's
expected URL set (per-env value, not a leaked default), and
(b) every backend fetch host matches a tight per-env allowlist —
anchored regexes pinned to the EXACT hosts captured from real
page-load inventories (Railway public domains for the staging
services, bare-domain copilotkit.ai hosts for prod, plus the
shared third-party allowlist for analytics/fonts/HubSpot/Reo/
REB2B/scarf/CDN).
An env-leak (someone re-bakes a URL into the artifact) shows up as
either the wrong `__SHOWCASE_CONFIG__` value OR a request to the
other-env's host — either branch fails the test.
This is the test referenced in plan-B B14 / spec §10 items 2,3,10.
Runs against LIVE deployments — no webServer block, fetches public
URLs. Gated to run in CI after the B15 Railway env-var wiring deploy
has settled.
Scoping notes:
- Spec file at `showcase/tests/env-routing.spec.ts`. The existing
`showcase/tests/playwright.config.ts` is the integrations smoke
harness (`testDir: ./e2e`); creating a separate, narrowly-scoped
config at `showcase/playwright.env-routing.config.ts` keeps the
two suites independent so neither can pull the other in by
accident under `playwright test`.
- `testMatch: /env-routing\.spec\.ts$/` belt-and-suspenders the
`testDir: ./tests` selection.
- Verified well-formed locally via `tsc --noEmit` against
`showcase/shell-dashboard/`'s @playwright/test + @types/node
install (the only place those deps are installed in the worktree;
`showcase/tests/` has no node_modules in this worktree because
npm install is symlinked-only by the blitz harness). Full
browser execution requires deployed shells and is gated to
post-deploy in CI.
Replay the Option B runtime-config spike as a vitest integration test
that guards the no-rebuild env switching property going forward:
- `next build` once with no per-env URL env vars (only a sentinel
OPS_BASE_URL so next.config.ts's rewrites() can validate; Next
evaluates rewrites at build, not only at start, so a placeholder
here is unavoidable — the assertions don't depend on it).
- `next start` twice, each on a fresh port and a DIFFERENT
POCKETBASE_URL / SHELL_URL / OPS_BASE_URL set.
- Fetch `/` on each boot and extract the inlined
`window.__SHOWCASE_CONFIG__={...}` JSON from the served HTML.
- Assert env-A URLs on the first boot and env-B URLs on the second
boot of the SAME built artifact. If anyone re-introduces a
build-time URL bake, the second boot's HTML still shows env-A
values and this test fails.
Test lives at `showcase/shell-dashboard/tests/runtime-env-switch.spike.test.ts`
and is picked up via a new vitest include for `tests/**/*.spike.test.ts`
(the default include is `src/**/*.test.{ts,tsx}` and the integration-
weight spike doesn't fit there). The `.spike.test.ts` suffix keeps
the include narrow so the visual snapshot suite under
`tests/visual/` stays out.
Total wall time locally: ~12s (build ~2s warm with prebuild data
generation already run; two boots ~10s combined). Heavy enough to
gate behind a `tests:integration` script in CI rather than running
on every push.
The client-side getRuntimeConfig() reader previously threw when window
was undefined, on the assumption that "use client" components only run
post-hydration. That assumption is wrong for the Next.js App Router:
"use client" component bodies ARE executed on the server during the
initial SSR pass (that's how the HTML stream is built before the JS
arrives), so throwing breaks SSR entirely — every page that uses a
client component which reads runtime config 500s on the first request
with `[runtime-config.client] getRuntimeConfig() called on the server`.
This was surfaced by the B13 spike-replay integration test (`next build`
once, `next start` twice with different POCKETBASE_URL/SHELL_URL/
OPS_BASE_URL): the inline `<script id="__showcase_config__">` injected
by the root layout never reaches the rendered DOM because the page
crashes during SSR and falls back to the Next.js error boundary, with
the would-be script content captured (and JSON-escaped) inside the
RSC streaming payload instead of as a real `<script>` tag.
Switch the SSR branch to return a sentinel RuntimeConfig with empty
strings rather than throwing. Client components see the placeholder
during the initial server render, then re-read post-hydration when
window.__SHOWCASE_CONFIG__ is populated. The post-hydration "config
missing" branch still throws so genuine wiring bugs (layout bypass,
empty injection) stay loud. Updated the matching unit test to assert
the new sentinel behavior in place of the old throw assertion.
Plan-B / Option-B migration moved every shell URL/analytics key off the
build-time NEXT_PUBLIC_* env channel and onto runtime config served via
__SHOWCASE_CONFIG__ + getRuntimeConfig(). To prevent a silent regression
where a future change reintroduces a direct process.env.NEXT_PUBLIC_*
read in shell code (which would re-freeze the value at build time and
break no-rebuild env switching), add a focused lint rule.
The rule (copilotkit/no-public-env-shell-read) is implemented as a
custom oxlint JS plugin rule in the existing copilotkit plugin and
enabled under shell-scoped overrides in .oxlintrc.json:
- Errors on process.env.NEXT_PUBLIC_<URL/ANALYTICS> reads in:
showcase/shell-dashboard/src/**, showcase/shell-docs/src/**,
showcase/shell/src/**, showcase/shell-dojo/src/**
- Banned keys: POCKETBASE_URL, SHELL_URL, BASE_URL, OPS_BASE_URL,
INTELLIGENCE_SIGNUP_URL, POSTHOG_KEY, POSTHOG_HOST, SCARF_PIXEL_ID,
GOOGLE_ANALYTICS_TRACKING_ID, REB2B_KEY, REO_KEY
- Intentionally allowed (NOT banned): NEXT_PUBLIC_COMMIT_SHA and
NEXT_PUBLIC_BRANCH (build-stamped artifact identifiers per B10/B11)
and NEXT_PUBLIC_LOCAL_BACKENDS (computed from shared/local-ports.json
at build, local-dev only).
- Excluded files (rule disabled via a follow-up override): MDX content
under shell-docs/src/content/**, runtime-config implementation files,
and *.test.{ts,tsx} / *.spec.{ts,tsx}. oxlint does not support
excludedFiles inside an override block, so the exclusion is expressed
as a later override that sets the rule to off.
Plan-B originally targeted oxlint's eslint/no-restricted-syntax with an
AST-selector regex. oxlint 1.x does not implement that rule (only
no-restricted-globals / no-restricted-imports), so the equivalent guard
is realized as a small custom rule in the existing copilotkit JS plugin
(meta.name=copilotkit), reusing the same plugin loader the repo already
has for require-cpk-prefix and no-single-arg-zod-record.
Verification (red-green): the rule fires on a fixture containing
process.env.NEXT_PUBLIC_POCKETBASE_URL and does NOT fire on a fixture
containing process.env.NEXT_PUBLIC_COMMIT_SHA. Test pins the config via
-c so it works inside git worktrees nested under .claude/worktrees/
where oxlint's automatic upward config search can miss the worktree's
own .oxlintrc.json.
All four shells lint clean: 0 errors of the new rule across
shell-dashboard (114 files), shell-docs (137), shell (29), shell-dojo (6).
Adds the inline <script id="__showcase_config__"> tag as the FIRST
child of <head> in shell-dojo's root layout (B6). Calls
getRuntimeConfig() server-side and serializes the result via the OWASP
recommended escape (< / U+2028 / U+2029) before injecting it as
window.__SHOWCASE_CONFIG__.
shell-dojo's RuntimeConfig is currently {}, so the injected value is
`window.__SHOWCASE_CONFIG__={};` — harmless but symmetric with the
other shells. The <script> sits ahead of the fonts <link> so it runs
before any other head-level script (including future next/script
beforeInteractive blocks).
Regex sources for U+2028 / U+2029 use ECMAScript escapes
(/\\u2028/g, /\\u2029/g) so the source compiles in TypeScript — a
literal codepoint in the regex source breaks tsc with TS1161
(unterminated regex literal). The regex engine resolves the escape at
runtime, so the substitution still targets the actual codepoint.
Mirror of the runtime-config pattern from shell-dashboard for shell-dojo
(B7). shell-dojo has no URL consumers today (B0 audit reported zero
process.env.NEXT_PUBLIC_* reads), so RuntimeConfig is an empty object
literal. Module exists to keep the runtime-config / layout-injection
pattern symmetric across all four shells; adding a URL later is a
single field addition.
- src/lib/runtime-config.ts: server reader with unstable_noStore()
opt-out (Node) and noStore-skip option (Edge wrapper not needed yet).
- src/lib/runtime-config.client.ts: client reader from
window.__SHOWCASE_CONFIG__ injected by root layout.
No tests included for shell-dojo: matches B7 file list (no test files
listed for shell-dojo) and reflects that there is no behavior to assert
on an empty config beyond the type contract.
Removes the env:{NEXT_PUBLIC_BASE_URL} entry that re-bakes the
build-time value of NEXT_PUBLIC_BASE_URL into every chunk (defeats
runtime injection). NEXT_PUBLIC_LOCAL_BACKENDS stays — it is computed
from shared/local-ports.json (a JSON file on disk, not an env var)
and only used in local-dev.
Refs plan-B §B10.3.
Replaces the module-load read of NEXT_PUBLIC_POSTHOG_HOST in
showcase/shell/src/middleware.ts (which Next inlines into the Edge
bundle at build time and freezes per artifact) with a per-request
read via getRuntimeConfigEdge().posthogHost. The Edge wrapper skips
unstable_noStore() — next/cache is not available in the Edge
runtime, and middleware always runs per-request so there is no
static cache to opt out of.
Refs plan-B §B9.6.
Adds a <head> element (shell previously had only <html> → <body>) and
emits an inline <script> as its first child that writes
window.__SHOWCASE_CONFIG__ from the server-side runtime config before
any client component mounts. The injection JSON is OWASP-escaped:
< → < (guards against </script> breakout from a hostile env
value), and U+2028 / U+2029 are escaped to / (line
separators are legal inside JSON strings but a syntax error inside a
JS string literal in pre-ES2019 engines / when parsed as
text/javascript).
The commit-sha overlay continues to read process.env.NEXT_PUBLIC_COMMIT_SHA
directly — COMMIT_SHA is build-stamped intentionally (identifies the
artifact, not the env).
Refs plan-B §B6.
Introduces showcase/shell/src/lib/runtime-config.ts (server-only —
imports next/cache and is read at request time by the root layout)
plus runtime-config.client.ts (reads window.__SHOWCASE_CONFIG__
injected by the layout). Shell's RuntimeConfig contains baseUrl and
posthogHost. getRuntimeConfigEdge() provides the Edge-runtime variant
for middleware (skips unstable_noStore).
Adds vitest config + dev deps to package.json and red-green tests for
both modules. Tests verify env-vs-fallback precedence, trailing-slash
stripping, no-module-load-freeze (live process.env reads per call),
and the Edge wrapper's noStore-skip behavior.
Refs plan-B §B7.
Two coupled changes that complete the Option B runtime-injection
switch for shell-docs:
- app/sitemap.ts + app/robots.ts: add `export const dynamic =
"force-dynamic"` so Next.js regenerates both routes per request.
Without this Next would statically prerender them at build time
and freeze whichever NEXT_PUBLIC_BASE_URL was set during
`next build` — the exact freeze that defeats the runtime-config
plumbing.
- next.config.ts: REMOVE the `next build` gates that threw when
NEXT_PUBLIC_BASE_URL or NEXT_PUBLIC_SHELL_URL were unset. Those
gates assumed build-time URL injection. Under Option B both URLs
are read per-request from process.env via getRuntimeConfig(), so
a single built artifact must be allowed to build with neither var
set — they're a deploy-time concern now. Missing-value surfacing
moves to `console.error` from runtime-config.ts at request time.
Migrate the remaining shell-docs consumers of NEXT_PUBLIC_* env vars
off of process.env reads and onto the runtime-config readers:
- lib/providers/posthog-provider.tsx (client): the PostHog key is
pulled inside PostHogProvider so the value reflects the current
deploy's NEXT_PUBLIC_POSTHOG_KEY; empty string disables analytics
via the existing truthiness gates around init/capture
- lib/providers/scarf-pixel.tsx (client): pixel id read at render
time; empty string returns null (existing no-op behavior preserved)
- lib/hooks/use-google-analytics.tsx (client): GA tracking id read
at hook invocation; empty string short-circuits via the existing
`if (!GA_ID) return` guard
- middleware.ts (Edge): POSTHOG_HOST is now resolved per-request via
getRuntimeConfigEdge() (the Edge wrapper skips unstable_noStore()
since next/cache is unavailable there and middleware always runs
per-request anyway). Server-side POSTHOG_KEY is unchanged — it's
not a NEXT_PUBLIC_* var.
Migrate every consumer of NEXT_PUBLIC_BASE_URL / NEXT_PUBLIC_SHELL_URL
/ NEXT_PUBLIC_INTELLIGENCE_SIGNUP_URL in the shell-docs tree off of
process.env reads and onto the runtime-config readers:
- lib/sitemap-helpers.ts (server): getBaseUrl() now delegates to
getRuntimeConfig().baseUrl
- components/search-modal.tsx (client): reads shellHost once per
render and threads it into normalizeHref() + the integration href
builder; normalizeHref is now parameterized rather than closing
over a module-scope SHELL_HOST
- components/integration-grid.tsx (client): reads shellHost after
the framework-scoped early return so we never touch the client
reader on no-op renders
- components/ai/page-actions.tsx (client): getClientBaseUrl()
delegates to getRuntimeConfig().baseUrl; the local inline copy is
retained so a "use client" file doesn't reach into the Node-only
sitemap-helpers module
- components/react/signup-link.tsx + ops-platform-cta.tsx (client):
buildHref() reads the signup URL lazily inside the function body
so the URL reflects the current deploy's env
Module-scope reads of process.env.NEXT_PUBLIC_* are eliminated from
all six files — every URL is now resolved at render time from the
runtime-config object the root layout injects.
Add inline <script> tag as the first child of <head> that populates
window.__SHOWCASE_CONFIG__ with values read from process.env at request
time via getRuntimeConfig(). This is the server side of the Option B
runtime-injection plumbing — every shell-docs client component will
read its URLs/analytics keys from this object instead of compiled-in
NEXT_PUBLIC_* values, so a single built artifact can serve staging
and prod by changing Railway env vars.
The serializer escapes < (XSS guard against </script> breakout) plus
U+2028 / U+2029 (legal in JSON, illegal in JS string literals when the
page is parsed as text/javascript) per OWASP guidance. The regex
sources are constructed via new RegExp(String.fromCharCode(...)) to
sidestep the TypeScript tokenizer treating the literal codepoints as
line terminators inside /regex/ literals.
Drop in NODE_ENV writes via Record<string,string> cast in the server
test — modern @types/node marks NODE_ENV read-only, and the runtime
reader only inspects the string value.
Adds the server (runtime-config.ts) and client
(runtime-config.client.ts) runtime config readers for shell-docs as
the foundation of Option B's per-request URL/key resolution. The
server module reads from process.env at request time (gated by
unstable_noStore so callers are not statically prerendered); a thin
getRuntimeConfigEdge wrapper skips the cache opt-out for middleware
(Edge runtime cannot import next/cache). The client module reads
window.__SHOWCASE_CONFIG__ which the root layout will inject in the
next commit.
Red-green: verified the test files fail without the modules
(module-not-found) and pass once the modules land — 10/10 green.
Part of plan-B (showcase per-env runtime URL injection).
Clarify that next.config.ts's rewrites() is evaluated once at process
start (not build time), matching the Option B runtime-injection model
where every NEXT_PUBLIC_* URL is read at request time from the Railway
env. The validation message now communicates start-time semantics so
operators know to set OPS_BASE_URL on the Railway service rather than
at image build time.
B8.1/B8.2/B8.3. Replaces every process.env.NEXT_PUBLIC_* read in
shell-dashboard consumer code with the runtime-config readers, all
in a single commit so the tree never sits red across a partial
rename.
B8.1 — src/lib/pb.ts: replace the eager module-load read with a
lazy getPb() getter. The previous `const resolvedUrl =
resolvePbUrl()` at module top froze the URL at import time, which
defeated runtime injection. Now the PB client is constructed on
first getPb() call from runtimeConfig.pocketbaseUrl. The
PocketBase instance is returned directly (NOT wrapped in a Proxy)
so `instanceof PocketBase`, detached methods, and this-sensitive
chains all work without surprise. pbIsMisconfigured is converted
from a const boolean to a function — every runtime-config-derived
export in the module is now a function call.
B8.2 — src/lib/ops-api.ts: resolveBaseUrl() now reads
runtimeConfig.opsBaseUrl on the client (gated on typeof window so
SSR / server tests still fall through to the explicit-param or
/api/ops fallback). The runtime-config throw is caught and
treated as 'no override' so a missing __SHOWCASE_CONFIG__ wiring
bug degrades to the safe same-origin rewrite path.
B8.3 — src/components/feature-grid.tsx: resolveShellUrl() body
collapses to `return getRuntimeConfig().shellUrl`. The sentinel
about:blank#shell-url-missing now lives in runtime-config.ts (its
single source of truth) instead of being re-implemented here.
Consumers and test mocks updated in the same commit:
- hooks: useBaseline.ts, useLiveStatus.ts, useLastTransition.ts —
swap `pb` import for `getPb`, capture `const pb = getPb()`
at hook/effect entry, change pbIsMisconfigured reads to
pbIsMisconfigured() calls.
- pb-auth-prompt.tsx — capture getPb() once at the top of the
submit handler.
- test mocks: useLiveStatus.test.tsx, useLastTransition.test.tsx,
__tests__/useBaseline.test.ts, cell-pieces.test.tsx — vi.mock
stubs return { getPb: () => pb, pbIsMisconfigured: () => false }
to match the new function-form API.
- use-probes.integration.test.tsx + ops-api.test.ts — swap the
env-var snapshot/clear pattern for the window.__SHOWCASE_CONFIG__
pattern that mirrors the production code path.
B5. Root server layout (shell-dashboard) now calls getRuntimeConfig()
once per request and injects window.__SHOWCASE_CONFIG__ via an
inline <script id="__showcase_config__"> as the FIRST child of
<head>, BEFORE the theme-init script and well before any client
component reads the global during hydration.
Serialization uses the OWASP-recommended escape for inline JSON in
HTML:
- < → \\u003c so a URL containing </script> in a hostile env value
cannot break out of the inline script tag.
- U+2028 / U+2029 → \\u2028 / \\u2029 because those codepoints are
legal in JSON strings but a syntax error inside JS string literals
when the page is parsed as text/javascript.
The regex sources use new RegExp() with \\u escape strings rather
than regex literals — U+2028 and U+2029 are line terminators that
prematurely close a regex literal in TypeScript / many JS engines.
Workstream B (Option B, runtime injection). Adds the shell-dashboard
runtime-config module pair plus their unit tests:
- src/lib/runtime-config.ts (server): reads URL env vars at REQUEST
time via unstable_noStore() so a single built artifact can serve
different values across staging vs prod by changing the Railway
service env vars. Sentinel fallbacks in production (visible
breakage) + console.error; localhost fallbacks + console.warn in
dev. getRuntimeConfigEdge() variant skips unstable_noStore for
Edge-runtime middleware.
- src/lib/runtime-config.client.ts (client): reads
window.__SHOWCASE_CONFIG__ injected by the root layout. Throws on
SSR (no window) and when the global is missing, so wiring bugs
surface loudly.
- runtime-config.test.ts + runtime-config.client.test.ts: 9 tests
covering env values, trailing-slash strip, dev defaults, sentinel
fallbacks + console.error, live-env-on-each-call (no module-load
freeze), and the Edge wrapper's noStore skip.
Implements plan-B B11. URL and analytics NEXT_PUBLIC_* values now reach
each shell at runtime via Option B (env-driven runtime-config), so the
GHA showcase_build.yml workflow no longer threads them through as Docker
build-args and the shell-dashboard/shell-docs Dockerfiles no longer
declare the matching ARG/ENV pairs.
- showcase_build.yml: shell-dashboard and shell-docs matrix entries lose
build_args_pb_url / build_args_shell_url / build_args_ops_url /
build_args_base_url / build_args_analytics; the 'Prepare build args'
step drops the corresponding env: keys and if-branches plus the five
analytics NEXT_PUBLIC_* secrets. COMMIT_SHA and BRANCH stay — they
identify the artifact.
- showcase/shell-dashboard/Dockerfile: remove ARG/ENV for
NEXT_PUBLIC_SHELL_URL, NEXT_PUBLIC_POCKETBASE_URL, OPS_BASE_URL plus
the explanatory comments. Update the runner-stage comment to point at
runtime-config.ts as the new source of truth.
- showcase/shell-docs/Dockerfile: remove ARG/ENV for
NEXT_PUBLIC_BASE_URL, NEXT_PUBLIC_SHELL_URL, NEXT_PUBLIC_POSTHOG_KEY,
NEXT_PUBLIC_REB2B_KEY, NEXT_PUBLIC_SCARF_PIXEL_ID, NEXT_PUBLIC_REO_KEY,
NEXT_PUBLIC_GOOGLE_ANALYTICS_TRACKING_ID. COMMIT_SHA / BRANCH retained.
shell/Dockerfile and shell-dojo/Dockerfile already only declare commit-sha
and branch ARGs — no changes needed there (per plan-B B11.4).
Cover four malformed-ref shapes the gate must reject:
* `:sha256-<hex>` (missing the @ separator — looks like a digest
pin but is actually a tag)
* `@sha256:<too-short-hex>` (truncated digest hex)
* The 2026-04-21 `...atest` corruption shape from the script
docstring (the original reason this gate exists)
* Non-ghcr.io registries on both envs
These match the canonical PROD_SHAPE / STAGING_SHAPE regexes in
verify-railway-image-refs.ts; the tests are the regression guard
that ensures a future "relax the regex" change cannot ship without
explicitly turning these red first.
Lock in shape behaviour for dashboard, docs, dojo, shell, and
harness with explicit per-service red-green cases:
* prod with :latest -> fail (must be @sha256)
* prod with @sha256 on the correct repo -> pass
* staging with :latest on the correct repo -> pass
* staging with @sha256 -> fail (must float on :latest)
* wrong GHCR repo name on either env -> fail
The validateImage body is unchanged — it was always shape-pure and
shape-correct. Before WS-C these five services were never exercised
through the gate at all (gateValidated:false). These tests are the
regression guard that ensures a future edit doesn't accidentally
re-introduce the Phase-2 carve-out without anyone noticing.
Flip dashboard, docs, dojo, shell, and harness from
gateValidated: false to gateValidated: true and simultaneously add
the corresponding repoNameOverride for both envs:
dashboard -> showcase-shell-dashboard
docs -> showcase-shell-docs
dojo -> showcase-shell-dojo
shell -> showcase-shell
harness -> showcase-harness
These two halves MUST land in the same commit. Flipping
gateValidated without the override would make the gate look up
ghcr.io/copilotkit/<railway-name> (e.g.
ghcr.io/copilotkit/dashboard:latest) which does not exist — the
gate would fail on first run. Adding the override without flipping
gateValidated is dead code: main() short-circuits unvalidated
services before consulting the override. Only the union is correct.
Also remove the Phase-2 deferral comments and refresh the
ServiceEntry.gateValidated JSDoc — there are no Phase-2 holdouts
left. All 27 services are now gate-validated; gateIgnore remains
the sole escape hatch and is unused by every current entry.
The Railway -> SSOT direction at verify-railway-image-refs.ts was
warn-and-continue: an out-of-band Railway service with a malformed
ref would not turn CI red as long as nobody read the warning line.
Replace that branch with a hard failure path that lists the
untracked service under its own failure class, with a clear remedy
in the error message (add to SSOT, or set gateIgnore on an existing
entry).
Refactor main() to call two pure helpers (findUntrackedServices,
summarizeFailures) so the policy is unit-testable without going
through Railway GraphQL. The SSOT -> Railway direction
(findMissingServices) is unchanged: it is already correct and is a
separate coverage class.
Adds red-green tests for the summarizeFailures shape, covering the
three failure classes (shape violations, SSOT->Railway drift,
Railway->SSOT drift) and the success path.
Widen `ServiceEntry` with an optional `gateIgnore: boolean` field
(default false / unset) so the image-ref gate can deliberately exclude
a Railway service from BOTH coverage-direction checks. No behavioural
change in this commit; the field is consumed by the next commit which
flips the Railway->SSOT direction from warn-only to hard-fail.
Stub-export `findUntrackedServices` from verify-railway-image-refs so
the unit test for the new direction can compile. Behavioural wiring
into main() lands in the next commit.
Adds verify-railway-image-refs.test.ts with the first unit tests
covering the new field surface and the helper contract.
Adds Domains, ProbeDriver, ProbeConfig types + domainFor() helper that
throws on unknown service/env. Populates domains.{staging,prod} and
probe.{staging,prod,driver} on every SERVICES entry. Adds
emit-railway-envs-json.ts to serialize the SSOT for the Ruby side and
the workflow consumers.
Refs spec §3 / §3a.
serviceForDispatchName iterates Object.entries(SERVICES) and returns the
first match — a silent dispatchName collision would route a redeploy to
whichever entry happens to iterate first. We now fail loud at module load
via assertDispatchNamesUnique(), and ship synthetic-input tests proving the
invariant fires on a real collision (and stays quiet for entries without a
dispatchName, which is legitimate for out-of-band services).
[BLITZ:L3-wf] E-7a/E-7b wiring.
Adds 'webhooks' as a workflow_dispatch choice in both showcase_build.yml
and showcase_deploy.yml so humans can redeploy/verify the webhooks service
on demand. webhooks' GHCR image (showcase-eval-webhook) is built by a
separate release workflow in the showcase-eval-webhook repo, so:
- paths-filter uses a sentinel that cannot match any in-tree path,
keeping push-driven runs from ever including webhooks.
- The build matrix entry carries skip_build: true; the Build and push
step skips the Depot build for that slot. The per-slot result still
publishes (job.status == success) so the redeploy path proceeds and
redeploy-env.ts picks the existing :latest from GHCR.
The SSOT entry in railway-envs.ts gains dispatchName: 'webhooks' so the
forward/reverse round-trip tests cover it. New tests pin that the SSOT
dispatchName is mirrored in both workflow files' dispatch choice lists
AND ALL_SERVICES JSON.
[BLITZ:L3-wf] E-6a/E-6b/E-6c/E-6d wiring.
redeploy-staging now also gates on aggregate-build-results.outputs.any_success
== 'true'. When every slot fails, the previous behavior silently kicked a
redeploy that just re-pulled the stale :latest and reported healthy. With this
gate, the redeploy is suppressed and a sibling notify-all-builds-failed job
marks the workflow red and posts to #oss-alerts so the all-broken state cannot
hide behind a green run.
The new notify-all-builds-failed job is distinct from the existing notify: job
(which fires on any build-slot failure); both can fire and that overlap is
intentional, per the Slack alert SOP.
[BLITZ:L3-wf] E-5c wiring.
Build matrix slots now publish per-slot build-result-<dispatch_name>
artifacts containing {service, status}. A new aggregate-build-results
job downloads every per-slot artifact, merges them via the shared
mergeBuildResultFiles helper, and uploads the canonical build-results
artifact (results.json) for cross-workflow consumption.
The deploy workflow's resolve-matrix step now downloads build-results
via gh api .../artifacts/<id>/zip and filters its verification matrix
from the structured JSON success set. No more gh api .../jobs calls,
no more capture("^build \\(") regex on job names — those silently
break on job-name renames. The contract (service + status enum) is
enforced in one place: showcase/scripts/lib/build-outputs.ts.
[BLITZ:L3-wf] E-4e/E-4f wiring.