The YAML had key_template "e2e-smoke:${name}" (hyphen) but every other
reference — DIMENSIONS constant, alert rules, kind: e2e_smoke — uses
the underscore form. This mismatch caused the producer to write
e2e-smoke:<slug> while the dashboard filtered for e2e_smoke, leaving
17 orphan hyphen rows and zero matches on the dashboard.
Changing the YAML key_template aligns the emitted probe key with the
dashboard filter. Driver code (deriveSlug) splits on ':' regardless
of prefix, so no driver/test changes are required.
The showcase shell's demo Docs tab was rendering "No documentation
available" for 25 of 31 langgraph-python demos because their demo
folders lacked a README.md. The bundler pulls each demo's Docs content
from README.md at the demo folder root.
This adds READMEs for the 24 manifest-listed demos that were missing
one, following the 3-section template used by the 6 pre-existing
READMEs (What This Demo Shows / How to Interact / Technical Details).
An orphan shared-state-read/README.md whose route doesn't match any
manifest demo is deleted.
Also extends the check-binaries pre-commit hook allowlist to cover
the shell-docs/shell-dojo copies of demo-content.json. The bundler
emits the same content to all three shells, but only the shell path
was previously allowlisted, so any bundle regen that grew past 1MB
got blocked.
As an incidental side effect of running the bundler, demo-content.json
also drops stale entries for langgraph-python::open-gen-ui and
langgraph-python::open-gen-ui-advanced — those folders exist in source
but are not listed in manifest.yaml.
The cli-start entry in each integration's demos[] is a copy-paste CLI
command, not a runnable demo, but the profile page rendered it as a
Live Demo tile whose drawer iframe loaded ${backend_url}undefined.
Split demos into liveDemos (runnable) and commandDemos (command-only)
and render commandDemos in a new "Get Started" section above the
Live Demos grid, mirroring how the dashboard already handles them.
## Summary
- Adds `showcase/aimock/RAILWAY.md` — reconstruction reference for the
Railway service
- Documents image, startCommand flags, env vars, deploy recipe, and the
legacy Dockerfile cleanup path
- Supports Phase 2 of the aimock wrapper-elimination work
## Why
Railway service config (image + startCommand + env) previously lived
only in Railway's state. If Railway loses it, the Notion plan was the
only reconstruction source. This persists an authoritative copy in-repo
so any reviewer can rebuild the service from scratch.
Values (project ID, environment ID, public domain) were queried live
from Railway GraphQL at doc-writing time.
## Test plan
- Docs-only, no runtime impact
- Verify every referenced URL renders on GitHub after merge
- Commitlint + CI green
The railway-services DiscoverySource wrapped its filter schema in an
outer {filter: FilterSchema} ConfigSchema, but the probe-invoker calls
source.enumerate(ctx, cfg.discovery.filter ?? {}) — it passes the filter
contents DIRECTLY, not a wrapped object. At runtime cfg.filter was
always undefined, so namePrefix AND nameExcludes silently defaulted to
undefined and ZERO filtering happened.
Impact: all 7 infra services declared in smoke.yml's nameExcludes
(showcase-ops, showcase-pocketbase, showcase-shell,
showcase-shell-dashboard, showcase-shell-docs, showcase-shell-dojo,
showcase-aimock) produced smoke:/health:/agent: ProbeResults on every
15-minute tick. PocketBase carried ~21 false-red rows (7 services × 3
dimensions) across smoke:*, health:*, agent:*; after this fix the
discovery source filters them out and only the 34 user-facing
package+starter services emit probe rows (102 legitimate rows).
Fix: drop the outer wrapper. ConfigSchema IS the filter block — parse
rawConfig against it directly. Matches the documented enumerate()
contract in probes/types.ts ("config is discovery.filter") and the
shape pnpm-packages has used all along.
Tests updated: the prior discovery tests passed a {filter: {...}}
wrapped argument that only parsed because passthrough() swallowed the
misshape; they now pass the flat filter object the invoker actually
hands the source. Added one regression test that enumerates two real
showcase services plus all 7 infra services and asserts the seven are
dropped before the per-service env fetch (saves seven variables
round-trips per tick as a side benefit).
## Summary
The `railway-services` DiscoverySource in `showcase/ops` passed
`environmentId` as an argument to `Service.serviceInstances` in its
project-level GraphQL query. Railway's current schema does not accept
that argument — every discovery tick 400'd with:
```
railway gql 400: Unknown argument "environmentId" on field "Service.serviceInstances".
```
Result: no probes ticked, PocketBase stopped receiving fresh rows, alert
pipeline starved of data.
Fix: drop the invalid field argument. The returned
`serviceInstances.edges` are already filtered by `environmentId`
client-side a few lines down (`svc.serviceInstances.edges.find((e) =>
e.node.environmentId === environmentId)`), so there is no semantic
change — the argument was dead weight that tripped GraphQL validation.
## Why
- Blocks `railway-services` from enumerating anything, which blocks
every downstream probe that fans out across Railway services
(image-drift, smoke, e2e-smoke, redirect-decom).
- Unblocks the fresh-redeploy of showcase-ops that carries the #4180
merge — that PR's fix is in the running image but never reachable
because no probes tick.
## Verification
Introspection confirmed the current shape of `Service.serviceInstances`:
- `Query { project(id: String!) }` exists and returns `Project`
- `Project.services` paginates to `ServiceEdge { node: Service }`
- `Service.serviceInstances` accepts no field arguments
- `ServiceInstance.environmentId` is present, so client-side filtering
is correct
Direct `curl` against `backboard.railway.com/graphql/v2` with the fixed
query returned 41 services for the CopilotKit showcase project, matching
the expected ~35 (17 starters + 17 packages + shell + aimock + a few
infra services).
## Test plan
New test: `omits environmentId arg from Service.serviceInstances
selection to match current Railway schema`
- Mocks a Railway endpoint that rejects any query whose
`serviceInstances` call includes parenthesised arguments with the
production 400 error body, and accepts the arg-less form.
- Red on pre-fix code: `DiscoverySourceBackendError: railway gql 400:
{"errors":[{"message":"Unknown argument \"environmentId\" on field
\"Service.serviceInstances\".",...}]}`
- Green after fix; full suite 743/743.
- [x] `pnpm --filter @copilotkit/showcase-ops typecheck`
- [x] `pnpm --filter @copilotkit/showcase-ops test` (743 passed, was 742
+ 1 new regression test)
- [x] `pnpm format`
- [x] `oxlint` clean on touched files (one pre-existing unused-import
warning left alone — not in this diff)
The `railway-services` DiscoverySource was passing
`environmentId` as an argument to `Service.serviceInstances`,
but Railway's current GraphQL schema doesn't accept it. Every
discovery tick 400'd with `Unknown argument "environmentId"
on field "Service.serviceInstances"` and no probes ran.
Drop the field argument; the returned instances are filtered
by `environmentId` client-side in the subsequent loop (existing
behaviour), so no semantic change — the bad arg was dead
weight that tripped validation.
Add a regression test that scripts Railway's exact 400 when
the old query shape is sent and asserts the source succeeds
against a schema-accurate mock. Verified against the live
Railway schema via introspection and a direct curl.
The Promise branch of applyFlushResult set flushedAt optimistically
before the promise resolved, and the .then handler never reverted it
when the flush was suppressed. Since onAggregationFlush is async, this
broke the A5 "keep bucket live for re-threshold" contract in production.
The sync branch test did not exercise the Promise path. Added
bucket_stays_live_when_flush_suppressed_async to lock the invariant.
probeOne and probeAgent catch-blocks both assigned
`errorDesc: sanitizeErrorDesc(errorDesc)` where the local `errorDesc`
had already been sanitized on the line above. Double-sanitize is
harmless today because sanitizeErrorDesc is idempotent on the token
set we emit, but it would turn into a silent escape-on-escape
regression the moment the sanitizer grew a non-idempotent rule
(e.g. entity-encoding `&` → `&`). Drop the outer wrap in both
sites.
validateTripleBrace previously scanned only rule.template.text and
rule.on_error.template.text. A rule with a malformed
`{{{firstSignal.unsafe}}}` reference in aggregation.template passed
load validation and rendered the unsafe value at runtime via
Mustache.render inside onAggregationFlush — same Slack mrkdwn-injection
surface R21-a closed for on_error.template.
Pull aggregation.template into the same scan:
- firstSignal.* and lastSignal.* are aliases for signal.* in the
aggregation render context; normalize both prefixes to `signal.`
before applying the existing per-dimension slackSafe allowlist.
- `count` and `services` are engine-injected scalars (number / string)
and are always loader-known-safe — allow their bare form for symmetry
with the existing `{{count}}` usage in fleet rules.
Tests: rule-loader.test.ts adds
"rejects aggregation.template with triple-brace on an unsafe signal.*
field" under the new "validateTripleBrace covers aggregation.template"
describe.
A2 — Run link guard
Every L1–L4 red-tick alert used to render the dashboard Run link
unconditionally, so a tick with no runId (invariant probes, cron ticks
without a propagated workflow run) emitted a broken
`<https://dashboard/.../runs/|Run>` in every alert. Wrap the token in
`{{#event.runId}}…{{/event.runId}}` across all four red-tick YAMLs
(smoke/agent/chat/tools) so the entire "· <…|Run>" chunk disappears when
runId is absent. The Showcase dashboard link stays unguarded (always
valid).
A3 — smoke endpoint link from signal.url
The smoke-red-tick template referenced `{{{signal.links.smoke}}}` and
`{{{signal.links.health}}}` section-guarded on those fields. The driver-
emitted SmokeDriverSignal (probes/drivers/smoke.ts) has no `links`
object — it exposes `url` instead, the URL that was actually probed. So
the guard sections were always empty and the link never rendered. Swap
the template to reference `{{{signal.url}}}` directly and emit a single
"endpoint" link when present. The other three red-tick YAMLs never
referenced `signal.links`; leave them untouched.
SMOKE_SLACK_SAFE_FIELDS updated: drop `links.smoke` / `links.health`,
add `url`. errorDesc stays safe.
Tests:
- render-red-tick.test.ts: parametrize across all four YAMLs, assert
renders_no_run_link_when_runId_missing (output contains no
`/runs/|` or `/runs/ |`) and renders_run_link_when_runId_present
(output contains `/runs/run-abc|Run`). Add
renders_endpoint_link_from_signal_url to cover the smoke endpoint
link swap directly.
- rule-loader.test.ts: update the smoke-red-tick full-YAML tests to
seed `signal.url` and assert `<…|endpoint>` instead of seeding
signal.links.
Alert engine (handleStatusChanged):
- A1: call resolveTriggers on aggregated-rule ingress and skip when the
triggered set is purely red_to_green. Prevents recovery events from
ingesting into the fleet bucket and firing <!channel> on recoveries.
error transitions also skip ingest — aggregated rules don't support
onError today.
- B3: dispatchCronAlert returns early when rule.aggregation is set. A
rule ever declaring both cron triggers and aggregation would otherwise
render the empty top-level template or bypass aggregation entirely.
Alert engine (onAggregationFlush + drain wiring):
- A5: return "suppressed" when the bootstrap / rate-limit gate short-
circuits dispatch, so the store keeps the bucket live. Pre-fix a
bootstrap-suppressed threshold flush was silently lost — the bucket's
flushedAt was marked and post-bootstrap ticks couldn't re-trigger.
- A4: drain() is now async. Register a SIGTERM handler that awaits
drain() so in-flight buckets actually dispatch before process exit;
keep beforeExit + stop() as best-effort fire-and-forget (Node
beforeExit cannot truly await). Remove both listeners in stop() for
test-cycle symmetry.
AggregationBucketStore (aggregation.ts):
- drain() returns Promise<void> and awaits each onFlush invocation via
Promise.allSettled.
- OnFlushCallback returns void | "suppressed" | Promise<…>. When the
callback returns "suppressed", flushedAt stays null so the bucket stays
live for re-ingestion after the engine gate reopens.
- A7: groupBy is optional; absent/empty collapses the bucket key to
rule.id alone. Schema updated so AggregationSchema.groupBy is
z.array(string).optional() (pre-fix .min(1) forced a decorative
[dimension] that always resolved to one bucket anyway).
- B4: intra-key field separator is NUL instead of "&" so a groupBy
value containing "&" + "=" cannot collide with a distinct grouping.
buildBucketKey and buildCompositeDedupeKey both match.
Schema (rules/schema.ts):
- B1: remove unused aggregation.targets field. The engine has always
used rule.targets for aggregation dispatch; the schema field was
silently dropped.
Tests (alert-engine.test.ts):
- A1 engine-level: red_to_green on an aggregated rule does NOT dispatch,
green_to_red DOES dispatch (minMatches: 1, empty groupBy).
- A4: drain_awaits_onflush_promises — async onFlush resolves before
drain's promise resolves.
- A5: bucket_stays_live_when_flush_suppressed — onFlush returning
"suppressed" keeps the bucket live; a follow-up ingestion re-fires.
- A6: fix TDZ in bootstrap_suppression_respected — declare rule before
the onFlush closure that captures it.
- A7: no_groupBy_uses_ruleid_only_as_bucket_scope — three signals
across distinct dimensions all collapse into one rule-id bucket when
groupBy is empty.
- B4: bucket_key_collision_safe_with_special_chars — concrete collision
{b:"y&c=z"} vs {b:"y",c:"z"} is rejected by the NUL separator.
Fleet aggregation rules previously listed red_to_green alongside green_to_red
and sustained_red, which meant any recovery event ingested into the fleet
bucket and could fire a <!channel> aggregation alert on recovery. Remove
red_to_green from all four fleet YAMLs (smoke/agent/chat/tools) — the engine
ingress gate that actually skips recovery-only matches is added in the
subsequent commit. While here, drop the decorative groupBy: [dimension] on
every fleet rule: signal.dimension already partitions traffic to one
dimension per rule, so the groupBy expanded to undefined → "" at runtime
and silently collapsed to a single bucket per rule. The new schema treats
absent/empty groupBy as an explicit "single bucket per rule" contract.
Four aggregated rules using groupBy=[dimension], windowMs=120s
(one 90s probe tick plus slack), minMatches=3. Each carries the
<!channel> escalation that moved off the per-service red-tick rules
so channel gets paged exactly once for a fleet-wide outage instead
of N times. rate_limit.window=10m prevents re-escalation across
consecutive aggregation windows.
Aggregated rules bypass per-match dispatch and ingress-side gates
(rate_limit, dedupe, bootstrap); the composite FLUSH callback still
honors bootstrap suppression and rate_limit against the composite
dedupe key so successive window expiries for the same groupValues
collapse to one dispatch.
Add errorDesc to SMOKE_SLACK_SAFE_FIELDS and introduce an inline
L1-L4 safe-field set (errorDesc only) for the agent/chat/tools
dimensions — they don't have dedicated probe modules, so the allow-
list is defined at the orchestrator's dimension-registration site.
rule-loader.test.ts loadRealRules() mirrors the orchestrator wiring
so the full-config-dir test path still validates successfully.
Migrate the smoke-red-tick "escalated block renders !channel"
assertion onto the new fleet rule — <!channel> now lives on
smoke-red-fleet's aggregation template, not on the per-service
red-tick template.
AggregationBucketStore collects matching signals keyed by rule-defined
groupBy, fires onFlush once the bucket reaches minMatches (threshold),
or at windowMs expiry when below threshold (timer); drains outstanding
buckets on beforeExit. buildCompositeDedupeKey stabilizes rate-limit
across consecutive windows with the same groupValues.
CompiledRule gains optional `aggregation` field; schema.ts adds a
strict AggregationSchema so a typoed aggregation key fails load rather
than silently arming a bucket with undefined.
Strip HTML tags, script/style bodies (including entity-encoded), Slack
mrkdwn controls (<!channel>, <!here>, <@U…>), backticks. Decode entities
first so entity-smuggled script payloads are caught. Cap at 120 chars
with U+2026 ellipsis. Apply at 8 call sites in smoke.ts.
Wrap each {{signal.links.<key>}} reference in a Mustache section guard
so absent links emit nothing. Switch {{signal.errorDesc}} to triple-brace
{{{signal.errorDesc}}}; the value is pre-sanitized in smoke.ts's
sanitizeErrorDesc. Drop <!channel> from per-service rules — fleet rules
(introduced in a following commit) carry it instead.
No Mustache partials are wired in the renderer (render/renderer.ts uses
Mustache.render directly with no `partials:` arg), so the guarded-link
pattern stays inline across four YAMLs rather than being extracted.
Cover all four red-tick alert YAMLs. Missing links should emit nothing
instead of empty <|label> brackets. Sanitized errorDesc should render
literal characters, not HTML-entity-escaped.
- Widen classifyShape package regex to accept hyphen-bearing multi-segment
names (showcase-langgraph-python, showcase-claude-sdk-typescript,
showcase-ms-agent-dotnet, showcase-crewai-crews) without firing the
name-shape-unknown warn on every tick.
- Broaden classifyShape warn trigger: any name that matches neither the
starter nor widened-package pattern now warns, including non-showcase-
prefixed names (my-random-service, copilotkit-cloud) and mixed-case
variants. Returns package as a safe default; warn is the audit trail.
- Drop NormalisedSmokeInput discovery-arm index signature so input.url in
a mode==="discovery" branch is a TS error rather than silently unknown.
- Drop the normaliseMode early-return arm that trusted raw.mode verbatim.
Field presence is the sole discriminator, keeping schema-parsed and
raw-input paths consistent. Bad-shape inputs now throw loud.
- Restore parse-time invariants on smokeInputSchema via a union-level
superRefine that rejects bare {key}, discovery-missing-publicUrl, and
mixed-mode {key,url,name,publicUrl} payloads that the discovery arm's
passthrough would otherwise absorb.
- Add regression tests for the primary-key rewrite in discovery mode
(smoke:showcase-ag2 -> smoke:ag2, smoke:showcase-starter-ag2 ->
smoke:starter-ag2) and verbatim pass-through in static mode.
- Add end-to-end package-shape test coverage for deriveSlug feeding
demosResolver (showcase-langgraph-python -> langgraph-python registry
lookup) and the throwing-resolver swallow behaviour as a regression
guard against the deferred fail-loud change.
Applies Round 2 code review findings to the starter-probe contract:
Critical:
- classifyShape now warns on every showcase-* name that fails both the
strict starter pattern and the single-segment package pattern; typos
(showcase-strater-*) and hyphen-bearing multi-segment names surface
audit warns instead of silently falling to package.
- C2 (add "skipped" state to the aggregate ProbeResult) deferred: the
aggregate state lives on the shared ProbeState type used by the
transition-detector, alert-engine, and status-writer; extending it to
a new literal ripples across the whole state machine and six files
outside scope, and would materially change alert semantics. Flagged
as follow-up.
Major:
- resolveShape consolidated into discovery/railway-services.ts; both
smoke and e2e-smoke drivers import it and thread ctx.logger so the
classifier's audit warn fires on the driver path.
- Silent fallback path (neither name nor shape supplied) now emits a
structured debug log (resolve-shape-fallback) so the assumption is
greppable if it breaks.
- E2eSmokeStarterSignal gains errorDesc?: never, turning the "no failure
reason on a skipped-green row" invariant into a compile error.
- smokeInputSchema refactored to a zod union over static + discovery
branches; runtime normaliseMode() maps legacy inputs onto a
NormalisedSmokeInput with a required mode discriminator, replacing
the non-null assertion on input.url and the ad-hoc refine()s.
- Tests and driver comments stripped of CR-process tags (C1:, R1
bundle, etc.); test-side two-GETs rationale shrunk to a one-line
pointer so the authoritative copy lives in the driver.
- Added classifyShape + resolveShape audit tests: warn fires on typo
names, not on well-formed roots or starters, and debug fires on the
name/shape-absent fallback.
Minor:
- Warn event key namespaced to discovery.railway-services.name-shape-unknown.
- deriveSlug JSDoc moved to sit directly above deriveSlug in e2e-smoke.
- SHOWCASE_SHAPES inlined into showcaseShapeSchema (no external consumers).
- Starter signal carries a structured skipReason: "starter-shape" enum
alongside the human-readable note.
- Removed dead "unreachable" discriminator asserts in the starter
contract test.
Critical
- smoke + e2e-smoke drivers resolve `shape` via the classifier when
`input.name` is present and throw on explicit-vs-classifier
disagreement. Silent-defaulting at the driver boundary inverted the
whole starter-shape fix; the driver now fails loud on wiring drift.
Majors
- E2eSmokeSignal is a discriminated union keyed on `shape`. Starter
variant carries `note` (never rendered as failure) instead of
`errorDesc`, so alert templates stop treating a green short-circuit
as a failure reason.
- Starter shape runs TWO independent GETs against `/api/health` (smoke
and health feed separate alert dimensions that dashboards correlate;
a single-GET reuse would make the rows byte-identical and defeat the
signal).
- smoke.test.ts uses a real `node:http` server to exercise genuine 308
redirects for the agent POST; prior tests accepted a synthetic 404
response and would have let a `redirect: "follow"` regression pass.
- New tests: starter 404 on /api/health, single-source shape enum
assertion on discovery+package and discovery+starter redirect option,
mixed-driver wiring via `.passthrough()` spread, shape-mismatch
throws on both drivers.
- e2e-smoke: starter short-circuit ordering locked in via a throwing
demos resolver; full starter signal contract asserted field-by-field.
Minors
- Single-source the shape enum: `SHOWCASE_SHAPES`, `ShowcaseServiceShape`,
`showcaseShapeSchema` exported from railway-services; both driver
schemas and `deriveUrls` now import the shared types.
- `classifyShape` emits a one-shot warning on umbrella-match-but-unusual
names (`showcase-`, `showcase-starter`) when a logger is supplied.
- Static-mode rejects `url` + `shape` combo via zod refine — shape
applies only to discovered `publicUrl`.
- Drop rot-prone counts (17, 34, 51), "CR A1", "Procedure 0" and
"previously"/"old contract" phrasing. Trim WHAT comments that restate
code; keep WHY that explains contract decisions.
Starters (17 Railway services showcase-starter-*) expose a different URL
surface than shell-based showcase packages, so the current discovery-driven
probe contract produces false-red alerts on every tick:
1. No /smoke route -> L1 primary probe 404s
2. Health lives at /api/health, not /health -> side-emit 404s
3. No /demos/* routing -> L3 navigates to /demos/agentic-chat and 404s;
L4 likewise 404s on /demos/tool-rendering
4. Registry lookup strips only the showcase- prefix, leaving starter-<slug>
which registry.json does not key on -> hasToolRendering silently false
5. L2 /api/copilotkit/ was accepted on a raw 308, masking regressions
where the redirect target is actually unmounted (404)
Fix via shape detection in the Railway discovery source. Classify each
service by name prefix:
- showcase-starter-* -> shape: "starter"
- showcase-* -> shape: "package"
Drivers branch on the shape:
drivers/smoke.ts:
- starter -> primary GET /api/health, reuses result for health
side-emit (same endpoint), POST /api/copilotkit/
- package -> legacy /smoke + /health + /api/copilotkit/
- both -> POST /api/copilotkit/ now follows 308 redirects and
judges the FINAL status, so an unmounted route (308->404)
correctly flips red instead of masquerading as green.
drivers/e2e-smoke.ts:
- starter -> short-circuit with l3: "skipped" + l4: "skipped"
BEFORE launching chromium; skips registry lookup
(starter-<slug> is not keyed). Dashboards see an
explicit skipped state rather than a silent red.
- package -> unchanged.
Classification is done by name prefix so adding a new starter requires no
YAML edit; the next tick picks it up automatically with the right contract.
Verified against three deployed starters (ag2, mastra, langgraph-python):
all return 200 on /api/health, 404 on /smoke / /health / /demos/*.
Before fix: next probe tick would fire 17x3=51 false red alerts across L1
(smoke + health 404s) plus 17x2=34 false reds on L3/L4 (chat + tools
404s on /demos/*). After fix: starters report green L1 on healthy
/api/health, L2 green via followed 308->400 runtime reply, L3/L4 cleanly
skipped with explicit status.
12 new driver + discovery tests cover the shape classifier, primary+health
+agent URL selection per shape, /smoke-never-called regression guard,
308-follow behavior for L2, and e2e-smoke short-circuit before chromium
launch. 716 tests pass (704 existing + 12 new).
30-minute schedule (half the smoke cadence — L3+L4 per tick pulls a
real chromium page so we trade freshness for memory + bandwidth
headroom). 180s outer timeout covers two 60s page timeouts plus
chromium launch overhead. max_concurrency 2 caps peak memory at ~600MB
(each chromium ~300MB RSS) to fit the orchestrator's Railway footprint.
Discovery uses the railway-services source with namePrefix + the new
nameExcludes filter to hit every user-facing showcase without the
7 infra services.
Runtime stage needs two things for e2e-smoke: chromium (headless) and
the shell registry so the driver can gate L4 on demos membership.
Build stage pulls the single registry.json file through rather than
fetching at probe-tick time, keeping the image hermetic. Runtime
stage is already on Debian-slim (needed by Playwright's glibc-linked
chromium); `playwright install --with-deps chromium` runs as root
before the USER node drop so apt-get succeeds. /ms-playwright is
chown'd to node so the runtime user can read the browser tree.
Wire the in-process Playwright driver so probe ticks resolving to
kind: e2e_smoke dispatch to it. Drops the 'deliberately NOT
registered yet' placeholder that flagged the driver as scaffold-only.
The e2e-smoke probe needs to discover every user-facing showcase
service without enumerating infra (aimock, ops, pocketbase, shell,
shell-{dashboard,docs,dojo}). Extend the discovery filter schema
with an optional nameExcludes array so a single namePrefix can pull
in all 34 services and explicitly drop the 7 that don't expose user
demos, rather than maintaining 27 separate namePrefix entries.
Replace the JSON-reporter scaffold with an in-process Playwright
driver. Each invocation launches one headless chromium, runs L3
against /demos/agentic-chat (any non-empty response) and, when the
service registry entry includes tool-rendering, runs L4 against
/demos/tool-rendering (weather vocabulary check). Primary return is
e2e-smoke:<slug>; chat:<slug> and tools:<slug> side-emit via
ctx.writer. AbortSignal is wired end-to-end so timeout teardown
kills the browser cleanly. Launcher + demos resolver are dep-injected
for unit testability.
Input schema accepts { key, backendUrl?, publicUrl?, name?, demos? }
with a zod refine requiring at least one URL. Slug derivation strips
the "showcase-" prefix so Railway service names align with registry
entries (bare slugs). The default demos resolver memoises a read of
/app/data/registry.json at first use, keeping probe-tick overhead at
a single JSON parse across the orchestrator lifetime.
Extends the smoke probe driver to emit a third side-emission per target:
`agent:<slug>` — POST `${backendUrl}/api/copilotkit/` with `{}` body,
green on non-404 response (runtime is mounted), red on 404 or transport
failure. Matches the `checkAgentEndpoint` contract in the e2e helpers
(integration-smoke.spec.ts L2 `@agent` assertion).
Switches `config/probes/smoke.yml` from a static 17-target list to the
`railway-services` discovery source with `namePrefix: "showcase-"` and
an explicit `nameExcludes` for infra services (aimock/ops/pocketbase/
shell*). New showcase services are picked up automatically on the next
tick without YAML edits.
- src/probes/discovery/railway-services.ts: add `nameExcludes: string[]`
filter field, applied AFTER the prefix check so excluded services skip
per-service env fetch entirely.
- src/probes/drivers/smoke.ts: accept both static (`{key,url}`) and
discovery (`{key,name,publicUrl,imageRef,env}`) input shapes. Derive
slug by stripping `showcase-` prefix from `name` (discovery) or
splitting key on `:` (static fallback). Rewrite primary result key to
`smoke:<slug>` in discovery mode so existing alerts keyed on
`smoke:ag2` / `smoke:starter-ag2` stay intact.
- Sequential probes kept for socket-count bound: at max_concurrency=6 *
34 services * 3 endpoints, parallel would push 612 inflight sockets.
- Coverage: smoke driver 98.55% lines, exceeds 95% gate.
Adds three new alert-rule YAMLs modeled on smoke-red-tick.yml so the
per-starter depth signals side-emitted by the smoke + e2e-smoke drivers
(agent:<slug>, chat:<slug>, tools:<slug>) produce Slack alerts on
green_to_red / sustained_red / red_to_green transitions with the same
!channel 1-hour escalation + 15-minute rate limit + 20-minute deploy
grace as the L1 smoke rule.
Also extends the DIMENSIONS enum with the three new closed values so
rule-side typos continue to fail at load with a listed-valid-members
Zod error instead of silently never matching any probe key.
No driver or discovery changes in this commit — pure rule YAML + the
single enum entry each rule's `signal.dimension` resolves against.
The multi-stage Docker refactor (#4147) moved TypeScript showcase services
to `node:22-slim` runtime. That base image does NOT include `curl`.
entrypoint.sh's watchdog runs `curl -sS --max-time 5 http://127.0.0.1:8124/ok`
every 5s to probe the liveness endpoint. Without curl the probe exits
rc=127 "command not found" every cycle. After 180s grace + 3x30s strikes
the watchdog kills the agent process. Railway auto-restarts. Kill-loop
forever.
Verified via `railway ssh --service showcase-langgraph-typescript`:
sh: 1: curl: not found
rc=127
while `ss -tlnp` on the same container confirmed BOTH 0.0.0.0:8123
(langgraph-api) and 0.0.0.0:8124 (liveness.mjs) LISTEN. The agent
servers were healthy the whole time — the probe tool was missing.
Fix: add `apt-get install -y --no-install-recommends curl` to the
runtime stage of:
- showcase/starters/template/dockerfiles/Dockerfile.typescript (source)
- showcase/packages/claude-sdk-typescript/Dockerfile
- showcase/packages/langgraph-typescript/Dockerfile
- showcase/packages/mastra/Dockerfile
3 starter Dockerfiles regenerated from the template via
`pnpm -C showcase/scripts generate-starters`.
Python/Java/.NET services already install curl (they need it for
NodeSource or similar) and are unaffected.
Smoke: local `docker build --platform linux/amd64` for both
starter + package variants succeeded; `docker run ... sh -c "which curl"`
prints /usr/bin/curl in both.
Railway deploys on and after main@5ed233f01 failed health within 600s with:
PermissionError: [Errno 13] Permission denied: '.langgraph_api'
First deploy of langgraph-fastapi since the b9bcf2e6b USER app change.
langgraph_runtime_inmem/store.py does `os.makedirs(".langgraph_api", exist_ok=True)`
at import time, resolving against CWD (/app). WORKDIR creates /app as root;
--chown=app:app on COPY only chowns contents. USER app (uid 1001) can't write
to /app, the agent subprocess crashes, Next.js stays alive so wait -n never
fires, and Railway / the watchdog restart loops until the 600s budget elapses.
Pre-create /app/.langgraph_api (and chown /app itself) so the first-import
makedirs is a no-op.
Verified locally:
- docker build + --load clean
- agent /ok → 200
- next /api/health → 200
- no PermissionError in logs
Non-navigable manifest entries (e.g. langgraph-python/cli-start, a CLI
copy-paste card) have no `route` field. The old code concatenated
`${backendUrl}${undefined}` producing URLs like
`...railway.appundefined/`, which Playwright chased until the 45s
response timeout, wasting CI budget and emitting a noisy
`ERR_NAME_NOT_RESOLVED` FAIL line per run.
Skip those entries with a visible `[SKIP]` log so they drop out of the
capture set entirely. Only effect on target composition: cli-start no
longer appears in the `Failed: N/158` tally.
Verified by running the script locally with `--slug langgraph-python
--demo cli-start` (prints `[SKIP]`, exits clean) and `--slug
langgraph-python --demo agentic-chat` (still captures the MP4).
showcase/scripts/package.json bumped vitest to ^4.1.5 and added
@vitest/coverage-v8 ^4.1.5, but package-lock.json still carried vitest
4.1.4 and was missing the coverage package entirely. Docker builder
stage runs `npm ci`, which refuses to reconcile drift and fails with:
Missing: @vitest/coverage-v8@4.1.5 from lock file
Invalid: lock file's vitest@4.1.4 does not satisfy vitest@4.1.5
Regenerated showcase/scripts/package-lock.json with
`npm install --package-lock-only` to bring it back in sync.
Also removed --silent from the builder-stage `npm ci` calls so the
next same-class lockfile drift surfaces the real error in CI instead
of a bare 'exit code: 1'. Kept --silent on the prod-deps stage (that
one wasn't masking anything here).
Verified: `docker build -f showcase/shell-dashboard/Dockerfile .`
completes end-to-end locally.
Two pre-existing type errors from commit 9b05ed41 surfaced on main's
Docker build:
1. `unscoped-docs-page.tsx` imports `findFrameworksWithCell` from
`@/lib/docs-render`, but the helper was only declared locally in the
two page.tsx routes with a 1-arg signature (and referenced an
undeclared `demos` in one case — dead code). Export a 3-arg shared
version from docs-render.tsx that accepts the integration slug list
and demo map as parameters (keeps the lib free of registry imports),
drop the dead local in `[[...slug]]/page.tsx`, and rewire the live
caller in `[framework]/[[...slug]]/page.tsx` through the shared
export.
2. The Step-2 section cards on the overview call `<SidebarLink>`
without the required `scope` prop. The prop was already ignored
internally (destructured as `_scope`) so relaxing it to optional is
the minimal fix and keeps the call-site intent documented.
`npm run build` in showcase/shell-docs now compiles cleanly.
Multi-stage Dockerfile built on bookworm-slim with Playwright chromium
pre-installed (e2e-smoke runner) + pnpm-workspace.yaml + every
workspace package.json COPYd so the runtime stage resolves workspace:*
deps. README organised operate-first / build-second so an on-call
operator can configure + inspect the service before needing to rebuild
it. Rotation drill doc captures the weekly / monthly schedule for
pin-drift + redirect-decommission audits.