build_snapshot enumerates every project service and queries each one's
serviceInstance. The only guard was `next if inst.nil?` — it handled a
NULL result but not a THROWN `GraphQL: ServiceInstance not found` error
(a half-deleted service that still appears in the project service list
but has no instance in the env). That error bubbled to Railway.run's
top-level `rescue GraphQL::Error` and aborted the ENTIRE promote with an
opaque exit 2 before any preflight/divergence logic ran (run 27144525566
killed the docs promote this way).
Scope the rescue narrowly to ONLY the per-service "ServiceInstance not
found" message — log+skip that one service exactly like the nil case —
so every other GraphQL failure (auth, rate-limit, schema drift) still
propagates fail-loud. Adds red-green coverage: a single thrown not-found
is skipped (healthy services still snapshot), while an unrelated GraphQL
error still raises.
The /strands/deploy-agentcore and /langgraph/deploy-agentcore pages
rendered only their title — the body was empty. <Content> resolved to a
dead stub in the MDX component registry that rendered nothing, despite
the content being authored in the shared agentcore partial.
- mdx-registry.tsx: replace the dead Content stub with a dedicated
component that renders the agentcore partial via PartialLoader and
threads the page's framework into MDX scope.
- mdx-registry-loader.tsx: PartialLoader accepts an optional scope,
forwarded to MDXRemote options.scope so partials can read bare scope
identifiers (next-mdx-remote binds scope as module identifiers, not as
the rendered component's props).
- agentcore/index.mdx: reference {framework} (bare scope var) instead of
props.framework so AgentCoreCommandTabs collapses to the single
relevant framework per page.
The icon geometry was defined via the Chrome-only CSS `d:` property, so
the spinner/checkmark rendered blank in Safari and Firefox. Move the
geometry into each path's `d` attribute and split the single morphing
path into two overlaid paths — a spinning arc that fades out and a
checkmark that draws itself in (stroke-dashoffset) upright.
Spin the arc with `transform-box: fill-box; transform-origin: center` so
it stays centered in WebKit (view-box is mis-resolved there; a SMIL
animateTransform stalls on first paint in Chrome). Both paths use
`pathLength="1"` so dashes read as fractions, and all motion is gated
behind `prefers-reduced-motion`.
## Summary
- **Cold-start retry before fast-fail (#71), gated to plain-fill turns**
— `conversation-runner` now performs a bounded turn-1 `page.reload()`
retry when an error banner appears on a cold start, recovering transient
boot flaps. The retry shares the single turn deadline (no ~2× budget
blowup, FF20) and is gated to plain-fill turns so it never masks a real
failure: a banner that survives the reload still fast-fails.
- **Fleet teardown surface-state honesty** — adds a
`worker-reclaimed-pending` comm-error kind, a control-plane SIGTERM
drain path that distinguishes graceful teardown from a crash (#70), and
pre-dispatch warm-up health pings (#72). Graceful Railway teardown is
now reported as pending, not as a red.
- **Dashboard pending-surface never masks a real red** — `cell-model`
decodes comm-errors severity-first, guards against stale `observedAt`,
and renders a dedicated pending chip. A pending surface can never
override a genuine red, while transient teardown noise resolves to
pending instead of flapping.
**Dashboard-green impact:** kills false flaps on Railway teardown
WITHOUT masking real reds, and bounds the cold-start retry so a
genuinely-broken cold start still surfaces.
## Test plan
- [x] harness `tsc --noEmit` (clean)
- [x] harness vitest — conversation-runner (46), fleet contracts +
control-plane/job-producer + queue-client (123 tests across 4 touched
suites, all passing)
- [x] shell-dashboard `tsc --noEmit` (clean)
- [x] shell-dashboard vitest — cell-model + depth-chip (148 tests, all
passing)
- [x] oxfmt `--check` clean on all changed files
## Known follow-ups
All pre-existing, tracked for a separate PR (not introduced by this
change):
- `priority` / `leaseSeconds` dead knobs in the fleet contracts
- `modelsEqual` does not compare `jobId`
- `DepthChip` switch is not exhaustiveness-checked
- unknown-state should map to a gray cell in `cell-model`
- `createRailwayAdapter` stale JSDoc
- no-reload-retry test-fake edge case
🤖 Generated with [Claude Code](https://claude.com/claude-code)
## Summary
Follow-ups to #5309 that further lower the d5/d6 dashboard red floor:
- **byoc browser-pool race:** guard `newContext()` against a
disconnected shared browser (the d5 contention source).
- **D4 probe:** `networkidle`→`load` — networkidle never settles on
CopilotKit's persistent-SSE demo pages, so D4 timed out locally and
gated every cell red on the local rig; staging unaffected but the fix is
correct and unblocks local==staging visual testing.
- **strands + spring-ai declarative-hashbrown/json-render:** dedicated
backend agents/controllers + tuned prompts + regenerated d6 fixtures +
manifest coherence (json-render is a real byoc-feature-type cell, at
parity with strands).
- **beautiful-chart:** port the pie/bar/scheduler generative-UI
renderers to built-in-agent + claude-sdk-python (backends already emit
the tool-calls).
- **cleanup:** remove the temporary x-diag-probe instrumentation (keeps
the permanent CVDIAG + the #5309 forwarding fixes).
Showcase-only; no package releases.
## Test plan
- [ ] CI rebuilds :latest; redeploy
- [ ] PB re-pivot:
byoc/declarative-hashbrown/json-render/beautiful-chart-chart cells green
on the fixed backends
- [ ] d5 contention sawtooth reduced (byoc guard)
## Deferred (separate follow-ups)
- Flap-band: decouple PB sampling from live run / detector cold-start
retry / warm+pool (the ±70 churn band)
- #67 backlog: harness accounting (d6 hung-teardown, d4 body-fallback
false-green), forwarding-shim hardening, spring-ai controller error-path
lifecycle (class-wide), built-in-agent @copilotkit/runtime pin
## Summary
Greens the staging d5/d6 showcase dashboard, which was red across ~450
cells. Root causes were a small set of genuine code/CI bugs (not a
single regression — the 06-01/06-06 flap edges were a harness
error-banner detector turning on + D5 starting to run the already-broken
D6 pill set):
- **ASSET-404 (multimodal):** CI checked out Git LFS files as pointer
stubs (`lfs:false` + no `git lfs pull`), so images served 130-byte
pointers instead of real assets. Now pulls LFS during build.
- **ROUTE-404 (beautiful-chat / declarative-hashbrown /
declarative-json-render):** 5 backends shipped the demo pages but no
per-pill API route → 404. Routes added (mirrored from same-language
siblings; injectA2UITool set per whether the backend already owns the
tool).
- **Gen-UI header forwarding:** the secondary `generate_a2ui` LLM call
dropped `x-aimock-context` — in Python because it runs in a thread-pool
executor that doesn't propagate the ContextVar (ported ag2's
`install_executor_contextvar_propagation` to
agno/ms-agent-python/pydantic-ai/strands), and in .NET because the AG-UI
SSE pump runs on a non-inheriting ExecutionContext (moved AsyncLocal →
HttpContext.Items, removed a racing finally-wipe). pydantic-ai
additionally sent a system-only prompt (StateDeps.copilotkit was always
None → now reads the real conversation). mastra's gen-a2ui threw
AI_UnsupportedModelVersionError (ai v4 vs @ai-sdk/openai v2 → bumped ai
to v5).
- **Harness measurement:** the fleet worker ignored the YAML
`timeout_ms` (used a hardcoded 10-min default → slow backends
false-aborted); now conveyed. Stopped rendering
`ms-agent-harness-dotnet`, which was flipped deployed:true but excluded
from all probes → perpetual stale red.
Showcase-only; no package releases.
## Test plan
- [ ] CI rebuilds `:latest` images with LFS assets; staging redeploys
- [ ] PB re-pivot: multimodal / beautiful-chat / declarative-* / gen-ui
cells green across backends
- [ ] aimock journal: secondary gen-ui calls carry `x-aimock-context`
(no ctx-absent 503s on real-pill traffic)
- [ ] ms-agent-dotnet d5 no longer abort-cascades
## Deferred to follow-ups (out of this PR's subject)
- strands + spring-ai declarative-hashbrown/json-render need backend
prompt specialization (separate PR)
- pre-existing bucket-(c)/(d): diag-hop labeling (reverts with the
temporary x-diag-probe), harness deploy-churn/total-invariant
accounting, .NET secondary-caller error-mapping, forwarding-shim
hardening, baseline-only stale-render partners
TEMPORARY DIAGNOSTIC (to be reverted after reading). Adds one ungated
header `x-diag-probe: thread=<name>;ctx=<present|EMPTY>;keys=<n>` on
every outbound LLM call (before the context gate) across the 7 backends
whose native tool calls drop x-aimock-context. The aimock journal
(HTTP-readable) will then show, per dropping call, which thread it ran
on and whether the forwarded-headers contextvar was empty — pinning the
context-losing run-path. Additive, try/except-wrapped (never throws), no
change to existing forwarding/injection/gating. py_compile clean.
## Summary
The `generateSandboxedUi` calculator (and the ping-host /
inline-evaluator pills) rendered static HTML with no interactivity — the
`=` and keypad buttons were no-ops — because the affected
per-integration fixtures omitted the `jsFunctions` key that the
`OpenGenerativeUIRenderer` injects into the iframe to wire click
handlers to the host bridge
(`Websandbox.connection.remote.evaluateExpression` / `notifyHost`).
This copies the `jsFunctions` VERBATIM from the canonical
langgraph-typescript fixture (apples-to-apples) into agno, crewai-crews,
langgraph-fastapi, and mastra, and creates the missing langgraph-python
advanced fixture (Calculator + Ping pills; the inline evaluator already
lives in `gen-ui-open.json`). Fixture-only change; the renderer is
correct.
## Verification
Proven in a prior session against the D6 path.
## Summary
The D6 cell `tool-rendering-custom-catchall` for langgraph-python
deterministically timed out (30s no-response). The probe sends two turns
in one thread: turn 1 "weather in Tokyo" -> `get_weather`, turn 2
"What's the price of AAPL?" -> `get_stock_price`. The AAPL first-leg
fixture was gated on `hasToolResult:false`, but aimock implements that
gate as `messages.some(m => m.role === "tool")` — i.e. ANY tool message
anywhere in the thread. Turn 1's `get_weather` leaves a tool result, so
on turn 2 `hasToolResult` is permanently true and the false gate can
never match -> `no_fixture_match` -> 503 -> 30s timeout.
This replaces the broken `hasToolResult:false` gate with
`toolName:"get_stock_price"` (the same pattern the sibling weather/chain
first-leg fixtures already use). A `toolName` gate fires whenever the
tool is registered, regardless of prior tool results; the `toolCallId`
follow-up fixture still wins on iteration 2.
Fixture-only change.
## Verification
Proven in a prior session via the D6 Docker path:
`tool-rendering-custom-catchall` now green (turn 1 `get_weather` + turn
2 `get_stock_price` both render, no 503/timeout);
`tool-rendering-default-catchall` still green (no regression).
## Summary
Diagnostic instrumentation to localize **where `x-aimock-context` is
dropped on the CV/D5 rung** (sustained fleet-wide aimock 503
`no_fixture_match`). Adds a uniform, durable, HTTP-readable per-hop
trace. **Instrumentation-only** — on non-diagnostic traffic every
forwarder is byte-identical to before (verified across all 22 breadcrumb
sites).
- **Shared contract + sink:** single-line `CVDIAG` log format (12-char
redacted header prefix), `x-diag-run-id`/`x-diag-hops` correlation +
breadcrumb headers, and a best-effort PocketBase `diag_events`
collection (anonymously HTTP-readable, so it's queryable mid-incident
without Railway logs).
- **Harness:** mint a run_id per D5/D6 run, inject the correlation
headers alongside `x-aimock-context`, and a timeout-bounded post-run
aimock-journal join that records a `cv-verdict` row tagged
`harness-d5`/`harness-d6` (localizing whether the header reached
aimock). Plus claim/worker_id/snapshot/pool-condition breadcrumbs and
surfaced previously-swallowed catches.
- **Per-framework forwarders** (LangGraph py/ts/fastapi, google-adk,
self-contained Node + 9 Python shims, spring-ai Java, ms-agent .NET ×2):
CVDIAG at each hop + a `route-<fw>`/`backend-<fw>` `x-diag-hops`
breadcrumb, gated on diagnostic-header presence. Surfaces
previously-silent forwarding misses (empty LangGraph `configurable`,
missing httpx event-hooks target, swallowed hook-install errors).
## Why
CV/D5 has been red ~42h fleet-wide; the aimock journal shows ~95% of
503s arrive with **no** `x-aimock-context`. Forwarding is divergent per
framework (LangGraph `configurable` SPOF; `claude-sdk-python` middleware
not wired; google-adk async path uses aiohttp, bypassing the httpx
hook). This makes the drop point visible per hop and durable, so the fix
can be made once — and it preserves apples-to-apples by adding a
*uniform* trace layer rather than more bespoke per-framework code.
## Test plan
- [x] Harness typecheck green; 2105/2105 harness tests pass
- [x] 7-agent CR converged (2 fix rounds); instrumentation-only
invariant verified across every outbound forwarder
- [ ] CI green (integration builds incl. Java/.NET)
- [ ] Post-deploy: read `/__aimock/journal` + PB `diag_events` to
localize the drop hop per framework and settle the open D5-vs-D6
question with data
🤖 Generated with [Claude Code](https://claude.com/claude-code)