Commit Graph

12110 Commits

Author SHA1 Message Date
Jordan Ritter 5cbc8cd6ac fix(showcase): correct turnIndex drift in agno/spring-ai/langgraph-fastapi d6 fixtures (#5140)
## Summary

The aimock matcher gates a fixture's `turnIndex` against the request's
assistant-message count: a fixture only fires when `assistantCount ===
turnIndex`. So `turnIndex` must equal the number of assistant turns
already in the conversation when that fixture should match — turn 1 is
`turnIndex: 0`, turn 2 is `turnIndex: 1`, etc. A multi-turn conversation
fixture that wants to match on *every* turn must **omit** `turnIndex`
entirely and disambiguate purely by `userMessage`, which is exactly what
the canonical clean pattern (langgraph-python and the other 15
frameworks) does.

The `agno`, `spring-ai`, and `langgraph-fastapi` `agentic-chat` fixtures
had baked `turnIndex: 0` onto **all three** goldfish conversation turns.
Turn 2+ carries one or more assistant messages, so `assistantCount !==
0` and the turn-2/turn-3 fixtures could never match. aimock returned
no-match, the request fell through to the proxy, and the d6 cell failed
— surfacing as the 503 → proxy → 502 cascade on staging.

This PR drops `turnIndex: 0` from the multi-turn goldfish turns in all
three files, mirroring the langgraph-python canonical pattern, and adds
clarifying `_comment` keys on each turn.

## Findings

- The 503/502 cascade had been diagnosed on two integrations (agno,
spring-ai). Auditing the full d6 fixture set turned up a **third**
affected integration, **langgraph-fastapi**, carrying the identical
`turnIndex: 0`-on-all-turns defect.
- **Single-turn** fixtures correctly keep `turnIndex: 0` (a single-turn
match where `assistantCount === 0` is correct) — those are unchanged.
- The other **15 frameworks** were already clean (no `turnIndex` on
multi-turn conversation turns).
- Local before/after evidence: turn-2 went 503 → 200 for all three
frameworks after the fix; turn-1 was unregressed.

## Test plan

- [x] All three JSON files parse (valid JSON).
- [x] `validate-fixture-tool-surface.ts` — green (no drift).
- [x] `validate-parity.ts` — exit 0.
- [x] `validate-pins.ts` ratchet — FAIL count and hash unchanged vs
`fail-baseline.json` (63, matching hash); none of the FAIL lines touch
these fixtures.
- [x] `vitest run` (build-pipeline tests) — 1670 passed.
- [ ] The real signal is a clean staging d6 run on these cells (agno /
spring-ai / langgraph-fastapi agentic-chat), verified post-merge.
2026-05-31 21:38:25 -07:00
Jordan Ritter df430ed3c0 fix(showcase): bake LANGGRAPH_HTTP configurable_headers into langgraph.json (#5139)
## Summary

Moves the D6 header-conveyance config from an env-only `LANGGRAPH_HTTP`
var to on-disk `langgraph.json`, so it rides image promotion instead of
drifting per-environment.

`langgraph.json` for `langgraph-python` and `langgraph-fastapi` now
declares:

```json
"http": { "configurable_headers": { "include": ["x-*"] } }
```

## Why

- **Required + uniform conveyance config belongs on-disk.** The env-only
`LANGGRAPH_HTTP` was missing on prod `langgraph-python` /
`langgraph-fastapi`, so header conveyance would 404 there. Baking it
into `langgraph.json` means the config travels with the image through
promotion and cannot drift out of sync between environments.
- **The env-only approach was latently unreliable even where set.**
`langgraph dev` actively pops the `LANGGRAPH_HTTP` env var when
`langgraph.json` lacks an `http` key, so relying on the env var alone
was fragile by design.

## Validation

- Schema key (`http.configurable_headers.include`) confirmed against
pinned `langgraph-cli` 0.4.21 / `langgraph-api` 0.7.101.
- `langgraph.json` is in each image's Docker build context (COPYed into
both images).
- Conveyance verified locally with the `LANGGRAPH_HTTP` env var
**unset**: both services forwarded `x-aimock-context`; negative control
without the key conveyed nothing.
- `langgraph-typescript` is unaffected — it does pure-code conveyance
and needs no `langgraph.json` change.
- Pre-commit hooks (full nx publint/attw suite) ran and passed at commit
time.

## Test plan

- [x] Local conveyance proof (env var unset, plus negative control)
- [ ] Real signal: a clean staging d6 run post-merge
2026-05-31 21:38:22 -07:00
Jordan Ritter dfeb06121b fix(showcase/d6): drop turnIndex:0 from multi-turn goldfish agentic-chat fixtures
The aimock matcher gates turnIndex against the request's assistant-message
count (assistantCount !== turnIndex → skip). The agno, spring-ai, and
langgraph-fastapi agentic-chat fixtures baked turnIndex:0 on all three
goldfish conversation turns. Turn 2+ carries >=1 assistant message, so
turnIndex:0 could never match those turns — aimock returned no-match and the
request fell through to proxy (503), failing the d6 cell.

Mirror the canonical clean pattern used by the other 15 frameworks
(langgraph-python et al.): omit turnIndex on the multi-turn conversation
turns and disambiguate purely by userMessage. Single-turn fixtures in the
same files keep turnIndex:0 (unchanged). Verified locally against the aimock
matcher: turn-2 goes 404→200 for all three frameworks, turn-1 unregressed.
2026-05-31 20:45:13 -07:00
Jordan Ritter 4dc8d465d4 fix(showcase): bake LANGGRAPH_HTTP configurable_headers into langgraph.json
Move the required, uniform D6 header-conveyance config on-disk so it rides
image promotion instead of depending on a per-environment env var. Prod was
missing the LANGGRAPH_HTTP configurable_headers env var, and baking the
`http.configurable_headers.include: ["x-*"]` setting directly into the
langgraph-python and langgraph-fastapi langgraph.json files removes the
promote-time drift gap (the config now travels with the image rather than
being re-supplied at each promotion).

langgraph-typescript needs no change — its header conveyance is pure-code.
2026-05-31 20:33:45 -07:00
Jordan Ritter 5f20887d77 test(showcase): delete orphan e2e specs with no backing demo
reasoning-default-render.spec.ts and shared-state-write.spec.ts navigate to
/demos/reasoning-default-render and /demos/shared-state-write respectively,
but neither demo directory exists in any integration (including the
langgraph-python gold reference) and neither is declared in any manifest.
They are stale, non-canonical leftovers from the #5127 page-mirror.

Removed from all baselines where present: reasoning-default-render from 8
(langgraph-fastapi never had it) and shared-state-write from all 9. The
parity validator stays green (0 fail) and langgraph-fastapi's prior
spec-under-coverage warning clears once its phantom spec is gone and the 5
canonical specs are added.
2026-05-31 20:28:02 -07:00
Jordan Ritter 28fee3dadf test(showcase): add 5 missing canonical e2e specs to baseline integrations
The 9 baseline integrations (ag2, agno, crewai-crews, langgraph-fastapi,
langroid, llamaindex, mastra, spring-ai, strands) were page-mirrored from
langgraph-python but the mirror omitted 5 LGP-canonical Playwright specs
whose demos are present on disk:

  - declarative-hashbrown
  - declarative-json-render
  - reasoning-custom
  - reasoning-default
  - threadid-frontend-tool-roundtrip

Copied each spec verbatim (byte-identical) from langgraph-python, which the
baselines mirror. All 5 backing demo directories exist in every baseline.
The specs are framework-agnostic (navigate by route + testid), so no
per-integration edits are needed. Restores apples-to-apples spec parity.
2026-05-31 20:25:57 -07:00
Jordan Ritter 078f97c3f4 fix(showcase/harness): serialize browser-pool launches (PID ceiling) (#5137)
## Summary

The staging D6 dashboard goes 0/18 because the harness's single shared
`BrowserPool` is contended across overlapping probe suites, and raising
the pool size to cover peak demand re-trips the container's **~1000-PID
ceiling** on the simultaneous chromium **launch burst**
(`pthread_create: Resource temporarily unavailable` / "Zygote could not
fork"). Pool sizing alone can't win: 10 starves under cross-probe
overlap (acquire timeouts), 16 exceeds the launch-burst thread ceiling.
The binding constraint is the *burst*, not the steady-state size.

This PR adds a **launch-serialization gate** to `BrowserPool`: every
chromium launch — init fill, recycle relaunch, reinit backstop, lazy
`relaunchPending` recovery — is funneled through one **concurrency-1**
gate with an env-tunable inter-launch stagger
(`BROWSER_LAUNCH_STAGGER_MS`, default **150ms**). This spaces chromium
spawns so the transient PID spike never exceeds the ceiling, **without
reducing the eventual pool size**. The caller's browser resolves the
instant its own launch settles (the stagger only gates the *next*
launch); a failed launch is swallowed on the chain so it can't poison
the queue, while still surfacing to its own caller.

With the burst removed, the pool can safely be sized to cover peak
cross-probe demand (~14) within the sustainable steady-state (~16 ≈ 800
PIDs). The shared pool *is* the global cross-probe cap, so no separate
budget mechanism is needed. **NOT** a "reduce concurrency to paper over
a race" change.

Full root-cause analysis (incl. the disproven pool=16 attempt and the
1000-PID discovery): the D6 BrowserPool Regression analysis doc, Parts
6-7.

## Commits

1. `serialize browser-pool chromium launches to prevent PID-ceiling
EAGAIN` — the gate (concurrency-1 + stagger), wrapping the single launch
seam so all paths inherit it.
2. `honor default stagger for negative/NaN launchStaggerMs arg` — CR
fix: a negative/NaN explicit arg fell back to **0** (silently disabling
the stagger), contradicting its own comment; now falls back to the
default like the env var.
3. `close partially-launched browsers when init() fill fails` — CR fix:
a mid-fill launch throw (the exact EAGAIN scenario) leaked
already-launched chromium; `init()` now closes them and resets state
before re-throwing (clears state before closing so the disconnect
handler can't re-enter recycle).

## Review

7-agent CR + a 7-agent confirmation round, converged to zero in-subject
findings; Procedure 3 bucket-(c) promotion audit returned 0 promotions.
Verified the gate's serialization stays well inside the 30s acquire
timeout: full-pool relaunch at poolSize 16 × 150ms ≈ ~10s, and a waiter
is served on the first successful relaunch (sub-second).

## Test plan

- [x] Red-green unit tests: launch concurrency === 1 + inter-launch
spacing honored; negative/NaN stagger falls back to default; `init()`
partial-fill closes already-launched browsers.
- [x] Full harness suite green (1675 tests; browser-pool 25).
- [x] typecheck / lint / format / build clean.
- [ ] **Staging validation (post-merge):** deploy to staging, set
`BROWSER_POOL_SIZE=14`, watch a real `:40` d6 cron run — confirm
acquire-timeouts gone AND no `pthread_create`/EAGAIN/Zygote (the
PID-ceiling fix can only be confirmed on the real container). The PID
ceiling is a Railway platform limit (~1000, unraisable), so keep
`poolSize × stagger` well under the 30s acquire timeout (≤16 × 150ms ≈
10s, safe).

## Follow-up (separate PR — out of this PR's subject)

CR surfaced a coherent cluster of **pre-existing** `BrowserPool` defects
unrelated to launch serialization (untouched by this diff), worth a
dedicated **"slot-publication / waiter-handoff hardening"** PR:
- `release()` double-publishes a slot to `available` on a double-release
(missing the `!available.includes(slot)` guard that `handOff` has) →
potential same-browser double hand-out.
- `handOff`/`release` deliver a browser to a waiter without the
`isConnected()` liveness check the available-scan path enforces → a
silently-disconnected browser can reach a probe.
- `track()` registers the raw promise in `inFlightRecycles` but removes
a derived wrapper, so `shutdown()`'s drain awaits a promise that settles
before its cleanup (currently mitigated by per-path `isShutdown`
re-checks).
- Plus observability/comment/test nits (inUse gauge divergence, stale
`slotIndex` logging, a mis-named recycle-warn test, gate
timing-tolerance flake risk).
2026-05-31 19:53:35 -07:00
Jordan Ritter a5612c21f0 fix(showcase/harness): close partially-launched browsers when init() fill fails
The BrowserPool init() fill loop launches browsers one at a time through the
serializing launch gate. If launch iteration N throws -- the exact PID-ceiling
failure (pthread_create EAGAIN / "Zygote could not fork") this file exists to
survive -- init() rejected and propagated, but the browsers already launched on
iterations 0..N-1 stayed live in this.slots/available/browserToSlot and were
never closed. The sole production caller (orchestrator.ts boot) catches the
rejection, marks the pool degraded, and does NOT call shutdown(), so those
chromium processes (~50 PIDs each) leaked permanently -- accelerating the very
PID exhaustion the launch gate is meant to prevent. The gate serializes the
fill, which makes a partial-then-throw fill MORE reachable.

Wrap the fill loop so a mid-fill launch failure resets the pool's internal
state and closes every browser already launched before re-throwing. State is
cleared BEFORE closing so the synchronous `disconnected` fire from a close
cannot re-enter the recycle path via the slot's disconnect handler (the
!this.slots.includes(slot) guard short-circuits it) -- the same ordering
intent shutdown() relies on, without adding an isShutdown re-check. init()'s
reject-on-failure contract is preserved.

Call sites: init() signature is unchanged. The only production caller
(showcase/harness/src/orchestrator.ts:292) and all test callers continue to
rely on reject-on-failure, which still holds.

Adds a multi-slot partial-fill test (pool size 4, launch 3 throws): asserts
init() rejects, the 2 browsers launched before the throw were closed, and
stats() reports an empty pool. The prior init-failure tests only covered
1-slot / first-launch-fails (nothing launched yet), which is why the leak was
missed.
2026-05-31 18:39:35 -07:00
Jordan Ritter 64f73d502c fix(showcase/harness): honor default stagger for negative/NaN launchStaggerMs arg
The browser-pool launch-stagger resolution honored the "negative or
non-numeric value falls back to the default rather than disabling the
stagger silently" contract only for the BROWSER_LAUNCH_STAGGER_MS env
var, NOT for an explicit launchStaggerMs constructor arg. A negative
explicit arg (e.g. -50) was taken by the `??` (non-null), skipped the
env/default branch, then got clamped to 0 by `resolvedStagger >= 0 ? ... : 0`
— silently DISABLING the stagger and reintroducing the launch-burst PID
spike (pthread_create EAGAIN / "Zygote could not fork") the gate exists
to prevent. A NaN explicit arg behaved the same way.

Validate the explicit arg the same way as the env var BEFORE it wins: a
valid explicit arg (>= 0, not NaN) wins; else a valid env value; else the
default (150ms). A negative/NaN explicit arg now falls back to the default,
not 0. An explicit 0 is still respected (tests rely on it to stay fast)
because `validExplicit ?? validEnv` keeps a literal 0.

Call sites: this only changes how this.launchStaggerMs is computed inside
the constructor. The sole consumer is launchBrowser()'s
`this.launchStaggerMs > 0 ? delay(...) : undefined` gate — its semantics
are unchanged (the field is still always a valid non-negative number). The
only production constructor (orchestrator.ts: `new BrowserPool(poolSize,
undefined, logger)`) passes no stagger arg, so its behavior (undefined →
env → default) is identical before and after.

Adds two red-green tests: a negative explicit arg and a NaN explicit arg
must each space launches by ~the default stagger (gap >= 140ms), proving
the stagger is not silently disabled.
2026-05-31 18:39:23 -07:00
Jordan Ritter 1ac16a4ed1 fix(showcase/harness): serialize browser-pool chromium launches to prevent PID-ceiling EAGAIN
The harness runs every e2e probe through one process-wide BrowserPool whose
chromium launches happen in BURSTS (initial fill, recycle relaunches, lazy
relaunchPending recovery, reinit backstop). On the Railway staging container
(~1000-PID ceiling, ~50 PIDs per headless chromium) a burst transiently
spikes PID demand past the ceiling and trips `pthread_create: Resource
temporarily unavailable` / "Zygote could not fork", so browsers fail to launch
and a d6 run goes 0/18.

Funnel EVERY launch through a single concurrency-1 serialization gate that
also waits BROWSER_LAUNCH_STAGGER_MS (default 150ms, env-tunable) after each
launch settles before the next may start. The caller still receives its
browser the instant the process is up; only the NEXT launch is gated. This
spaces process spawns so the transient PID spike never exceeds the ceiling
WITHOUT reducing the eventual pool size. All launch paths (init fill,
acquire-time/recycle relaunch, reinit, relaunchPending recovery) now route
through the gated `launchBrowser` wrapper over the raw launcher.

Adapts the two deterministic-race guard tests that previously required two
recovery launches to be in flight simultaneously — incompatible with the
serialization invariant — to assert the same re-entry guards under serialized
launches (verified by mutation: disabling the in-loop guard still fails the
test). Existing tests pin stagger to 0 to stay fast.
2026-05-31 18:26:08 -07:00
Jordan Ritter a195cc28d4 Revert "fix(showcase/harness): isolate d6 feature-timeout cascade" (#5134) (#5135)
Reverts #5134.

Post-deploy verification on the 23:40Z staging d6 run showed #5134 did
NOT isolate the cascade AND introduced a regression: 12/18 services fail
with `BrowserPool acquire timeout` before running any feature (pool
starvation from the recycle-between-features path +
FEATURE_CONCURRENCY_D6 4→2).

Reverting to restore the #5133 state where services at least execute
features. The BrowserPool cascade needs deeper investigation; tracking
separately.
2026-05-31 17:04:17 -07:00
Jordan Ritter 8839bf354a Revert "fix(showcase/harness): isolate d6 feature-timeout cascade (#5134)"
This reverts commit 470f6f0687, reversing
changes made to 6cd630919f.
2026-05-31 16:57:35 -07:00
Jordan Ritter 470f6f0687 fix(showcase/harness): isolate d6 feature-timeout cascade (#5134)
## Root cause

In `d6-all-pills-e2e`, each service acquires ONE pooled browser and runs
`FEATURE_CONCURRENCY_D6` (=4) feature workers concurrently via
`browser.newContext()`. When a feature's prompt has no matching aimock
fixture, the agent hangs/loops until the 300s `featureTimeoutMs`. Four
such concurrent hung 5-minute contexts (each holding a loaded page +
accumulating SSE/DOM state) create enough memory pressure to OOM-kill
the shared Chromium subprocess. After Chromium dies, every subsequent
`browser.newContext()` / `newPage()` on that handle fails with `"Target
page, context or browser has been closed"` — wiping the rest of that
service's ~40 features and destroying per-feature D6 signal.

The per-feature timeout path (`d6-all-pills.ts`) already aborts the hung
feature but does NOT recycle the shared browser, so the cascade
propagates.

## Fix

1. **Recycle the pooled browser between features when disconnected /
after timeout.** The per-service feature loop now checks
`browser.isConnected()` before each feature's `newContext()` call and
single-flight recycles (close + relaunch) when dead. It also recycles
proactively after a `feature-timeout` result so the next feature on this
worker does not inherit the pressure. Capped at
`MAX_BROWSER_RECYCLES_PER_SERVICE = 5` to prevent infinite-loop on a
poisoned pool.

2. **Lower `FEATURE_CONCURRENCY_D6` from 4 to 2** to roughly halve the
simultaneous in-flight contexts per service. Worst-case wall-clock per
service rises modestly (~10-14 min vs ~7-10 min) but per-feature signal
quality — the actual ask of D6 — matters more here. Now env-overridable
via `FEATURE_CONCURRENCY_D6`.

## Out of scope

Does NOT change `max_concurrency`, `BROWSER_POOL_SIZE`, or
`DEFAULT_FEATURE_TIMEOUT_MS`. The pool's own dead-slot recovery (PR
#5133) handles the per-launcher zombie-acquire case; this PR adds the
orthogonal between-features-in-one-service recovery.

## Test plan

- [x] `pnpm typecheck` clean
- [x] `pnpm vitest run src/probes/drivers/d6-all-pills
src/probes/helpers/browser-pool` — 43 tests pass, including two new
focused tests for the between-features recycle path (one verifies the
swap to a fresh browser, one verifies the recycle cap)
- [x] Full harness `pnpm vitest run` — 1670 tests pass
- [x] `oxfmt --write` + `oxlint` on touched files (no new warnings)
- [ ] Production verification: deploy to harness service, observe a D6
run where a fixture-miss happens and confirm subsequent features in the
same service still run (`probe.e2e-full.between-features-recycle` log +
fresh per-feature results)
2026-05-31 16:16:30 -07:00
Jordan Ritter 12dc633302 fix(showcase/harness): recycle browser between d6 features after timeout + lower FEATURE_CONCURRENCY_D6 to isolate fixture-miss cascade 2026-05-31 16:10:13 -07:00
Jordan Ritter 6cd630919f fix(showcase/harness): prevent d6 BrowserPool PID-exhaustion crash (#5133)
## Summary

The staging \`d6-all-pills-e2e\` probe was aborting at T+2s with every
per-feature \`browser.newContext()\` failing with \"Target page, context
or browser has been closed.\"

### Root cause

The probe-invoker's worker fan-out launches \`max_concurrency\` workers
in the SAME JS tick (a 4ms thundering herd). On d6 (\`max_concurrency:
8\`, ~8 services) each worker's first call drives \`pool.acquire()\` → a
Chromium launch (~50 threads/procs each) on top of the already-warm
10-browser pool. The container hits its PID/thread ceiling, the kernel
returns EAGAIN on fork/pthread_create, fresh Chromium processes die, and
per-feature \`browser.newContext()\` then fails on every single feature
— wiping out every service's full ~40-feature matrix. Not OOM, not a
code race in the pool's locking.

### Fix

1. **Stagger service startup** in the invoker's worker fan-out
(\`probe-invoker.ts\`): worker \`i\` sleeps \`i *
SERVICE_STARTUP_STAGGER_MS\` before its FIRST pull, so per-Chromium
thread-spawn bursts (~150ms each in practice) overlap rather than
collide. Subsequent iterations of each worker run at full speed —
wall-clock throughput is preserved. Configurable via
\`SERVICE_STARTUP_STAGGER_MS\` env (default 300ms, set 0 to disable).

2. **Defensive re-acquire** in \`createPooledE2eFullLauncher\`
(\`d6-all-pills.ts\`): after \`pool.acquire()\`, check
\`browser.isConnected()\`; if false, \`pool.release(browser)\` +
re-acquire once. Covers the narrow window where a browser dies AFTER
acquire returns but BEFORE the caller hands it to \`newContext()\` — so
a dead browser doesn't doom an entire service's ~40 features.

\`BROWSER_POOL_SIZE\`, \`max_concurrency\`, and
\`FEATURE_CONCURRENCY_D6\` defaults are unchanged — concurrency is
preserved; the stagger only spreads the initial-acquire burst.

## Test plan

- [x] \`pnpm typecheck\` passes for \`showcase/harness\`
- [x] Targeted vitest run (\`d6-all-pills.test.ts\`,
\`browser-pool.test.ts\`, \`probe-invoker.test.ts\`) passes — added
focused coverage for stagger ordering, stagger=0 opt-out, and
re-acquire-on-disconnected
- [x] Full harness vitest (1668 tests) passes
- [x] \`oxfmt --write\` + \`oxlint\` on touched files; no new warnings
introduced by this change
- [ ] Watch CI to green
- [ ] Verify in staging after merge: d6 runs complete past T+2s and emit
non-empty per-service aggregates
2026-05-31 15:22:21 -07:00
Jordan Ritter 4dab96569c fix(showcase/harness): stagger d6 service startup + re-acquire disconnected browsers to avoid PID-exhaustion crashes 2026-05-31 15:16:33 -07:00
Jordan Ritter 68fb0bdfe9 chore(showcase): remove unused QA-to-Notion sync (#5131)
## Summary

Removes the unused QA-to-Notion sync feature from the showcase:

- Deleted `.github/workflows/showcase_qa-sync.yml` (the "Showcase: Sync
QA to Notion" workflow).
- Deleted `showcase/scripts/sync-qa-to-notion.ts` (the sync script).
- Removed the `sync-qa` and `sync-qa:dry` npm script entries from
`showcase/scripts/package.json`.

A reference sweep across the repo (excluding `node_modules`) for
`sync-qa`, `sync_qa`, `syncQa`, `sync-qa-to-notion`, `showcase_qa-sync`,
and `qa-sync` found no other references after these removals.

## Rationale

Unused per repo owner. The `qa/*.md` integration content under
`showcase/integrations/*/qa/**` is intentionally retained
(validate-parity counts it); only the sync machinery is removed.

## Orphaned secret

`SHOWCASE_NOTION_API_KEY` was referenced only by the removed
`showcase_qa-sync.yml` workflow. It is now orphaned and can be deleted
from the repo's GitHub Actions secrets if it has no other use.

## Test plan

- [ ] Confirm CI is green on the branch.
- [ ] Confirm no other workflow references the removed job.
- [ ] (Optional) Delete the orphaned `SHOWCASE_NOTION_API_KEY` GitHub
Actions secret.
2026-05-31 12:49:45 -07:00
Jordan Ritter cfe1b1ae95 ci: fix zizmor ref-version-mismatch version comments on 3 pinned actions 2026-05-31 12:36:13 -07:00
Jordan Ritter 803f3d8d01 chore(showcase): remove unused QA-to-Notion sync workflow and script 2026-05-31 11:43:49 -07:00
Jordan Ritter 734148b9fd showcase(D6): reclassify manifest not_supported_features as skipped-incapable (stop counting incapable demos as red) (#5130)
## Summary

Stops the D6 \`e2e-full\` harness driver from counting
framework-incapable demos as RED. When an integration's
\`manifest.yaml\` lists a feature under \`not_supported_features\`
(NSF), that feature is now reclassified at probe time as
\`skipped-incapable\` (a green side-row), not red.

### Reclassification mechanism

- \`showcase/harness/src/probes/drivers/d6-all-pills.ts\` partitions
\`requestedFeatures\` into capable + incapable sets BEFORE script
resolution / runnable filtering.
- Incapable features emit a side row with \`state: green\` and
\`errorClass: \"skipped-incapable\"\`, and surface in the aggregate
signal's \`skipped[]\` array plus a new \`incapable[]\` field.
- NSF flows in through two paths:
- CLI: \`cli/targets.ts\` reads \`not_supported_features\` from the
manifest and passes it through \`buildFullInputs\` as
\`notSupportedFeatures\` on \`FullInput\`.
- Discovery: \`probes/discovery/railway-services.ts\` extracts NSF from
\`registry.json\` per integration via a new \`RegistryIntegrationInfo\`
shape and emits it on each \`RailwayServiceInfo\` record.

### Red-green test

New \`describe(\"NSF (not_supported_features) reclassification\")\`
block in \`d6-all-pills.test.ts\` (3 tests) asserts NSF features without
a registered script emit \`state=green\`. RED phase was confirmed by
stashing the driver changes (test failed with \`expected 'red' to be
'green'\`); GREEN phase: all 21 driver tests pass, full suite 1664/1664
green, typecheck clean.

### Manifests edited

Promoted PARITY_NOTES-documented incapabilities into NSF, moving each
entry out of \`features[]\`:

- **mastra**: \`gen-ui-interrupt\`, \`interrupt-headless\`,
\`agentic-chat-reasoning\`, \`reasoning-default-render\`,
\`tool-rendering-reasoning-chain\`
- **langroid**: \`mcp-apps\`, \`tool-rendering-reasoning-chain\`
- **ag2**: \`gen-ui-interrupt\`, \`interrupt-headless\`
- **crewai-crews**: \`gen-ui-interrupt\`, \`interrupt-headless\`,
\`mcp-apps\`
- **llamaindex**: \`gen-ui-interrupt\`, \`interrupt-headless\`,
\`hitl-in-chat-booking\`
- **spring-ai**: \`byoc-json-render\`

\`npm run validate-manifests\` passes (no features/NSF overlap).

## DO NOT MERGE

Draft per request — for review only. \`strands\` and \`agno\` already
declare their interrupt skips and were verified unchanged.
\`validate-pins\` and \`examples\` are untouched.

## Test plan

- [ ] Harness \`npm run typecheck\` clean
- [ ] Harness \`npm test\` 1664/1664 green
- [ ] \`scripts/generate-registry.ts --validate-only\` succeeds across
all 19 integrations
- [ ] Confirm CI publishes \`incapable[]\` array on D6 aggregate signals
2026-05-31 11:28:07 -07:00
Jordan Ritter dd0e94b7f3 showcase(integrations): promote PARITY_NOTES incapabilities into manifest not_supported_features
Move framework-incapable features from features[] into not_supported_features[]
so the D6 harness reclassifies them as skipped-incapable instead of red:

- mastra: gen-ui-interrupt, interrupt-headless, agentic-chat-reasoning,
  reasoning-default-render, tool-rendering-reasoning-chain
- langroid: mcp-apps, tool-rendering-reasoning-chain
- ag2: gen-ui-interrupt, interrupt-headless
- crewai-crews: gen-ui-interrupt, interrupt-headless, mcp-apps
- llamaindex: gen-ui-interrupt, interrupt-headless, hitl-in-chat-booking
- spring-ai: byoc-json-render

All entries already documented as incapable in each integration's PARITY_NOTES.md.
Validated via 'npm run validate-manifests' (no features/NSF overlap).
2026-05-31 11:21:58 -07:00
Jordan Ritter 66e003f986 showcase(D6): wire not_supported_features through CLI + Railway discovery paths
- cli/targets.ts: read not_supported_features from manifest YAML, pass through
  buildFullInputs as notSupportedFeatures on FullInput
- probes/discovery/railway-services.ts: extract not_supported_features from
  registry.json per integration via new RegistryIntegrationInfo type and emit
  on RailwayServiceInfo records so the D6 driver receives it at probe time
2026-05-31 11:21:58 -07:00
Jordan Ritter 2fcb7fea3e showcase(D6): reclassify NSF features as skipped-incapable in d6-all-pills driver
When an integration's manifest lists a feature in not_supported_features (NSF),
the D6 e2e-full driver now partitions requestedFeatures into capable + incapable
sets BEFORE script resolution. Incapable features get a green side-row with
errorClass='skipped-incapable' and surface in the aggregate skipped[] list plus
a new incapable[] field. They no longer count as red.

Includes red-green test that asserts NSF features without a registered script
emit state=green (was red prior to this change).
2026-05-31 11:21:58 -07:00
Jordan Ritter a14959b876 fix(showcase): align tool-rendering/frontend-tools testids to gold for claude-sdk-python + built-in-agent (#5129)
## Summary

Aligns testids in two non-mirrored showcase integrations with the
langgraph-python gold reference so D6 probes pass.

- `claude-sdk-python`:
- `tool-rendering-custom-catchall/custom-catchall-renderer.tsx`:
`custom-catchall-{card,tool-name,args,result,status}` ->
`custom-wildcard-{card,tool-name,args,result,status}`
- `frontend-tools/page.tsx`: `background-container` ->
`frontend-tools-background`
- `built-in-agent`:
- `tool-rendering-custom-catchall/custom-catchall-renderer.tsx`:
`custom-catchall-*` -> `custom-wildcard-*`
- `frontend-tools` already on the new testid
(`frontend-tools-background`) — skipped.

Pure testid conveyance: no markup/behavior/styling changes, no dep
changes.

## Test plan

- [ ] D6 probes for `tool-rendering-custom-catchall` and
`frontend-tools` now pass for `claude-sdk-python` and `built-in-agent`
- [ ] No other integrations affected
2026-05-31 11:20:59 -07:00
Jordan Ritter b4f1984cf6 fix(showcase): align tool-rendering/frontend-tools testids to gold for claude-sdk-python + built-in-agent 2026-05-31 11:13:13 -07:00
Jordan Ritter 5f4d8764ae showcase(D6): structural LGP frontend parity for 9 baseline integrations (#5127)
## Summary (DRAFT — for scope review)

Mirrors the langgraph-python (LGP) gold-reference **frontend** onto the
9 baseline integrations — **structural parity**, not full D6-green (see
Scope below).

**Integrations:** mastra, agno, langroid, strands, spring-ai, ag2,
crewai-crews, llamaindex, langgraph-fastapi

**What this does:**
- Mirrors LGP's `src/app/demos/**` (+ `_shared/`),
`src/components/ui/**` (25 shadcn primitives), `src/lib/utils.ts`, and
the manifest-driven homepage `src/app/page.tsx` verbatim into each
baseline.
- Aligns `package.json` deps to LGP; pins `@copilotkit/*` to exact
**1.59.2**.
- **Bumps the Dojo reference** (`examples/integrations/*`)
`@copilotkit/*` to 1.59.2 so `showcase == Dojo` and validate-pins drift
retires properly (not a baseline hash-bump).
- Fixes the manifest.yaml SSG Docker build
(strands/spring-ai/llamaindex) — same fix as pydantic-ai in #5125.
- **Backend / API layer preserved untouched** (route.ts, agent servers,
Java/Mastra runtimes).

## Scope — what this is NOT

This is **structural** parity. It does **not** make the baselines
D6-green. Measured on langroid: most features pass, but ~7 stay red
(`gen-ui-agent`, `subagents`, `frontend-tools`, `gen-ui-declarative`,
`gen-ui-headless-complete`, `tool-rendering-custom-catchall`,
`beautiful-chat`). These are **backend-behavior reds** — the mirrored
LGP pages expect agent-emitted state/tool/gen-ui events that each
baseline's (different-framework) backend doesn't implement. Greening
them requires per-framework backend work (and some demos aren't feasible
per framework) — intentionally **out of scope** here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-05-31 11:11:59 -07:00
Jordan Ritter 8f8b7857e3 fix(showcase): stub agents.gen_ui_agent in forwarded-props test fixtures 2026-05-31 11:05:52 -07:00
Jordan Ritter e599da1409 fix(showcase): remove stale manifest highlights and orphan demo entries from mirrored baselines 2026-05-31 11:05:48 -07:00
Jordan Ritter 6b9b13d334 fix(showcase): copy src/components + src/lib into frontend Docker stage for mirrored baselines 2026-05-31 11:05:42 -07:00
Jordan Ritter fd0c3ed051 showcase: exact-pin mastra openai to satisfy canonical-pin validator 2026-05-31 10:39:37 -07:00
Jordan Ritter 76874f06dc showcase(D6): add gen-ui-agent set_steps backend to 8 baselines
Wire the set_steps state-card tool into each framework's existing
state-snapshot primitive so the (already-mirrored) gen-ui-agent demo can
drive [data-testid=agent-state-card] + agent-step rows:
- agno: state-aware AGUI route (StateSnapshotEvent)
- ag2: ContextVariables + ReplyResult via AGUIStream (dedicated sub-app)
- llamaindex: get_ag_ui_workflow_router backend_tools + StateSnapshotWorkflowEvent
- strands: ag_ui_strands ToolBehavior(state_from_args)
- crewai-crews: dedicated GenUiAgentFlow + copilotkit_emit_state
- mastra: working-memory STATE_SNAPSHOT (genUiAgent + setStepsTool)
- langroid: raw-OpenAI loop emitting TOOL_CALL_* + STATE_SNAPSHOT
- spring-ai: GenUiAgentController set_steps FunctionToolCallback + route override
Frontend (mirrored) untouched; route.ts slug overridden per integration.
2026-05-31 10:39:34 -07:00
Jordan Ritter b4a2b279d4 fix(showcase): copy manifest.yaml into frontend Docker stage for strands/spring-ai/llamaindex
The mirrored LGP homepage prerenders / and reads manifest.yaml at build
time; these 3 baselines use an explicit COPY of config files that omitted
manifest.yaml, breaking SSG export (same fix as pydantic-ai in #5125). The
other 6 baselines copy the dir wholesale and already include it.
2026-05-31 10:39:25 -07:00
Jordan Ritter 06e795b8e2 showcase(D6): mirror LGP frontend to 9 baseline integrations
Page-mirror the langgraph-python gold reference (demo pages + _shared/ +
src/components/ui shadcn primitives + src/lib/utils.ts + manifest-driven
homepage) into mastra, agno, langroid, strands, spring-ai, ag2,
crewai-crews, llamaindex, langgraph-fastapi. Align deps to LGP and pin
@copilotkit/* to exact 1.59.2. Backend / API-layer (route.ts, agent
servers, Java/Mastra runtimes) preserved untouched.

Known follow-up: agent-slug conveyance gaps (mirrored LGP demos reference
slugs each integration's route.ts registers under different names) — to be
measured + addressed per integration.
2026-05-31 10:39:21 -07:00
Jordan Ritter e11eb7de13 showcase: isolate pin-validation to showcase-internal (decouple from external Dojo) (#5128)
## Summary

Phase 1 of the showcase-internal pin-validation spec. Decouples
`showcase/integrations/<slug>` pin validation from the external
`examples/integrations/<source>` Dojo.

- Introduces `showcase/scripts/showcase-canonical-pins.json` as the
single source of truth for `@copilotkit/*` pinning (canonical `1.59.2`
plus per-slug per-dep overrides for `built-in-agent` and
`ms-agent-harness-dotnet`).
- Rewrites `showcase/scripts/validate-pins.ts` to enforce a
showcase-internal invariant:
  - Every `@copilotkit/*` dep pins to canonical OR its declared override
  - Every other framework dep is an exact pin (no `^/~/>=/latest/next`)
  - Workspace refs skip, not fail
  - Non-framework deps (react, lodash, etc.) are unconstrained
- Preserves the `[FAIL]` line format so the drift-ratchet
(sorted-unique-hash + count baseline) continues to work.
- Bumps every `showcase/integrations/*/package.json` `@copilotkit/*` to
canonical/override; regenerates lockfiles via `npm install
--package-lock-only --legacy-peer-deps`.
- Re-ratchets both `showcase/scripts/fail-baseline.json` and
`showcase/harness/test/fixtures/pin-drift/fail-baseline.json` to the new
output (`count=63`, single shared hash). Regenerates the harness
`cli-baseline-{stdout,stderr}.txt` fixtures.

Validator no longer reads `examples/integrations/` — that cross-product
comparison is retired.

## Test plan

- [x] `vitest showcase/scripts/__tests__/validate-pins` — 115 tests, 5
files, all pass
- [x] `tsx showcase/scripts/validate-pins.ts` — exits 1 with 63 FAILs
(baseline matches)
- [x] Drift-ratchet hash stable across multiple runs and post-lockfile
regen
- [ ] CI: `showcase_validate.yml` ratchet step passes (drift FAIL count
+ hash match committed baseline)
- [ ] CI: harness pin-drift probe parses new stderr fixture correctly

## Out of scope (future phases)

The 63 remaining FAILs are framework-side deps using ranges/dist-tags
(e.g. `@mastra/core "beta"`, `openai "^5.9.0"`, `langroid ">=0.53.0"`).
They are captured by the baseline hash for future ratchet-down.

**DRAFT** — DO NOT MERGE without orchestrator sign-off.
2026-05-31 10:23:43 -07:00
Jordan Ritter c4ce9781f6 showcase: regen nested langgraph-typescript agent lockfile for canonical 1.59.2 2026-05-31 10:13:21 -07:00
Jordan Ritter d15d2637a2 showcase: re-ratchet validate-pins baselines under new invariant 2026-05-31 10:13:21 -07:00
Jordan Ritter 03c9b91382 showcase: bump @copilotkit/* in integrations to canonical 1.59.2 2026-05-31 10:13:21 -07:00
Jordan Ritter 4ee78247ef showcase: rewrite validate-pins as showcase-internal canonical-pin validator 2026-05-31 10:13:13 -07:00
Jordan Ritter b75563fd9a showcase(D6): CST/pydantic page-mirror + MAF toolCallId parity + pydantic build fix (#5125)
## Summary

D6 showcase-parity wave for the priority bloc. Brings four integrations
closer to the langgraph-python (LGP) gold reference and unblocks
pydantic's production Docker build.

| Integration | D6 result | Change |
|---|---|---|
| **claude-sdk-typescript** | **130/55** (was 79/106, **+51**) | Full
LGP demo-page mirror + shadcn primitives + `copilotkit-beautiful-chat`
route (was 404) |
| **ms-agent-dotnet** | **180/2** (was 177/5) | toolCallId-strip via
universal `createAgent` `FunctionMiddleware` + reasoning-chain injection
+ fixtures |
| **ms-agent-python** | toolCallId-strip failures resolved | Ported the
universal strip middleware + snake_case `tool_calls` / nested
`function.tool_call_id` to match dotnet. Residual non-toolCallId
failures remain (interrupt timeouts, multimodal endpoint, declarative
pie-chart) — out of scope here |
| **pydantic-ai** | **Docker build fixed** (SSG export) + page-mirror |
Aligned deps to LGP; mirrored demo pages; **Dockerfile now copies
`manifest.yaml`** into the frontend stage (the homepage prerenders `/`
and reads it at build time, previously crashing the SSG export with
ENOENT). Residual gen-ui/timeout failures remain |

## Notes

- **Scope:** showcase integrations only (no shipped SDK code). CST and
dotnet are clear wins; pydantic build-fix is essential; python lands
correct middleware parity with documented residuals on the
de-prioritized Python bloc.
- **Page-mirrors** are verbatim copies of the LGP gold reference (demo
pages, `components/ui/*`, `lib/utils.ts`, the LGP dep set) with only the
backend/runtime swap retained at the API layer.
- **Code review:** 7-agent CR loop run against `origin/main`; **zero
functional (bucket-a) findings**. Flagged items were verbatim-LGP
dep-pin parity, pre-existing foundation machinery (not in this diff), or
internal-tool-bar non-functional (security/style) — deferred.
- **CI is the authoritative package-test gate** (the isolated worktree's
pnpm/lefthook SDK suite can't run locally). Integration builds verified
green via local docker builds.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-05-31 07:17:35 -07:00
Jordan Ritter 5f224c5af3 fix(showcase): sync validate-pins baseline hash to CI canonical value
Count is correct at 104 (down from 106 — net improvement from the
@copilotkit 1.59.2 exact-pin + manifest highlight fixes). The prior hash
was computed in a local env whose FAIL-line set differed from CI's
canonical pipeline; sync to CI's printed actual hash.
2026-05-31 03:11:35 -07:00
Jordan Ritter 2b13cbab38 fix(showcase): pin @copilotkit/* to exact 1.59.2 + re-ratchet validate-pins baseline
Previously pydantic-ai and claude-sdk-typescript used 'latest' for all
@copilotkit/* dependencies (and the @copilotkit/web-inspector pnpm
overrides on @copilotkit/core), which the showcase validate-pins ratchet
counts as non-exact pin drift. Intended design is to pin to exact 1.59.2.

Changes:
- integrations/pydantic-ai/package.json: @copilotkit/{a2ui-renderer,react-core,runtime,shared,voice} 'latest' -> '1.59.2'; npm + pnpm overrides on @copilotkit/web-inspector>@copilotkit/core 'latest' -> '1.59.2'
- integrations/claude-sdk-typescript/package.json: same set as above
- Regenerated both package-lock.json files via npm install --legacy-peer-deps --package-lock-only
- scripts/fail-baseline.json: ratcheted DOWN validatePinsFailCount 106 -> 104; updated validatePinsFailHash to dde7950e8d691de5a7b2c0c16ca64b3e550221cb6072d2c29e24dcb497515cf6 (matches local sort-uniq + shasum-256 of the new [FAIL] set)

Verified locally: npx tsx validate-pins.ts reports Summary FAIL=104
(2 fewer than baseline because pydantic-ai's @copilotkit/react-core and
@copilotkit/runtime moved from 'latest' (non-exact) to '1.59.2' (exact);
CST already had unrelated FAILs that remain). Non-@copilotkit deps
(lucide-react, cmdk, openai, @ag-ui/*) intentionally left unchanged.
2026-05-31 03:02:53 -07:00
Jordan Ritter d630bfa2a0 fix(showcase): resolve CST stale highlight paths in manifest
Four stale highlight paths in claude-sdk-typescript/manifest.yaml referenced
files that don't exist in the demo folders, breaking the shell-docs
bundle-demo-content step:

- chat-slots: custom-welcome-screen.tsx -> slot-wrappers.tsx (mirrors LGP layout)
- gen-ui-tool-based: agent.ts -> bar-chart.tsx (mirrors LGP layout)
- gen-ui-interrupt: time-picker-card.tsx -> _components/time-picker-card.tsx
- headless-complete: tool-renderers.tsx + use-rendered-messages.tsx -> chat/chat.tsx + hooks/use-tool-renderers.tsx (mirrors LGP layout)

Verified locally: scripts/bundle-demo-content.ts now bundles all 695 demos
with 0 failures.
2026-05-31 03:02:40 -07:00
Jordan Ritter b7620f1d26 fix(showcase): unify dashboard staleness and fix false-green chips (#5124)
## Summary

Unifies how the showcase status dashboard decides whether a demo's data
is "too old to trust," and removes two ways a chip/badge could show
green when the underlying data was stale or absent.

- **Shared staleness helper** (`lib/staleness.ts`): one `isStale` +
per-driver windows (e2e 6h, D4 1h, liveness 45m), so the chip, badges,
depth, and filters all apply the same freshness rule instead of three
private copies.
- **Stale-green → amber**: `resolveCell` and `resolveD5Row` downgrade
frozen-green rows so a driver that stopped reporting no longer reads as
healthy.
- **Honest D5 coverage**: `resolveD5Row` now returns `null` for
unmapped/empty-map features, matching the chip (`resolveD5`) and depth
(`isD5Green`). Previously an unmapped feature with a stray
`d5:<slug>/<featureId>` row rendered a green badge while the chip showed
gray — a visible contradiction.
- **No phantom regressions**: `isRegression` requires emitted data on
the rung above the achieved depth before flagging a regression.
- **D4 worst-state-wins**: `deriveDepth` folds chat/tools to the worst
state rather than OR-ing green.
- **One clock per render**: a single `now` (memoized on `liveStatus`) is
threaded through the cell-matrix render path so the chip, the
regressions/gaps filter, and the badges all judge staleness against the
same instant; previously the render path used a fresh `Date.now()` per
cell and could disagree across a window boundary.

## Test plan

- [x] `npm run typecheck` — clean
- [x] `npm run test` — 681 passed, 1 skipped
- [x] `npm run build` — clean
- [x] New red-green tests: stale-green downgrade (order-independent),
STRICT missing-sub-row, unmapped-feature → gray (not green), and
shared-`now` agreement across a staleness-window boundary
2026-05-31 02:38:23 -07:00
Jordan Ritter 1eefbc3d97 fix(showcase): copy manifest.yaml into pydantic-ai frontend Docker stage
The homepage prerenders / at build time and reads manifest.yaml via the
filesystem, but the frontend stage only copied next.config/tsconfig/postcss.
Without the manifest the SSG export of /page failed with ENOENT, breaking the
Docker build. Copy manifest.yaml alongside the other config files.
2026-05-31 02:38:10 -07:00
Jordan Ritter b1da0e836c Merge branch 'blitz/d6-parity/fix-pyd-deps' into blitz/d6-parity/integration 2026-05-30 21:37:13 -07:00
Jordan Ritter f356a56cf7 Merge branch 'blitz/d6-parity/fix-mafpy-mw' into blitz/d6-parity/integration 2026-05-30 21:37:12 -07:00
Jordan Ritter 44b57c5f25 fix(showcase): align pydantic-ai deps to LGP reference (unblock D6 prod build)
Aligns showcase/integrations/pydantic-ai/package.json to the
langgraph-python reference so the mirrored D6 demo pages can resolve
their shared-UI imports during `next build`.

- @copilotkit/*: 1.59.0 (exact-pinned) -> latest (matches LGP/CST)
- Add overrides + pnpm.overrides for @copilotkit/web-inspector>@copilotkit/core
- lucide-react: ^0.469.0 -> ^1.14.0
- openai: ^4.0.0 -> ^5.9.0
- tailwind-merge: ^2.6.0 -> ^3.5.0
- @playwright/test: ^1.50.0 -> ^1.59.1
- Add cmdk ^0.2.1, embla-carousel-react ^8.6.0, react-markdown ^10.1.0, remark-gfm ^4.0.1
- @ag-ui/client + @ag-ui/core remain at ^0.0.52 (satisfies @copilotkit/shared peer)
- Regenerate package-lock.json

Verified locally: npm install --legacy-peer-deps + npm run build ->
Compiled successfully in 6.9s, Generating static pages (56/56).
2026-05-30 21:36:40 -07:00
Jordan Ritter 6e934a79b7 fix(showcase): add CST copilotkit-beautiful-chat route (was 404)
The claude-sdk-typescript beautiful-chat demo page sets
runtimeUrl="/api/copilotkit-beautiful-chat" but no such Next.js route
existed, so the cell 404'd on every request.

Add a dedicated runtime mirroring pydantic-ai's beautiful-chat route:
openGenerativeUI + a2ui (injectA2UITool: false) + mcpApps middleware
configured together (the canonical LGP combined-runtime shape). Wire
the HttpAgent to CST's pass-through / mount since agent_server.ts has
no dedicated beautiful_chat backend graph — CST is frontend +
middleware driven, matching how mcp-apps is wired.

Also defensively register beautiful-chat in the shared /api/copilotkit
agentNames list so probe requests against the default runtime resolve
cleanly, matching the pattern used by the other dedicated-runtime
demos.
2026-05-30 21:35:07 -07:00
Jordan Ritter 2b05f9edc7 fix(showcase): port universal strip middleware + snake_case tool_calls to ms-agent-python route (D6 fixture parity) 2026-05-30 21:35:04 -07:00
Jordan Ritter f7203f3368 fix(showcase): thread shared now through cell-matrix render path
Hoist a single now (memoized on liveStatus) and thread it into the
render-path buildCellModel so the chip, the regressions/gaps filter, and
the badges all judge staleness against one instant. Previously the render
path defaulted to a fresh Date.now() per cell, so a cell the filter
included could render a different staleness state across a window boundary.
2026-05-30 21:34:48 -07:00