Two D5 conversation scripts that probe gen-UI rendering against the
showcase's `headless-simple` and `gen-ui-tool-based` demos.
`d5-gen-ui-headless` mirrors the recorded fixture against
`/demos/headless-simple` (frontend `show_card` via useComponent).
Assertion: a custom-rendered component is present in the DOM (not
just text), the matched node is structurally non-trivial, and the
assistant's follow-up narration references the rendered card.
`d5-gen-ui-custom` mirrors the recorded fixture against
`/demos/gen-ui-tool-based` (frontend `render_pie_chart`). Stricter
than the headless tier: also asserts a STRUCTURAL match — an `<svg>`
with multiple `<circle>` (or `<path>`) drawing children, which is
what distinguishes "custom" from "headless".
Shared selector cascade and DOM walkers live in `_gen-ui-shared.ts`
(leading underscore so the D5 driver's `^d5-.*\.(js|ts)$` loader
skips it). Selector cascade matches CopilotKit testids first, falls
back to `[data-tool-name]`, `.copilotkit-render-component`, and
`[role="article"] svg` so showcases that don't yet wire dedicated
testids still get probed.
Tests cover registration shape, buildTurns matches the fixture's
userMessage, and assertions catch missing component / empty wrapper /
wrong SVG shape / missing narration tokens. Robustness: the custom
assertion accepts both `<circle>` and `<path>` slices so a future
showcase switching chart libraries doesn't false-fail.
Adds the D5 tool-rendering script under
showcase/ops/src/probes/scripts/d5-tool-rendering.ts. Drives the
/demos/tool-rendering page with a single 'weather in Tokyo' user
turn (matching the recorded LGP fixture verbatim) and asserts on
DOM CARD STRUCTURE rather than just response text — the heart of
D5 for this feature is verifying the per-tool renderer mounted a
structured weather card.
The assertion uses a 4-selector cascade so the script works across
the fleet without per-integration overrides:
1. [data-testid=weather-card] — LGP canonical
2. [data-tool-name=get_weather] — alternative convention
3. .copilotkit-tool-render — class-based fallback
4. [data-testid=copilot-tool-render] — generic CK testid
Fail modes carry distinct messages so triage can distinguish
framework regression (no card at all) from content regression
(card present but missing temperature/city/inner elements).
Tests cover registration, buildTurns shape, the cascade order, and
all four assertion failure paths plus two green paths (canonical
selector and fallback selector). Run via pnpm exec vitest run
src/probes/scripts/d5-tool-rendering.test.ts (14/14 pass).
One D5 script claims both `mcp-apps` and `subagents` feature types and
routes both to `/demos/subagents` via `preNavigateRoute`. The fixture
(`mcp-subagents.json`) was recorded against the LangGraph Python
supervisor agent which chains research → writing → critique sub-agents;
coupling `/demos/mcp-apps` to a public Excalidraw MCP server would have
made replay depend on an external service, so we reuse the subagents
chain for both.
Per-turn assertion scrapes the rendered reply and asserts five fragments
unique to each sub-agent's contribution (research facts, writing draft
language, critique framing). If any sub-agent had been skipped its
fragment would not have made it into the critique's final draft, so the
fixture-driven text check is sufficient signal that the entire chain
executed.
Tests cover registration under both feature types, fixture filename,
buildTurns shape (single turn matching fixture's userMessage matcher
verbatim), preNavigateRoute returning `/demos/subagents` for both feature
types, and the chained-reply assertion's pass/fail behaviour against
scripted page text.
Authors the agentic-chat probe script under
showcase/ops/src/probes/scripts/d5-agentic-chat.ts. Self-registers with
the D5 registry on import; the e2e-deep driver's dynamic loader picks
it up automatically.
Three-turn conversation mirrors the recorded fixture verbatim
(showcase/ops/fixtures/d5/agentic-chat.json) so aimock's first-match-
wins userMessage matcher routes each turn to the correct canned
response:
1. greeting / first ask — "good name for a goldfish"
2. follow-up referencing prior turn — "name for its tank"
3. context verification — "what we named the goldfish"
Per-turn assertions:
- turn 1: assistant non-empty.
- turn 2: assistant mentions a tank-name fragment (anyOf list to
survive plausible model rephrasings).
- turn 3: assistant recalls "Bubbles" from turn 1 — the actual
context-retention check (case-insensitive).
Tests cover registration-on-import, fixture/turn alignment with
cross-turn substring isolation, and turn-3's context-retention check
in both pass and fail modes.
Adds the D5 (e2e-deep multi-turn conversation) probe driver as the
runtime entry point for the D5/D6 complex-interact parity probes.
- helpers/d5-registry.ts: keyed-by-featureType registry with
registerD5Script() entry point. Single registered script may claim
multiple feature types; double-registration throws so script-file
collisions surface at boot.
- drivers/e2e-deep.ts: kind:"e2e_deep" ProbeDriver. Fans out per
declared feature type, runs runConversation() from
helpers/conversation-runner.ts on a fresh Playwright context per
feature, emits aggregate e2e-deep:<slug> + per-feature
d5:<slug>/<featureType> rows. Loads scripts dynamically from
src/probes/scripts/d5-*.{js,ts} so Wave 2b script authors can ship
in parallel without touching this file.
- config/probes/e2e-deep.yml: cadence */60, max_concurrency 2,
timeout 300_000ms per spec.
- types/index.ts: adds e2e_deep + d5 to DIMENSIONS so the loader and
rule schema accept the new kind / dimension at load time.
- Driver runs cleanly with an empty scripts/ directory: warns and
emits zero rows rather than throwing, so this lands ahead of the
Wave 2b script PRs without a chicken-and-egg problem.
Hand-authored against LangGraph Python (LGP) source as the reference
implementation for the eight D5 (complex interact) feature types:
agentic-chat, tool-rendering, shared-state, hitl-approve-deny,
hitl-text-input, gen-ui-headless, gen-ui-custom, mcp-subagents
(/demos/subagents — chosen over /demos/mcp-apps to avoid the external
Excalidraw MCP dependency).
aimock's match model is single-shot, so multi-turn behavior is expressed
as sibling fixtures: userMessage substring matches for distinct user
turns, plus toolCallId-routed fixtures for mid-loop re-invocations after
a tool result is appended. toolCallId fixtures appear above their
userMessage counterparts in each file so first-match-wins doesn't
re-route the agent loop into the original tool-call fixture and create
an infinite loop.
Per-feature-type-against-LGP-only — the same fixture is replayed across
all 17 integrations via aimock's fixture pool. Per-integration overrides
are deferred until D5 surfaces real divergences.
All eight files validate cleanly via @copilotkit/aimock loadFixtureFile +
validateFixtures (0 errors / 0 warnings, 21 fixtures total). All 20
conversation legs across the eight files replay-pass against a booted
aimock instance (chat completions short-circuit to the expected text or
tool_calls — no escape to a live LLM).
Notes are not picked up by the existing showcase aimock-fixtures.test.ts
(its globs do not include showcase/ops/fixtures/d5); extending those
globs is left for the PR that lands the e2e-deep driver.
Captures the three orthogonal axes the parity-vs-reference probe needs
from the runtime SSE stream: ordered tool-call names, stream timing
profile (TTFT, inter-chunk gaps, p50, total chunks, duration), and JSON
contract shape (field path to JS-native type).
Uses CDP Network.dataReceived for per-wire-chunk timestamps without
intercepting the response, then retrieves the full payload via
Network.getResponseBody once loadingFinished fires. Playwright route
APIs cannot stream cleanly — route.fulfill takes a complete body and
route.fetch buffers upstream, both of which collapse chunk-arrival
timing to a single point. CDP is the only path that preserves real
per-chunk arrival timestamps. The response itself flows naturally to
the page; we observe, we do not intercept.
Parser tolerates SSE multi-line data records, comments, CRLF endings,
and malformed half-line chunks (returns a non-json record rather than
throwing). Tool-call extraction walks the four likely field shapes
(toolCallName, name, tool_call.name, tool_use.name) so the parity
engine can compare against reference implementations that have not
migrated to ag-ui canonical naming.
28 unit tests cover the parser, contract walker, timing computation,
and capture assembly using captured raw SSE samples — no live browser.
Adds `runConversation(page, turns, opts)` — a multi-turn Playwright
helper that drives a CopilotKit chat surface through a sequence of
assertions. Returns a structured `ConversationResult` describing how far
the conversation got, where it failed, and per-turn wall-clock cost.
Behaviour:
- Resolves the chat input via the same 5-selector cascade used in
`e2e-demos.ts` (canonical testid → default placeholder → generic
text-input fallbacks). Cascade is short-circuited when the caller
supplies `chatInputSelector`.
- Detects "response complete" by polling the assistant-message DOM
count via `page.evaluate` and waiting for `assistantSettleMs`
(default 1500 ms) of no growth past the previous-turn baseline.
Uses `evaluate` rather than `textContent(selector)` to avoid
Playwright's auto-wait machinery defeating the poll cadence —
same rationale documented in `e2e-smoke.ts`.
- On any failure (selector miss, fill/press throw, response timeout,
assertion throw) records `failure_turn` (1-indexed) + `error` and
returns immediately — subsequent turns are NOT executed.
Tests inject scripted fakes for the Page surface so no real browser is
required. Coverage on the helper is 94% lines / 97% branches; the
uncovered tail is the browser-context `evaluate` callback that can't be
exercised without chromium (matches the pattern in `e2e-smoke.ts`).
The `Page` type is a structural minimal interface (NOT the playwright
`Page`) so consumers don't have to acquire DOM lib types just to call
the helper. Real `playwright.Page` satisfies it structurally.
Pure-logic comparator for the four D6 parity axes (DOM, tools, stream,
contract). Takes a reference + captured ParitySnapshot plus optional
tolerances and returns a per-axis verdict, aggregate verdict, and
structured details for the dashboard / Slack writer to render.
No I/O — Playwright, fetch, and capture logic live elsewhere; this
module is the canonical scoring implementation shared between the
reference-capture script and the live D6 driver. Default stream
tolerances (TTFT 2x, P50 chunk 3x) are pulled verbatim from the spec
and flagged in code as needing empirical calibration post-merge.
Includes 27 unit tests covering pass/fail/edge cases for each axis
(testId-vs-class fallback, double-claim prevention, ratio overrides,
zero-divisor + NaN guards, mixed-axis aggregate counting) and three
JSON fixtures under test/fixtures/parity-snapshots/ exercising the
all-pass and all-fail end-to-end paths.
## Summary
The e2e-demos probe writes PocketBase rows as `e2e:<slug>/<featureId>`
but the dashboard looked up `e2e_smoke:<slug>/<featureId>`. Keys never
matched, so per-demo depth was always null and cells were stuck at D2.
One-line fix in `resolveCell` (change dimension from `e2e_smoke` to
`e2e`), plus corresponding updates to all tests and the
cell-pieces/feature-grid components that referenced the old dimension
name.
## Test plan
- [x] All 130 unit tests pass (`npm test` in shell-dashboard)
- [ ] After deploy, dashboard cells show D3/D4 for demos the probe has
already visited
## Summary
Replace the hardcoded 200-line INTEGRATIONS array in
`integration-smoke.spec.ts` with a 6-line derivation from
`registry.json`. New demos automatically appear in smoke tests when
manifests are updated.
- Delete hardcoded Integration interface + 17-entry array
- Derive slug, name, backendUrl, deployed, hasToolRendering, demos from
registry
- Add unit tests verifying derivation correctness + regression guard
against re-hardcoding
## Test plan
- [x] New vitest tests pass (7/7)
- [x] Existing showcase build pipeline tests pass (1183/1183 across 26
files)
- [ ] CI green
The e2e-demos probe writes per-demo rows keyed as e2e:<slug>/<featureId>
but resolveCell looked up e2e_smoke:<slug>/<featureId>. The mismatch
caused every per-demo depth chip to show gray (no data), leaving cells
stuck at D2. Fix: read the e2e dimension the probe actually writes.
All 15 framework packages now have demo routes wired after the omnibus
parity merge. Update comments to reflect full-fleet scope — namePrefix
and max_concurrency were already set correctly in the config values.
Add missing feature IDs (hitl, hitl-in-chat-booking) to langgraph-python
manifest so it reclaims reference status from langgraph-fastapi. Update
catalog test expectations and validate-pins fail-baseline.json hash/count
after PR #4287 dependency changes.
Packages that gained beautiful-chat or byoc-json-render demos via the
omnibus parity merge were missing direct dependencies that work locally
(hoisted from root node_modules) but fail in Docker's isolated npm
install.
Added:
- @radix-ui/react-checkbox to crewai-crews, langgraph-fastapi,
ms-agent-dotnet, ms-agent-python
- @json-render/core + @json-render/react to ms-agent-python, strands
- @copilotkit/voice + openai to ms-agent-python
Replace the 200-line hardcoded INTEGRATIONS array in integration-smoke.spec.ts
with a 6-line derivation from registry.json. New demos automatically appear in
smoke tests when manifests are updated — no manual maintenance needed.
18 per-framework MDX pages had the closing half of a multi-line
import statement left at the top of the file (the orphaned
"Symbol1, Symbol2, } from '...'" tail with the opening "import {"
line missing). The current sync-docs-from-main import stripper
handles multi-line imports correctly, so this was a historical
artifact from an earlier stripper version that misparsed them —
every affected file predates that fix.
MDX treats the orphan as plain text, so pages rendered the
identifier list and "} from ..." string as visible garbage at the
top before the real content started. Strip the orphan blocks; future
syncs won't re-introduce them.
The MDX registry had TailoredContent and TailoredContentOption
stubbed as passthrough <div>{children}</div> wrappers. MDX pages
author these as a variant-switcher (e.g. on Readables:
"Custom graph" vs "Prebuilt agent" paths for LangGraph setup), so
the stub rendered both option paths stacked — the same useAgentContext
example, the same steps, effectively duplicating multi-hundred-line
sections on every page that uses the component.
Swap the stubs for the real implementation at
components/react/tailored-content.tsx, which has been in the repo
since the shell-docs move but was never wired in. The real component
renders only the selected option and persists the choice in a URL
search param (?impl=graph / ?impl=prebuilt).
Registry slugs don't always match the integrations/<folder>/ name on
disk. Three LangChain/LangGraph variants (langgraph-python, langgraph-
typescript, langgraph-fastapi) read from the single langgraph/ tree,
ms-agent-dotnet and ms-agent-python share microsoft-agent-framework/,
and google-adk/strands are legacy renames that point at adk/ and
aws-strands/ respectively.
Before: the framework-scoped router, the sidebar-override nav
builder, and the "not available for this framework" fallback all
used the URL slug directly as a folder name, so any of the seven
mismatched slugs showed empty sidebars, 404s on framework-unique
pages (/langgraph-python/auth, /ms-agent-dotnet/auth), and missing
'available in other integrations' matches.
After: lib/registry exposes getDocsFolder(slug) backed by a small
DOCS_FOLDER_OVERRIDES table. Callers resolve the URL slug to its
actual folder before touching disk; findFrameworksWithPage takes the
resolver as a parameter so docs-render stays registry-free.
Per-page variant selectors authored as <Tabs groupId="..." default="Python">
now open with the URL-matching tab preselected instead of the author's
hardcoded default. getTabDefault(slug, groupId) reads
TAB_DEFAULTS_BY_SLUG; a wrapper in DocsPageView's MDX components map
injects the resolved value into <Tabs> via a 'default' prop alias.
/langgraph-typescript/configurable opens TypeScript, /ms-agent-dotnet/
auth opens .NET, /langgraph-fastapi/deep-agents opens FastAPI. Slugs
and groupIds without a mapping fall through to the existing behavior
(author default, then first items label).
Every framework under integrations/<fw>/ shipped its own copy of
contributing/ + telemetry/ content that was byte-identical (or
trivially divergent — a stray "cd" path, a legacy CLI name) to the
canonical copy at root (other)/. The duplicates were stale sync
artifacts with no framework-specific content: "how to contribute to
CopilotKit" and "how to configure telemetry" don't vary by agent
framework.
Beyond disk clutter (~2.5k lines across 8 frameworks), the duplicates
surfaced as a "Other" group nested under the framework-scoped
sidebar section whenever the merged nav built from meta.json — a
second copy of root's own "Other" section at the bottom of the
sidebar. Deleting the trees makes that UI bug disappear without any
filter patching.
Scope:
- Remove integrations/<fw>/(other)/ trees for ag2, agno, aws-strands,
crewai-flows, langgraph, llamaindex, mastra, microsoft-agent-framework.
- Drop ---Other--- + ...(other) entries from each framework's meta.json.
- Add a path-exclusion filter in sync-docs-from-main.ts so the next
sync run doesn't resurrect the subtrees when upstream edits touch
them. Upstream keeps its copies (removing them there means touching
all 13 parallel framework trees, out of scope for this branch).
The docs landing page and sidebar framework-selector both gated a
grayed-out card state and a "soon" label on integration.deployed.
That flag tracks whether a live showcase demo with tagged cells
exists — a showcase concern, not a docs concern. Built-in Agent has
ready docs even though no showcase package is published yet, so the
card was incorrectly rendered as "coming soon" and stayed visually
inert after clearing the stored framework.
Every integration that ships docs should be pickable from these
surfaces with the same visual weight. The deployed flag continues to
drive the showcase app, router-pivot filtering, and snippet cell
assertions — this change only touches the two docs-picker call sites.
Before this change, visiting /<framework>/<slug> for a topic that only
exists under integrations/<other-framework>/ returned a bare 404 — for
example, /mastra/advanced-configuration (the page only lives under
integrations/built-in-agent/).
Now the router checks whether the slug exists in any other integration
and renders a framework-scoped fallback page inside the docs shell:
sidebar and framework switcher stay intact, and the body lists the
integrations where the topic does exist with direct links. Genuine
unknown slugs still 404.
New helper findFrameworksWithPage walks integrations/<slug>/ for each
registered framework. NotAvailableForFrameworkPage renders the shell
around the fallback body. Nav tree build is hoisted so both the happy
path and the fallback share one source.
Per-framework meta.json files (mastra, langgraph, llamaindex, etc.)
mirror the root tree's section names ("Getting Started", "Basics").
Passing them through buildFrameworkOverridesNav into the merged
sidebar caused duplicate React keys when the merge ran — every root
section collided with the override's copy of the same title.
The override block is already wrapped in a single "{frameworkName}"
section by mergeFrameworkNav, so nested section headers added no
information anyway. Drop them at the filter step.
Registers Built-in Agent as a framework in the registry and wires up
a router + sidebar-nav pattern so its content can live at /built-in-agent/*
without needing a dedicated per-framework content tree for every topic.
Content model:
- Root MDX pages (/quickstart, /frontend-tools, /shared-state, etc.) are
the canonical home for framework-agnostic topics. Rendered at
/built-in-agent/<slug> via the existing framework-override mechanism.
- integrations/built-in-agent/*.mdx is the escape hatch for topics that
are genuinely BIA-specific (copilot-runtime, server-tools, mcp-servers,
model-selection, advanced-configuration, custom-agent). The router
falls back to these when no root equivalent exists.
- Root wins when both exist.
Changes:
- shared/manifest.schema.json: add 'built-in' to the category enum.
- shared/packages.json: register built-in-agent slug.
- packages/built-in-agent/manifest.yaml: new. deployed:false (showcase
package TBD in a follow-up), sort_order:0, category:popular so it
appears at the top of the framework dropdown.
- public/logos/built-in-agent.svg: new logo asset (extracted from the
inline CopilotKit mark in brand-nav.tsx).
- shell-docs/src/app/[framework]/[[...slug]]/page.tsx: router gains a
fallback to integrations/<framework>/<slug>.mdx when the root file
doesn't exist. Sidebar nav merges in per-framework overrides as a
labeled section positioned after 'App Control' (mirrors upstream's
integrations/built-in-agent/meta.json ordering).
- shell-docs/src/components/docs-page-view.tsx: new optional
contentSlugPath prop lets the router thread through the override
content path without changing the URL-slug used for breadcrumbs and
active-link detection.
- shell-docs/src/lib/docs-render.tsx: new buildFrameworkOverridesNav
helper that walks integrations/<framework>/* and filters out pages
that already exist at root.