Aimock fixtures match by userMessage substring, which collides across
integrations whose pills share the same prompt template. The validator
correctly catches that the d5 probe fixtures my PR added would dangle
on integrations whose tool surfaces don't include the emitted tool.
Fix by adding `match.toolName` constraints to the affected fixtures so
aimock only fires them when the agent emits that specific tool — and
the validator skips them for demos that don't declare the tool.
Also delete the obsolete `Visualize pitch, yaw, and roll` fixture
(F3 changed the d5-gen-ui-open-advanced probe to use the
`Inline expression evaluator` pill instead).
Locally `pnpm exec tsx validate-fixture-tool-surface.ts` returns
0 drift after this change.
Two small hygiene fixes:
- frontend-tools cosmic-gradient pill matched bare token 'navy' which
could collide with any prompt mentioning navy. Tighten to
'navy → magenta cosmic gradient' — verbatim substring of the
pill prompt in frontend-tools/suggestions.ts. Mirror into d5-all.json.
- d5-all.json header comment claimed the first 7 entries are the
open-gen-ui pills, but ordering shifts as features land (headless-
simple/complete pills currently sit ahead of open-gen-ui after
recent merges). Rewrite the header to describe the ordering
convention generically: high-priority verbatim-prompt fixtures
appear first, first-match-wins; per-fixture _comments and the
per-feature source files are the source of truth.
The per-feature source file owned only 2 fixtures (one matcher,
'project planning') while the bundled d5-all.json had 6 fixtures
(3 pills × 2 turns each: project-planning, auth, reading). Because
d5-all.json is auto-merged from the per-feature sources, any future
regen would wipe the auth and reading fixtures from the bundle and
silently regress the frontend-tools-async probe.
Copy all 6 fixture entries from d5-all.json into
frontend-tools-async.json so the source is authoritative. No change
to d5-all.json — the bundle already has these entries.
The render_a2ui matcher emitted the same Card+Metric payload for every
pill, so the d5-gen-ui-declarative probe went red on the second pill —
its per-pill expectedTestIds map demands declarative-pie-chart for the
pie-chart pill, declarative-bar-chart for the bar-chart pill, and
declarative-status-badge for the status-report pill, none of which
were rendered.
Branch the render_a2ui response by combining userMessage substring with
toolName so each pill emits the catalog component its probe expects:
- KPI dashboard → Card + 3 Metric children
- pie chart → PieChart with regional sales data
- bar chart → BarChart with quarterly revenue data
- status report → Card + 3 StatusBadge children
Keep the bare toolName-only matcher at the bottom as a deterministic
fallback so unforeseen pills still render something instead of erroring.
Mirror all five fixture entries into the bundled d5-all.json.
Two pill mismatches were silently routing the headless-complete Stock and
Highlight pills to the showcase-assistant catch-all in feature-parity.json:
- The Stock fixture matched 'AAPL trading' but the SuggestionBar pill
configured by use-headless-suggestions.ts sends 'What's the price of
AAPL right now?' — substring 'AAPL trading' is not in that prompt.
- The Highlight fixture matched 'Highlight \'meeting at 3pm\'' but neither
the empty-state pill ('Highlight: ship the demo on Friday') nor the
SuggestionBar pill ('Highlight this note for me: ...ship the demo on
Friday...') contains that substring.
Switch both matchers to short distinctive substrings ('AAPL' and
'ship the demo on Friday') that appear verbatim in BOTH the empty-state
and SuggestionBar prompts. The tool-rendering AAPL fixture (matcher
'What\'s the current price of AAPL?') stays at higher priority via
array order — first-match-wins keeps it pinned to the tool-rendering
pill, so substring 'AAPL' here only catches headless-complete pills.
Update narration response text to reference the correct highlighted
phrase.
Two related issues in the per-pill assertions:
1. The status-report pill listed BOTH `declarative-status-badge` and
`declarative-card` as expected testids and used `.some()` to pass.
Pill 1 (kpi-dashboard) already mounts `declarative-card`, so the
status-report pill trivially passed via leftover even when it
produced no status badge at all.
2. All pills used `.some()` (any-of), so mounting a single satisfying
testid was enough — even one that was leftover from an earlier pill.
Switch the per-pill expected-testid check to `.every()` (conjunctive)
and require the status-report pill's distinguishing testid
(`declarative-status-badge`) explicitly. Add a per-pill baseline
captured via the conversation-runner's `preFill` hook so the assertion
also requires at least ONE expected testid to be NEWLY mounted (not
just present in the DOM as cross-pill leftover).
`[data-testid="agent-step"]` rows accumulate across pills — pill N's
rows persist into pill N+1's DOM. The previous absolute-count check
(`stepCount >= 2`) trivially passed on pill 2/3 even when the pill
produced no new rows; the cumulative-text duplication check was the
only signal catching cross-pill regressions.
Use the conversation-runner's `preFill` hook to capture the agent-step
DOM state BEFORE each pill is sent, then assert the count grew by at
least 2 NEW rows (delta-based, not absolute). The cumulative-text
fingerprint check is kept as a secondary signal to catch fixtures that
return the same canned step list across prompts.
`[data-testid="document-content"]` is sticky DOM that carries text from
pill N into pill N+1's pre-fill state. The previous assertion only
checked the absolute final char count, so pill 2/3 trivially passed on
pill 1's leftover document content even when the stream regressed.
Use the conversation-runner's `preFill` hook to capture the document
content baseline (text + char count) BEFORE each pill is sent, then
after settle assert either:
- delta from baseline >= STREAMING_MIN_FINAL_CHARS (substantive new
content appended), OR
- text differs from baseline AND final char count >= threshold
(document was REPLACED with substantive content of similar size).
Drop the misleading "guards against single-chunk emission" claim from
the docstring — the assertion runs after the runner's settle window,
so mid-stream chunking is NOT observable here. Documented as a
follow-up.
The probe input was "hello" with a misleading comment claiming it was a
no-op. The MCP-apps demo only mounts a sandboxed iframe AFTER an MCP
tool call returns a UI resource — `MCPAppsActivityRenderer` subscribes
to the runtime's MCP activity event and dynamically appends the iframe;
"hello" never triggers that path, so the iframe assertion was either
trivially red or trivially green via DOM leftovers.
Switch the input to the verbatim suggestion-pill prompt
"Open Excalidraw and sketch a system diagram with a client, server,
and database." (from `langgraph-python/src/app/demos/mcp-apps/suggestions.ts`)
and document honestly that the canned aimock fixture currently keyed to
"Use Excalidraw to sketch" returns content-only (no MCP tool call), so
the iframe assertion can only pass under a live agent + reachable MCP
endpoint until the fixture is updated. F4 owns the fixture file — TODO
left in the comment.
The probe sent "render the advanced gen-ui sandbox" — a string with no
fixture entry in `showcase/aimock/d5-all.json` or
`showcase/harness/fixtures/d5/gen-ui-open.json`, so the iframe never
rendered through a deterministic tool call. Switch to the verbatim
suggestion-pill prompt "Inline expression evaluator" which is keyed to
a `generateSandboxedUi` tool call in d5-all.json (line 77).
Also tighten the iframe selector cascade: the bare `iframe` fallback
matched ANY iframe on the page (e.g. third-party analytics frames).
Replace with `iframe[sandbox*="allow-scripts"]`, matching the advanced
demo's actual sandbox spec.
The probe returned at the first snapshot with critic count == 1, so a
regression that briefly shows 1 critic card before the supervisor
re-invokes critique_agent (count grows to 2) would silently pass.
After the validator passes, hold for CHAIN_SETTLE_DWELL_MS (3s) and
re-snapshot. If the second snapshot fails validation, including a
critic count > 1 or the new sub-agent-empty sentinel appearing in
the page text, throw with a clear destabilised-during-dwell message
that surfaces both initial and post-dwell counts.
Adds the empty-sub-agent sentinel from subagents.py to the probe
BOILERPLATE_MARKERS so an empty sub-agent result fails the test
instead of rendering as an empty card.
[first.status, second.status].sort() defaults to lexicographic
ordering — it returned [200, 429] correctly only by accident and
would mis-order [200, 1000]. Pass an explicit (a,b)=>a-b comparator
so the assertion holds for any future status-code pair.
_invoke_sub_agent's last-resort branch silently returned '' or a
Python repr like \"[{'type': 'text', ...}]\" for block-list content,
which the UI would render as a blank/garbled card. Return a stable
SUB_AGENT_EMPTY_SENTINEL ('<sub-agent produced no output>') instead
so the d5-subagents probe can match it against its boilerplate-marker
list and fail the genuine-pass test loudly when a sub-agent produces
no usable output.
The get_stock_price tool returned randint-based price/change every
call, so e2e specs asserting $338.37 / -2.96% would only pass when
aimock short-circuited the entire call. The Python tool body still
runs server-side under aimock — only the LLM call is mocked.
Mirror the deterministic-`value` pattern on roll_d20: accept optional
price_usd / change_pct arguments and echo them back when present.
Defaults to random mock data when the args are omitted, preserving
the legacy live-LLM behaviour.
The probe sent "show your reasoning step by step" and accepted any
assistant transcript containing "reasoning", "step", or "thinking".
Those tokens are typical assistant acknowledgements of the prompt
itself, so the keyword fallback always matched regardless of
whether the framework actually surfaced REASONING_MESSAGE_* events
to the frontend — making the assertion non-genuine.
Strict-only behavior now: the assertion requires one of the four
stable role/testid markers (reasoning-block / reasoning-content /
reasoning-default / [data-message-role="reasoning"]) to render
within the timeout. Integrations that rely on CopilotKit's default
reasoning slot without a stable testid will fail this probe — by
design.
Replace REASONING_KEYWORDS export with REASONING_SELECTORS, drop
the readAssistantTranscript helper (no longer needed), and update
the test to exercise the testid-presence success path and the
timeout failure path.
The HELLO_LEADING phrase was the showcase-assistant catch-all
boilerplate ('I can help you with weather lookups...') that other
tests in this PR explicitly guard AGAINST. The dedicated d5-all.json
fixture for 'Say hello in one short sentence' now returns a distinct
non-boilerplate greeting; the spec asserts that distinct phrase, so a
fixture-priority misroute fails loudly instead of passing by accident.
PILL_GRADIENT_HINTS contained naked hex prefixes (e.g. "#0", "#1",
"#2", "#ff", "#fb", "#3"-"#9") that accidentally cross-matched
across families — cosmic's #1e3a8a matches forest's "#1", so a
regression returning the same gradient for every pill could have
silently passed.
Drop ALL naked-hex prefix hints. Keep only color-name keywords:
- sunset: sunset, orange, red, rose, pink, amber, coral, peach
- forest: forest, green, emerald, lime, olive, teal
- cosmic: cosmic, space, navy, magenta, purple, violet, indigo,
fuchsia
If full-hex matching is ever needed, use full 6-digit codes from
the actual fixture rather than 1-2 character prefixes.
Also fix the broken "background did not change off baseline" test:
the previous assertion built a baseline-only page mock and then
discarded it via `void page`. Replaced with a fake-timer-driven
invocation that actually triggers the timeout branch, plus a
separate test for the wrong-family path.
The length-mode (responseLength) assertion compared the value-B
delta against the FULL CUMULATIVE pre-A transcript, so on the third
knob pair the math was strongly negative and the threshold check
always failed. Same shape mismatch existed on text-mode.
Thread a per-knob snapshot through each pair so both sides of the
comparison are per-turn deltas:
- Capture priorCumulative BEFORE turn A (chained from the previous
pair's value-B assertion via chainPriorCapture; pair 0 starts
with empty since the page is blank pre-run).
- Snapshot assertion stores aOnly = postA - priorCumulative and
postACumulative for the next subtraction.
- Diff assertion isolates onlyB = postB - postACumulative and
compares aOnly vs onlyB (text equality or length delta).
Tests updated to construct KnobSnapshot bundles via a snapshotFor
helper and exercise the cumulative-prior third-pair scenario for
length-mode.
The default-catchall probe queried `tool-status-{pending,executing,complete}`
but `DefaultToolCallRenderer` emits a single `copilot-tool-render-status`
pill with the lifecycle state on the container's `data-status` attribute.
Collapse the testid array to that single id and update the selector.
The custom-catchall probe queried `custom-catchall-render`, but the
production `CustomCatchallRenderer` emits `custom-wildcard-card`. Update
the constant + selector. Also swap the unmatched `"AAPL stock price"`
prompt for `"What's the current price of AAPL?"`, the verbatim key
already wired in `showcase/aimock/d5-all.json`, so the fixture
deterministically resolves instead of falling through to the live model.
Tests in both .test.ts files updated to match the new constants and
prompt; both files green locally.
The blitz removed the explicit useDefaultRenderTool() call expecting a
framework-level fallback to handle zero-hook registrations, but the
integration uses the published @copilotkit/react-core@1.56.5 which does
not yet ship that fallback. Without the call, useRenderToolCall has no
'*' renderer and tool calls render invisibly — the user only sees the
agent's final text summary instead of the OOTB tool card.
Restore the explicit useDefaultRenderTool() invocation. The framework
fallback (committed in this PR but inert until react-core publishes a
release that includes it) becomes a no-op once that ships.
## Summary
Seven commits, batched into one PR for the May 12 docs cutover. All
under shell-docs.
- `3961a7e34` Normalize V2 canonical form across canonical V2 docs (137
files): `<CopilotKitProvider>` to `<CopilotKit>`, imports from
`@copilotkit/react-core/v2` to `@copilotkit/react-core` (root), styles
from `@copilotkit/react-ui/v2/styles.css` to
`@copilotkit/react-core/styles.css`, drop `@copilotkit/react-ui` from
npm install. Also replaces stale V1 hook references in prose:
`useCopilotAction` to `useFrontendTool` in `agentic-protocols/a2a.mdx`,
`useCopilotReadable` to `useAgentContext` in `backend/custom-agent.mdx`.
Excludes intentional V1 references in migrate guides, V1 reference tree,
migration callouts, V1 release notes.
- `7cc939103` Consolidate `custom-agent` onto
`backend/custom-agent.mdx`. Two divergent shell-docs copies of Factory
Mode content existed: `backend/custom-agent.mdx` (508 lines,
structurally complete) is now canonical;
`integrations/built-in-agent/custom-agent.mdx` (240 lines, missing 5
sections) is deleted. Adds 301 redirects from both historical paths.
Five inbound link retargets landed in the V2 normalization commit.
- `211785135` Add `IframeSwitcher` component at
`src/components/content/iframe-switcher.tsx`. Five MDX files import
`IframeSwitcher` from `@/components/content` but the component file was
missing, causing those tags to silently render nothing. Component wraps
the existing `<Tabs>` with two iframes (Demo + Code).
- `8cc79c26e` `IframeSwitcher` `id` prop forwarding (was declared but
never consumed) + V2 SDK prose fix in `concepts/oss-vs-enterprise.mdx`
(Frontend SDK bullet now reflects that components ship from
`react-core`).
- `23c9b6f7c` Replace the stub `IframeSwitcher` in `mdx-registry.tsx`
with the real component. The stub took `src`/`title` props (old shape)
and silently dropped `exampleUrl`/`codeUrl`/`exampleLabel`/`codeLabel`
from MDX consumers, so all `<IframeSwitcher>` tags were rendering empty
containers despite the component being added.
- `31995933d` Render Demo + Code tabs on `<InlineDemo>`. Pages using
`<InlineDemo demo="..." />` previously showed only the demo iframe; they
now render a Tabs strip with the demo iframe and a code iframe
side-by-side. Code iframe pulls from
`feature-viewer.copilotkit.ai/<framework>/feature/<demo>?view=code&sidebar=false&codeLayout=tabs`.
Framework slug translates from the registry slug (e.g.
`langgraph-python`) to the upstream-style slug (`langgraph`) via
`getDocsFolder()`.
- `447751567` Translate `InlineDemo` demo slug from dash-form
(`agentic-chat`, registry) to underscore-form (`agentic_chat`,
feature-viewer) for the Code iframe URL. Without this, the Code tab
404'd.
## Test plan
- [ ] `npm run build` from `showcase/shell-docs/` clean.
- [ ] Visit a V2-canonical page (e.g. `/agent-config`,
`/backend/copilot-runtime`, `/threads`). Imports show `<CopilotKit>`
from `@copilotkit/react-core` (root), styles from
`@copilotkit/react-core/styles.css`. No `/v2` or `react-ui` paths in
canonical V2 docs.
- [ ] `/backend/custom-agent` renders the full 508-line Factory Mode
content (Tabs UI, all upstream sections). `/built-in-agent/custom-agent`
and `/integrations/built-in-agent/custom-agent` 301 to
`/backend/custom-agent`.
- [ ] On a page using `<IframeSwitcher>` (e.g.
`/langgraph-python/human-in-the-loop/interrupt-flow`), confirm the Demo
and Code tabs both load iframes from `feature-viewer.copilotkit.ai`.
- [ ] On a page using `<InlineDemo>` (e.g.
`/langgraph-python/prebuilt-components/chat`, `/agentic-chat-ui`,
`/voice`, `/shared-state`, `/auth`), confirm the embedded demo now has a
Demo / Code tab strip. Demo iframe loads from the integration backend,
Code iframe loads from
`feature-viewer.copilotkit.ai/<upstream-framework>/feature/<demo-with-underscores>?view=code...`.
- [ ] Spot-check the 5 inbound links that retargeted to
`/backend/custom-agent`: `backend/copilot-runtime.mdx`,
`integrations/built-in-agent/{index,model-selection,server-tools,advanced-configuration}.mdx`.
- [ ] Grep `showcase/shell-docs/src/content/` for `react-core/v2`,
`react-ui/v2`, `CopilotKitProvider`. Only matches in V1 migrate guides
and V1 reference docs.
The earlier V2 canonical sweep misread the brand-name guidance ("always
CopilotKit") as a directive on import paths and stripped /v2 from
@copilotkit/react-core specifiers. Inline review feedback clarified that
guidance applied only to the component name. This restores
"@copilotkit/react-core/v2" and "@copilotkit/react-core/v2/styles.css"
across the docs sweep scope (now including 8 framework quickstarts
inherited via rebase onto main); the <CopilotKit> rename and the drop
of @copilotkit/react-ui from install commands are kept.
## Release monorepo v1.57.1
**Scope:** `monorepo` | **Bump:** `patch`
---
### How this release process works
1. **This PR was created automatically** by the "release / create-pr"
workflow.
It bumped the `monorepo` packages to `1.57.1`
and generated AI-enhanced release notes.
2. **CI runs on this PR** — the full test suite (unit tests, lint, type
checks, build)
must pass before merging. This is the review gate.
3. **Review the release notes** in `release-notes.md` in this PR.
If a Notion draft was created, you can edit the release notes there
before merging.
4. **When this PR is merged**, the `release / publish` workflow
automatically:
- Builds all packages
- Publishes the `monorepo` packages to npm at version `1.57.1`
- Creates git tag `monorepo/v1.57.1`
- Creates a GitHub Release with the final release notes
### Before merging
- [ ] CI is green (tests, lint, types, build)
- [ ] Version bumps look correct
- [ ] Release notes are accurate (edit in Notion if a draft was created)
---
> **Do not merge until CI is fully green.** The full test suite runs
automatically on this PR.
Registry demo slugs are dash-form (`agentic-chat`); feature-viewer.copilotkit.ai
serves them at underscore-form (`agentic_chat`). The Code tab iframe was
404'ing because the slug was passed through verbatim. Replace `-` with `_`
when building the code URL.
Wraps the InlineDemo iframe in a Demo / Code tab strip mirroring the
IframeSwitcher component, so pages using <InlineDemo demo="..." /> now
expose both the live demo (integration backend) and a code view. The
code iframe points at feature-viewer.copilotkit.ai/<framework>/feature/
<demo>?view=code&sidebar=false&codeLayout=tabs, where <framework> is
translated through getDocsFolder() to map registry slugs like
langgraph-python down to their upstream folder name (langgraph) used by
the feature viewer.
The mdx-registry shipped a stub IframeSwitcher that took `src`/`title`
props and rendered a single iframe. MDX consumers actually pass
`exampleUrl`/`codeUrl`/`exampleLabel`/`codeLabel`, so the stub was
silently dropping those props and rendering an empty container. The
result: pages using `<IframeSwitcher>` showed no demo+code tabs.
Replace the stub with the real component from `@/components/content`,
which matches the props shape consumers actually use.
IframeSwitcher: forward the `id` prop to a wrapping `<div id={id}>` so MDX
authors can deep-link to a specific switcher instance. The prop was
declared but never consumed.
oss-vs-enterprise.mdx: fix the Frontend SDK bullet to reflect canonical
V2 — both hooks and prebuilt components ship from `@copilotkit/react-core`
(the standalone `@copilotkit/react-ui` package is V1-era).
Several MDX files import `IframeSwitcher` from `@/components/content`
(prebuilt-components, frontend-tools, interactive, tool-rendering, plus
integration overrides), but the component file was missing. This adds it
as a thin wrapper around the existing `<Tabs>` component, rendering a
demo iframe and a code iframe in a Tabs strip.
Props: `exampleUrl`, `codeUrl`, `exampleLabel` (default "Demo"),
`codeLabel` (default "Code"), `height` (default "600px"). Both iframes
are sandboxed and lazy-loaded.
Matches the upstream IframeSwitcher pattern used on docs.copilotkit.ai
to render embedded feature-viewer demos alongside their backing code.
backend/custom-agent.mdx (508 lines, structurally complete) is the canonical
Factory Mode page. The integrations/built-in-agent/custom-agent.mdx copy
(240 lines, missing 5 sections) was retired:
- Add 301 redirects in next.config.ts for the two historical paths
(/built-in-agent/custom-agent and /integrations/built-in-agent/custom-agent
→ /backend/custom-agent).
- Inbound link retargets to /backend/custom-agent landed in the previous
V2 normalization commit (5 files).
- Delete the divergent 240-line copy.
Apply the canonical V2 import form across all V2 docs:
- `<CopilotKit>` (not `<CopilotKitProvider>`)
- imports from `@copilotkit/react-core` (root, not `/v2`)
- styles from `@copilotkit/react-core/styles.css` (not `react-ui/v2/styles.css`)
- drop `@copilotkit/react-ui` from npm install commands
Fix V2 leaks in canonical pages: replace stale `useCopilotAction` and
`useCopilotReadable` references in agentic-protocols/a2a.mdx and
backend/custom-agent.mdx with their V2 equivalents (`useFrontendTool`,
`useAgentContext`).
Excludes intentional V1 references in migrate guides, V1 reference tree,
migration callouts, and V1 release notes.
## Summary
Six commits, batched into one PR for the May 12 docs cutover.
- `2c10c735d` Port upstream CLI walkthrough verbatim onto 8 framework
quickstarts (adk, agno, aws-strands, langgraph, llamaindex, mastra,
microsoft-agent-framework, pydantic-ai). Initial byte-perfect import
from upstream `docs/content/docs/integrations/<fw>/quickstart.mdx`.
- `f17c86867` Add `sample-audio-button` region marker on the
built-in-agent voice demo. Closes a B-cell snippet gap; the bundler now
reports 4 regions on `built-in-agent::voice`.
- `cebca9216` Remove AG-UI brand tab from desktop header for visual
parity with `docs.copilotkit.ai` (canonical has zero AG-UI elements in
the header).
- `abeb2792b` Also remove AG-UI link from mobile slide-out menu and drop
now-unused helpers (`AgUiIcon`, `AG_UI_LINKS`, `AG_UI_PREFIXES`, `Brand`
type, `activeBrandFromPath`, `usePathname`). Net 82-line cleanup of
`brand-nav.tsx`.
- `99d8ac7a5` Intermediate rewrite of the 8 quickstarts to use `npx
copilotkit@latest create --framework <id>` (after CLI accuracy audit
found `--intelligence` and `--framework` are mutually exclusive in
`copilotkit@2.0.3`). Also applies V2 canonical imports to the same 8
files: `<CopilotKit>` from `@copilotkit/react-core` (root), styles from
`@copilotkit/react-core/styles.css`, drop `@copilotkit/react-ui` from
npm install.
- `59a8a71a6` Final CLI-section shape: each quickstart now leads with
the upstream interactive walkthrough (`npx copilotkit@latest create`
plus a bulleted explanation of Project name / Enterprise Intelligence
Platform with sign-up link / Framework picker), followed by an "Or skip
the prompts and pin a framework directly" `--framework <id>` block.
LangGraph and Microsoft Agent Framework show both language variants in
the flag block.
## Test plan
- [ ] Visit each of the 8 framework quickstarts on the dev server.
Verify the CLI section shows the interactive walkthrough with the
EIP/sign-up bullet AND the `--framework <id>` flag block below it.
- [ ] LangGraph quickstart shows both `langgraph-py` and `langgraph-js`
flag commands. Microsoft Agent Framework shows both
`microsoft-agent-framework-dotnet` and `microsoft-agent-framework-py`.
- [ ] All 8 quickstarts: imports show `@copilotkit/react-core` (root),
styles from `@copilotkit/react-core/styles.css`. No `/v2` paths.
- [ ] Inspect shell-docs header at desktop and mobile (narrow viewport).
No AG-UI brand tab anywhere.
- [ ] Visit a built-in-agent voice page that consumes the
`sample-audio-button` region. The snippet renders.
## Summary
- Bump `@copilotkit/license-verifier` from `0.2.0` → `0.4.0` in
`packages/runtime`, `packages/shared`, and the root pnpm `overrides`
block
- Refresh `pnpm-lock.yaml` accordingly
## Test plan
- [ ] CI green
Resolved d5-all.json conflict by appending B6's headless-simple/complete fixtures (Say hello, joke, fun fact, Highlight, chart) before the tool-rendering and open-gen-ui entries already on integration.
Rewrites 10 D5 smoke probes from keyword-match-on-transcript checks
to side-effect / DOM-mount / network-payload assertions per
Phase-2B of .claude/specs/lgp-test-genuine-pass.md:
- d5-frontend-tools: 3 per-pill turns (Sunset/Forest/Cosmic) asserting
the page background data-attribute mutated to the pill's family-
specific gradient and differs from the prior pill's gradient.
- d5-frontend-tools-async: notes-card testid mount + settled-state
detection (list items OR empty-state copy).
- d5-agent-config: 3 knob pairs (tone/expertise/responseLength) sent
as 6 turns; assertion compares value-A vs value-B response text or
character-count delta. Aimock fixtures key on prompt-encoded knob
tokens since JSON fixtures don't expose context-keyed predicates.
- d5-gen-ui-agent: 3 per-pill turns asserting agent-state-card mounts
with ≥ 2 agent-step rows whose text fingerprint differs across
pills (catches a regression where fixtures stop differentiating).
- d5-gen-ui-declarative: 4 per-pill turns asserting at least one
catalog testid (declarative-card / metric / status-badge /
pie-chart / bar-chart) mounted.
- d5-gen-ui-a2ui-fixed: pill prompt + a2ui-fixed-card testid mount
assertion.
- d5-gen-ui-interrupt: 2 per-pill turns; click first time-picker-slot
after time-picker-card mounts; assert time-picker-picked appears.
- d5-gen-ui-open: pill prompt; iframe[srcdoc] mounts with srcdoc
length ≥ 100 chars.
- d5-shared-state-streaming: 3 per-pill turns asserting
document-content text length ≥ STREAMING_MIN_FINAL_CHARS.
- d5-readonly-state-context: seeds [data-testid=ctx-name] with a
sentinel; installs page.route() interceptor on /api/copilotkit;
asserts the captured outgoing request body contains the sentinel.
Adds _genuine-shared.ts with the structural Page extension
(GenuinePage), pill-click helper, testid waiter, and runtime guard
that mirrors the _beautiful-chat-shared.ts pattern.
All 10 *.test.ts files updated with focused assertions over fakes —
none rely on the production page being live, all 63 unit tests pass.
Adds per-pill aimock fixtures backing the Phase-2B genuine D5 probes:
- agent-config: 6 fixtures (3 knob pairs) so concise vs detailed
responses differ deterministically; satisfies the new probe's
text-diff + length-diff assertions.
- frontend-tools: 3 per-pill fixtures (sunset/forest/cosmic) emitting
change_background tool calls with family-specific gradient hexes.
- frontend-tools-async: query_notes tool call so the NotesCard mounts.
- gen-ui-agent: 3 per-pill set_steps tool calls with distinct step
content so the per-pill content-fingerprint assertion catches
fixture-drift.
- gen-ui-declarative: 4 per-pill generate_a2ui calls + render_a2ui
fixture for the secondary LLM call that paints the catalog.
- gen-ui-a2ui-fixed: SFO/JFK display_flight tool call.
- gen-ui-interrupt: 2 per-pill schedule_meeting tool calls with
distinct time-slot payloads.
- gen-ui-open: generateSandboxedUi tool call with a non-trivial
HTML payload so the iframe[srcdoc]-mount assertion has ≥ 100 chars
to observe.
- shared-state-streaming: 3 per-pill write_document tool calls with
substantive content payloads (≥ 100 chars each).
- readonly-state-context: pill-prompt fixture; the probe's network-
payload assertion checks the request body, not the response.
d5-all.json is updated with the new entries prepended so first-match
precedence routes specific pill prompts to their per-pill fixtures
ahead of the generic adjacent matches.