The shell embeds each demo in a cross-origin iframe whose `allow`
attribute only granted clipboard access. Browsers block
`getUserMedia({ audio: true })` at the Permissions Policy layer in
cross-origin frames unless the parent grants `microphone` via `allow`,
so every voice demo across every integration threw "Microphone
permission denied" before any user prompt was shown.
Add `microphone` to the iframe `allow` in all three places that embed
demo previews — the per-demo viewer, the standalone preview route, and
the demo drawer — so voice demos work uniformly across all 18
integrations. No other demo type uses getUserMedia / getDisplayMedia
/ geolocation, so no other Permissions Policy features are needed.
## Summary
Fixes D5 voice test across all 18 showcase integrations. Voice was
previously untested at the D5 level.
### What was fixed
- **V1 to V2 voice routes**: 9 integrations had voice routes using V1
CopilotRuntime (single-route mode rejects multipart/form-data with 415).
Ported to V2 createCopilotRuntimeHandler (multi-route).
- **Tool-free voice agents**: strands, llamaindex, ms-agent-python had
agents with tools registered -- aimock returned tool calls that the
adapter didn't loop on. Added dedicated /voice endpoints with tools=[].
- **Langgraph-typescript**: Added sample_agent to langgraph.json (was
only in dev-mode config, not the graphSpec used at runtime).
- **Google ADK**: Added SampleAudioButton, voice route, and sample.wav.
- **Missing assets**: Added sample.wav to agno, ms-agent-python,
ms-agent-dotnet.
- **Aimock fixtures**: Mastra get-weather hyphen variant (Mastra
registers as get-weather, not get_weather).
- **Infra**: OPS_BASE_URL dashboard build arg for local docker-compose.
### Locally verified
All 18 integrations pass D5 voice test locally.
## Test plan
- [x] 18/18 integrations voice D5 green locally
- [ ] CI green
- [ ] Production harness D5 probe green after deploy
The validator was not considering the aimock match.toolName field when
checking for drift. When a fixture has toolName set, aimock only fires
it for agents that register that tool -- skip demos that don't have it.
The dashboard needs OPS_BASE_URL at build time to connect to the local
harness. Without this, the dashboard container builds with no API URL
and cannot display probe results locally.
- Add dedicated tool-free voice agents for strands, llamaindex,
ms-agent-python (aimock returns tool calls when tools are registered,
which the adapters don't loop on)
- Add sample_agent alias to langgraph-typescript langgraph.json
(was only in dev-mode config)
- Add SampleAudioButton and voice route to google-adk
- Add sample.wav to agno, ms-agent-dotnet, ms-agent-python, google-adk
V1 CopilotRuntime in single-route mode rejects multipart/form-data
with 415 Unsupported Media Type. Port all 9 integrations to V2
createCopilotRuntimeHandler which handles the /voice sub-route
natively.
Integrations: claude-sdk-python, claude-sdk-typescript, crewai-crews,
llamaindex, ms-agent-dotnet, ms-agent-python, pydantic-ai, spring-ai,
strands
When a Railway service deployed within the last 120 seconds, the D5
probe driver now skips all features for that service with green side
rows (note: "skipped: deploy in progress (Ns ago)") instead of
launching a browser and producing false reds from deploy churn.
The skip fires before the script loader and browser launch so recently
deployed services cost zero probe resources.
Discovery source: added deployedAt (latestDeployment.createdAt) to
RailwayServiceInfo and the GraphQL query.
Driver: added deployedAt to the input schema and DEPLOY_CHURN_GRACE_MS
constant (120_000ms). When deployedAt is within the grace window, all
features short-circuit to green with a descriptive note.
Tests: 8 new tests covering the skip path, boundary conditions
(exactly at grace window), backwards compat (absent/empty/unparseable
deployedAt), and the 0s-age edge case.
## Summary
- Never mutate originals -- temp overlay in $TMPDIR instead of in-place
.iso-bak
- Slot-based port allocation via atomic mkdir for parallel support (slot
0 = +200, slot 1 = +400, etc.)
- --project-name scoping prevents Docker container/network collisions
- Stale slot cleanup via PID check + 2-hour age fallback
- LOCAL_PORTS_FILE env var for TS harness to read offset ports
## Test plan
- [ ] Originals untouched after isolated run (md5 verified)
- [ ] CI green
- [ ] Two parallel --isolate runs don't conflict
The COMPOSE_CMD in apply_isolation was missing --project-name, so Docker
Compose would infer the project name from the directory and collide with
the base showcase stack (and other isolated runs). Adding --project-name
ensures containers, networks, and volumes are fully scoped to the
isolation slot.
apply_isolation previously mutated docker-compose.local.yml and
local-ports.json in-place with .iso-bak backups. If the process crashed
the originals stayed corrupted with +200 port offsets, breaking all
subsequent showcase commands.
Now writes modified copies to a temp directory and overrides
COMPOSE_FILE/PORTS_FILE shell variables so downstream code reads from
the overlay. Originals are never touched. restore_isolation just removes
the temp dir.
Also replaces hardcoded +200 port offset with atomic mkdir-based slot
allocation. Two parallel --isolate runs now get different port ranges
(slot 0 = +200, slot 1 = +400, etc.) instead of colliding on the same
ports. Container names include the slot number for collision-free Docker
naming. Stale slots from crashed runs are reclaimed via PID liveness
checks and a 2-hour age fallback.
TS harness files (config.ts, lifecycle.ts, doctor.ts) honor
LOCAL_PORTS_FILE env var so they read offset ports from the temp overlay.
D3/D4 was blue (accent), now amber/yellow to signal "not yet D5".
D1/D2 was amber, now red to signal "needs work".
Removed D6 references — D6 no longer exists.
## Summary
Both spring-ai and mastra tool-rendering D5 failures had the same root
cause: aimock couldn't match `toolName: "get_weather"` because the tool
name wasn't present in the OpenAI API request.
**spring-ai:** `StreamingToolAgent.streamFirstTurn()` omitted tool
definitions from the streaming request. Added `toolCallbacks` with
`internalToolExecutionEnabled=false` so Spring AI sends tool schemas
without auto-executing.
**mastra:** JS object shorthand `{ weatherTool }` expanded to function
name `"weatherTool"`, not `"get_weather"`. Used explicit keys matching
expected names.
## Test plan
- [x] spring-ai: 31/32 D5 green locally (voice is separate workstream)
- [x] mastra: 32/32 D5 green locally
spring-ai: StreamingToolAgent Phase 1 streaming was not including tool
definitions in the OpenAI API request. Without tool schemas, aimock
could not match the fixture requiring toolName:"get_weather" and fell
through to a generic text-only fixture. The tool call was never
emitted, so useRenderTool never triggered and the weather card never
rendered.
Attach toolCallbacks to the Phase 1 streaming request while keeping
internalToolExecutionEnabled=false. The LLM/aimock sees the tool
schemas and can return tool_calls, but Spring AI does not auto-execute
them. Phase 2 handles execution as before.
mastra: Mastra Agent class uses the object KEY in the tools config as
the OpenAI function name, not the createTool id. The weatherAgent
registered tools as { weatherTool, queryDataTool, ... } which sent
function names "weatherTool", "queryDataTool" etc. to aimock. The D5
fixture expects toolName:"get_weather", so it never matched. Also
aligned the createTool id from "get-weather" to "get_weather".
Use explicit keys matching the expected function names:
{ get_weather: weatherTool, query_data: queryDataTool, ... }
Verified locally: spring-ai 31/32 green (voice pre-existing), mastra
32/32 fully green.
## Summary
When the harness process dies mid-probe (Railway redeploy, OOM),
PocketBase rows stay in `state=running` forever. These zombies block the
API run list, show stale partial results, and paint the dashboard blue.
`sweepStaleRuns()` marks any `running` row older than 15 minutes as
`failed`. Called once at orchestrator boot.
Also includes the D5 flap diagnostics instrumentation (warn-level
logging on every failure — hydration timing, page state, console errors,
request failures).
## Test plan
- [x] `npx tsc --noEmit` clean
- [ ] On next harness redeploy, zombie runs should be swept
When the harness process dies mid-probe (Railway redeploy, OOM),
PocketBase rows stay in state=running forever. These zombies block
the API's run list, show stale results, and confuse the dashboard.
sweepStaleRuns() marks any running row older than 15 minutes as
failed. Called once at orchestrator boot.
## Summary
Two changes:
1. **Fix referenceCount ReferenceError** — PR #4562 removed the
`referenceCount` variable but left a log line using it, crashing
`generate-registry.ts` at build time. The dashboard never rebuilt with
langgraph-python as REF because the build crashed silently. Fixed the
log to use `referenceWiredFeatures.size`.
2. **Add D5 flap diagnostics** — warn-level logging on every probe
failure capturing hydration timing, page body text, console errors,
request failures, and error boundary state. Zero functional changes.
This data identifies whether production flaps are from hydration
timeouts, error boundaries, aimock mismatches, or network failures.
Locally verified: `generate-registry.ts` outputs `reference integration
= langgraph-python (39 wired features)`.
## Test plan
- [x] `npx tsx showcase/scripts/generate-registry.ts` → `reference =
langgraph-python`
- [x] `catalog.json` shows `"reference": "langgraph-python"` in both
shell + shell-dashboard
- [x] `npx tsc --noEmit` clean on harness
- [ ] Dashboard shows langgraph-python as REF after deploy
PR #4562 removed the referenceCount variable but left a log line
referencing it, causing ReferenceError at build time. The dashboard
never rebuilt with langgraph-python as REF because the build crashed.
## Summary
The reference integration was auto-detected as whichever had the most
wired features, ties broken alphabetically. This caused ag2 to appear as
the reference when it matched langgraph-python's feature count.
langgraph-python is always the gold standard reference. Pinned
explicitly.
## Test plan
- [ ] Dashboard shows langgraph-python as REF (first column)
The auto-detect logic picked whichever integration had the most wired
features, with ties broken alphabetically. This caused ag2 to appear
as the reference when it matched langgraph-python's feature count.
langgraph-python is always the gold standard reference.
## Summary
Adds a D5 voice test that exercises the voice transcription flow via the
sample audio button. Deterministic: aimock intercepts the Whisper
transcription call and returns a canned response.
### What's new
- **D5 voice script** (`d5-voice.ts`) — clicks sample audio button,
waits for transcription to fill textarea, sends, asserts weather
response
- **skipFill in conversation runner** — new `skipFill: true` option on
ConversationTurn that skips `page.fill()` when preFill already populated
the textarea (9 new tests)
- **Voice in D5 registry** — `voice` added to D5FeatureType enum +
REGISTRY_TO_D5 mapping
- **Aimock transcription fixture** — returns "What is the weather in
Tokyo?" for any `/v1/audio/transcriptions` request
- **Voice runtime docs** — confirmed OPENAI_BASE_URL routes Whisper
calls through aimock automatically
### How it works
1. preFill clicks `[data-testid="voice-sample-audio-button"]`
2. Aimock returns canned transcription → textarea fills with "What is
the weather in Tokyo?"
3. skipFill sends without overwriting → aimock handles the chat
completion
4. Assertion checks for weather/Tokyo in the assistant response
### Test plan
- [x] 1490 harness tests pass (91 files)
- [ ] CI green
- [ ] Local `showcase test langgraph-python --d5` with voice feature
Add D5 voice test that exercises sample-audio transcription via aimock.
Infrastructure: voice in D5 feature type registry + mapping, skipFill
support in conversation runner (9 new tests), inputValue forwarding
in e2e-deep Page wrappers, aimock transcription fixture, tool-free
weather fallback fixture for agents without tools. Verified locally:
D5 suite passes green on langgraph-python (60.4s).
## Summary
Adds Strategy 9 to DEBUGGING.md: pulling PocketBase D5 history to
distinguish real bugs from production flapping. Includes ready-to-copy
curl+python commands for error categorization, deploy correlation, and
per-service flapping rate analysis.
## Test plan
- [x] Documentation only — no code changes
Strategy 9: pull all D5 status records from PocketBase in one request,
categorize errors, cross-reference with deploy history, and compute
per-service flapping rates to distinguish real bugs from transient
production issues.
## Summary
The D5 probe's chat input selector cascade gave each candidate only 2s
to appear (`SELECTOR_PROBE_TIMEOUT_MS = 2_000`). Under Railway
production load, React hydration takes longer than 2s, causing the probe
to miss the input and mark the feature red. This was the root cause of
production D5 flapping — all 18 integrations pass locally but only
2-3/18 were stable green in production.
All 7 currently-red PocketBase D5 records showed the same error:
`page.fill: Timeout 30000ms exceeded — waiting for locator
'[data-testid=...]'`.
One-line fix: `SELECTOR_PROBE_TIMEOUT_MS` from `2_000` to `5_000`.
## Test plan
- [x] langgraph-python 31/31 D5 green locally after change
- [ ] Production D5 stable green rate improves from ~15% to ~90%+ over
next few probe cycles
The chat input selector cascade gave each candidate only 2s to appear.
Under Railway production load, React hydration takes longer than 2s,
causing the probe to miss the input and mark the feature red. This was
the root cause of D5 flapping — all 18 integrations pass locally but
only 2-3 are stable green in production.
All 7 currently-red PocketBase D5 records showed the same error:
page.fill timeout waiting for the chat input selector.
Mechanical cleanup of the perf commit's three repeated patterns. No
behavioral change.
- Three `_xxxTplCache` fields collapse into a single
`_panelTplCache: Map<ThreadDetailsTab, { key, tpl }>` with a shared
`cachedPanelTpl(slot, key, build)` helper. Cache key is now a tuple
compared element-wise by reference, so each panel passes everything
the template depends on (conversation passes
`[_conversation, _expandedTools, _expandedMessages]`) without
duplicating the cache-check shape three times.
- Three sibling tab-content `<div>` blocks in render() collapse into one
`TAB_LIST.map(...)` driven by a new `renderTabContent(id)` dispatcher.
- Tab-button click handler extracts to `activateTab(id)` plus
`maybeFetchTabData(id)`, keeping the rAF-and-spinner dance and the
lazy-fetch decision off the inline lambda.
- Two-rAF first-activation collapses to a single rAF: Lit batches the
`_activatedTabs` add and the `_panelInitializing = false` clear into
one update, so the second rAF was redundant.
- `highlightedJsonImpl` was a pass-through layer split out only to host
the WeakMap memo; inline it back into `highlightedJson`.
- `ReturnType<typeof html>` swaps to the canonical `TemplateResult`
type imported from lit.
- Class renames `CpkThreadDetails` → `ɵCpkThreadDetails` (already
exported for tests; the prefix keeps the internal-API hint).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three follow-ups to the perf commit:
1. The conversation TemplateResult cache was keyed only on `_conversation`,
so toggling a tool-call expand or "Show more" on a long message — both
of which mutate `_expandedTools` / `_expandedMessages` without touching
the conversation array — returned the pre-toggle template and the
disclosure appeared broken. Widen the cache key to include both expand
sets; production toggles already replace the Set instance, so reference
equality flips correctly.
2. Wrapping each tab in a panel <div> for the keep-mounted approach broke
the `gap` flow that previously cascaded from `.cpk-td__content > *`.
Add a `.cpk-td__panel` class with `display:flex; flex-direction:column;
gap:12px` so conversation items and event rows have breathing room.
3. Export the `ɵCpkThreadDetails` class so unit tests can pin down the
per-panel cache-invalidation contract. Add four tests in
web-inspector.spec.ts covering: threadId change drops all three caches;
conversation cache invalidates on `_conversation` reassignment;
conversation cache invalidates on expand-state change (regression
guard); state and events caches invalidate on their fetched data
reassignment.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Three layered fixes for the tab-switch jank on threads with many AG-UI
events. Symptoms: clicking back to a previously-opened tab took roughly
a second per switch on a thread with several hundred recorded events,
even though the underlying data was already cached.
1. Keep activated tab panels mounted. The render conditional swapped
between renderConversation/renderState/renderEvents based on `_tab`,
so Lit tore down the previous panel's DOM and rebuilt the next one
from scratch on every switch. Now once a tab is activated, its panel
stays mounted and inactive panels are hidden via `display:none`.
Activated set resets on threadId change.
2. Memoize per-panel TemplateResults by data reference. Even with the
panel mounted, render() still re-evaluated the template on every
parent update, allocating fresh nested TemplateResults for every
event row. Each render now returns the cached TemplateResult when
`_conversation` / `_fetchedState` / events array references haven't
changed; Lit then short-circuits the entire diff.
3. Defer layout for off-screen events with `content-visibility: auto`
plus a `contain-intrinsic-size` hint. The cached-data switch back to
the events panel still triggered a full layout pass over every
recorded event, which on a 600-event thread shows up as a seconds-
long freeze when the panel becomes visible. The browser now skips
layout/paint for off-screen rows entirely.
Also adds a WeakMap memo around `highlightedJson` so identical event
payloads don't re-run JSON.stringify + the syntax-highlight regex pass
on every render.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Frontend-rendered generative-UI tools (charts, custom UI) never produce
a `role: tool` result message because they execute client-side, so the
prior `item.result ? DONE : PENDING` rule rendered them as PENDING
forever even after the run finished and the chart was on screen.
The args block being populated is itself the resolution signal for these
tools, so flip the condition: any tool call with parsed arguments shows
DONE. The badge stays PENDING only for the brief window where args have
not yet streamed in.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
## Summary
Two multimodal D5 fixes, both locally verified green.
**ms-agent-python (31/31 green):** `_MultimodalAgent.run()` override
used `*args/**kwargs` but `AgentFrameworkAgent.run()` expects
`input_data: dict`. Changed to match the base class signature and yield
events from the base generator.
**langroid (31/31 green):** `_normalize_part()` checked
`isinstance(part, dict)` but Pydantic deserializes AG-UI content parts
as model instances, not dicts. All multimodal parts were silently
dropped. Added Pydantic model detection via `model_dump()`.
## Test plan
- [x] ms-agent-python: 31/31 D5 green locally
- [x] langroid: 31/31 D5 green locally (both image + PDF turns pass)
The _normalize_part function in multimodal_agent.py checked
isinstance(part, dict) to gate all content-part processing.
When RunAgentInput is deserialized via Pydantic, ag_ui.core
types (TextInputContent, ImageInputContent, DocumentInputContent)
are model instances — not dicts — so every multimodal content
part was silently dropped.
This caused the D5 multimodal probe's PDF turn to fail: the user
message text was lost, the aimock matched the stale turn-1 fixture
instead of the turn-2 fixture, and the assertion saw the image
response where it expected the document response.
Convert Pydantic models to dicts via model_dump(by_alias=True)
before processing, preserving camelCase field names (mimeType)
that the rest of the function relies on.
## Summary
The `_MultimodalAgent.run()` override used `*args/**kwargs` but
`AgentFrameworkAgent.run()` expects `input_data: dict`. The mismatch
caused `TypeError: got an unexpected keyword argument 'messages'` at
runtime, making the multimodal endpoint return RUN_ERROR on every
request.
Fixed to match the base class signature and yield events from the base
generator.
Locally verified: ms-agent-python 31/31 D5 green.
## Test plan
- [x] ms-agent-python passes 31/31 D5 features locally
- [ ] CI passes
The _MultimodalAgent.run() override used *args/**kwargs but
AgentFrameworkAgent.run() expects input_data: dict. The mismatch
caused TypeError at runtime. Changed to match the base signature
and yield events from the base generator.
## Summary
- Switch `gen-ui-interrupt` and `interrupt-headless` demo pages in
built-in-agent from V1 `CopilotKit` (with named `agent=` prop) to V2
`CopilotKitProvider` with `useSingleEndpoint`
- Remove `agentId` prop from `CopilotChat` in both pages
## Why
The built-in-agent runtime (`/api/copilotkit/route.ts`) uses
CopilotRuntime V2 with `mode: "single-route"` and only registers a
single `default` agent. The two interrupt pages were using V1
`CopilotKit` with `agent="gen-ui-interrupt"` and
`agent="interrupt-headless"` respectively, which caused the runtime to
return agent-not-found errors since those named agents don't exist in
the V2 registry.
All other built-in-agent demo pages already use V2 `CopilotKitProvider`
with `useSingleEndpoint`. This PR aligns the two interrupt pages with
that pattern.
## Test plan
- [ ] D5 probe for `gen-ui-interrupt` on built-in-agent turns green
- [ ] D5 probe for `interrupt-headless` on built-in-agent turns green
- [ ] Other built-in-agent D5 features remain unaffected
gen-ui-interrupt and interrupt-headless pages were using V1 CopilotKit
with named agents (agent="gen-ui-interrupt", agent="interrupt-headless")
but the built-in-agent runtime only registers a single "default" agent
via CopilotRuntime V2. This caused 404/agent-not-found errors at D5.
Switch both pages to CopilotKitProvider with useSingleEndpoint and
remove agentId from CopilotChat, matching the pattern used by all
other built-in-agent demo pages.
## Summary
- Routes PostHog analytics through `docs.copilotkit.ai/ingest/*`
(rewritten to `eu.i.posthog.com`) so requests bypass ad blockers and
tracking-protection that target the PostHog hostname directly.
- Mirrors the existing setup on the marketing website (`../website`).
- Drops the `NEXT_PUBLIC_POSTHOG_HOST` dependency — host is now
hardcoded since it's tied to the proxy path.
## Changes
- `docs/next.config.mjs` — added two `beforeFiles` rewrites:
`/ingest/static/:path*` → `eu-assets.i.posthog.com/static/*`,
`/ingest/:path*` → `eu.i.posthog.com/*`
- `docs/lib/providers/posthog-provider.tsx` — `api_host: '/ingest'`,
`ui_host: 'https://eu.posthog.com'`, removed `POSTHOG_HOST` env-var
check
- `docs/middleware.ts` — excluded `ingest` from the redirect-middleware
matcher so proxy traffic skips it
## Test plan
- [ ] Deploy preview — confirm DevTools shows PostHog requests going to
`docs.copilotkit.ai/ingest/*` instead of `eu.i.posthog.com`
- [ ] Verify `$pageview` events appear in the PostHog dashboard (EU
project)
- [ ] Test with uBlock Origin / Brave Shields enabled — events should
still capture
- [ ] Verify session recordings still load (uses `/ingest/static/*` for
assets)
🤖 Generated with [Claude Code](https://claude.com/claude-code)