Commit Graph

8565 Commits

Author SHA1 Message Date
Benjamin Taylor 789fc98bdb fix(docs): reverse-proxy PostHog through /ingest
Routes PostHog analytics through docs.copilotkit.ai/ingest/* (rewrites to
eu.i.posthog.com) so requests bypass ad blockers and tracking-protection
that target the PostHog hostname directly. Mirrors the existing setup in
the marketing website.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 20:08:40 -05:00
github-actions[bot] 56524a7282 style: auto-fix formatting 2026-05-01 01:08:19 +00:00
Martha Schumann a258e8c7c7 revert: drop throwaway langgraph-python-threads workspace wiring
Reverts the example-app integration that was used for local end-to-end
testing of the inspector against an intelligence-backed runtime. The
example app is not part of the pnpm workspace, ships its own
package-lock.json, and pulls @copilotkit/* from the npm registry.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 18:05:27 -07:00
Martha Schumann 0721414fbb Merge remote-tracking branch 'origin/main' into feat/CPK-7193-inspector-threads-clean
# Conflicts:
#	examples/integrations/langgraph-python-threads/apps/app/package.json
#	examples/integrations/langgraph-python-threads/apps/bff/package.json
#	examples/integrations/langgraph-python-threads/package-lock.json
#	pnpm-lock.yaml
2026-04-30 18:01:51 -07:00
Jordan Ritter fbe882eca0 fix(showcase): restore langgraph-python declarative-gen-ui to a2ui_dynamic graph (#4553)
## Summary

PR #4542 incorrectly changed the langgraph-python declarative-gen-ui
route from `graphId: "a2ui_dynamic"` to `graphId: "sample_agent"` and
removed `injectA2UITool: false`. This caused a regression from 31/31
green to red.

The `a2ui_dynamic` graph owns the `generate_a2ui` tool itself — the
runtime must NOT auto-inject its own A2UI tool on top (`injectA2UITool:
false`). The `sample_agent` graph is a generic chat agent with no tools,
which can never produce A2UI surfaces.

## Test plan

- [ ] langgraph-python returns to D5 green (31/31)
2026-04-30 17:51:46 -07:00
Jordan Ritter 738a85cfe9 fix(showcase): restore langgraph-python declarative-gen-ui to a2ui_dynamic graph
PR #4542 incorrectly changed graphId from "a2ui_dynamic" to
"sample_agent" and removed injectA2UITool: false. The a2ui_dynamic
graph owns the generate_a2ui tool itself — the runtime must NOT
auto-inject. This caused langgraph-python to regress from 31/31 to red.
2026-04-30 17:50:23 -07:00
Jordan Ritter fc854c0436 fix: use custom stream converter for BYOC agents to fix D5 timeout (#4552)
## Summary

- Switch byoc-hashbrown and byoc-json-render agent factories from `type:
"tanstack"` to `type: "custom"` with a dedicated stream converter,
matching the proven pattern used by the main built-in-agent factory
(`tanstack-factory.ts`)
- Remove invalid `response_format: { type: "json_object" }` from
`modelOptions` -- TanStack AI v0.8.x uses the OpenAI Responses API which
does not support this Chat Completions parameter
- System prompts already enforce JSON-only output; `temperature: 0.2` is
retained as a valid Responses API parameter

The `type: "tanstack"` path routes through the runtime's
`convertTanStackStream` which has a `runFinished` flag (PR #4476) that
blocks all events after the first `RUN_FINISHED`. Combined with the
invalid `response_format` parameter being silently rejected by the
Responses API, this prevented text events from reaching the frontend,
causing the D5 probe to timeout waiting for
`[data-testid="copilot-assistant-message"]`.

## Test plan

- [ ] Verify byoc-hashbrown D5 passes locally with `bin/showcase test
built-in-agent --d5`
- [ ] Verify all other 27 built-in-agent features still pass D5
- [ ] Verify byoc-json-render D5 passes (same fix pattern)
2026-04-30 17:35:46 -07:00
Jordan Ritter 0c270e5eb1 fix: use custom stream converter for BYOC agents to fix D5 timeout
The byoc-hashbrown and byoc-json-render agents used `type: "tanstack"`
which routes through the runtime's `convertTanStackStream`. That
converter has a `runFinished` flag (PR #4476) that blocks all events
after the first RUN_FINISHED, which can prevent text events from
reaching the frontend.

Additionally, both agents passed `response_format: { type: "json_object" }`
via `modelOptions`. TanStack AI's OpenAI adapter v0.8.x uses the
Responses API (`client.responses.create()`), not Chat Completions. The
Responses API does not support `response_format` (it uses `text.format`
instead), so this parameter was silently causing failures.

Fix both issues by:
- Switching from `type: "tanstack"` to `type: "custom"` with a dedicated
  stream converter that skips RUN_FINISHED and forwards text events,
  matching the proven pattern in tanstack-factory.ts
- Removing the invalid `response_format` from modelOptions (the system
  prompt already enforces JSON-only output)
- Keeping `temperature: 0.2` which is valid for the Responses API
2026-04-30 17:33:58 -07:00
Jordan Ritter 3ec233fd26 fix(showcase): fix 3 llamaindex D5 timeouts caused by deferred annotations (#4551)
## Summary

- Remove `from __future__ import annotations` from 3 llamaindex agents
that were timing out at D5
- The deferred annotations import causes Pydantic to fail with
"class-not-fully-defined" when resolving `Annotated[str, "..."]` tool
parameters, producing a `RUN_ERROR` SSE event instead of streaming text
- The D5 conversation runner cannot detect `RUN_ERROR` as an assistant
response, so it times out after 30s

## Affected agents

- `a2ui_fixed.py` (gen-ui-a2ui-fixed / a2ui-fixed-schema feature)
- `a2ui_dynamic.py` (gen-ui-declarative / declarative-gen-ui feature)
- `tool_rendering_reasoning_chain_agent.py`
(tool-rendering-reasoning-chain feature)

## Root cause

`from __future__ import annotations` (PEP 563) defers evaluation of all
annotations to strings. When the LlamaIndex `AGUIChatWorkflow` passes
`backend_tools` to Pydantic for schema generation, `Annotated[str,
"Origin airport code"]` is stored as the string `"Annotated[str, 'Origin
airport code']"` instead of the actual type. Pydantic cannot resolve
this and raises `PydanticUserError: class-not-fully-defined`.

Agents without `backend_tools` (like `reasoning_agent.py`) were
unaffected because the code path that triggers Pydantic schema
validation is never reached with an empty tool list.

## Test plan

- [x] Built Docker image locally and curled all 3 endpoints directly
- [x] Before fix: all 3 returned `RUN_ERROR` with Pydantic
class-not-fully-defined
- [x] After fix: all 3 stream `TEXT_MESSAGE_CHUNK` events and finish
with `RUN_FINISHED`
- [ ] CI green
- [ ] D5 probe passes for all 3 features on production
2026-04-30 17:26:57 -07:00
Jordan Ritter 3dc802a927 fix(showcase): remove from __future__ import annotations from 3 llamaindex agents
The `from __future__ import annotations` import causes all type
annotations to be stored as strings rather than evaluated at
definition time. When the LlamaIndex AGUIChatWorkflow validates
backend_tools via Pydantic, `Annotated[str, "..."]` parameters
fail with "class-not-fully-defined" because Pydantic cannot
resolve the deferred string annotations.

This caused RUN_ERROR on every request to these three agents,
which the D5 conversation runner cannot detect as a response,
leading to the 30s timeout.

Affected agents:
- tool_rendering_reasoning_chain_agent.py (4 backend tools)
- a2ui_fixed.py (display_flight backend tool)
- a2ui_dynamic.py (generate_a2ui backend tool)

Verified locally: all three endpoints now stream
TEXT_MESSAGE_CHUNK events and finish with RUN_FINISHED.
2026-04-30 17:24:59 -07:00
Jordan Ritter 9cec7d7b70 fix(showcase): render CLI Start command in docs-only cells (#4550)
## Summary

The docs-only early return in `ComposedCellInner` only rendered
`DocsLayer`, skipping `LinksLayer` which renders `CommandCell` for demos
with a `command` field. The `npx copilotkit@latest init` starter command
disappeared from the matrix.

Fix: render `LinksLayer` alongside `DocsLayer` when links overlay is
active.

## Test plan

- [ ] CLI Start Command row shows the `npx` command text (not just docs
links)
- [ ] Other docs-only features unaffected
2026-04-30 17:20:34 -07:00
Jordan Ritter 0d41702f68 fix(showcase): wait for auth banner state before post-signout probe send (#4549)
## Summary

- Fix race condition in the D5 auth probe where the post-signout chat
message was sent before React's `useEffect` flushed `setHeaders()`,
causing the request to go out with valid auth headers (no 401, no error
surface, 8s timeout)
- After clicking sign-out, the probe now waits for the auth banner to
flip to `data-authenticated="false"` plus a 500ms settle delay before
triggering the probe send
- Add unit test covering the new "banner doesn't flip" failure mode

## Test plan

- [x] All 7 d5-auth unit tests pass (including new banner-flip test)
- [x] TypeScript compiles clean
- [x] Verified all 18 integrations already have
`data-testid="auth-banner"` with `data-authenticated` attribute
- [ ] CI green
2026-04-30 17:19:39 -07:00
Jordan Ritter 7d7386edc6 fix(showcase): add missing module stubs in crewai-crews forwarded props tests (#4548)
## Summary

- Add missing `agents.interrupt_crew` and `agents.tool_rendering` stubs
to the `_stub_agent_server_deps` fixture in `test_forwarded_props.py`
- The interrupt_crew import was added to `agent_server.py` in d6b784ee9
(interrupt demos PR) but the test fixture was not updated, causing
`ModuleNotFoundError` on all 5 tests that `import agent_server`

## Test plan

- [x] All 8 tests in `test_forwarded_props.py` pass locally (Python 3.9)
- [ ] CI: Python unit tests (3.10) pass
- [ ] CI: Python unit tests (3.12) pass
2026-04-30 17:19:36 -07:00
Jordan Ritter 59f8a54c80 fix(showcase): render CLI Start command in docs-only cells
The docs-only early return only rendered DocsLayer, skipping LinksLayer
which is responsible for rendering CommandCell when a demo has a
command field. The npx starter command disappeared from the matrix.
2026-04-30 17:18:39 -07:00
Jordan Ritter 38d1f1a09f fix(showcase): wait for auth banner state before post-signout probe send
After clicking sign-out, React's useEffect that calls setHeaders() runs
async (after paint). The probe was immediately filling the textarea and
sending a message before useEffect flushed, so the request went out with
valid auth headers — no 401 occurred, no error surface rendered, and the
probe timed out at 8s.

Add a waitForSelector gate on the auth banner's data-authenticated="false"
attribute plus a 500ms settle delay to ensure setHeaders() has flushed
before the probe triggers the post-signout chat send.
2026-04-30 17:16:44 -07:00
Jordan Ritter 3cd2520c5a fix(showcase): add missing module stubs in crewai-crews forwarded props tests
The _stub_agent_server_deps fixture was missing stubs for
agents.interrupt_crew and agents.tool_rendering, which were added to
agent_server.py in d6b784ee9 (interrupt demos). The missing
interrupt_crew stub caused ModuleNotFoundError on import, failing 5
tests on both Python 3.10 and 3.12 in the Showcase: Validate workflow.
2026-04-30 17:15:53 -07:00
Jordan Ritter aa76994f5f fix(showcase): get claude-sdk-python D5 to all green (#4541)
## Summary

- Add missing `data-testid="copilot-assistant-message"` and
`data-message-role="assistant"` to claude-sdk-python's BYOC hashbrown
renderer, matching langgraph-python and claude-sdk-typescript
- Fix D5 auth probe to use JS-level `element.click()` instead of
Playwright `force: true`, which `<cpk-web-inspector>` silently
intercepts

## Details

**byoc**: The custom `AssistantMessageRenderer` in the hashbrown
renderer was missing the harness testid attributes. Without them the
conversation runner sees 0 assistant messages and times out at 30s.

**auth**: Playwright's `page.click(sel, { force: true })` dispatches
pointer events that the `<cpk-web-inspector>` overlay absorbs before
React's synthetic event system fires the `onClick`. Switching to
`page.evaluate(() => document.querySelector(sel).click())` triggers the
DOM click directly, reliably flipping the auth state so the 401 error
surface appears.

## Test plan

- [x] `showcase test claude-sdk-python --d5 --verbose` — 29/29 green on
two consecutive runs
- [x] No changes to demo functionality — both fixes are test
infrastructure / harness plumbing
2026-04-30 17:06:47 -07:00
Jordan Ritter 1e97f5ecb7 fix(showcase): D5 probe fixes across 16 integrations (#4542)
## Summary

Fixes D5 (deep conversation) probe failures across 16 showcase
integrations, addressing 39 of 42 failing features. After this PR, we
expect 15-16/18 integrations at D5 green (up from 2/18).

**Root causes identified and fixed:**

- **byoc-hashbrown testid** (12 integrations): The custom HashBrown
renderer overrides CopilotChat's `assistantMessage` slot, dropping the
`data-testid="copilot-assistant-message"` attribute that the D5
conversation runner uses to count responses
- **langgraph-typescript server.mjs** (14 features): Only 5 of 25 graphs
were registered in the production LangGraph server — all unregistered
graph endpoints returned 404
- **Agent-not-found errors** (8 integrations): V1→V2 import mismatches,
missing agent registrations, per-request runtime creation race
conditions, and stale AgentConfig subclasses incompatible with LangGraph
0.6.0+
- **byoc-hashbrown backend wiring** (5 integrations): Agent name
mismatches, missing default aliases, wrong agent (mastra used
weatherAgent instead of a dedicated hashbrown agent)
- **Agno reasoning**: Stock Agno handler emits STEP_STARTED/FINISHED
(ignored by CopilotKit); replaced with custom handler emitting
REASONING_MESSAGE AG-UI events
- **ms-agent-python**: Missing chat-slots components, unregistered
interrupt agents, missing public/demo-files/ assets
- **mastra subagents**: e2e test expected nonexistent UI elements

**Remaining (3 llamaindex timeouts):** tool-rendering-reasoning-chain,
gen-ui-declarative, gen-ui-a2ui-fixed — exhaustive static analysis found
no code bug; hypothesis is a LlamaIndex `astream_chat_with_tools`
streaming issue when tool definitions are present but the response is
text-only. Needs runtime testing.

## Test plan

- [ ] CI passes
- [ ] Deploy to Railway (auto-deploy on merge)
- [ ] Wait for next D5 probe cycle (~15 min)
- [ ] Verify 15+ integrations at D5 green on the showcase dashboard
- [ ] Investigate remaining llamaindex timeouts with local Docker
container
2026-04-30 17:05:12 -07:00
Jordan Ritter ef1bf442eb fix(showcase): revert custom agno reasoning handler, use stock AGUI
The custom _run_reasoning_agent handler had a bug where text messages
weren't rendered by the frontend despite the backend emitting correct
AG-UI events. The stock AGUI handler works with reasoning=False and
aimock fixtures — the D5 probe checks for reasoning keywords in the
transcript, not for REASONING_MESSAGE events specifically.

Locally verified: 28/29 D5 features pass (only auth fails — pre-existing
auth gate regression unrelated to this change).
2026-04-30 17:04:57 -07:00
Jordan Ritter 73833a368b fix(showcase): show docs links for docs-only features under any active overlay
The CLI Start Command feature (kind: "docs-only") was rendering empty
cells in the dashboard grid because the docs row only displayed when the
"docs" overlay was explicitly toggled on. Since docs are the only content
for docs-only features, show the docs row whenever any content-producing
overlay is active (links, depth, health, or docs), not just docs alone.
2026-04-30 17:04:57 -07:00
Jordan Ritter aea980526e fix(showcase): use repo root as dashboard build context in docker-compose
The dashboard Dockerfile uses COPY paths prefixed with `showcase/...`
which expect the repo root as build context. The compose entry was
using `build: ./shell-dashboard` which set context to
showcase/shell-dashboard/, causing COPY failures in worktrees.

Split into explicit context (../ relative to compose file = repo root)
and dockerfile path so COPY paths resolve correctly.
2026-04-30 17:04:56 -07:00
Jordan Ritter bf98500bd6 fix(showcase): agno reasoning handler must emit text message for D5 probes
The _run_reasoning_agent handler's fallback path (for aimock fixtures
that return plain text with "Reasoning:" prefix) was setting
answer_text="" which skipped emitting any TEXT_MESSAGE events.
CopilotKit requires a text message to render an assistant bubble in
the conversation view -- reasoning events alone produce no visible
DOM element that the D5 probe selectors can match, causing both
reasoning-display and tool-rendering-reasoning-chain to timeout
with 0 assistant messages.

Fix: set answer_text = full_text so the response is emitted as both
a reasoning message (for the ReasoningBlock slot) and a text message
(for the conversation transcript the probe reads).
2026-04-30 17:04:56 -07:00
Jordan Ritter f02c6fa906 chore(showcase): ratchet validate-pins baseline 134→135
New pin drift from ms-agent-dotnet V1→V2 import change.
2026-04-30 17:04:56 -07:00
github-actions[bot] 751eb7d389 style: auto-fix formatting 2026-04-30 17:04:56 -07:00
Jordan Ritter db1d7d05cb fix(showcase): agno reasoning, ms-agent-python slots/multimodal, mastra subagents
- agno: custom _run_reasoning_agent handler emitting proper
  REASONING_MESSAGE AG-UI events (Agno's stock handler only emits
  STEP_STARTED/FINISHED which CopilotKit ignores); disable reasoning=True
  to avoid multi-call CoT loop that breaks aimock fixtures
- ms-agent-python: wire chat-slots assistantMessage + disclaimer overrides;
  add missing public/demo-files/ (sample.png, sample.pdf)
- mastra: register byocHashbrownAgent in main route; rewrite subagents
  e2e test to match actual page structure
2026-04-30 17:04:56 -07:00
Jordan Ritter 1db0bd7042 fix(showcase): resolve agent-not-found errors across integrations
- ms-agent-dotnet auth: V1→V2 CopilotKit import for proper agent discovery
- ms-agent-python: register interrupt agents (array declared but never iterated)
- claude-sdk-python: register hitl-in-chat-booking agent + fix stale dates
- ag2 + langgraph-python: declarative-gen-ui routes use default agent with
  runtime auto-injection instead of custom backend a2ui agents
- google-adk: hoist copilotRuntimeNextJSAppRouterEndpoint to module scope
  (per-request invocation caused race condition in agent Promise chain)
- langgraph-fastapi: remove AgentConfigLangGraphAgent that caused HTTP 400
  with LangGraph 0.6.0+; add default alias for open-gen-ui
2026-04-30 17:04:55 -07:00
Jordan Ritter 60d8139d3d fix(showcase): register all 25 graphs in langgraph-typescript server
The production LangGraph server only registered 5 of 25 graphs from
langgraph.json, causing 404s for all unregistered graph endpoints.
Also adds 6 missing agent registrations to the main route and
normalizes deploymentUrl trailing slashes across dedicated routes.
2026-04-30 17:04:40 -07:00
Jordan Ritter faac42c313 fix(showcase): wire byoc-hashbrown backend agents correctly
- agno: add default agent alias + per-request runtime
- langgraph-fastapi: add default agent alias
- llamaindex: fix agent name mismatch (byoc_hashbrown → byoc-hashbrown-demo)
- mastra: create dedicated byocHashbrownAgent with hashbrown system prompt
  (was using weatherAgent which produced plain text instead of JSON)
- ms-agent-dotnet: upgrade byoc page to V2 CopilotKit import
2026-04-30 17:04:40 -07:00
Jordan Ritter d36660ba24 fix(showcase): add D5 probe testid to byoc-hashbrown across all integrations
The D5 conversation runner detects assistant responses via
data-testid="copilot-assistant-message". The byoc-hashbrown demo
overrides the assistantMessage slot with a custom HashBrown renderer,
which dropped that attribute. Without it the harness sees 0 messages
and times out.
2026-04-30 17:04:39 -07:00
Jordan Ritter a7d63e4000 feat(showcase): add --isolate flag for parallel-safe local testing (#4546)
## Summary

Adds `--isolate [name]` flag to `showcase test` so multiple
agents/sessions can run D5 probes without Docker container conflicts.

- `showcase test agno --d5 --isolate` — auto-names project
`isolate-<PID>`
- `showcase test agno --d5 --isolate d5verify` — explicit project name

When active, `apply_isolation()` offsets all host ports by +200, renames
containers from `showcase-*` to `<name>-*`, and sets
`COMPOSE_PROJECT_NAME`. The EXIT trap tears down the isolated group and
restores both `local-ports.json` and `docker-compose.local.yml` on exit.

## Test plan

- [x] Verified `--isolate d5verify` shows correct offset ports (3309 =
3109 + 200)
- [x] Containers start with `d5verify-*` names, don't conflict with
existing `showcase-*` group
- [x] Cleanup restores files and tears down containers on exit
- [ ] CI passes
2026-04-30 16:34:56 -07:00
Jordan Ritter c233bb7ac7 revert: perf(showcase-dashboard) — broke live status (#4547)
Reverts #4504. The parallelized fetch change broke the dashboard's live
status connection, causing all integrations to show 'offline' and the
Feature Matrix to show 'OFFLINE'.
2026-04-30 16:32:00 -07:00
Jordan Ritter d439b61c68 Revert "perf(showcase-dashboard): yield during fetch + parallelize initial pages (#4504)"
This reverts commit d17ea911e0, reversing
changes made to a0770ea0cf.
2026-04-30 16:31:37 -07:00
Jordan Ritter f9578b2ab7 feat(showcase): add --isolate flag for parallel-safe local testing
When passed to `showcase test`, creates an isolated Docker Compose
project with offset ports (+200) and renamed containers, allowing
multiple agents/sessions to run showcase tests simultaneously without
container conflicts.

Usage:
  showcase test agno --d5 --isolate           # auto-names isolate-<PID>
  showcase test agno --d5 --isolate d5verify  # explicit name
2026-04-30 16:31:12 -07:00
Jordan Ritter d17ea911e0 perf(showcase-dashboard): yield during fetch + parallelize initial pages (#4504)
## Summary

Follow-up to #4502. With matrix cells memoized and SSE deltas coalesced,
the dashboard is much smoother — but the **initial dashboard load**
still freezes for a few seconds because:

1. **Initial PB fetch is 10 sequential round-trips.** `fetchInitial`
walks `getList(page=1, 200) → getList(page=2, 200) → …` up to 10 pages,
each awaiting the previous. Network wall time stacks.
2. **First commit with real data is unavoidably a full-matrix
re-render.** The empty-map → populated-map transition invalidates
per-key memo checks on every cell — that's 720 cell renders in one
synchronous React commit, blocking the main thread.

## Changes

- **Parallelize the initial fetch.** Pull page 1 sequentially to learn
`totalItems`, then fire pages 2..N concurrently via `Promise.all`. Wall
time drops from `sum(rtt_per_page)` to roughly `max(rtt_per_page)`
modulo network parallelism. PocketBase reads are independent so this is
safe.
- **`startTransition` around the initial `setRows(initial)`.** The first
big commit is unavoidable, but marking it as a transition lets React 19
yield to user input mid-walk instead of blocking for the entire render.
`setStatus` stays urgent so the "connecting → live" indicator still
flips immediately.
- **`startTransition` around the SSE flush.** Bursts that the 16ms
coalescer can't fully absorb (large reconnect replays, simultaneous
probe completions) still need to yield rather than block.

## Test plan

- [x] `useLiveStatus.test.tsx` 15/15 pass with parallelized fetch (the
mock handles arbitrary page numbers).
- [x] `composed-cell.test.tsx` 9/10 pass (1 pre-existing failure, same
as #4502).
- [x] No new TypeScript errors in `useLiveStatus.ts`.
- [ ] **Manual perf trace before/after on the live dashboard.**
Expected: initial-load wall time drops sharply (sequential → parallel
pages), and any remaining heavy commit no longer blocks the main thread
(transition lets React yield).
- [ ] **Functional sanity** — confirm rows still arrive correctly when
pages return out of order, and the "connecting → live" status transition
still fires before the heavy render.
2026-04-30 16:02:43 -07:00
Jordan Ritter a0770ea0cf feat(showcase): close 28 unsupported gaps — 93.5% → 97.4% coverage (#4545)
## Summary

Closes 28 of 29 "unsupported" showcase cells, moving the dashboard from
673 to 701 wired features (93.5% → 97.4% coverage).

### Interrupt demos (26 cells, 13 integrations)
- Added gen-ui-interrupt + interrupt-headless to: ag2, agno,
built-in-agent, claude-sdk-python, claude-sdk-typescript, crewai-crews,
google-adk, langroid, llamaindex, mastra, pydantic-ai, spring-ai,
strands
- Uses MS Agent Python "Strategy B": useFrontendTool + async Promise
handler — no native interrupt primitive needed
- Backend agents use system prompt + tools=[] — CopilotKit runtime
routes tool calls to frontend

### Spring AI additional features (2 cells)
- **byoc-json-render**: zero-tool agent + @json-render frontend (routine
port, PARITY_NOTES justification was incorrect)
- **shared-state-streaming**: per-token STATE_SNAPSHOT emission via
tool-call argument interception in a dedicated Spring controller

### What remains unsupported (1 cell)
- hitl on built-in-agent — genuinely unsupported (TanStack AI
chat-completions has no graph to pause). Equivalent UX ships as
hitl-in-chat.

## Test plan
- [x] showcase-harness: 1480 tests pass (91 files)
- [ ] CI green
- [ ] Registry regeneration validates 701 wired / 1 unsupported
2026-04-30 16:00:47 -07:00
Jordan Ritter 7d04e0ca37 fix(showcase): add --ci flag to eval workflow command (#4544)
## Summary

- The eval workflow command was missing `--ci`, causing it to try Docker
Compose lifecycle in CI — which immediately fails because
`showcase/.env` doesn't exist in the runner environment.
- This is why the auto-triggered eval on PR #4541 failed with exit code
1.
- The `--ci` flag tells the eval orchestrator to skip Docker lifecycle
and assume services are already running (or use native execution via
`ci-native-eval.sh`).

## Test plan

- [ ] Trigger eval on a PR, verify it no longer fails with "env file not
found"
2026-04-30 16:00:27 -07:00
Jordan Ritter be01722df4 feat(showcase): close all spring-ai unsupported gaps
Add interrupt demos (Strategy B), byoc-json-render demo (zero-tool
agent + @json-render frontend), and shared-state-streaming (per-token
STATE_SNAPSHOT emission via tool-call argument interception).
Spring AI moves from 4 unsupported features to 0.
2026-04-30 15:59:07 -07:00
Jordan Ritter d6b784ee9a feat(showcase): add interrupt demos to 12 integrations via Strategy B
Replace gen-ui-interrupt and interrupt-headless "not supported" stubs
with working demos using useFrontendTool + async Promise pattern.
Backend agents use system prompt + tools=[] — CopilotKit runtime
routes tool calls to the frontend handler. Pattern proven by
ms-agent-python/dotnet, now extended to ag2, agno, built-in-agent,
claude-sdk-python, claude-sdk-typescript, crewai-crews, google-adk,
langroid, llamaindex, mastra, pydantic-ai, strands.
2026-04-30 15:59:00 -07:00
Jordan Ritter 25c7fa368b fix(showcase): add --ci flag to eval workflow command
The eval workflow was missing --ci, causing it to try Docker Compose
lifecycle in CI (which fails because there's no .env file). The --ci
flag skips Docker and assumes services are already running or uses
native execution.
2026-04-30 15:58:42 -07:00
Jordan Ritter 3eeb32216f fix(showcase): prevent eval auto-trigger by requiring confirmation click (#4543)
## Summary

- The "Run Evaluation" link in PR comments was a GET endpoint that
dispatched the eval workflow directly. GitHub's link unfurling/preview
bot fetches URLs embedded in comments, which caused evals to
auto-trigger on every PR push without anyone clicking the link.
- Changed the GET endpoint to render a confirmation page (no side
effects). The actual dispatch now requires a POST via the form button
click.
- After dispatch, the response opens in a new tab (via `target="_blank"`
on the form) and redirects to the Actions workflow run page where
progress can be monitored — instead of redirecting back to the PR.

## Test plan

- [ ] Deploy eval-webhook to Railway
- [ ] Open a test PR, verify the bot comment link shows the confirmation
page (no auto-dispatch)
- [ ] Click "Run Evaluation" on the confirmation page, verify eval
dispatches and new tab opens to Actions run
- [ ] Verify the original tab preserves navigation history
2026-04-30 15:52:18 -07:00
Jordan Ritter 6303cd6b88 fix(showcase): prevent eval auto-trigger by requiring confirmation click
Split GET /trigger/eval into a two-step flow: GET renders a confirmation
page with a "Run Evaluation" button, POST performs the actual dispatch.
This prevents GitHub's link unfurling bot from auto-triggering evals
when it fetches the URL from PR comments.

The POST handler polls for the Actions run URL after dispatch and opens
it in a new tab via window.open, with a fallback link if the run isn't
found within 5 seconds.
2026-04-30 15:50:57 -07:00
Martha Schumann 29be068dfb chore(example): wire langgraph-python-threads into the pnpm workspace
The intelligence-backed example was a standalone npm-managed monorepo
pinned to published @copilotkit/runtime 1.56.3, so iterating on the
runtime against an Intelligence platform required publishing a release
or hand-linking. Pulling its two JS apps into the pnpm workspace lets
them resolve workspace:* and use whatever the local packages currently
build to.

Specific changes:

- Adds the example's app and bff to pnpm-workspace.yaml.
- Switches their @copilotkit/* deps to workspace:*; adds
  @copilotkit/web-inspector as a workspace dep on the frontend so the
  inspector can mount alongside the chat.
- Adds a small <Inspector /> React component that side-effect-imports
  @copilotkit/web-inspector to register the custom element and assigns
  the CopilotKit core via a ref (Lit elements expect properties, not
  attributes, for complex values).
- Bumps the docker-compose image to the rc.16 composite, which is the
  earliest published tag containing CopilotKit/Intelligence#144's
  /api/_inspect/threads/:id/{events,state} endpoints.
- Drops the example's package-lock.json so npm and pnpm don't fight
  over lockfiles in the workspace.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 15:44:56 -07:00
Martha Schumann 762eb79d74 fix(inspector): lazy-load events/state + spinner so tab clicks feel instant
The threads-detail sub-tabs (Conversation, Agent State, AG-UI Events)
all fetched eagerly on threadId change. Against an Intelligence-backed
runtime, the AG-UI events response can be many MB; the JSON.parse alone
blocks the main thread for several seconds while the user is still on
the conversation tab. Any sub-tab click queued during that window
couldn't fire until the parse finished, so the tab itself appeared
unresponsive — the user perceived 15s of frozen UI before the panel
swapped, and the active-tab highlight didn't paint either.

Two fixes:

1. Defer the events / state fetches to first sub-tab click. Conversation
   stays eager because it's the default tab and visible immediately.
   When the user clicks an Agent State or AG-UI Events sub-tab the fetch
   kicks off then — so the heavy JSON.parse blocks AFTER the click has
   registered, not before.

2. Add a `_panelInitializing` flag set on tab change and cleared in the
   next animation frame. The render switches to a generic "Loading…"
   placeholder while the flag is true, so the active-tab highlight and
   spinner paint before the heavy per-tab render runs.

Total wait for events to display is unchanged; perceived responsiveness
of the click is now immediate.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 15:43:29 -07:00
github-actions[bot] ee50e1f416 style: auto-fix formatting 2026-04-30 22:26:18 +00:00
Jordan Ritter b76a19bc55 fix(showcase): ratchet validate-pins baseline to 135
Pre-existing drift on main (strands ag-ui-protocol pin) bumped the
count from 134 to 135. Update baseline hash to unblock CI.
2026-04-30 15:23:47 -07:00
Jordan Ritter 90d478c75f fix(showcase): update d5-auth test fakes for evaluate-based click
The defaultClick implementation now uses page.evaluate() instead of
page.click(), so the test fakes need to inject the click via the
buildAuthAssertion({ click }) option rather than relying on a page.click()
method that the assertion no longer calls.
2026-04-30 15:20:42 -07:00
Jordan Ritter e9501a58f5 fix(showcase): use JS-level click in D5 auth probe to bypass cpk-web-inspector
The <cpk-web-inspector> overlay intercepts Playwright pointer events even with
force: true, preventing the sign-out button's React onClick from firing. Switch
from page.click(selector, { force: true }) to page.evaluate() with a
document.querySelector(sel).click() call, which triggers the DOM click event
directly without pointer dispatch. This bypasses the overlay and reliably flips
the auth state.

Without this fix the auth D5 probe times out waiting for the error surface
that only appears after a successful sign-out + 401 request cycle.
2026-04-30 15:18:37 -07:00
Jordan Ritter 9a6ce11716 fix(showcase): add harness testids to claude-sdk-python BYOC hashbrown renderer
The AssistantMessageRenderer was missing data-testid="copilot-assistant-message"
and data-message-role="assistant" attributes. Without these, the e2e-deep
conversation runner cannot detect assistant responses (it counts elements by
these selectors), causing a 30s timeout and D5 failure for the byoc feature.

Mirrors the same fix already applied to langgraph-python and
claude-sdk-typescript.
2026-04-30 15:18:29 -07:00
Jordan Ritter 2fdbb1caed fix(showcase): reorder legend — L1-L4 first, add D3, remove D0-D4 (#4540)
Legend now reads: L1-L4 Strip, D2 API, D3 Page Load, D4 Round Trip, D5
Conversation, then indicators.
2026-04-30 14:51:02 -07:00
Jordan Ritter 8e36bb8b2d fix(showcase): reorder legend — L1-L4 first, add D3, remove D0-D4 2026-04-30 14:50:39 -07:00