Commit Graph

8565 Commits

Author SHA1 Message Date
Jordan Ritter 81b2d9a53e fix(showcase): add harness testids to claude-sdk-typescript BYOC renderers (#4539)
## Summary

- Added `data-testid="copilot-assistant-message"` and
`data-message-role="assistant"` to the BYOC hashbrown and json-render
custom assistant message renderers in claude-sdk-typescript
- These attributes were already present in the langgraph-python
gold-standard; this aligns the TS port

## Why

The byoc-hashbrown and byoc-json-render demos override CopilotChat's
`messageView.assistantMessage` slot with custom renderers. The e2e-deep
conversation runner counts assistant messages via
`[data-testid="copilot-assistant-message"]` to detect "response
settled." Without that attribute on the custom renderers, the count
stayed at 0 forever and the byoc feature timed out -- the only D5
failure blocking claude-sdk-typescript from green.

## Test plan

- `showcase test claude-sdk-typescript --d5 --verbose` goes from 28/29
(byoc timeout) to 29/29 green
- No demo functionality changed -- only test-harness attributes added to
wrapper divs
2026-04-30 14:44:38 -07:00
Jordan Ritter a1e27b330e fix(showcase): add harness testids to claude-sdk-typescript BYOC renderers
The byoc-hashbrown and byoc-json-render demos override the CopilotChat
assistantMessage slot with custom renderers, but the overrides were
missing the data-testid="copilot-assistant-message" attribute that the
e2e-deep conversation runner uses to detect settled assistant responses.
Without it, the runner's readMessageCount always returned 0 and the
byoc feature timed out at D5.

The langgraph-python gold-standard already had these attributes; this
aligns claude-sdk-typescript to match.
2026-04-30 14:44:10 -07:00
Jordan Ritter 0c4ad13460 fix(showcase): exclude docs-only features from catalog metadata counts (#4538)
## Summary

- docs-only features (e.g. cli-start) exist purely for documentation
coverage tracking and have no route, depth probes, or health signals
- The catalog metadata was counting their 18 stub cells in the headline
wired/stub/unshipped/unsupported breakdown, inflating the total and
making the stats bar misleading
- Exclude docs-only cells from the headline counts; add a new
`docs_only` field to track them separately

**Before:** total_cells=720 wired=673 stub=18 unsupported=29
**After:** total_cells=702 wired=673 stub=0 unsupported=29 docs_only=18

## Test plan

- [x] `showcase/scripts` test suite: all 574 tests pass (including 13
catalog tests)
- [x] `showcase/shell-dashboard` test suite: same 3 pre-existing
failures, no regressions
- [x] TypeScript check clean (no new errors)
- [x] Invariant holds: wired + stub + unshipped + unsupported +
docs_only == cells.length
2026-04-30 14:39:50 -07:00
Jordan Ritter 79c10dbce2 fix(showcase): exclude docs-only features from catalog metadata counts
docs-only features (e.g. cli-start) exist purely for documentation
coverage tracking -- they have no route, no depth probes, and no
health signals. The catalog metadata was counting their 18 stub cells
in the headline wired/stub/unshipped/unsupported breakdown, inflating
the total and making the stats bar misleading.

Exclude docs-only cells from the headline counts. A new docs_only
field tracks the excluded count separately so the invariant
(wired + stub + unshipped + unsupported + docs_only == cells.length)
holds.

Before: total_cells=720 wired=673 stub=18 unsupported=29
After:  total_cells=702 wired=673 stub=0  unsupported=29 docs_only=18
2026-04-30 14:38:18 -07:00
Jordan Ritter 6d2b6ba16b fix(showcase): wire all three slot overrides in llamaindex chat-slots demo (#4537)
## Summary

- Add `CustomAssistantMessage` and `CustomDisclaimer` components to
llamaindex chat-slots demo, matching langgraph-python gold standard
- Wire `input.disclaimer` and `messageView.assistantMessage` slot props
into `<CopilotChat>` alongside the existing `welcomeScreen` slot
- Update e2e spec to D5 coverage (welcome screen, suggestions, assistant
message slot, disclaimer slot, multi-turn slot persistence)
- Update manifest.yaml highlight list with new component files

## Why

The llamaindex chat-slots demo only had the `welcomeScreen` slot
override, while the D5 probe (`d5-chat-slots.ts`) asserts
`data-testid="custom-assistant-message"` is present after the assistant
responds. This left the feature stuck at D4.

## Test plan

- [x] `showcase build llamaindex` -- builds clean
- [x] `showcase test llamaindex --d5 --verbose` -- chat-slots passes
(was the only feature stuck at D4)
- [x] Verified Docker container renders all three slot testids
2026-04-30 14:33:03 -07:00
Martha Schumann c9e1bd8e34 feat(runtime): wire intelligence /_inspect/threads/:id/{events,state} endpoints
Replaces the 501 stubs in handleGetThreadEvents and handleGetThreadState
with real delegation to CopilotKitIntelligence. Adds two new client
methods (`getThreadEvents`, `getThreadState`) on intelligence-platform/
client.ts that hit the new `/api/_inspect/threads/:id/{events,state}`
routes shipped in Intelligence PR #144 (CPK-7453). Auth flows through
the existing `resolveIntelligenceUser` path; threadId scoping happens
server-side via the API key's org/project resolution.

Wire shapes match the in-memory branch so the inspector consumes both
runtimes identically:
- events: `{ events }` (platform-internal `decodeErrorRowIds` and
  `truncated` flags are stripped at the runtime boundary)
- state: `{ state }` where the platform's discriminated `ThreadStateResult`
  flattens to the snapshot value for `kind: "snapshot"` and to `null` for
  both `no-snapshot` and `snapshot-decode-error`

Replaces the 501 regression-protection tests with real delegation tests
that mock the platform methods, assert call args, and exercise the
no-snapshot / decode-error / throw paths.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 14:33:02 -07:00
Jordan Ritter 7db463a208 fix(showcase): wire all three slot overrides in llamaindex chat-slots demo
The llamaindex chat-slots demo only had the welcomeScreen slot,
missing the input.disclaimer and messageView.assistantMessage
overrides that langgraph-python (gold standard) provides. The D5
probe checks for data-testid="custom-assistant-message" after the
assistant responds, so the feature was stuck at D4.

Add CustomAssistantMessage and CustomDisclaimer components (matching
langgraph-python), wire them into CopilotChat props, update the e2e
spec to cover all three slots, and update manifest highlights.
2026-04-30 14:31:35 -07:00
Martha Schumann 099700721c fix(inspector): silent re-fetch for live conversation updates
The first reactivity fix triggered a `_loadingMessages` flicker between
streaming chunks because every live re-fetch toggled the loading state
and replaced the conversation array wholesale. The silent flag suppresses
both: live re-fetches keep the loading indicator off and preserve the
last-good conversation on transient fetch errors. Initial threadId-change
fetches still show a real loading state.

Adds the same staleness guard the other tab fetches use (skip applying
results if `threadId` changed mid-flight) so a quick thread switch can't
leave the wrong thread's data on screen.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 14:27:57 -07:00
Jordan Ritter 3fc4c9900f fix(showcase): remove D6 from depth stats, tighten grid spacing (#4536)
Remove D6 from depth distribution counts and DepthDistribution
interface. Reduce whitespace between stats bar and feature matrix
header.
2026-04-30 14:24:16 -07:00
Jordan Ritter 9ec8b90516 fix(showcase): remove D6 from depth stats, tighten grid spacing 2026-04-30 14:23:58 -07:00
Jordan Ritter 6abb87ee2b fix(showcase): improve cell drilldown dialog UX (#4535)
## Summary

- Widen dialog from `w-72` (288px) to `w-[480px]` so error details are
not cramped
- Extract key signal fields (`errorDesc`, `error`, `failureSummary`,
`backendUrl`, `apiRequestCount`, `step`) as readable key-value pairs
instead of raw JSON
- Deduplicate `errorDesc` and `error` when both present (first match
wins)
- Make raw signal payload collapsible ("Raw Signal" toggle) for
debugging
- Remove duplicate tooltip text that repeated status/timing info already
conveyed by badge color and failure metadata
- Compact fail_count and first_failure_at into a single inline row

## Test plan

- [x] All 11 cell-drilldown tests pass (4 new tests added for extracted
fields, collapsible toggle, deduplication, and wider width)
- [x] No new TS errors introduced
- [x] Pre-existing test/TS failures in other files are unrelated
2026-04-30 14:17:33 -07:00
Jordan Ritter 7556b0d244 fix(showcase): improve cell drilldown dialog width, signal readability, and deduplication
Widen dialog from w-72 to w-[480px], extract key signal fields
(errorDesc, error, failureSummary, backendUrl) as readable key-value
pairs instead of buried JSON, make raw signal collapsible, remove
duplicate tooltip text that repeated badge color/status info.
2026-04-30 14:15:46 -07:00
Jordan Ritter 763aeb7662 fix(showcase): hide depth chip for docs-only features, add depth distribution to stats bar (#4534)
## Summary

- Suppress depth chip rendering for docs-only features in the RefDepth
column (parity overlay) -- show `--` instead, since docs-only features
have no probe coverage and should not display a depth level
- Add depth distribution counts (D6, D5, D4, D3, D2, D1) to the
AdaptiveStatsBar when the depth overlay is active, giving operators
at-a-glance visibility into how many wired cells sit at each depth level

## Test plan

- [ ] Verify docs-only features (e.g. `cli-start`) show `--` in the Ref
Depth column instead of a depth chip
- [ ] Verify non-docs-only features still show their correct depth chip
in Ref Depth
- [ ] Toggle the depth overlay on and confirm the depth distribution
section appears in the stats bar
- [ ] Toggle depth overlay off and confirm the distribution section
disappears
2026-04-30 14:04:28 -07:00
Jordan Ritter accf9b5050 fix(showcase): hide depth chip for docs-only features, add depth distribution to stats bar
Suppress RefDepthCell for docs-only features in the parity overlay,
showing '--' instead of a depth chip since docs-only features have no
probe coverage. Add depth distribution counts (D6..D1) to the
AdaptiveStatsBar when the depth overlay is active, giving operators
at-a-glance visibility into how many wired cells are at each depth.
2026-04-30 14:02:57 -07:00
Jordan Ritter dffe3e25d8 fix(showcase): add favicon and og-image to shell-dashboard (#4533)
Copy icon.svg and og-image.png from shell to shell-dashboard. Add
metadata to layout.tsx.
2026-04-30 13:55:03 -07:00
Jordan Ritter 1d817577f2 fix(showcase): add favicon and og-image to shell-dashboard 2026-04-30 13:53:57 -07:00
Martha Schumann 94305c0927 fix(inspector): re-fetch conversation when active agent emits new messages
The conversation view in cpk-thread-details only re-fetched on threadId
change, so live agent output during streaming wasn't visible until the
user switched threads (or unmounted/remounted the element by clicking
out and back in). Restoring the original conversationOverride channel
isn't viable because the parent's agentMessages map is keyed by agentId
(not threadId) and ConversationItem mapping lives in the child — passing
mapped messages would either leak across threads or duplicate the
runtime's AG-UI → ConversationItem conversion in the parent.

Instead, the parent now tracks each agent's currently-running threadId
(from `onAgentRunStarted`) and ticks a per-thread liveMessageVersion
counter every time `syncAgentMessages` fires for that agent. The counter
is passed to cpk-thread-details, which watches it in `updated()` and
re-fetches `/threads/:id/messages` when it changes for the same threadId.
The runtime endpoint is the single source of truth for conversation
shape, so streaming output flows in without any client-side mapping.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 13:51:27 -07:00
Jordan Ritter 35586633db feat(showcase/dashboard): add docs-only feature kind (#4532)
## Summary

- Adds `docs-only` as a new `FeatureKind` alongside `primary` and
`testing`
- Features with this kind show only the docs row in the dashboard -- no
Dx badges (API/RT/CV), no links, no depth chips
- Row label is muted/italic with a "docs-only" tag, matching the visual
treatment of "testing" rows
- `cli-start` is the first feature changed to `docs-only` since it has
no runnable demo

## Changes

| File | What |
|------|------|
| `showcase/shared/feature-registry.json` | `cli-start` kind: `primary`
-> `docs-only` |
| `showcase/shell-dashboard/src/lib/registry.ts` | `FeatureKind` union
extended |
| `showcase/shell-dashboard/src/components/cell-pieces.tsx` |
`CellStatus` returns null for docs-only |
| `showcase/shell-dashboard/src/components/composed-cell.tsx` | Only
renders DocsLayer for docs-only |
| `showcase/shell-dashboard/src/components/feature-grid.tsx` | Muted row
styling + tag |
| `cell-pieces.test.tsx` | New test for docs-only hiding all badges |
| `composed-cell.test.tsx` | 3 new tests for docs-only overlay behavior
|

## Test plan

- [x] New unit tests pass for all three behavioral changes
- [x] Pre-existing test failures confirmed unrelated (useOverlays URL
hash, composed-cell docs-dedup, missing generated data files)
2026-04-30 13:49:37 -07:00
Jordan Ritter 03aa2f3756 feat(showcase/dashboard): add docs-only feature kind
Features with kind "docs-only" show only the docs row in the
dashboard -- no Dx badges, no status badges, no links or depth
layers. The row label is muted and tagged like "testing" rows.

- feature-registry.json: cli-start changed from primary to docs-only
- registry.ts: FeatureKind union extended with "docs-only"
- cell-pieces.tsx: CellStatus returns null for docs-only (hides all
  badges)
- composed-cell.tsx: docs-only features render only DocsLayer, skip
  links/depth/health
- feature-grid.tsx: docs-only rows use muted italic styling with a
  "docs-only" tag
- Tests added for all three behavioral changes
2026-04-30 13:49:05 -07:00
Jordan Ritter 003da0d9d0 fix(showcase/dashboard): unsupported cells show indicator instead of numeric depth (#4531)
## Summary

- `deriveDepth()` was computing D2 for `not_supported_features` because
D1/D2 are integration-scoped
- Added `unsupported` boolean to `DepthResult`, guard in `deriveDepth`,
updated consumer components
- 24/24 depth-utils tests pass

## Test plan

- [ ] Verify dashboard renders "—" instead of numeric depth for
unsupported cells
- [ ] Confirm 24/24 depth-utils tests pass (`nx run
@copilotkit/showcase-shell-dashboard:test`)
2026-04-30 13:41:02 -07:00
Jordan Ritter 607c5567d8 fix(showcase/google-adk): D5 green — 6 demo fixes verified locally (#4530)
## Summary

- **29/29 google-adk features pass D5 probes locally**
- Bug fixes across 6 demos:
- **chat-slots**: agent name dashes→underscores, added
custom-assistant-message + custom-disclaimer slots
  - **prebuilt-popup**: added `defaultOpen` prop for D5 probe visibility
- **auth**: demo defaults to unauthenticated state so D5 can exercise
the full auth flow
- **byoc**: added missing `data-testid="assistant-message"` to hashbrown
and json-render renderers
- **multimodal**: added `SampleAttachmentButtons` component + sample
PDF/PNG demo files
- **chat-customization-css**: aligned CSS custom properties and page
layout with D5 reference theme

## Test plan

- [ ] `bin/showcase test google-adk` — all 29 features green
- [ ] Visual spot-check: prebuilt-popup opens by default, auth starts
unauthenticated, multimodal shows attachment buttons
- [ ] No other integrations affected (google-adk files only)
2026-04-30 13:40:23 -07:00
Jordan Ritter d8d73d6b38 feat(showcase): one-click eval trigger via signed bot comment link (#4527)
## Summary

Adds a clickable "Run Evaluation" link in a bot comment on every PR. One
click triggers the showcase eval — no Checks tab navigation needed.

### How it works
1. `showcase_eval_check.yml` creates the Check Run AND posts a bot
comment with a signed link
2. The link points to
`eval-webhook-production.up.railway.app/trigger/eval?pr=N&check_run_id=M&sig=HMAC`
3. User clicks → webhook service verifies signature, dispatches eval
workflow, updates Check Run to in_progress, redirects user back to the
PR
4. Eval results appear in the Check Run (already implemented in #4517)

### Security
- Link is HMAC-signed using the webhook secret — can't be forged
- Signature covers `pr:check_run_id` — changing either invalidates the
link
- `EVAL_WEBHOOK_SECRET` added as repo secret for the signing step

### Changes
- `showcase/eval-webhook/src/server.ts` — new `GET /trigger/eval` route
with signature verification + redirect UX
- `.github/workflows/showcase_eval_check.yml` — posts bot comment with
signed link (updates on push, idempotent via HTML marker)

## Test plan
- [ ] Open a test PR → verify bot comment appears with "Run Evaluation"
link
- [ ] Click the link → verify redirect page shows, eval dispatches,
Check Run updates
- [ ] Tamper with the sig parameter → verify 403 response
2026-04-30 13:35:08 -07:00
Jordan Ritter 1541cbc96f fix(showcase/dashboard): unsupported cells show indicator instead of numeric depth
deriveDepth() was computing D2 for not_supported_features because D1/D2
are integration-scoped. Add unsupported boolean to DepthResult, guard in
deriveDepth, update consumer components. 24/24 depth-utils tests pass.
2026-04-30 13:34:57 -07:00
github-actions[bot] 4629573f32 style: auto-fix formatting 2026-04-30 20:33:59 +00:00
Jordan Ritter a9193b2d7f fix(showcase/google-adk): align chat-customization-css theme with D5 reference
Update CSS custom properties and page layout to match the D5 probe
reference theme expectations.
2026-04-30 13:32:21 -07:00
Jordan Ritter b4361d8ea7 fix(showcase/google-adk): add SampleAttachmentButtons and multimodal attachment pipeline
Add SampleAttachmentButtons component with PDF/image sample files,
wire multimodal demo page to support file attachments matching D5
probe expectations.
2026-04-30 13:32:14 -07:00
Jordan Ritter ebcbdc5774 fix(showcase/google-adk): demo wiring fixes for D5 probe compatibility
Fix chat-slots agent name (dashes to underscores), prebuilt-popup
defaultOpen prop, auth demo defaults to unauthenticated state, and
add missing data-testid attributes to byoc renderers.
2026-04-30 13:32:07 -07:00
Jordan Ritter f11e3c60dd feat(showcase): use hooks.showcase.copilotkit.ai custom domain for eval trigger 2026-04-30 13:21:18 -07:00
Jordan Ritter 6159c1885c fix(showcase/ag2): D5 green — 15 demo fixes verified locally (#4529)
## Summary

Fixes 15 ag2 showcase files to bring all 29 ag2 features to D5-green
(verified locally with the D5 probe runner).

### Backend fixes
- **Starlette mount order**: `/open-gen-ui-advanced` must mount before
`/open-gen-ui` — Starlette resolves mounts via prefix matching in
registration order, so the shorter prefix was shadowing the longer one
- **`@tool` -> `@tool()` decorator**: bare `@tool` in
`agent_config_agent.py` silently fails to register the function as a
CopilotKit tool; needs the call form `@tool()`
- **headless-complete route wiring**: wired the headless-complete agent
into the `copilotkit-mcp-apps` runtime route so MCP Apps activity
messages render in the headless cell
- **Auth route cleanup**: removed stale `@ts-ignore`
- **Manifest update**: corrected headless-complete's route reference
from `copilotkit` to `copilotkit-mcp-apps`

### Demo page fixes
- **agent-config**: replaced broken `useAgent`/`setState` pattern with
CopilotKit `properties` prop (the correct way to pass config to AG2
ContextVariables)
- **multimodal**: rewrote the legacy-binary-shape converter to a content
flattener — AG2's AGUIStream validates message content as a plain string
and rejects content-part arrays with a 400
- **headless-complete**: changed `runtimeUrl` from `/api/copilotkit` to
`/api/copilotkit-mcp-apps`
- **byoc-hashbrown, byoc-json-render**: re-attached
`data-testid="copilot-assistant-message"` on custom slot overrides
(harness conversation runner needs this to count settled responses)
- **agentic-chat-reasoning, beautiful-chat**: added demo-level
`data-testid` wrappers
- **tool-rendering-reasoning-chain, voice**: fixed CopilotKit import
paths

### Railway rebuild note
Several demos (chat-slots, shared-state-*, subagents) were already
code-correct on `main` but showed D5 failures due to a stale Railway
image. A Railway rebuild (no code changes) will bring those green as
well.

## Test plan
- [ ] 29/29 ag2 features pass D5 probes locally (verified before PR)
- [ ] CI passes
- [ ] After merge + Railway rebuild, ag2 column should be fully D5-green
2026-04-30 13:18:47 -07:00
Jordan Ritter 69ca0c6afc fix(showcase/ag2): demo page fixes — agent-config, testids, imports, wiring
- agent-config: replace useAgent/setState pattern with CopilotKit properties
  prop (the correct way to pass config to AG2 ContextVariables)
- headless-complete: change runtimeUrl from /api/copilotkit to
  /api/copilotkit-mcp-apps so MCP Apps activity messages render
- byoc-hashbrown, byoc-json-render: re-attach data-testid="copilot-assistant-message"
  on custom assistantMessage slot overrides (harness conversation runner
  needs this to detect settled responses)
- agentic-chat-reasoning, beautiful-chat: add demo-level data-testid wrappers
- tool-rendering-reasoning-chain, voice: fix CopilotKit import path
  (import from @copilotkit/react-core, not /v2, for consistency)
2026-04-30 13:13:51 -07:00
Jordan Ritter 0f6942c5d0 fix(showcase/ag2): multimodal content flattener + attachment pipeline
Rewrite the multimodal demo's message converter from a legacy-binary-shape
rewriter to a content flattener. AG2's AGUIStream validates message content
as a plain string and rejects arrays of content parts with a 400. The new
ContentFlattenerShim extracts text from multipart user messages before the
AG-UI run dispatches them to the AG2 backend.

Also adds defensive validation to sample-attachment-buttons: magic-byte
checks, LFS pointer detection, and actionable error messages.
2026-04-30 13:13:38 -07:00
Jordan Ritter 0aef10f9a2 fix(showcase/ag2): backend wiring — mount order, @tool() decorator, route fixes
- Fix Starlette mount order: /open-gen-ui-advanced before /open-gen-ui to
  prevent prefix shadowing (open-gen-ui-advanced was unreachable)
- Fix @tool -> @tool() decorator in agent_config_agent.py (bare @tool
  silently fails to register the function as a CopilotKit tool)
- Wire headless-complete agent into copilotkit-mcp-apps route so the
  headless-complete demo can exercise MCP Apps rendering
- Update manifest to reference copilotkit-mcp-apps route for headless-complete
- Remove stale @ts-ignore in auth route
2026-04-30 13:13:27 -07:00
Jordan Ritter c4e1650871 docs(showcase): D5 debugging strategies (#4528)
Replace conclusions-heavy D5 section with 8 investigation strategies:
binary classification, same-endpoint differential, curl-the-backend,
message-count-N-to-0, gold-standard-first, custom-renderer-testids,
full-event-chain-trace, production-parity-env-vars.
2026-04-30 13:10:57 -07:00
Jordan Ritter cb6e26880b docs(showcase): replace D5 conclusions with investigation strategies
8 reusable debugging strategies distilled from the D5 all-green push.
Techniques for learning what's wrong, not just cataloging past bugs.
2026-04-30 13:09:48 -07:00
Jordan Ritter 3b4d1eb7c8 feat(showcase): signed trigger link for eval button in PR comments
Add GET /trigger/eval route to eval-webhook with HMAC-signed URLs.
The showcase_eval_check.yml workflow now posts a bot comment with a
clickable "Run Evaluation" link. Clicking triggers the eval and
redirects back to the PR. The link is signed so it can't be forged.
2026-04-30 13:09:02 -07:00
Tyler Slaton 38ac534cfc chore: release monorepo v1.56.5 (#4421) v1.56.5 2026-04-30 13:00:59 -07:00
Jordan Ritter 2c32b48dda fix(showcase): fix spring-ai D5 failures — all 28 features green (#4525)
## Summary

- Fix TOOL_CALL_RESULT events reusing parent assistant message ID across
StreamingToolAgent, SharedStateReadWriteController, and
SubagentsController —
React's deduplicateMessages() Map overwrote assistant messages with tool
  messages when they shared the same ID
- Add `data-testid="copilot-assistant-message"` to BYOC hashbrown
renderer so
  the D5 test harness can detect assistant messages
- Register AG-UI Jackson mixins via JacksonConfig postConfigurer,
normalize
array-format content via ContentNormalizingModule, remove ChatMemory
beans
  that interfered with aimock fixture matching, inject Spring-managed
  ObjectMapper into AgentConfigController

## Test plan

- [x] All 28/28 spring-ai D5 features pass locally (verified stable
across
  multiple consecutive runs)
- [x] Docker build succeeds
- [x] No regressions — previously passing 22 features remain green
2026-04-30 12:55:21 -07:00
Jordan Ritter 81a0fbe080 fix(showcase): fix spring-ai D5 failures — all 28 features green
Root cause: TOOL_CALL_RESULT events reused the parent assistant
message ID, causing React deduplicateMessages() to overwrite
assistant messages with tool messages (Map keyed by ID). Fix uses
unique UUIDs for each tool result message across all three
controllers. Also adds data-testid to the BYOC hashbrown renderer
so the D5 harness can detect assistant messages, registers AG-UI
Jackson mixins via JacksonConfig postConfigurer, normalizes
array-format content via ContentNormalizingModule, removes ChatMemory
beans that interfered with aimock fixture matching, and injects
Spring-managed ObjectMapper into AgentConfigController.
2026-04-30 12:54:52 -07:00
Martha Schumann 2e498341ad test(runtime): pin per-thread routes, GET-only enforcement, and agentId tagging
Three new test areas covering surfaces this PR introduced or relies on:

- fetch-router: matchRoute tests for `/threads/:id/events`,
  `/threads/:id/state`, and `/threads/clear` (with and without URL
  encoding). Critically pins that "/threads/clear" resolves to
  `threads/clear` and does NOT fall through to the more permissive
  `threads/update` arm with threadId="clear" — the explicit guard in
  the router exists for this reason.
- fetch-handler validation: 405 enforcement for POST/PATCH/DELETE on
  the read-only `/threads/:id/events` and `/threads/:id/state`
  endpoints, with `Allow: GET` header. Complementary positive case
  asserts GET is NOT a 405.
- handle-run: end-to-end test that handle-run.ts:40
  (`agent.agentId = agentId`) propagates the registry key onto historic
  runs. Without this stamp, InMemoryAgentRunner falls back to "default"
  and the agentId filter on `GET /threads?agentId=...` breaks for the
  local-dev fallback.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 12:52:35 -07:00
Martha Schumann cf51841d18 test(core): cover onAgentsChanged auto-unregister + seed previousAgentIds from constructor
Pins three integration cases for the thread-store auto-unregister branch
in CopilotKitCore's internal onAgentsChanged subscriber:

1. agent removed from agents → store IS unregistered; subscriber receives
   the previous store via prevStore.
2. FIRST onAgentsChanged({ agents: {} }) on a published-style core where
   the store was registered before any agents arrived → store SURVIVES.
   Reproduces the race that the "previously had" guard exists to
   prevent.
3. add → register → remove cycle → store IS unregistered. Complements (2)
   by exercising the same code path's positive branch.

Test (1) surfaced a real bug in the seed: agentRegistry.initialize does
NOT emit onAgentsChanged, so the constructor's internal subscriber never
saw the agents__unsafe_dev_only set, and `previousAgentIds` stayed empty.
A later removeAgent__unsafe_dev_only call would then be guarded into a
no-op because the agentId looked "new". Fixed by seeding
`previousAgentIds = new Set(Object.keys(agents__unsafe_dev_only))` right
before wiring the subscriber.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 12:52:02 -07:00
Martha Schumann 3d6166611e refactor(runtime,inspector): tighten types and align JSDoc with reality
Type quality:
- intelligence/threads.ts handleGetThreadMessages: switch on the Message
  discriminant (role) and read narrowed fields directly. Removes
  `as Record<string, unknown>` laundering and chained `as` casts on
  toolCalls/function/arguments. AssistantMessage's toolCalls always have
  `function: { name, arguments }`, so the prior fallbacks (`tc.name`,
  `tc.args`) were dead branches.
- in-memory.ts getThreadState: import StateSnapshotEvent and use it
  instead of `(event as { snapshot?: unknown }).snapshot`.

Comment / API alignment:
- intelligence/threads.ts handleClearThreads JSDoc no longer claims the
  inspector calls this; the actual caller is the demo button.
- in-memory.ts clearThreads JSDoc updated to match.
- in-memory.ts getThreadEvents JSDoc no longer references a SQLite
  runner that does not exist; just describes the compaction logic.
- web-inspector lint fix: rename unused `changed` parameter to
  `_changed` per oxc no-unused-vars.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 12:45:07 -07:00
Jordan Ritter b6c8cdc10d docs(showcase): D5 failure classification and framework debugging guide (#4523)
Adds D5-specific debugging sections to DEBUGGING.md based on lessons
from the all-green push.
2026-04-30 12:44:49 -07:00
Jordan Ritter e2649af7a0 docs(showcase): add D5 failure classification and framework debugging guide
Covers: failure triage (text-only vs tool features, probe text extraction,
message disappearing, chatMemory pollution), framework-specific gotchas
(LlamaIndex, Spring AI, Built-in Agent, MS Agent, AG2), and production
vs local parity checklist.
2026-04-30 12:44:32 -07:00
Martha Schumann bd9fe0169a fix(core): guard thread-store auto-unregister against initial empty agents
CopilotKitCore subscribes to onAgentsChanged and unregisters thread
stores for any agentId not in the new agents map. For published cores,
core.agents is asynchronously populated, so the FIRST
onAgentsChanged({ agents: {} }) notification fires BEFORE published
agents are merged in. Without a guard, that empty notification rips out
a thread store that a consumer (e.g. useThreads) just registered.

Track previousAgentIds and only unregister an agentId that was present
in the previous snapshot AND missing from the new one. The first
empty-agents notification (where the agentId was never previously
present) becomes a no-op.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 12:44:13 -07:00
Martha Schumann daf52a2068 fix(inspector): plug data-staleness, error-swallowing, and parse silent-fail
Five behavioural fixes on cpk-thread-details / WebInspectorElement:

- fetchEvents/fetchState: mirror the AbortController pattern fetchMessages
  already uses. Without this, switching threads quickly (A→B) can leave
  the user looking at thread B with thread A's events/state when A's
  request resolves last.
- mapMessages: when JSON.parse fails on tool-call args or tool result
  content, log via console.error and attach __parseError + __raw on the
  parsed object instead of silently substituting `{}`. The inspector is a
  debugging surface; hiding malformed payloads defeats its purpose.
- fetchAnnouncement: keep the captured error and console.warn it. The
  prior `catch {}` swallowed Malformed-payload throws, JSON parse
  failures, and convertMarkdownToHtml exceptions silently.
- subscribeToThreadStore: also subscribe to ɵselectThreadsError, store
  per-agent in `_threadsErrorByAgent`, and surface in renderThreadsView /
  cpk-thread-list as an error state branch alongside empty/no-results.
  Previously a thread-store load failure (REST list rejection, Phoenix
  subscribe failure, retry exhaustion) left the user with stale data and
  no indication of failure.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 12:43:50 -07:00
Martha Schumann 324d0dfabb fix(angular-demo): use POST /threads/clear for clear-threads button
The button hit DELETE /threads, which the runtime router resolves to
"threads/list" (a GET-only route) and returns 405. The intended endpoint
is POST /threads/clear (defined at fetch-router.ts:175). Add try/catch
and surface non-ok HTTP status via console.error so the click no longer
silently appears to succeed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 12:43:01 -07:00
Tyler Slaton 52f8030f82 Merge branch 'main' into release/publish/monorepo/v1.56.5 2026-04-30 12:42:10 -07:00
Tyler Slaton bca17d954c chore: run pnpm format (#4522)
Previous PR's didn't run the formatter.
2026-04-30 12:41:37 -07:00
Tyler Slaton 2d616543ec Merge branch 'main' into tyler/fix-formatting 2026-04-30 12:41:17 -07:00
Jordan Ritter 330b4eb065 feat(showcase): Check Run eval button + native CI execution (#4517)
## Summary

Replaces the `/eval` PR comment trigger with a professional GitHub Check
Run button UX, and adds native (no-Docker) CI execution infrastructure.

### Check Run Button (Phase 1)
- **showcase_eval_check.yml** — creates a Check Run with "Run Showcase
Eval" action button on every PR open/push
- **showcase/eval-webhook/** — tiny Hono relay service on Railway that
receives `check_run.requested_action` webhooks, authenticates as devops
bot, and dispatches `showcase_eval.yml` via `workflow_dispatch`
- **showcase_eval.yml** — adds `workflow_dispatch` trigger with
`dispatch-gate` job, Check Run lifecycle updates (neutral → in_progress
→ completed with results), devops bot token for Checks API

### Native Execution (Phase 2)
- **`--ci` flag** on eval orchestrator — skips Docker lifecycle
(`up/down/isRunning/docker-inspect`), assumes services already running
- **ci-native-eval.sh** — standalone helper that installs deps, starts
`next dev --turbopack` + Python agents natively, health-waits, runs
`showcase eval --ci`
- **test_e2e-showcase-on-demand.yml** — fixed and extended with
`langgraph-python` support, agent-type detection (uvicorn vs
langgraph_cli), relaxed aimock_toggle.py requirement

### Setup required (post-merge)
1. Add `checks:write` permission to the copilotkit-devops-bot GitHub App
2. Deploy eval-webhook to Railway with app secrets
3. Configure GitHub App webhook URL to point to the Railway service

## Test plan
- [x] showcase-harness: 1480 tests pass
- [ ] CI green
- [ ] Deploy eval-webhook, verify `/health` endpoint
- [ ] Open test PR, verify Check Run appears with button
- [ ] Click button, verify eval triggers and Check Run updates
2026-04-30 12:36:09 -07:00