Commit Graph

8565 Commits

Author SHA1 Message Date
Ran Shemtov 992e6f884b Merge branch 'main' into release/publish/monorepo/v1.56.5 2026-04-30 18:14:24 +02:00
Jordan Ritter e28295dd3d fix(showcase): spring-ai HITL D5 features — 6/11 green (#4505)
## Summary

- Spring-ai StreamingToolAgent now classifies tool calls as frontend vs
backend, emitting TOOL_CALL events without TOOL_CALL_RESULT for frontend
tools so the CopilotKit runtime's processAgentResult handles them
- On HITL re-invocation (tool result sent back), sends full AG-UI
message history to aimock/LLM via convertMessages() for correct fixture
matching
- Fixes 3 HITL features: hitl-text-input, hitl-approve-deny, hitl-steps
(spring-ai now 6/11 D5 green, up from 3/11)

## Remaining 5 failures (pre-existing, not caused by this PR)

| Feature | Root cause |
|---------|-----------|
| tool-rendering | CopilotKit frontend re-render drops assistant
messages (count goes 1→0) |
| shared-state-read | SharedStateReadWriteController — separate
controller, not StreamingToolAgent |
| shared-state-write | Same as shared-state-read |
| subagents | SubagentsController — aimock returns "No fixture matched"
(404) |
| mcp-apps | Depends on external MCP server (excalidraw.com) unreachable
from Docker |

All 5 fail identically on baseline (main) without this PR's changes.

## Test plan

- [x] `showcase test spring-ai --d5` — 6 green, 5 red (same 5 red on
baseline)
- [x] Direct curl to Java backend confirms correct AG-UI event sequences
for all HITL flows
- [x] No regressions: agentic-chat, gen-ui-custom, gen-ui-headless still
pass
2026-04-30 09:02:43 -07:00
Ran Shemtov e8061b707b Merge branch 'main' into release/publish/monorepo/v1.56.5 2026-04-30 17:59:09 +02:00
Sam Julien 9099aa5d70 fix(shell-docs): retarget residual broken cross-links uncovered by visual QA
Follow-up to the docs cross-link sweep — visual QA surfaced four more sets
of broken links that the framework-scope middleware was rewriting into 404s
(or that resolved against the wrong path due to relative-link ambiguity).

- agentic-protocols/index.mdx: switch ./ag-ui, ./mcp, and ./a2a to absolute
  /agentic-protocols/* paths so they resolve regardless of trailing slash.
- faq.mdx: drop link wrapping on the three "Rich agentic experiences" row
  labels (Deep support for LangChain, Human-in-the-loop, Shared state).
  The /langgraph/* slugs were docs.copilotkit.ai legacy paths that don't
  exist in shell-docs; matches the playbook used for the V1 reference rows.
- integrations/langgraph/agent-app-context.mdx: retarget the "Frontend Data
  documentation" Callout link from /langgraph/agent-app-context to
  /langgraph-python/agent-app-context (the working slug).
- inspector.mdx: retarget the useAgentContext link in the Context row from
  /langgraph/agent-app-context to /langgraph-python/agent-app-context.
2026-04-30 08:56:08 -07:00
Sam Julien 30b68b12c2 fix(shell-docs): correct premium walkthrough on opt-in migrations Job
The Helm chart's `migrations.enabled` value defaults to `false`, but
the install walkthrough and overview prose both implied the
pre-install migrations Job always runs. Realign the docs with what the
chart actually does:

- intelligence-platform: rewrite the "Install" prose to describe the
  Job as conditional on `migrations.enabled: true`.
- self-hosting: add a new step in the install walkthrough between
  "Create secrets" and "Install the chart" that explains how to opt
  into migrations and when to leave them disabled. Update the
  "Verify the install" prose so it only claims a Completed Job if
  the reader opted in, and adjust the `--timeout` rationale to match.
2026-04-30 08:56:08 -07:00
Sam Julien 4be277055a fix(shell-docs): replace rogue gpt-5.2* with gpt-5.4* and extend CI validator
The model-name allowlist (`docs/model-allowlist.json`) ships
`gpt-5.4` and `gpt-5.4-mini` but never `gpt-5.2*` — the latter slipped
in during a model-bump cycle and was never caught because the CI
validator only scanned the legacy `docs/` tree.

- Sweep replace `gpt-5.2-mini` -> `gpt-5.4-mini` and `gpt-5.2` ->
  `gpt-5.4` across `showcase/shell-docs/src/content/` (~17 files).
- Extend `scripts/validate-doc-model-names.ts` with an
  `EXTRA_DOCS_DIRS` list so the validator now scans the shell-docs
  content tree alongside the legacy Nextra tree under `docs/`,
  preventing the same drift in future.
2026-04-30 08:56:08 -07:00
Sam Julien 61c7d8c55b fix(shell-docs): repair broken cross-page links
Sweep of dead links surfaced in the QA triage:

- agentic-protocols/index: retarget AG-UI / MCP / A2A links to the
  actual sibling slugs (`./ag-ui`, `./mcp`, `./a2a`); update the
  Generative UI table to point at the real `/generative-ui/*`
  pages and drop the link to the open-json-ui spec page (now hidden).
- multimodal-attachments: drop the orphaned migration Callout — the
  `/migration-guides/migrate-attachments` page does not exist.
- faq: remove `/reference/v1/*` links from the "What's available?"
  table; engineering direction is V2-only outside the reference area.
- agent-app-context: retarget the cross-link in the LangGraph guide
  and the Inspector reference to `/langgraph/agent-app-context` so
  they resolve.
- troubleshooting/common-issues: fix two broken `../X` links to
  point at the real `/backend/copilot-runtime` and
  `/built-in-agent/model-selection` slugs.
2026-04-30 08:56:02 -07:00
Sam Julien 4f795a13e3 fix(shell-docs): restore import lines on broken code blocks
Quickstart pages and a handful of high-traffic guides had their
`import { ... }` headers stripped, leaving bare identifiers above an
orphan `} from "..."` line — copy-pasting the snippets failed to
compile. Restore the missing import statements and add full imports
to layout.tsx / page.tsx blocks that were previously empty so each
block is independently copy-pasteable.

Also fixes the BIA-family quickstart V1 import path: `CopilotKit` is
exported only from the `/v2` entrypoint, so swap
`@copilotkit/react-core` for `@copilotkit/react-core/v2` in the BIA,
agent-spec, and Microsoft Agent Framework quickstarts.
2026-04-30 08:56:02 -07:00
Sam Julien ca95475607 fix(shell-docs): IA, sidebar, and HITL cleanup
Sidebar / IA / HITL cleanup from the shell-docs QA triage:

- Add `absent` mode to `<WhenFrameworkHas>` so MDX pages can declare a
  fallback branch for frameworks where a flag is null/missing, instead
  of collapsing to an empty middle.
- Use the new `absent` branch on `useInterrupt.mdx` and `headless.mdx`
  to point readers without `interrupt_pattern` at `useHumanInTheLoop`.
- Wrap the `useHeadlessInterrupt`-using "Driving it from plain UI"
  section in `headless.mdx` inside the native gate where the symbol is
  actually defined.
- Add `multi-agent/meta.json` so breadcrumbs / section labelling for
  `/multi-agent/subagents` use the explicit "Multi-Agent" title.
- Add a shared lead-in between `<InlineDemo>` and the gated branches in
  `agent-config.mdx`.
- Move `ag-ui-middleware.mdx` into `agentic-protocols/`, register it in
  the section's `meta.json`, link to the upstream AG-UI guide, and add
  a 302 from the old `/ag-ui-middleware` path.
2026-04-30 08:55:49 -07:00
Sam Julien f2b3c96c45 fix(shell-docs): reference accuracy + observability V2 translation + contributor docs rewrite (#4495)
## Summary

Updates incorrect/outdated documentation about real APIs. Reference
pages aligned with current package source; observability page
mechanically translated V1->V2; contributor onboarding rewritten to
point at shell-docs/Fumadocs/Nx instead of the legacy Nextra tree.

**Items addressed (from [triage
plan](https://app.notion.com/p/3523aa3818528128bcb0ee9e137cfff0)):**
- 9.3 — `LangChainAdapter.mdx`: `gpt-5.4` -> `gpt-4o`, drop the
misleading "auto-generated" header comment
- 19.1 — V2 hook reference pages: add missing `threadId` on `useAgent`,
fix `throttleMs` default cascade, add `lastRunAt` on `useThreads` Thread
shape
- 5.1 — `observability-connectors.mdx`: mechanical V1->V2 translation
(`<CopilotKit>` -> `<CopilotKitProvider>`, V2 import path, updated
`onError` event shape, optional server-side `CopilotObservabilityConfig`
section)
- 20.1 — `docs-contributions.mdx`: rewrite for shell-docs / Fumadocs /
Nx with real dev commands and ports

## Visual inspection

1. Start the dev server:
   ```bash
   nx run shell-docs:dev
   ```
   Open http://localhost:3003.

2. **Reference pages — accuracy.** Visit each page and verify the
documented shape matches the source:
- `/reference/v1/classes/llm-adapters/LangChainAdapter` — `## Example`
should show `model: "gpt-4o"`, not `gpt-5.4`. Page header should no
longer claim it's auto-generated.
- `/reference/v2/hooks/useAgent` — Parameters section should now list
`threadId`. `throttleMs` description should mention the provider-cascade
default.
- `/reference/v2/hooks/useThreads` — `threads` Return Value's Thread
shape should now list `lastRunAt`.

3. **Observability page — V2 translation.** Visit
`/troubleshooting/observability-connectors`:
   - Code blocks should use `<CopilotKitProvider>`, not `<CopilotKit>`.
   - Imports should be from `@copilotkit/react-core/v2`.
- The `onError` event shape should be `{ error, code, context }`, not
the V1 `CopilotErrorEvent` fields.
- `publicApiKey` and `publicLicenseKey` props should still appear (they
carry over to V2).
- A server-side section was added referencing
`CopilotObservabilityConfig` from the runtime.

4. **Contributor docs — read-through.** Visit the rendered
"Documentation Contributions" page (under the `(other)/contributing`
group). Walk through as if you're a new contributor:
- Clone instructions should point at `CopilotKit` root with Nx commands
run from there.
   - Dev port should be 3003.
   - Should mention Fumadocs (not Nextra).
   - Should mention `<Snippet>`-region authoring at a high level.
   - Should mention the pre-commit hook expectation.

5. **Anti-checks (should NOT have changed):**
   - The actual JSDoc source in `packages/` is unchanged.
- Other reference pages (e.g. `/reference/v2/hooks/useCapabilities`) are
unchanged.
   - The legacy `docs/` tree (the Nextra one) is unchanged.
2026-04-30 08:51:20 -07:00
Sam Julien 4793017a12 chore(shell-docs): hide and clean up stale pages (#4494)
## Summary

Hide and clean up stale shell-docs pages. Drops orphaned/broken/AI-slop
pages from nav, adds 302 redirects, deletes dead per-framework
overrides.

**Items addressed (from [triage
plan](https://app.notion.com/p/3523aa3818528128bcb0ee9e137cfff0)):**
- 1.1 — Tutorials (broken end-to-end; rewrite post-launch)
- 6.1 — `coding-agent-setup.mdx` (rename straggler)
- 10.1 — `copilot-suggestions.mdx` (orphaned broken stub)
- 11.1 — `generative-ui/open-json-ui.mdx` (AI-slop placeholder; rewrite
post-launch)
- 21.1 — `migrate/1.10.X.mdx` (~1-year-old migration target, no longer
supported)
- 3.4 — Legacy per-framework HITL overrides — see Concerns
- 16.1 — 3 orphan `state-inputs-outputs` / `workflow-execution` files in
adk/langgraph/llamaindex

All redirects are 302 (not 301) — we'll restore at the same URLs when
the affected pages are properly authored post-launch.

## Concerns

**3.4 (HITL overrides) intentionally NOT executed in this PR.** Audit
found that every per-framework HITL override carries substantive
framework-specific content beyond the canonical (e.g. `IframeSwitcher`
blocks pointing at fw-specific feature-viewer URLs, `CTACards` linking
to fw-specific sub-pages like
`/mastra/human-in-the-loop/interrupt-flow`,
`/crewai-flows/human-in-the-loop/flow`, etc.). Per the task's "be
conservative" guidance, none were deleted. If we want to land 3.4 the
right way, it likely needs a follow-up that either (a) merges the
fw-specific iframe demos into the canonical via per-fw branching, or (b)
explicitly designates these as legitimate per-fw overrides and accepts
them. Flagging for a separate decision.

## Visual inspection

1. Start the dev server:
   ```bash
   nx run shell-docs:dev
   ```
   Open http://localhost:3003.

2. **Sidebar disappearance.** Confirm the following are no longer in the
left sidebar:
   - "Tutorials" section (entire dividing block + sub-pages)
   - `Open-JSON-UI` under Generative UI > Declarative
   - `Migrate to 1.10.X` under Migrate
- `Coding Agent Setup` and `Chat Suggestions` (these were already not in
nav — visual check just confirms nothing changed)

3. **Redirect check.** Visit these URLs directly; each should 302 to the
destination:
   - `http://localhost:3003/tutorials/ai-todo-app/overview` → `/`
   - `http://localhost:3003/tutorials/multi-conversation-chat` → `/`
   - `http://localhost:3003/coding-agent-setup` → `/coding-agents`
   - `http://localhost:3003/copilot-suggestions` → `/`
- `http://localhost:3003/generative-ui/open-json-ui` → `/generative-ui`
   - `http://localhost:3003/migrate/1.10.X` → `/migrate`

4. **Shared-state per-fw check.** For adk, langgraph, llamaindex:
navigate to `/shared-state` (or the corresponding shared-state sidebar
entries). Confirm only the wired version (per `meta.json`) appears in
nav. The orphan URLs should 404 cleanly.

5. **Anti-checks (these should NOT have changed):**
   - All other docs pages render normally.
   - The canonical `human-in-the-loop` content (root) is unchanged.
   - Existing redirect rules in `next.config.ts` still work.
2026-04-30 08:45:45 -07:00
Jordan Ritter 78deb8b488 fix(showcase): spring-ai HITL features — frontend tool detection + multi-turn history
StreamingToolAgent now classifies tool calls as frontend vs backend by
comparing input.tools() (CopilotKit-injected frontend tools) against
the registered toolCallbacks (backend tools). Frontend-only tool calls
emit TOOL_CALL_START/ARGS/END without TOOL_CALL_RESULT so the CopilotKit
runtime's processAgentResult detects the missing result and executes the
frontend handler (useHumanInTheLoop, useFrontendTool).

On re-invocation (when the runtime sends back the tool result), the
agent now sends the full AG-UI message history to aimock/LLM via
convertMessages() so the fixture matcher sees the tool result and
returns a follow-up text response instead of repeating the tool call.

Fixes 3 HITL D5 features: hitl-text-input, hitl-approve-deny,
hitl-steps. Spring-ai now passes 6/11 D5 features locally (up from 3).
The remaining 5 failures (tool-rendering, shared-state-read/write,
subagents, mcp-apps) are pre-existing and unrelated to StreamingToolAgent
— they involve separate controllers, missing aimock fixtures, external
MCP servers, and CopilotKit frontend rendering issues.
2026-04-30 08:40:47 -07:00
Alem Tuzlak c34a28acf8 feat(showcase): D5 coverage for LangGraph Python (24 features, 20 new D5FeatureType literals) (#4491)
## Summary

Expands D5 (e2e-deep multi-turn) coverage for the LangGraph Python (LGP)
showcase integration from the existing 11 `D5FeatureType` literals to
**31 total** (20 new). 24 new D5 scripts cover features that previously
max'd out at D4 on the dashboard.

The plan and per-feature design lives in
[`.claude/specs/lgp-d5-coverage.md`](https://github.com/CopilotKit/CopilotKit/blob/blitz/lgp-d5-coverage-design/integration/.claude/specs/lgp-d5-coverage.md)
— start there for the bigger picture and per-feature turns sketches.

### What landed (8 commits)

- `docs:` — design plan with §1 summary, §2 new literal taxonomy, §3
per-feature plan table, §4 wave proposal, §5 open design questions, §6
anti-scope
- `S0` — scaffold 20 new `D5FeatureType` literals in
`harness/src/probes/helpers/d5-registry.ts`, mirror entries in
`d5-feature-mapping.ts` (`REGISTRY_TO_D5`) and
`shell-dashboard/src/lib/live-status.ts` (`CATALOG_TO_D5_KEY`). Adds
`beautiful-chat` and `shared-state-read` as additional registry-id
mappings to existing literals.
- `B1` chat-surface family — `chat-slots`, `chat-css`,
`prebuilt-sidebar`, `prebuilt-popup`
- `B2` platform family — `auth` (sign-out flow), `multimodal` (image+PDF
via sample buttons), `agent-config`
- `B3` frontend-tools + reasoning — `frontend-tools`,
`frontend-tools-async`, `reasoning-display` (covers
`agentic-chat-reasoning` + `reasoning-default-render`),
`tool-rendering-reasoning-chain`
- `B4` state — `shared-state-streaming`, `readonly-state-context`
- `B5` gen-UI — `gen-ui-declarative`, `gen-ui-a2ui-fixed`, `gen-ui-open`
(covers `open-gen-ui` + `open-gen-ui-advanced`), `gen-ui-agent`
- `B6` interrupt + BYOC — `interrupt-headless`, `gen-ui-interrupt`,
`byoc` (covers `byoc-hashbrown` + `byoc-json-render`)
- `F` — regenerated `aimock/d5-all.json` bundle (29 → 52 fixtures),
DOM-typing fix on chat-css probe

User-locked semantics encoded verbatim:
- **`auth`** — turn 1 chat works → click
`[data-testid=auth-sign-out-button]` → assert
`[data-testid=auth-demo-error]` or
`[data-testid=auth-demo-chat-boundary]` appears
- **`multimodal`** — click "Try with sample image/PDF" buttons → ask
"what is in this?" → assert assistant references attachment content
- **`chat-slots`** — send message → assert
`[data-testid=custom-assistant-message]` slot rendered
- **`prebuilt-sidebar`** / **`prebuilt-popup`** — surface root visible
(`.copilotKitSidebar` / `.copilotKitPopup`) → message → response inside
surface
- **`chat-customization-css`** — assert user-bubble bg contains `255, 0,
110` (hot pink) AND assistant-bubble bg `253, 224, 71` (amber)

`voice` is excluded (mic input not aimockable, deferred to a future
probe family).

## What's verified

- ✅ All 188 harness unit tests pass (30 test files including 24 new
ones)
- ✅ `pnpm typecheck` clean on `showcase/harness/`
- ✅ `aimock/d5-all.json` bundle regenerated cleanly with all 52 fixtures
- ✅ Pre-commit hooks pass (lint, format, package tests, commitlint)
- ✅ Each new D5 script self-registers via `registerD5Script()` and is
picked up by the e2e-deep driver's dynamic loader at boot

## What's NOT verified yet (verification gap — please read before
merging)

- ❌ **End-to-end run against the LGP integration was NOT executed.** A
full `showcase test langgraph-python --d5 --verbose --cycle` requires
the local docker compose to be brought up against this branch's
`aimock/d5-all.json` bundle, which would interrupt other running
services. We expect the first e2e pass to surface real issues — the most
likely failure modes per the design plan §5:
- `multimodal`: pre-fill hook for sample-button click is folded into the
assertion, which races the runner's automatic fill on turn 1. Likely
needs a runner-level `preFill` hook OR the demo to grow a query-string
attach trigger.
- `auth`: the post-sign-out 401 surfaces via `/info` refetch — if the
refetch isn't auto-triggered, the assertion will time out.
- `chat-css`: computed-style `background` shorthand resolution varies by
browser; the gradient parse may not flatten to the expected RGB
substring on all engines.
- `agent-config`: the `tone`/`expertise`/`responselength` keyword check
assumes the canned response surfaces them all — actual agent output may
differ.
- All transcript-keyword assertions are loose by design and will pass on
any response containing the keywords. They detect "no response" / "wrong
agent" but won't catch subtle behavioral regressions.
- ❌ **`cr-loop` was NOT run.** Sandbox blocks subagent file writes in
this environment, which would stall the cr-loop fix cycle. Recommend
running `cr-loop` in a fresh local session OR accepting reviewer
feedback on this PR directly.
- ❌ **8 open design questions** in `.claude/specs/lgp-d5-coverage.md` §5
— most are family-collapse vs split decisions (Q1
`beautiful-chat`/`agentic-chat`, Q3 `gen-ui-open` vs split, Q4 `byoc` vs
split, Q5 `reasoning-display` vs split). Defaults applied; flag in
review if any need reversing.

## Test plan (for reviewer / merger)

- [ ] Pull this branch locally
- [ ] `cd showcase && ./bin/showcase down && ./bin/showcase up`
(rebuilds aimock with the new `d5-all.json`)
- [ ] `./bin/showcase test langgraph-python --d5 --verbose --cycle` —
expect failures on first run; iterate on script/fixture/agent fixes
- [ ] Verify dashboard: features that were strikethrough `~~D5~~` should
now show D5 chips after the next 15-min probe tick on Railway

## Anti-scope

- D6 parity coverage is separate (post-D5).
- Other framework integrations (CrewAI, Mastra, LangGraph-TS,
Pydantic-AI) — not in this PR. The new `D5FeatureType` literals are
registry-wide; per-framework implementation is each framework's own
work.
- Voice probe family — needs a separate non-aimock probe shape,
deferred.
2026-04-30 17:23:10 +02:00
Alem Tuzlak fa4cc06b18 perf(showcase-dashboard): yield during fetch + parallelize initial pages
Follow-up to the matrix memoization PR. The initial dashboard load still
freezes for a few seconds because (a) the initial PB fetch is 10
sequential getList round-trips before any data arrives, and (b) when it
does arrive, the empty-map → populated-map transition forces every cell
in the matrix to re-render in one synchronous commit.

- Parallelize initial fetch: pull page 1 sequentially to learn
  totalItems, then fire pages 2..N concurrently via Promise.all. Wall
  time drops from `sum(rtt_per_page)` to roughly `max(rtt_per_page)`
  modulo network parallelism. PocketBase reads are independent so this
  is safe.
- Wrap initial setRows in startTransition. The first commit with real
  data is unavoidably a full-matrix re-render (per-key memo checks
  invalidate on every cell when the map flips from empty to populated);
  marking it as a transition lets React 19 yield to user input mid-walk
  instead of blocking the main thread for the duration. setStatus stays
  urgent so the "connecting → live" indicator still flips immediately.
- Wrap the SSE flush setRows in startTransition for the same reason —
  bursts that the 16ms coalescer can't fully absorb still need to yield.
2026-04-30 17:18:11 +02:00
Jordan Ritter a3bbe3effc perf(showcase-dashboard): memoize matrix cells + coalesce SSE deltas (#4502)
## Summary

The shell-dashboard matrix tab can spike CPU and freeze the browser on
load and during update bursts. Two compounding causes:

1. **Every PB SSE delta forces a full grid re-render.** The matrix
renders ~720 cells (18 integrations × ~40 features), and nothing in the
cell tree was memoized — a single status row update propagated a new
`ctx.liveStatus` Map identity through every cell, and identity-based
React.memo couldn't help because the Map ref changed.
2. **PocketBase fires the subscribe callback once per record.** A probe
finishing dozens of services, or the initial-state replay on reconnect,
produced N consecutive `setRows` calls and N React commits across
separate microtasks (auto-batching only catches synchronous bursts).

## Changes

- **`ComposedCell`** wrapped in `React.memo` with a custom
`arePropsEqual`. Compares `overlays`/`catalogCell` refs, ctx scalars,
and — when the `liveStatus` Map identity changes — only the 5 row keys
this cell actually reads (`health:slug`, `e2e:slug/feature`,
`smoke:slug`, `d5:slug/feature`, `d6:slug/feature`). Because
`upsertByKey` preserves row identity for unchanged keys, deltas that
don't touch a cell's `slug/featureId` short-circuit at the memo
boundary.
- **`useLiveStatus`** buffers SSE callbacks into a per-key `Map<string,
PendingOp>` and flushes via a single 16ms `setTimeout`. A burst of N
deltas now produces 1 React commit instead of N. Last-write-wins per
key. Buffer is cleared on effect teardown and on reconnect kickoff so
post-reconnect initial fetches never land on top of stale buffered rows.

## Why these two together

Throttling alone would still re-render every cell on each flush.
Memoization alone would still get invalidated on every delta because
each delta produces a new Map. Together: a burst of unrelated row
updates → 1 React commit, and within that commit only cells whose
specific rows changed actually re-render.

## What I have not changed

- The three independent `useNowTick` 1s timers on the Ops tab. They only
matter if the freeze reproduces on Ops, not Matrix; happy to consolidate
as a follow-up if needed.
- `OverlayColumnHeader` (rendered 18× in the header, calls `LevelStrip`
which does 4 Map lookups each). Same memo treatment would apply if these
still show up after this lands.
- No virtualization. Should not be necessary if the memo + throttle
combo works as expected.

## Test plan

- [x] Existing unit tests in the touched files pass —
`useLiveStatus.test.tsx` 15/15, `composed-cell.test.tsx` 9/10 (the 1
failure is a pre-existing test-vs-impl drift in the "renders 3 layers
when all active — docs deduped by health" case, verified by stashing my
changes and re-running).
- [x] No new TypeScript errors in the touched files.
- [ ] **Manual perf trace before/after on the live dashboard** — Chrome
DevTools → Performance tab → record while loading the dashboard against
a real PB feed. Expected: scripting time during initial load and during
probe-finish bursts drops sharply.
- [ ] **Functional sanity** — confirm cells still update on real status
changes (e.g. trigger a probe and watch the matching cell flip tone
within ~1 frame).
2026-04-30 08:00:25 -07:00
github-actions[bot] 8b5c6d2d6d style: auto-fix formatting 2026-04-30 14:43:44 +00:00
Alem Tuzlak 797e4177f2 Merge branch 'main' into blitz/lgp-d5-coverage-design/integration 2026-04-30 16:42:48 +02:00
Alem Tuzlak 2f649075b1 style: oxfmt 6 d5 probe files 2026-04-30 16:41:49 +02:00
Alem Tuzlak 61b8c663e3 perf(showcase-dashboard): memoize matrix cells + coalesce SSE deltas
The matrix tab can freeze the browser on load and during update bursts
because every PB SSE delta triggers a full re-render of all ~720 cells
(18 integrations × ~40 features), and PocketBase fires the subscribe
callback once per record — a probe finishing dozens of services or an
initial-state replay produces N consecutive React commits.

- ComposedCell: wrap in React.memo with a custom equality check that
  compares overlays/catalogCell refs, ctx scalars, and (when the
  liveStatus Map identity changes) only the 5 row keys this cell
  actually reads. Because upsertByKey preserves row identity for
  unchanged keys, deltas that don't touch a cell's slug/featureId
  short-circuit at the memo boundary.

- useLiveStatus: buffer SSE callbacks into a per-key Map and flush via
  a single 16ms setTimeout. A burst of N deltas now produces 1 React
  commit instead of N. Last-write-wins per key. Buffer is cleared on
  teardown and on reconnect kickoff so post-reconnect initial fetches
  never land on top of stale buffered rows.
2026-04-30 16:40:19 +02:00
Jordan Ritter 839aa1fb9d chore(showcase): bump @copilotkit/aimock to 1.16.4 (#4501)
## Summary

Bumps `@copilotkit/aimock` to 1.16.4 to pick up the router fix from
<https://github.com/CopilotKit/aimock/pull/148>. After this lands and
Railway rebuilds `ghcr.io/copilotkit/aimock:latest`, the showcase demos
pick up the fix automatically — no fixture changes required.

**The bug it fixes:** `match.toolCallId` was scanning the entire
conversation for the most recent `tool` message, so once any prior tool
result was in history, every subsequent request still had a "last tool
message" buried in the array. A stale `toolCallId` fixture could win and
shadow `userMessage` matchers for new user turns.

**User-visible symptom:** in `beautiful-chat`, clicking the first
suggestion (e.g. pie chart) worked, but clicking any subsequent
suggestion replayed the previous chart's "Pie chart rendered above —
Electronics is the largest slice…" content fixture instead of producing
a new tool call. Looked broken to anyone clicking through demos.

**Upstream fix:** the matcher now requires the tool message to be the
**last** message in the request — the only state in which the LLM is
being asked to respond to a tool result. Two regression tests in aimock
cover the "new user turn after tool" and "assistant content reply after
tool" cases.

## Changes

- `pnpm-lock.yaml` — refresh resolutions for `@copilotkit/runtime`'s
`@copilotkit/aimock` devDep (`1.16.2` → `1.16.4`). Mostly mechanical;
lockfile is slightly more compact than before.
- `showcase/scripts/package-lock.json` — refresh to `1.16.4` (was
`1.14.3`).
- `.github/workflows/test_e2e-showcase-on-demand.yml` — bump the
known-good floor from `@copilotkit/aimock@^1.14.3` to
`@copilotkit/aimock@^1.16.4` so the `/test-aimock <slug>` PR-comment
workflow always installs a build that contains the fix.

`packages/runtime/package.json` and `showcase/scripts/package.json` keep
their `"latest"` specifier per existing convention; the lockfile pins
are the reproducibility layer.

## Test plan

- [x] `pnpm nx run @copilotkit/runtime:test` — 1412/1412 pass (covers
`LLMock` / `MCPMock` imports from `@copilotkit/aimock` in the v2 MCP
integration tests)
- [x] `npm test -- aimock-fixtures` in `showcase/scripts` — 18/18 pass
(loadFixtureFile + validateFixtures schema validation)
- [x] Lefthook pre-commit (`check-binaries`, `sync-lockfile`,
`lint-fix`, `test-and-check-packages`) green
- [x] commitlint conventional-commit format green
- [ ] After merge: confirm Railway picks up the new
`ghcr.io/copilotkit/aimock:latest` image on next service restart and
beautiful-chat suggestions work end-to-end

The 71 unrelated test failures in `showcase/scripts` are pre-existing
Windows path-separator issues on `main` (audit CLI subprocess tests,
integration registry tests) — verified identical counts on origin/main
without this branch's changes. Out of scope for this bump.

## Related

- aimock PR: <https://github.com/CopilotKit/aimock/pull/148>
- aimock release:
<https://www.npmjs.com/package/@copilotkit/aimock/v/1.16.4>
2026-04-30 07:34:16 -07:00
Alem Tuzlak 4722d4f4da chore(showcase): bump @copilotkit/aimock to 1.16.4
Picks up the router fix from CopilotKit/aimock#148 — `toolCallId` matchers
now only fire when the tool message is the *last* message in the request,
preventing stale tool_call_ids from history shadowing `userMessage`
matchers on new user turns.

Surfaced as: in beautiful-chat, clicking a second suggestion replayed the
prior chart's "Pie chart rendered above…" content fixture instead of
producing a new tool call. Once Railway rebuilds `ghcr.io/copilotkit/aimock:latest`
and restarts the service, demos will pick up the fix automatically.

- Refresh `pnpm-lock.yaml` resolutions (workspace `@copilotkit/runtime` devDep)
- Refresh `showcase/scripts/package-lock.json` to 1.16.4
- Bump the floor in `test_e2e-showcase-on-demand.yml` from `^1.14.3` → `^1.16.4`
  so the `/test-aimock` PR-comment workflow always installs a build that
  contains the fix
2026-04-30 16:28:08 +02:00
Jordan Ritter 1a2f388863 fix(showcase): get built-in-agent to all-D5-green (10/10) (#4500)
## Summary

- Fixes all 10 D5 features for the built-in-agent TanStack integration
(was 6/10)
- Custom stream converter bypasses the runtime's `runFinished` blocking
(PR #4476) that breaks TanStack's multi-turn agent loop for server tools
- Forwards AG-UI frontend tools (useHumanInTheLoop, useRenderTool) as
TanStack definition-only declarations
- Fixes HITL TimePickerCard status case mismatch (CopilotKit v2 uses
lowercase "executing" not "Executing")
- Migrates hitl-in-chat demo to CopilotKit v2 with proper agentId wiring

## Test plan

- [x] `showcase/bin/showcase test built-in-agent --d5 --verbose` passes
10/10 (2 consecutive runs)
- [x] All 10 features: agentic-chat, gen-ui-headless, gen-ui-custom,
tool-rendering, shared-state-read, shared-state-write, hitl-text-input,
hitl-approve-deny, subagents, mcp-apps
2026-04-30 07:24:41 -07:00
Jordan Ritter 4107ab486b fix(showcase): get built-in-agent to all-D5-green (10/10)
Three fixes for the built-in-agent TanStack integration:

1. **Custom stream converter** — the runtime's `convertTanStackStream`
   (PR #4476) blocks all events after the first `RUN_FINISHED`, which
   breaks TanStack's multi-turn agent loop. Server tools like
   `get_weather` and `set_notes` need the loop to execute the tool,
   emit TOOL_CALL_RESULT, and re-prompt for the text response. Switch
   to `type: "custom"` with a local converter that skips RUN_FINISHED
   without blocking subsequent events, and deduplicates tool-call
   events (TanStack's buildToolResultChunks re-emits START/ARGS/END
   for server tool results).

2. **Frontend tool forwarding** — register AG-UI frontend tools
   (useHumanInTheLoop, useRenderTool, useFrontendTool) as TanStack
   definition-only declarations so the LLM can call them. Without
   this, tools like `book_call` were unknown to TanStack and silently
   dropped.

3. **HITL status case mismatch** — CopilotKit v2's ToolCallStatus uses
   lowercase strings ("executing") but the TimePickerCard checked for
   PascalCase ("Executing"). Normalize with toLowerCase() so buttons
   enable correctly.

Also migrates hitl-in-chat demo from CopilotKitProvider (v1) to
CopilotKit (v2) with proper agentId wiring.
2026-04-30 07:24:09 -07:00
Jordan Ritter d65b4cdbef fix(showcase): unblock dashboard D3 — pool zombie detection + e2e-demos abort release + deploy webhook schema (#4498) 2026-04-30 06:45:09 -07:00
Jordan Ritter 3db447c92e fix(showcase): llamaindex D5 all 11 features green (#4499)
## Summary

- Register `request_user_approval` FunctionTool stub in the llamaindex
`hitl_in_app_agent.py` so that `AGUIChatWorkflow` emits the necessary
`TOOL_CALL_CHUNK` AG-UI event for CopilotKit to intercept the tool call
and open the approval dialog
- Same pattern as the existing `book_call` stub in
`hitl_in_chat_agent.py`

## Why

`AGUIChatWorkflow` only emits AG-UI tool-call events for tools in its
`frontend_tools` registry. Without the stub, the aimock's
`request_user_approval` tool call was silently dropped by the workflow,
so CopilotKit never opened the approval dialog and the
`hitl-approve-deny` D5 test timed out.

## Test plan

- [x] `showcase/bin/showcase build llamaindex` succeeds
- [x] `showcase/bin/showcase test llamaindex --d5 --verbose` — all 11
features green (was 10/11)
- [x] Verified the fix follows the same pattern as `book_call` in
`hitl_in_chat_agent.py`
2026-04-30 06:41:57 -07:00
Jordan Ritter 63878c3b12 fix(showcase): register request_user_approval stub in llamaindex hitl-in-app agent
AGUIChatWorkflow only emits TOOL_CALL_CHUNK events for tools registered
via the frontend_tools constructor argument. Without a backend stub for
request_user_approval the workflow silently dropped the tool call from
the aimock response, so CopilotKit never intercepted it and the approval
dialog never opened.

Add a FunctionTool stub (same pattern as book_call in hitl_in_chat_agent)
and pass it in frontend_tools so the AG-UI event chain fires correctly.

All 11 llamaindex D5 features now pass locally.
2026-04-30 06:41:34 -07:00
Alem Tuzlak f9be5d56ba fix(showcase): add preFill hook to conversation runner; multimodal probe attaches before send
Adds an optional `preFill` callback to `ConversationTurn` that runs
before the runner fills the chat input and presses Enter for each turn.
Failure semantics mirror the post-settle `assertions` callback: a thrown
error records the turn as failed and stops the conversation.

Rewires `d5-multimodal.ts` to use `preFill` to click the sample image /
PDF buttons before each turn — fixing the false-green where the
multimodal probe's transcript-keyword assertion passed without the
attachment ever being sent.

Adds unit tests covering preFill ordering, failure mode, and a
no-preFill regression guard, plus tests verifying the multimodal
script wires `preFill` to the right sample-button selectors.
2026-04-30 15:36:05 +02:00
Alem Tuzlak 25902561c6 fix(showcase): unblock dashboard D3 — pool zombie detection + e2e-demos abort release + deploy webhook schema
Three changes that together restore D3 (e2e-readiness per-cell) emission for
half the dashboard cells after the pool sat on dead chromium instances and
the deploy webhook started 400-ing:

1. BrowserPool now detects dead browsers proactively. Once a chromium
   process died (OOM, crash, network blip), the dead Browser instance
   stayed in `available[]` and every probe that drew the slot failed with
   "browser.newContext: Target page, context or browser has been closed"
   — for hours/days, until the harness restarted or 100 release-cycles
   tripped contextCount-based recycle. Adds a `disconnected` listener per
   slot (registered in init() and recycleSlot()), an isConnected() check
   at the top of acquire() that skips zombies and recycles them, a
   release-time isConnected() check that catches the disconnect-event-
   pending race, and a per-slot recyclingSlots guard preventing double-
   relaunch when both paths fire concurrently. New unit tests use
   fake browsers via the existing test-injection point on the constructor.

2. createPooledE2eDemosLauncher now honours the driver's abort signal,
   mirroring the createPooledE2eDeepLauncher fix from ed0933e5c. Without
   this, an outer-timeout kept the pooled browser held until the orphaned
   driver promise drained all remaining demos — pool starvation across
   ticks. Tracks open contexts so abort closes them before releasing,
   uses a forceReleased flag so the driver's normal `browser.close()` in
   the finally block doesn't double-release, and wires the logger
   through the orchestrator registration.

3. /webhooks/deploy now accepts buildRunId / buildRunUrl. PR #4471 split
   build and deploy into separate workflows and started co-sending these
   fields, but `deployPayloadSchema.strict()` rejected them as unknown
   keys, returning 400. Every Showcase: Verify Deploy run has 400'd
   since. Adds both as optional, mirrors runUrl's http(s)-only
   refinement on buildRunUrl, and propagates them onto DeployResultEvent
   so downstream consumers can link a red deploy back to the build that
   produced its images.

Pre-push gates: oxfmt --check clean; tsc --noEmit clean; harness vitest
suite passes the 1367 platform-portable tests (the 19 Windows-specific
pre-existing failures — path separators in test fixtures + boot-wiring
test timeouts — are present on origin/main with the same shape, untouched
by this change); tsc -p tsconfig.build.json clean.
2026-04-30 15:31:44 +02:00
Jordan Ritter 17d68e5a0d fix(showcase): test built-in-agent with runtime PR #4482 pkg-pr-new build (#4485)
## Summary
Install pkg-pr-new build of @copilotkit/runtime (from PR #4482) in
built-in-agent, and align all demo page wiring with the LangGraph-Python
D5 fixtures.

**Runtime fix** — The `@copilotkit/runtime` dependency points to a
`pkg.pr.new` URL instead of the `next` npm tag. This MUST be reverted to
`"next"` once a proper @copilotkit/runtime release includes the TanStack
RUN_FINISHED fix from PR #4482.

**D5 wiring fixes** — Seven D5 features failed because built-in-agent
demo pages used different tool names and hook types than the LGP
reference agent:

- **tool-rendering**: switched from `useComponent("weather")` to
`useRenderTool("get_weather")` with correct `{parameters, result,
status}` render shape
- **gen-ui-tool-based**: renamed `useComponent("haiku")` to
`useComponent("generate_haiku")`, rewrote HaikuCard to accept args as
direct props
- **hitl-in-chat**: created `/demos/hitl-in-chat` route (was missing);
reuses TimePickerCard from hitl-in-chat-booking with `book_call` tool
- **shared-state**: added `set_notes` server tool so the fixture's
`set_notes` call has a backend handler
- **subagents**: renamed `delegate_to_planner`/`delegate_to_researcher`
to `research_agent`/`writing_agent`/`critique_agent` matching LGP
fixture tool names
- **subagents page**: updated `useComponent` registrations and
DelegationCard for the renamed tools

## Results
- Before: 3/10 D5 features pass
- After: targeting 10/10 (pending D5 probe verification)

## Test plan
- [x] Docker image builds
- [x] TypeScript check — no new errors (only pre-existing Zod version
mismatch)
- [x] All 6 files modified/created verified
- [ ] D5 probe run against deployed built-in-agent
- [ ] Team review of runtime change scope
- [ ] Revert runtime dep to `"next"` after proper release
2026-04-30 06:26:29 -07:00
Jordan Ritter dadd89e083 fix(showcase): fix llamaindex D5 failures (tool-rendering, gen-ui-headless, hitl-steps) (#4488)
## Summary

Fixes three D5 (e2e-deep) probe failures for the llamaindex integration:

- **tool-rendering**: `get_weather` was registered as a backend tool, so
the LlamaIndex AG-UI workflow emitted `TOOL_CALL_CHUNK` but never
`TOOL_CALL_RESULT`. CopilotKit's `useRenderTool` stayed stuck in loading
state and the WeatherCard never rendered with
`data-testid="weather-card"`. Fix: move `get_weather` to
`frontend_tools`, add `ToolCallResultWorkflowEvent` (subclasses
`ToolCallEndWorkflowEvent` to pass the AG_UI_EVENTS isinstance filter),
override `aggregate_tool_calls` to emit it for render-only tools, and
make WeatherCard always render its testid wrapper even during loading.

- **gen-ui-headless**: the shared agent was missing a `show_card`
frontend tool stub, so the workflow never emitted `TOOL_CALL_CHUNK` for
it. Fix: add `show_card` stub and register in `frontend_tools`.

- **hitl-steps**: `human_in_the_loop` was routed to the hitl-in-chat
specialized agent (which only has `book_call`), not the shared agent
(which has `generate_task_steps`). Fix: move `human_in_the_loop` to
`sharedAgentNames`. Introduce `render_only_tool_names` set to
distinguish render-only tools from interactive tools -- only render-only
tools get premature `TOOL_CALL_RESULT`; interactive tools let CopilotKit
manage the result lifecycle.

Also fixes two harness probe bugs:
- esbuild's `keepNames` transform injected `__name()` wrappers in
`page.evaluate()` browser context, causing `ReferenceError: __name is
not defined`. Replaced with string-based `new Function()` construction.
- `querySelector` only returned the first match; on headless chat pages
the first assistant message is an empty wrapper. Switched to
`querySelectorAll` with iteration.

## Test plan

- [x] `showcase/bin/showcase test llamaindex --d5 --verbose` passes
10/11 (hitl-approve-deny is pre-existing)
- [x] Two consecutive stable runs confirm no flapping
- [x] LangGraph-Python regression check confirms no breakage from
harness changes
2026-04-30 06:26:26 -07:00
Jordan Ritter cd8b46eeab fix(showcase/spring-ai): prevent streaming crash on frontend tool calls (#4487)
## Summary

- **StreamingToolAgent Phase 1 (streaming)**: Set
`internalToolExecutionEnabled=false` via `OpenAiChatOptions` so Spring
AI detects `tool_calls` in the stream without attempting execution
through the global `ToolCallingManager`. Previously,
`OpenAiChatModel.internalStream()` would auto-execute tool calls, and
when the model called frontend tools (like `generate_task_steps`,
`show_card`) that aren't registered on the Java backend,
`StaticToolCallbackResolver` returned null causing
`DefaultToolCallingManager` to throw `IllegalStateException`.
- **BoundedToolCallingManagerConfig**: Wrap `StaticToolCallbackResolver`
with `LenientToolCallbackResolver` that returns a
`FrontendToolPlaceholder` for unknown tools instead of letting the
resolver return null. This prevents crashes during Phase 2 (`.call()`)
when the model calls frontend tools injected by the CopilotKit runtime.
- **AG-UI event ordering**: Reorder event emission so tool call events
are emitted BEFORE `textMessageEnd`, which is required for the
frontend's `useRenderTool` to see them while the message is still open.

## Why

PR #4475 introduced `StreamingToolAgent` which uses Spring AI's
`.stream()` for real-time text delivery. However, Spring AI's
`OpenAiChatModel.internalStream()` auto-executes tool calls through the
global `ToolCallingManager`, which only knows about backend tools.
Frontend tools injected by the CopilotKit runtime caused
`IllegalStateException` crashes, breaking all D5 tests (0/11).

This fix recovers agentic-chat, gen-ui-custom, and gen-ui-headless D5
tests (3/11). The remaining 8 failures are pre-existing issues in other
controllers (subagents, shared-state, hitl) unrelated to the streaming
regression.

## Test plan

- [ ] `showcase test spring-ai --d5 --verbose` passes agentic-chat,
gen-ui-custom, gen-ui-headless
- [ ] Docker build succeeds locally
- [ ] No `IllegalStateException` in Spring AI container logs during D5
runs
- [ ] Frontend tool calls (generate_task_steps, show_card) handled
gracefully via placeholder
2026-04-30 06:26:22 -07:00
Jordan Ritter dca50b7dc7 fix(showcase): remove trailing slash from ms-agent-python hitl-in-app agent URL (#4486)
## Summary

- Remove trailing slash from the `hitl-in-app` agent URL in
ms-agent-python's CopilotKit route handler

## Why

The `hitl-in-app` agent was the **only** agent registered with a
trailing slash in the URL (`/hitl-in-app/`). The FastAPI backend mounts
the endpoint at `/hitl-in-app` (no slash). FastAPI's default
`redirect_slashes=True` returns a **307 redirect** for POST requests to
the trailing-slash variant, and the AG-UI `HttpAgent` does not follow
POST redirects during streaming. This caused the agent to appear
completely unresponsive — the D5 `hitl-approve-deny` probe timed out at
60s with `baseline=0, current=0` (zero assistant messages).

Verified via container: `POST /hitl-in-app/` returns 307 → `POST
/hitl-in-app` returns 422 (correct routing, body validation).

The fix uses the shared `createAgent("/hitl-in-app")` helper (which does
not append a trailing slash) for consistency with every other agent
registration in the file.

## Test plan

- [ ] D5 `hitl-approve-deny` passes for ms-agent-python (`showcase test
ms-agent-python --d5`)
- [ ] No regression in other ms-agent-python D5 features (10/11 → 11/11)
2026-04-30 06:26:18 -07:00
github-actions[bot] e6235fe304 style: auto-fix formatting 2026-04-30 13:22:12 +00:00
Alem Tuzlak bf93745f19 style: oxfmt d5-byoc.ts 2026-04-30 15:20:45 +02:00
Sam Julien 5151548235 fix(shell-docs): reference page accuracy + observability V2 translation + contributor docs rewrite 2026-04-30 06:19:30 -07:00
Sam Julien 7e0853daef chore(shell-docs): hide and clean up stale pages
Drop orphaned/broken/AI-slop pages from nav, add 302 redirects, and
delete dead per-framework override stragglers. Items addressed:

- 1.1: Tutorials section hidden (broken end-to-end; rewrite post-launch)
- 6.1: coding-agent-setup.mdx (rename straggler -> /coding-agents)
- 10.1: copilot-suggestions.mdx (orphaned broken stub)
- 11.1: generative-ui/open-json-ui.mdx (AI-slop placeholder)
- 21.1: migrate/1.10.X.mdx (~1yr-old migration target)
- 16.1: 3 orphan shared-state files in adk/langgraph/llamaindex
  (each meta.json wires only one of state-inputs-outputs vs
  workflow-execution; the other was a dead duplicate)

All redirects use permanent: false (302) so URLs can be restored at
the same paths once the affected pages are properly authored.
2026-04-30 06:17:28 -07:00
Alem Tuzlak 6cdc863815 Merge branch 'blitz/lgp-d5-coverage-design/integration' of https://github.com/CopilotKit/CopilotKit into blitz/lgp-d5-coverage-design/integration 2026-04-30 15:11:49 +02:00
Alem Tuzlak 0eb81b1406 fix(showcase): unbreak agent-config + byoc D5, reaching 31/31 green
agent-config: drop the AgentConfigLangGraphAgent subclass and use plain
LangGraphAgent. The subclass repacked CopilotKit provider properties
into forwardedProps.config.configurable.properties so the Python graph
could read them via RunnableConfig.configurable.properties — but
@ag-ui/langgraph@0.0.31 builds the LangGraph SDK request as
{ ..., config, context: { ...input.context, ...config.configurable } }
which merges configurable INTO context. LangGraph 0.6.0+ then rejects
with HTTP 400 'Cannot specify both configurable and context' on every
chat round-trip. Net effect: chat sent the user message, runtime 400'd,
no assistant response ever rendered. Removing the subclass unbreaks
the round-trip; the Python agent falls back to its DEFAULT_* constants
so the demo's frontend toggles no longer steer the system prompt
(known regression, tracked separately pending @ag-ui/langgraph fix
that decouples context from configurable).

byoc:
- D5 probe now sends the 'Sales dashboard' pill prompt (matches the
  fixtures added in main:f0a89b843 in feature-parity.json) instead of
  the previous generic 'render a byoc hashbrown' prompt that had no
  matching JSON-shaped fixture. Removed the now-obsolete byoc.json D5
  fixture file and regenerated the d5-all.json bundle (52 -> 50
  fixtures).
- Added data-testid='copilot-assistant-message' + data-message-role=
  'assistant' to the byoc-hashbrown and byoc-json-render renderer
  wrapper divs. The CopilotChat default assistantMessage slot includes
  these markers; overriding the slot with a custom JSON-rendering
  component dropped them, so the e2e-deep conversation runner's
  settle-detection cascade (which counts these selectors) never saw
  the response and timed out at 30s. Re-attaching the markers is a
  purely additive change that doesn't affect the renderers'
  behavior.
- D5 byoc assertion now waits for [data-testid='metric-card'] AND a
  chart (bar-chart or pie-chart) to render — a structural check on
  the BYOC contract output, not a transcript-keyword check that the
  custom renderer would never produce.

E2E status: 31/31 passing locally against
./bin/showcase up langgraph-python aimock with this branch's bundle.
2026-04-30 15:11:19 +02:00
github-actions[bot] 56b9cfd887 style: auto-fix formatting 2026-04-30 12:45:01 +00:00
Alem Tuzlak c9b513a948 fix(showcase): D5 e2e route + click + evaluate-body fixes (24->29 of 31)
Brings local D5 pass rate for langgraph-python from 24/31 to 29/31.

- preNavigateRoute added to chat-css, gen-ui-declarative,
  gen-ui-a2ui-fixed, readonly-state-context — featureType literals
  differ from registry IDs so default /demos/<featureType> 404'd.
- auth: <cpk-web-inspector> intercepts the sign-out click — added
  force: true. The demo doesn't auto-refetch /info on header change,
  so the post-sign-out 401 surface only appears after a probe send
  — assertion now triggers one before polling.
- chat-css: inline-only page.evaluate body to avoid esbuild's __name
  helper emit (undefined in the browser).

Remaining 2/31 known issues (documented in PR body for follow-up):
- byoc: agent runs but demo's hashbrown renderer expects streaming
  JSON; aimock plain-text canned response leaves nothing to show.
- agent-config: runtime route never forwards to LangGraph (no
  graph_id=agent_config_agent runs in the backend logs).
2026-04-30 14:43:31 +02:00
Alem Tuzlak 721cd98405 Merge branch 'blitz/lgp-d5-coverage-design/integration' of https://github.com/CopilotKit/CopilotKit into blitz/lgp-d5-coverage-design/integration 2026-04-30 14:10:39 +02:00
Alem Tuzlak a0abd81c6b Merge main into blitz/lgp-d5-coverage-design/integration 2026-04-30 14:09:58 +02:00
Alem Tuzlak b1c708887b fix(showcase): add aimock fixtures for byoc-json-render + byoc-hashbrown pills (#4492)
## Summary

- All six BYOC suggestion pills (3 per demo on `byoc-json-render` and
`byoc-hashbrown`) were getting eaten by older generic substring fixtures
in `showcase/aimock/feature-parity.json`. The `dashboard` rule swallowed
the Sales pills with placeholder text, and the `revenue by category as a
pie chart` / `monthly expenses as a bar chart` rules redirected the
json-render Revenue/Expense pills to a `render_pie_chart` /
`render_bar_chart` tool call from a different demo — leaving the BYOC
renderers staring at either canned text or empty content.
- Adds one fixture per pill, keyed on the **full pill prompt sentence**,
placed before the conflicting generic rules so first-fixture-wins picks
the right one. Each fixture returns the exact content shape the demo's
renderer expects: `{root, elements}` flat spec for json-render and
`{ui:[{<tag>:{props:{...}}}]}` for hashbrown.
- Removes two pre-existing duplicate fixtures from the bottom of the
file that had the right content but were placed after the conflicting
rules and never matched.

## Pill → fixture map

| Demo / pill | Match (full prompt) | Response shape |
|---|---|---|
| json-render — Sales dashboard | `Show me the sales dashboard with
metrics and a revenue chart` | flat spec, `MetricCard` with nested
`BarChart` |
| json-render — Revenue by category | `Break down revenue by category as
a pie chart` | flat spec, `PieChart` (4 segments) |
| json-render — Expense trend | `Show me monthly expenses as a bar
chart` | flat spec, `BarChart` (3 months) |
| hashbrown — Sales dashboard | `Show me a Q4 sales dashboard. Include a
total-revenue metric card, …` | ui kit: Markdown + metric + pieChart +
barChart |
| hashbrown — Revenue by category | `Break down Q4 revenue by product
category as a pie chart. …` | ui kit: pieChart, 4 segments |
| hashbrown — Expense trend | `Show me monthly operating expenses for
the last six months as a bar chart …` | ui kit: barChart, 6 months |

## Test plan

- [ ]
`https://showcase.copilotkit.ai/integrations/langgraph-python/byoc-json-render/preview`
— click each of `Sales dashboard`, `Revenue by category`, `Expense
trend`. Assert a chart/metric renders inside the assistant bubble (not
empty, not "Working on your request…").
- [ ]
`https://showcase.copilotkit.ai/integrations/langgraph-python/byoc-hashbrown/preview`
— same three pills. Sales should now render the multi-component
dashboard (Markdown + metric + pie + bar). Revenue/Expense should render
their single chart.
- [ ] Console: no `No renderer for component type:` warnings from
`@json-render/react`.
- [ ] `node -e "require('./showcase/aimock/feature-parity.json')"`
parses; no duplicate `userMessage` warnings from the aimock fixture
loader.
2026-04-30 14:08:00 +02:00
Alem Tuzlak 8e67d42a66 Merge branch 'main' into fix/showcase-byoc-aimock-fixtures 2026-04-30 14:07:33 +02:00
Alem Tuzlak a966999a21 fix(showcase): stop infinite tool-call loop in beautiful-chat + restore brand styling (#4490)
## Summary

Two related Beautiful Chat bugs in the integration showcases:

### Bug 1 — Infinite tool-call loop on every suggestion click

In `showcase/aimock/feature-parity.json`, 10 tool-calling fixtures
(`pieChart`, `barChart`, `render_pie_chart`×3, `render_bar_chart`×2,
`scheduleTime`, `search_flights`, `toggleTheme`) lacked an explicit `id`
on the `toolCall` and lacked a paired `toolCallId` followup fixture.

aimock's matcher (`router.js:36-46`) matches `userMessage` by
**substring** against the **last** user message. After the agent ran the
returned toolCall and re-prompted the LLM with the tool result attached,
the last user message was unchanged → the same fixture matched → the
same toolCall was returned forever. Repro: click any of the Pie Chart /
Bar Chart / Schedule Meeting / Search Flights / Toggle Theme suggestions
on `langgraph-python`'s `beautiful-chat` demo and watch the chart
re-render until manually stopped.

Fix follows the existing convention in the same file (compare the `Ada
Lovelace` / `show_card` pair):

- Added explicit `id` to each broken `toolCall`
(`call_fp_pie_chart_001`, `call_fp_bar_chart_001`, …)
- Added 11 new `toolCallId` followup fixtures at the top of the file
that match each id and respond with a brief `content` summary instead of
another toolCall

This shared file is loaded by every integration's `beautiful-chat` demo,
so the fix benefits all of them.

### Bug 2 — Black/white background split + broken CopilotKit logo

8 integrations use the full `ExampleLayout` pattern with `ThemeProvider`
and an `<img src="/copilotkit-logo.svg" />` (crewai-crews,
langgraph-fastapi, langgraph-python, langgraph-typescript, mastra,
ms-agent-dotnet, ms-agent-python, pydantic-ai). Two issues:

1. `globals.css` hardcoded `body { background: #fafaf9 }` and never
defined the brand tokens (`--background`, `--foreground`, `--card`,
`--primary`, `--secondary`, `--muted`, `--accent`, `--destructive`,
`--border`, `--input`, `--ring`, `--radius`) that the layout,
mode-toggle, todo card/column, and chart components reference via
Tailwind 4 arbitrary values. `ThemeProvider` was also adding `dark` to
`<html>` from system preference, so CopilotKit's chat went dark via
`[data-copilotkit] { background: var(--background) }` while body stayed
cream — that's the visible split the user reported.
2. `example-layout/index.tsx` renders `<img src="/copilotkit-logo.svg"
/>` but the file did not exist in any integration's `public/`.

Fix per integration:

- Added the full token set (light + dark) under `:root` and `:root.dark,
.dark`, registered the Tailwind 4 dark variant with `@custom-variant
dark (&:where(.dark, .dark *));`, switched `body` to
`var(--background)`/`var(--foreground)` so it tracks the same `dark`
class CopilotKit's chat is already responding to.
- Copied `copilotkit-logo.svg` and `copilotkit-logo-mark.svg` into each
integration's `public/` from
`examples/integrations/langgraph-python/public/`.

The 10 simplified-layout integrations (ag2, agno, built-in-agent,
claude-sdk-*, google-adk, langroid, llamaindex, spring-ai, strands)
don't render the logo or canvas split, so they don't need either fix.

## Verification

- `feature-parity.json`, `d5-all.json`, `smoke.json` parse cleanly
- `pnpm tsx showcase/scripts/validate-fixture-tool-surface.ts` → `104
fixtures × 603 demos — no drift`
- Programmatic check: every `toolCalls[].id` in `feature-parity.json`
now has a paired `toolCallId` followup fixture (no orphans)
- All 8 updated `globals.css` files have balanced braces and the
`:root.dark` block

## Test plan

- [ ] Click each of the 9 suggestion pills in `langgraph-python`
`beautiful-chat`; confirm tool-using suggestions no longer loop and end
with a text summary
- [ ] Toggle system theme → confirm chat side and canvas/body side stay
in sync (both light or both dark) on `langgraph-python` `beautiful-chat`
- [ ] CopilotKit logo renders in the chat header (no broken-image
placeholder)
- [ ] Spot-check one Pattern A integration (`langgraph-fastapi` or
`pydantic-ai`) and one Pattern B (`mastra` or `crewai-crews`) for the
same behavior
2026-04-30 14:05:20 +02:00
Alem Tuzlak 5c945b6b09 chore(showcase): unify unsupported demo convention (#4489)
## Summary

Aligns `langroid` and `claude-sdk-typescript` with the rest of the
showcase fleet for how they expose unsupported HITL/interrupt demos on
the dashboard matrix.

Today these two integrations list `gen-ui-interrupt` and
`interrupt-headless` in BOTH `not_supported_features` AND `demos[]`
(with stub pages that explain why the demo doesn't work). Every other
integration that marks these features unsupported (ag2, agno,
built-in-agent, claude-sdk-python, crewai-crews, google-adk, llamaindex,
mastra, pydantic-ai, spring-ai, strands) lists them only in
`not_supported_features` — the dashboard then renders a single no-entry
icon with a hover tooltip and no demo link.

This PR moves langroid and claude-sdk-typescript onto that same
convention.

## Changes

- `showcase/integrations/langroid/manifest.yaml` — drop
`gen-ui-interrupt` and `interrupt-headless` entries from `demos:` (kept
in `not_supported_features`)
- `showcase/integrations/claude-sdk-typescript/manifest.yaml` — same
- Delete the orphan stub `page.tsx` + `README.md` files under
`src/app/demos/gen-ui-interrupt/` and
`src/app/demos/interrupt-headless/` for both integrations (8 files
total)

## Test plan

- [x] `pnpm run validate-manifests` — all 18 integrations pass
- [x] `pnpm exec oxfmt --check <both manifests>` — clean
- [x] `pnpm exec nx run @copilotkit/react-core:test` — passes (one flaky
retry from the pre-commit hook resolved on second run, per Nx flaky-task
warning)
- [ ] Verify on the dashboard matrix that langroid and
claude-sdk-typescript render an icon-only cell with hover tooltip for
`In-Chat HITL (useInterrupt — low-level primitive)` and `Interrupt
(Headless)`, matching the other 12 integrations
2026-04-30 14:04:56 +02:00
Alem Tuzlak d748bc6a83 chore(showcase): ratchet validate-pins hash baseline
Pre-existing drift on main: count held at 134, hash shifted (one
FAIL healed, another regressed). Captured the new sorted-FAIL hash
locally with the same algorithm CI uses (sort -u | shasum -a 256)
and updated showcase/scripts/fail-baseline.json to match.
2026-04-30 14:03:12 +02:00
Alem Tuzlak c3f3631e58 fix(showcase/langroid): drop DEBUG warning leak + update stale tool-result tests
Two bugs introduced in 534cd1efa (D5 integration fixes) when the
langroid adapter switched backend tool results from TEXT_MESSAGE_*
triples to TOOL_CALL_RESULT events:

1. Two leftover `logger.warning("DEBUG ...")` calls in agui_adapter.py
   that fired on every chat turn — exactly the noise pattern the
   `test_plain_text_turn_does_not_warn` test was written to prevent.

2. Three stale tests in test_agui_adapter.py (only run on 3.11+, so
   the failure was 3.12-only):
   - test_backend_tool_execution_happy_path: still expected the old
     TEXT_MESSAGE_START/CONTENT/END inner triple instead of a single
     TOOL_CALL_RESULT event after TOOL_CALL_END.
   - test_backend_tool_exception_returns_sanitized_error: still read
     the sanitized error JSON from TEXT_MESSAGE_CONTENT.delta — it
     now rides on TOOL_CALL_RESULT.content.
   - test_plain_text_turn_does_not_warn: caught the DEBUG leak above.

Tests are skipped on Python 3.10 (langroid imports `typing.Self`),
which is why this only shows up on the 3.12 matrix entry.
2026-04-30 14:03:03 +02:00
Alem Tuzlak f0a89b8430 fix(showcase): add aimock fixtures for byoc-json-render + byoc-hashbrown pills
Both BYOC demo pills were getting eaten by older generic substring
fixtures ("dashboard", "revenue by category as a pie chart", "monthly
expenses as a bar chart") — the json-render pills came back as foreign
tool calls (render_pie_chart) and the hashbrown sales pill came back as
the "Working on your request..." placeholder, leaving both renderers
empty.

Add a dedicated fixture per pill keyed on the full pill prompt sentence,
placed before the conflicting generic rules so first-fixture-wins lands
on the right one. Each fixture returns the exact content shape the
demo's renderer expects:

  byoc-json-render -> {root, elements} flat spec for <Renderer />
  byoc-hashbrown   -> {ui:[{<tag>:{props:{...}}}]} for useJsonParser

Removes two pre-existing duplicates that had been added at the bottom of
the file but never matched because the conflicting generic rules came
first.
2026-04-30 13:54:29 +02:00
github-actions[bot] e1fc123332 style: auto-fix formatting 2026-04-30 11:36:20 +00:00