Commit Graph

228 Commits

Author SHA1 Message Date
Alem Tuzlak 7747033d68 fix(showcase/langgraph-python): mock d20 as 5 deterministic rolls
Replace the random roll_dice tool with roll_d20(value), which echoes
the LLM-supplied value back as the result. Aimock fixtures script the
five sequential calls returning [7, 14, 3, 19, 20] so the e2e suite can
assert exact values rather than rolling until 20 lands.

Update SYSTEM_PROMPT to allow multi-tool chaining when the user
explicitly asks for it (Chain tools pill emits 3 tool calls in one
turn).
2026-05-07 17:54:26 +02:00
Alem Tuzlak 6ce43eeaf6 test(showcase/langgraph-python): rewrite hitl-in-app and frontend-tools-async to genuine assertions
hitl-in-app — 7 explicit tests (1 skip):

- Page-load and pill-render tests retained.
- New refund #12345 approve/reject pair asserts the deterministic
  fixture leading phrases ("I am processing the $50 refund" vs
  "refund request was not approved").
- New escalate #12347 approve/reject pair asserts "Escalated ticket
  #12347" vs "Not escalated".
- The describe block runs in serial mode so the approve test
  (sequenceIndex 0 in the fixture) always runs before the reject test
  (sequenceIndex 1) for each pill.
- Downgrade #12346 stays skipped per spec (broken upstream as of
  2026-05-07).

frontend-tools-async — 4 explicit tests:

- Page-load test asserts composer + 3 pills.
- Project-planning, auth, and reading pills each click and assert the
  Notes DB card renders with the correct keyword heading and the
  per-note testid rows that the async handler returned. Anti-regression
  assertions catch the previous fixture-priority bugs (generic-plan
  boilerplate, showcase-assistant catch-all).
- Reading pill locks the full canonical shape per spec test #4: keyword,
  match count, note title, content lines, tag chip, and the assistant
  narration leading phrase.

No production-code testid changes — the existing dialog and notes-card
testids cover every assertion. Fixture work lives in d5-all.json (prior
commit).
2026-05-07 17:52:30 +02:00
Alem Tuzlak 390ab6fbbc test(showcase/langgraph-python): rewrite subagents to 5 deterministic 3-card tests
- Drop the 8 stale tests that asserted travel-planner shapes
  (Current Itinerary / supervisor-indicator / .bg-gray-50). They
  predated the supervisor + 3-subagent rewrite and could not catch
  the 3 production bugs.
- New suite (5 tests, 0 skipped):
    1) page loads with composer + 3 verbatim suggestion pills + 3
       subagent role indicators (testid-based).
    2-4) one test per pill (Write a blog post / Explain a topic /
       Summarize a topic). Each clicks the pill, waits for all
       three role-scoped subagent cards to reach data-status=
       "complete", then asserts each card's subagent-result is
       non-empty AND does not contain the showcase-assistant
       boilerplate fragments. Test 4 also serves as a regression
       gate on the delegations reducer fix — without it, that pill
       returns HTTP 400 and the cards never reach complete.
    5) clicks any pill, waits for terminal state, then asserts the
       critic-card count is exactly 1 and stays at 1 with status
       complete across a 5s dwell — catches any return of the
       supervisor -> critic loop.
- Aimock fixtures: 3 verbatim pill chains (cold exposure training,
  LLM tool calling, reusable rockets) added to both d5-all.json and
  the harness source d5/mcp-subagents.json. Each chain drives the
  full supervisor flow (turnIndex 0..3) plus three nested sub-agent
  fixtures so every Researcher/Writer/Critic card surfaces real
  prose instead of showcase boilerplate.
2026-05-07 17:52:02 +02:00
Alem Tuzlak a6f1076cf5 test(showcase/langgraph-python): rewrite open-gen-ui and open-gen-ui-advanced to iframe-presence assertions
Drop the 5 cross-origin contentFrame() / page.on('console', ...) skipped
assertions across the two specs — sandbox=allow-scripts only blocks host
introspection of the iframe DOM, and console-spying on the host page
catches no inner-iframe logs. Replace with iframe-presence assertions:
each pill click must produce iframe[sandbox*='allow-scripts'] with a
non-empty srcdoc (or src) attribute. That is the load-bearing signal
that the open-generative-ui pipeline mounted SOMETHING.

Rewrite each suggestion message string as a short verbatim label that
doubles as a deterministic aimock fixture key (paired with the new
fixtures in showcase/aimock/d5-all.json). Drop pill-title parentheticals
per the cosmetic note in lgp-test-genuine-pass.md so titles read as
natural human prompts; keep the message field aligned with the fixture
key.

Final test counts: 5 minimal (page-load + 4 pill-iframe), 4 advanced
(page-load + 3 pill-iframe). All .skip() removed. The sandbox-function
round-trip (evaluateExpression / notifyHost) is intentionally not
asserted here — that requires a same-origin sandbox option or a
host-side spy on the runtime's sandbox-function-call event, both
deferred to a follow-up.
2026-05-07 17:51:44 +02:00
Alem Tuzlak 9b3d64dda4 feat(showcase/langgraph-python): add subagent-card and subagent-result testids
- subagent-activity-card: emit data-testid="subagent-card-<role>" on
  each card wrapper (researcher | writer | critic), data-testid=
  "subagent-result" on the Result content div, and data-testid=
  "subagent-status" on the status pill. The previous testids
  (subagent-activity-card, subagent-activity-result) collapsed across
  roles, so the e2e suite couldn't count or content-assert per role.
- delegation-log: render a fixed row of 3 always-visible role
  indicators (data-testid="subagent-indicator-<role>") so the page
  exposes a stable hook for the load-state assertion regardless of
  whether the supervisor has delegated yet.
2026-05-07 17:51:40 +02:00
Alem Tuzlak 527e46c5d9 fix(showcase/langgraph-python): fix subagents delegations reducer, single-critic cap, and Result-echo boilerplate
- Annotate AgentState.delegations with operator.add reducer so concurrent
  sub-agent emissions in one supervisor step accumulate instead of
  raising INVALID_CONCURRENT_GRAPH_UPDATE (HTTP 400 on the Summarize pill).
- Update _delegation_update to return only the new entry (the reducer
  concatenates) instead of echoing the full prior list, which would
  duplicate entries each step under operator.add.
- Cap supervisor -> critique_agent loop at _MAX_CRITIQUE_ITERATIONS
  (default 1). Re-entrant critique calls short-circuit with a finish-now
  ToolMessage and do not append a second delegation, so the UI shows
  exactly one critic card per supervisor run.
- _invoke_sub_agent now walks messages newest-first and returns the
  first non-empty AIMessage content (handles list-of-content-blocks
  shape too). Prevents the previous failure mode where a final empty
  AIMessage made the card Result blank or echoed the showcase-assistant
  intro.
- Strengthen supervisor system prompt: each sub-agent must be called
  exactly once, with no further calls after critique returns.
2026-05-07 17:51:19 +02:00
Alem Tuzlak c7b65e9ec7 test(showcase/langgraph-python): rewrite headless-simple and headless-complete to deterministic pill-driven tests
Rewrites both headless specs to the pill-driven plan from
.claude/specs/lgp-test-genuine-pass.md (Family-1F).

headless-simple (4 tests):
1. page loads with custom composer + 3 pills visible
2. hello pill -> assistant bubble starts with the greeting leading phrase
3. joke pill -> assistant bubble contains the deterministic joke
4. fun fact pill -> assistant bubble contains 'Honey never spoils!'

headless-complete (5 tests):
1. page loads with custom composer + 4 pills visible
2. weather pill -> WeatherCard with Tokyo / Sunny / 68F + narration
3. AAPL pill -> StockCard with AAPL / $189.42 / +1.27% + narration
4. highlight pill -> HighlightNote with 'ship the demo on Friday' + narration
5. revenue chart pill -> ChartCard with 'Quarterly revenue', subtitle,
   month labels Jan-Jun + narration

Each test asserts on the headless-specific testids introduced in the
preceding commit, so a regression that demotes the headless surface
back to the default CopilotChat surface fails every tool test. Each
pill exercises a different render-hook path so regressions surface
test-by-test.

No .skip() — all 9 tests are live.
2026-05-07 17:46:47 +02:00
Alem Tuzlak cb1b8bfaf6 feat(showcase/langgraph-python): add headless surface testids
Adds stable data-testid hooks on the hand-rolled headless chat surface
shared by headless-simple and headless-complete so e2e specs can target
the headless surface (and not the default CopilotChat surface) by
selector.

Shared:
- 'headless-message-assistant' on the custom assistant bubble
- 'headless-message-user' on the custom user bubble
- 'headless-composer' on the composer container

headless-complete only:
- 'headless-weather-card' on the WeatherCard rendered via useRenderTool
- 'headless-stock-card' on the StockCard
- 'headless-highlight-card' on the HighlightNote rendered via useComponent
- 'headless-revenue-chart' on the ChartCard

If the headless surface ever silently regresses to the default
CopilotChat surface, the headless-specific testids are absent and the
spec fails. Each tool-card testid is scoped per-component so a regression
in a single render hook fails only that test.
2026-05-07 17:46:25 +02:00
Alem Tuzlak c3d8fdc14f test(showcase/langgraph-python): rewrite readonly-state-agent-context to 5 deterministic tests
Replace the previous 4-test (2 active, 2 skipped) suite with a
5-test deterministic plan keyed off the demo's actual published
context defaults and pill verbatim prompts.

Tests:
1. page loads — context-card + composer render
2. edits propagate to JSON — type into name/timezone, JSON updates
3. "Who am I?" pill — assistant reply leads with "I see you're Atai"
   and the identity card has name=Atai, timezone=America/Los_Angeles,
   avatar text=A (defaults from page.tsx)
4. activity checkboxes default-checked — "viewed the pricing page"
   and "watched the product demo video" are checked on first paint
5. "Suggest next steps" pill — assistant reply leads with "Since
   you recently viewed the pricing page and watched the product
   demo video"

Tests 3 and 5 rely on the deterministic aimock fixtures added in
the previous commit. Tests 4 uses the new activity-<slug> testids.
0 .skip() remaining.
2026-05-07 17:46:11 +02:00
Alem Tuzlak f1d55a40f5 feat(showcase/langgraph-python): add identity testids and aimock fixtures for readonly-state-agent-context
Add four production-code testids on demo-layout.tsx:
- identity-name on the user-name display
- identity-timezone on the timezone display
- identity-avatar on the first-letter avatar circle
- activity-<slug> on each activity <label>, kebab-case of activity name

These let Playwright lock the identity card against the published
context defaults (Atai, America/Los_Angeles, A) and assert the
default-checked activity rows ("Viewed the pricing page", "Watched
the product demo video") without depending on assistant text.

Add two verbatim-prompt aimock fixtures to showcase/aimock/d5-all.json
for the demo's suggestion pills:

- "What do you know about me from my context?" leading phrase
  "I see you're Atai, and you're in the America/Los_Angeles timezone.
  Recently, you viewed the pricing page and watched the product demo
  video."
- "Based on my recent activity, what should I try next?" leading phrase
  "Since you recently viewed the pricing page and watched the product
  demo video, ..."

Pinned to the verbatim message bodies so they don't contend with the
showcase-assistant catch-all fixture.
2026-05-07 17:45:34 +02:00
Alem Tuzlak 265c1567bf chore(showcase/langgraph-python): Phase 0 cleanup of orphans and dead probes
- delete orphan e2e specs (hitl, interrupt-headless,
  tool-rendering-reasoning-chain)
- rename reasoning-default-render → reasoning-default (route + describe)
- delete dead D5 probes (hitl-steps, interrupt-headless,
  tool-rendering-reasoning-chain)
- update d5-feature-mapping: rename agentic-chat-reasoning →
  reasoning-custom and reasoning-default-render → reasoning-default;
  repoint hitl + hitl-in-chat-booking to hitl-text-input (still used by
  ag2/agno/built-in-agent)
- update d5-registry: drop hitl-steps, interrupt-headless,
  tool-rendering-reasoning-chain feature types
- update d5-reasoning-display preNavigateRoute branch logic and tests
- update probes.test.ts and e2e-deep.test.ts to use surviving feature
  types
2026-05-07 17:18:21 +02:00
Alem Tuzlak bbe128ca65 fix(showcase): repair langgraph-python manifest highlight paths and bump test snapshots
Three highlight paths in langgraph-python's manifest pointed at files
that don't exist after the PR #4694 reorganization:
  - hitl-in-chat → src/agents/hitl_in_chat.py (actually hitl_in_chat_agent.py)
  - chat-slots → custom-welcome-screen.tsx (file doesn't exist; use slot-wrappers.tsx)
  - mcp-apps → copilotkit-mcp-apps/route.ts (actually .../[[...slug]]/route.ts)

The bundler walks every highlight at build time; one missing path aborts
the whole CI step. Fix all three.

Test snapshot counts in generate-catalog and generate-registry hardcoded
40 features / 720 cells / 702 total. With the two new feature IDs added
to the registry (reasoning-default + reasoning-custom), counts shift to
42 / 756 / 738; the LGP-specific cell distribution moved from
39 wired + 1 stub + 0 unshipped to 35 wired + 1 stub + 6 unshipped, and
the registry-side LGP feature/demo count drops to 36 (PR #4694 trimmed
4 items from the manifest's features list).
2026-05-07 14:30:45 +02:00
Alem Tuzlak 621c4e3bd1 fix(showcase/langgraph-python): copy manifest.yaml into runtime image
The demos layout's `generateMetadata` calls `headers()` (forces dynamic
rendering) and reads `manifest.yaml` at request time for per-demo titles.
The Dockerfile's runner stage didn't include the manifest, so every
`/demos/*` route in production threw a Server Components render error
(ENOENT on /app/manifest.yaml). The home page was unaffected because it
has no dynamic APIs and gets statically prerendered at build time.

Add a single COPY of `manifest.yaml` into the runner stage.
2026-05-07 13:55:39 +02:00
github-actions[bot] f9cef5048d style: auto-fix formatting 2026-05-07 08:35:24 +00:00
Tyler Slaton f85c1f8328 chore(showcase/langgraph-python): prune dead dependencies
The branch's iteration arc went through several headless-demo
implementations (CSS Modules, AI Elements, prompt-kit, shadcn) and a
streaming-markdown experiment with `streamdown`. After the spec
narrowed to shadcn-only and `react-markdown` for assistant rendering,
those component libraries and their satellite packages were left
installed but unused.

Removed (zero importers across `src/`):
- @copilotkit/react-ui (v1 UI; the showcase consumes /v2 hooks only)
- @rive-app/react-webgl2 (animations, never wired)
- @streamdown/cjk, /code, /math, /mermaid + streamdown (replaced by
  react-markdown for assistant content)
- @xyflow/react (graph viz, never wired)
- ai (Vercel AI SDK, was the AI Elements demo's transport)
- ansi-to-react (terminal output, never wired)
- marked (alternative markdown parser, never wired)
- media-chrome (media player UI, never wired)
- motion (animation library, was a prompt-kit demo dep)
- nanoid (we use crypto.randomUUID())
- react-jsx-parser (never wired)
- remark-breaks (markdown extension, never wired)
- shiki (syntax highlighting, never wired)
- tokenlens (never wired)
- use-stick-to-bottom (was an AI Elements dep)
- @radix-ui/react-use-controllable-state (no importers)

Kept everything actually imported: shadcn primitive surface intact
(class-variance-authority, cmdk, embla-carousel-react, lucide-react,
radix-ui umbrella + per-package @radix-ui/react-checkbox /
react-separator for beautiful-chat's local components), CopilotKit v2,
Hashbrown + json-render BYOC catalogs, recharts, react-markdown +
remark-gfm, openai (voice route), yaml (manifest parsing), zod.

Build still green at 54 routes. Lockfile updated. tsconfig.json
`jsx: "preserve"` — Next 15 reset this from `react-jsx` automatically
on build.
2026-05-06 23:32:47 -07:00
Tyler Slaton ebad989855 chore(showcase/langgraph-python): manifest, runtime, e2e, cleanup
Cross-cutting changes that don't belong with any one demo: manifest +
landing-page tags, runtime route adjustments, e2e + QA notes that
follow the demo renames, and a few small cleanups.

Manifest (manifest.yaml + src/app/page.tsx tag labels):
- Naming convention: every demo uses `Thing: Subthing` (Generative UI:
  Tool Rendering - Default / Custom Default / Specific; Open Generative
  UI: Default / Advanced; Shared State: Streaming / Read + Write;
  Reasoning: Default / Custom; Frontend Tools: In-App Actions / Async;
  Human in the Loop: In-chat / In-App / Interrupt based; Chat
  Customization: CSS / Slots; Headless UI: Simple / Complete).
- Retags: Auth → `platform`; HITL Step Selection + Interrupt-based →
  `interactivity`; Reasoning Default + Custom → `chat-ui`; Generative
  UI: Tools → `generative-ui`.
- Renames: Readonly State (Agent Context) → Frontend Context Sharing.
- HITL slot points at /demos/hitl-in-chat (working
  useHumanInTheLoop+interrupt path) instead of the previous
  /demos/hitl that had no backend `interrupt()` calls.
- Highlight paths corrected for the rebuilt headless demos (root-level
  paths replaced with hooks/, chat/, tools/, attachments/ subdirs).
- Descriptions rewritten where they had drifted from the implementation
  (gen-ui-agent: dropped useCoAgentStateRender claim;
  headless-simple: shadcn primitives, not raw Tailwind;
  headless-complete: enumerates the actual hooks wired).

Runtime / route:
- src/app/api/copilotkit/route.ts — 30 agents registered (incl. the
  reasoning-custom rename from agentic-chat-reasoning).
- copilotkit-mcp-apps/route.ts replaced with [[...slug]]/route.ts so v2
  subpath POSTs (/v2/agent/run) resolve.
- src/app/api/copilotkit-voice/[[...slug]]/route.ts — env var standardized
  (was `AGENT_URL || LANGGRAPH_DEPLOYMENT_URL`, now matches the rest
  of the showcase with just LANGGRAPH_DEPLOYMENT_URL); trailing `/`
  removed from deploymentUrl.

Tests / QA:
- e2e specs renamed and paths updated for the demo renames.
- qa notes for a2ui-fixed-schema (booked-state checklist removed) and
  byoc-json-render (Wave 4a residue removed).
- docs-links.json key renamed for reasoning-custom.

Cleanup:
- Removed remaining stub agent.py files in demo dirs (real graphs in
  src/agents/); removed dead beautiful-chat/components/headless-chat.tsx
  (zero importers); removed [A2UI-DEBUG] / [A2UI-RESPONSE] print
  statements from beautiful_chat.py; gpt-5.4-mini → gpt-5-mini typo
  fix in beautiful_chat.py:249 (would have 4xx'd every model call);
  stripped iframe-restriction LLM-prompt copy bleed from
  open-gen-ui-advanced suggestion titles.

The convention pass that ran across ~28 demos earlier in this branch is
already reflected in their per-demo commits — every page.tsx reads as
imports + provider + suggestions hook + JSX, with `useConfigureSuggestions`
extracted to a sibling suggestions.ts.
2026-05-06 23:19:57 -07:00
Tyler Slaton 7a380ecbeb feat(showcase/langgraph-python): Platform — Auth + Agent Config
Two demos that showcase how a platform team integrates CopilotKit:
auth gating and runtime config injection.

Auth — defaults UNAUTHENTICATED. First paint is a centered shadcn Card
with the demo token visible (`demo-token-123`) and a "Sign in" button.
<CopilotKit> doesn't mount until the user signs in; clicking the button
stores the token in localStorage and triggers a re-render that mounts
the chat with `Authorization: Bearer <token>` header attached. Reload
preserves the token; sign-out clears it and returns to the card.

Earlier drafts of this demo defaulted to authenticated with a
ChatErrorBoundary + onError-driven banner to recover from <CopilotChat>
's 401-on-mount crash. Both have been removed — gating the chat behind
`isAuthenticated` means it never mounts with bad creds, so the boundary
has no purpose.

Backend route uses `createCopilotRuntimeHandler` from
@copilotkit/runtime/v2 directly because the Next.js adapter does not
forward `hooks`. The `onRequest` hook validates the bearer token and
throws a Response(401) on missing/wrong tokens.

Agent Config — typed knobs (tone / expertise / responseLength) that
change the agent's behavior per turn. Pivoted to `useAgentContext` from
the original `<CopilotKit properties={...}>` transport, which silently
dropped the values in @ag-ui/langgraph 0.0.31 — those payloads landed
at the top level of the LangGraph stream and weren't routed into
RunnableConfig["configurable"]. A prior workaround that repacked them
there triggered LangGraph 0.6's "cannot specify both configurable and
context" 400.

useAgentContext is the supported LangGraph 0.6+ path for "frontend →
agent runtime context." A small ConfigContextRelay component sits inside
the provider and publishes the live toggles. The Python graph collapses
to a single static system prompt with three rulebooks; CopilotKitMiddleware
injects the context entry into the model's prompt automatically.
2026-05-06 23:19:16 -07:00
Tyler Slaton a78e6bddc7 feat(showcase/langgraph-python): Shared State + Sub-Agents
Four state-flow demos plus the multi-agent demo, all sharing the
page-as-entry-point convention with extracted suggestions.

- shared-state-streaming (Shared State: Streaming) —
  StateStreamingMiddleware(state_key="document", tool="write_document",
  tool_argument="document"). The argument name MUST match the state_key
  for the partial-JSON streamer to index correctly. Fixed in this pass
  (previous version had a name mismatch). Frontend renders `LIVE` badge
  + char counter so per-token streaming is visible.
- shared-state-read-write (Shared State: Read + Write) — bidirectional.
  UI writes preferences via `agent.setState`; agent writes notes via a
  `set_notes` tool that returns Command(update={...}). PreferencesInjector
  middleware reads the state on every turn and injects it as a system
  message. CopilotPopup layout, 2-col card UI, "Agent Scratch pad"
  copy.
- shared-state-read (deprecated route stub) — kept for back-compat with
  any external links; the canonical demo is shared-state-read-write.
- readonly-state-agent-context (Frontend Context Sharing) — frontend
  publishes read-only context via useAgentContext (the LangGraph 0.6+
  idiom). Backend has tools=[] and only CopilotKitMiddleware; read-only
  is enforced by the absence of any state-write tool, not a flag.

- subagents (Sub-Agents) — supervisor + research / writer / critic
  sub-agents. Per-tool useRenderTool registrations surface delegation
  events to the chat. State has a delegations[] log; the Python side
  only writes status="completed" so the type was tightened (dropped
  unused "running" / "failed" legs from both Python TypedDict and the
  TypeScript shape). Frontend infers active sub-agent from in-flight
  tool calls via a defensive structural probe over agent.messages.
2026-05-06 23:18:51 -07:00
Tyler Slaton a7a6a5c6fd feat(showcase/langgraph-python): Declarative UI + BYOC + Open Generative UI
A single commit for the "agent-authored UI" cluster — five distinct
strategies, all sharing a common shape (declare a catalog, let the
agent pick + populate components):

- declarative-gen-ui (Declarative UI: A2UI) — A2UI dynamic schema. The
  agent calls `generate_a2ui` (not the runtime's auto-injected
  `render_a2ui`) which secondary-binds an internal render tool with
  forced tool_choice, then returns operations via `a2ui.render(...)`.
  Custom catalog (Card / StatusBadge / Metric / InfoRow / PrimaryButton
  / PieChart / BarChart) wired via `a2ui.catalog` on the provider.
- a2ui-fixed-schema (Declarative UI: A2UI Fixed Schema) — fixed
  server-side schema. The "Book flight" button is an inert label; the
  earlier draft tried a schema swap to a booked-confirmation but the
  SDK doesn't yet expose `action_handlers` from Python. Removed
  BOOKED_SCHEMA + booked_schema.json since they were dead weight.
  Cleaned 8 (props as Record<string, any>) casts down to a single
  shared `s()` helper.
- mcp-apps — MCP server-driven UI via activity renderers. The runtime's
  `mcpApps.servers` config wires Excalidraw; agent has tools=[] and
  the middleware emits activity events that the built-in
  MCPAppsActivityRenderer auto-mounts as a sandboxed iframe.
- byoc-hashbrown (Declarative UI: Hashbrown) — streaming structured
  output via @hashbrownai/react. Agent prompt locks output to JSON via
  `response_format: json_object` with a `{ ui: [{ tag: { props } }] }`
  contract. Custom slot override on messageView.assistantMessage parses
  the streaming JSON.
- byoc-json-render (Declarative UI: json-render) — streaming hierarchical
  JSON UI spec via @json-render/react with a Zod-validated catalog.
  Catalog (defineCatalog) + registry (defineRegistry) split keeps the
  schema as the single source of truth.
- open-gen-ui (Open Generative UI: Default) — runtime's
  `openGenerativeUI` config injects a sandboxed UI tool; design-skill
  override steers the agent toward educational visualizations. Built-in
  OpenGenerativeUIActivityRenderer auto-mounts.
- open-gen-ui-advanced (Open Generative UI: Advanced) — adds frontend
  sandboxFunctions registered on the provider; each Zod-typed handler
  is exposed to the iframe via the host bridge. Suggestion titles read
  as normal user prompts (no iframe-restriction LLM-prompt copy bleed).

Backends sit at src/agents/{a2ui_fixed,byoc_hashbrown_agent,
byoc_json_render_agent}.py and the MCP runtime at
src/app/api/copilotkit-byoc-hashbrown/route.ts.
2026-05-06 23:18:18 -07:00
Tyler Slaton 8ad38cfdf2 feat(showcase/langgraph-python): Generative UI tool-rendering family
Five demos exercising the per-tool / catch-all / agent-state rendering
patterns. The three tool-rendering cells share the tool_rendering_agent
graph; they differ only in how the frontend renders the same tool
calls.

- tool-rendering (Tool Rendering - Specific) — per-tool useRenderTool
  for get_weather + search_flights, plus a useDefaultRenderTool wildcard
  for everything else.
- tool-rendering-default-catchall (Tool Rendering - Default) — single
  shadcn-styled wildcard via useDefaultRenderTool. Without registering
  *some* renderer the runtime has no `*` entry and tool calls render
  invisibly; this demo shows the minimum-viable shape.
- tool-rendering-custom-catchall (Tool Rendering - Custom Default) —
  same single-wildcard shape, branded with a custom card.

Backend system prompt (src/agents/tool_rendering_agent.py) defaults to
ONE tool per user question. Chaining is opt-in via a "Chain tools"
suggestion that triggers an explicit-ask exception in the prompt — the
previous default-on-chaining generated extra unsolicited tool-call
cards on every turn.

- gen-ui-tool-based (Generative UI: useComponent) — useComponent for
  render_bar_chart + render_pie_chart with Zod schemas; backend has
  tools=[] and the runtime injects the tools.
- gen-ui-agent (Generative UI: Agent State) — agent-state-driven step
  list. The Python graph plans steps via a `set_steps` tool that
  returns Command(update={"steps": …}); the frontend reads via
  useAgent({updates: [OnStateChanged]}) + a custom MessageList that
  renders steps inside CopilotChat's messageView.children slot.

Removed: src/app/demos/{tool-rendering,gen-ui-agent}/agent.py — TODO
stubs; real graphs live in src/agents/.
2026-05-06 23:17:45 -07:00
Tyler Slaton 9e40b1bf64 feat(showcase/langgraph-python): Frontend Tools + HITL family
Frontend Tools (in-app + async) demonstrate the spectrum of agent →
client tool calls:

- frontend-tools — useFrontendTool with a synchronous handler that
  mutates page state. The agent calls `change_background` with any CSS
  background value and the canvas re-paints. <CopilotSidebar /> layout;
  default background is solid indigo so the canvas reads as a clean
  start.
- frontend-tools-async — useFrontendTool with an async handler. The
  agent calls `query_notes`, the handler awaits a 500ms simulated DB
  query, and the agent uses the returned notes in its reply. Pure
  frontend tool — backend has tools=[].

HITL family — three patterns for human-in-the-loop, all keyed on the
useHumanInTheLoop / useInterrupt primitives:

- hitl — step-feedback variant rendering inside the chat via
  useHumanInTheLoop + useInterrupt (the v2 replacement for
  useLangGraphInterrupt, which is not exported in v2).
- hitl-in-chat — time-slot picker variant. Backend graph (hitl_in_chat)
  calls `interrupt({slots, …})` and resumes when the user picks. This
  is the canonical "in-chat HITL" demo; the manifest entry points here.
- hitl-in-app — async useFrontendTool with an app-LEVEL approval modal
  (rendered via createPortal OUTSIDE the chat). The completion callback
  resolves the pending tool Promise with the user's decision. This is
  HITL where the human surface is your app's UI, not the chat.
- gen-ui-interrupt — the lower-level useInterrupt primitive. Backend
  (src/agents/interrupt_agent.py) has a real `schedule_meeting` tool
  that emits an interrupt with topic / attendee / slots, and the
  frontend renders an inline TimePickerCard.
2026-05-06 23:17:21 -07:00
Tyler Slaton 3e157a1e02 feat(showcase/langgraph-python): Multimodal + Voice + Beautiful Chat
Three "production-feel" chat demos that exercise the same convention
(page.tsx as entry point + suggestions extracted) and add their own
specialized hooks:

- multimodal — image + PDF uploads via CopilotChat attachments. Includes
  a LegacyConverterShim (until @ag-ui/langgraph ships an updated
  converter), magic-byte + LFS-pointer guard for safe content sniffing,
  and sample-attachment-buttons that inject test images via DataTransfer
  + dispatch `change` event so screenshots / Playwright reproduce.
- voice — speech-to-text via @copilotkit/voice. Uses the V2 runtime
  directly with [[...slug]] catch-all because `transcriptionService` is
  V2-only. A guarded sample-audio-button injects deterministic sample
  text via the textarea's native value setter (CopilotChat has no
  controlled-input prop today). useSingleEndpoint={false} opts into the
  V2 multi-endpoint protocol.
- beautiful-chat — flagship polished starter chat with brand fonts,
  theme tokens, suggestion pills, generative-UI charts, and
  enableAppMode / enableChatMode tools. Backend
  (src/agents/beautiful_chat.py) wires query_data, manage_todos,
  search_flights, and an A2UI dynamic generator (generate_a2ui) that
  hits a secondary LLM for schema design. Model: gpt-5-mini.

The convention pass extracted suggestions into separate files for each
demo and slimmed page.tsx down to imports + provider + render.
2026-05-06 23:17:00 -07:00
Tyler Slaton 4737eb46f1 feat(showcase/langgraph-python): Prebuilt — CopilotChat / Sidebar / Popup
Three demos showing the prebuilt component formats. All three share the
neutral assistant graph and follow the page-as-entry-point convention:
each page.tsx slimmed to imports + provider + the suggestions hook
mount + JSX, with `useConfigureSuggestions` extracted to its own file.

- agentic-chat (Prebuilt: CopilotChat) — full-page CopilotChat. Plain
  text suggestions (joke / fun fact / limerick) — earlier drafts had a
  "Weather in Paris" prompt but the agent has no weather tool; trimmed
  to non-tool prompts so suggestions actually work.
- prebuilt-sidebar (Pre-Built: Sidebar) — <CopilotSidebar /> docked to
  the edge of the viewport. Page content centered in mx-auto max-w-2xl
  column with icon + heading + paragraph. The sidebar genuinely PUSHES
  the page now — see the body width fix in src/app/globals.css from the
  scaffolding commit.
- prebuilt-popup (Pre-Built: Popup) — <CopilotPopup /> with a floating
  launcher. Same centered content shape as the sidebar demo.

Removed: src/app/demos/agentic-chat/agent.py — TODO stub with a
misleading docstring; the real graph is at src/agents/agentic_chat.py.
2026-05-06 23:16:40 -07:00
Tyler Slaton ac89b11efd feat(showcase/langgraph-python): Reasoning - Default + Custom
A pair of demos that exercise the same backend reasoning graph but
differ only in whether the frontend overrides the
`messageView.reasoningMessage` slot.

Backend (src/agents/reasoning_agent.py): uses a reasoning-capable OpenAI
model (gpt-5-mini by default, override via OPENAI_REASONING_MODEL) routed
through the Responses API so the model's chain-of-thought streams as
AG-UI REASONING_MESSAGE_* events with `role: "reasoning"`. The prompt
asks for a concrete physics answer, which reliably triggers reasoning;
meta-prompts like "show your reasoning step by step" produce no
reasoning summary because the model recognizes those as a request to
reveal chain-of-thought (which it refuses).

Frontend:
- reasoning-default/ — no slot override; built-in
  CopilotChatReasoningMessage renders the "Thinking… / Thought for X"
  header with an expandable content region.
- reasoning-custom/ — overrides `messageView.reasoningMessage` with a
  ReasoningBlock (amber banner with `data-testid="reasoning-block"`).
  The label flips from "Thinking…" while streaming to "Agent reasoning"
  once the stream settles.

Suggestions live in their own files (per the page-as-entry-point
convention). Both demos share `agent="reasoning-default"` /
`agent="reasoning-custom"` against the same `reasoning_agent` graph,
registered in api/copilotkit/route.ts.

Removed:
- src/app/demos/agentic-chat-reasoning/ — replaced by reasoning-custom/
  for naming clarity.
- src/app/demos/reasoning-default-render/ — earlier draft of the Default
  demo with a slightly different page name.
- tests/e2e/agentic-chat-reasoning.spec.ts — replaced by
  reasoning-custom.spec.ts.
2026-05-06 23:16:13 -07:00
Tyler Slaton 4ff8dd608c feat(showcase/langgraph-python): chat customization (CSS + Slots)
Two paired demos showing the spectrum of "change the look" without
rewriting components:

CSS theming (chat-customization-css/) — HALCYON, a warm-paper editorial
brand. Two layers do the work:
  1. v2 token overrides on `[data-copilotkit]` recolor every Tailwind
     utility (cpk:bg-muted, cpk:text-foreground, …) the runtime renders.
  2. Class-targeted styling on .copilotKitChat, .copilotKitMessage*,
     .copilotKitInput, suggestions, scrollbar, welcome screen for the
     editorial details that CSS variables alone can't express.
Aesthetic: cream parchment surface, sharp 90° corners, copper-ember
accents, italic display serif (Instrument Serif) + Fraunces body +
JetBrains Mono dispatch, paper-grain noise via inline SVG, mono masthead
pinned under the top edge. All selectors namespaced under
`.chat-css-demo-scope` — no leakage.

Slot atlas (chat-slots/) — every overrideable slot wrapped in a
dashed-outline marker. Markers nest correctly: each shows ONLY its own
label on hover via `:has(.slot-marker:hover)` rather than lighting up
every nested label as the cursor enters the outermost one. Each label is
a click-to-copy button (✓ flash on success) that copies the slot's
PascalCase component path (`Input.TextArea`,
`MessageView.AssistantMessage`, …) so a developer can paste straight
into IDE search. SuggestionPill, ScrollToBottomButton, and Feather slots
are all wired. CustomFeather has its own `FeatherCopyLabel` because the
default Feather uses position:absolute and can't share SlotMarker.

Both demos use the neutral assistant graph (chat-customization-css and
chat-slots are entries in the route.ts neutralAssistantCells list).
2026-05-06 23:15:49 -07:00
Tyler Slaton bd9da67c48 feat(showcase/langgraph-python): rebuild Headless UI: Complete (modular)
Full headless surface — a hand-rolled CopilotChat replacement that wires
every render hook on top of shadcn/ui primitives. Visual chrome matches
Headless: Simple so the two read as a paired sibling demo.

Architecture is progressive-disclosure: the entry file is a 30-line
HeadlessCompleteRoot that enumerates capabilities, each registered via
a focused hook module:

  page.tsx
  hooks/
    use-tool-renderers.tsx     — useRenderTool x3, useDefaultRenderTool
    use-frontend-components.ts — useComponent (highlight_note)
    use-headless-suggestions.ts — useConfigureSuggestions
  chat/
    chat.tsx, header.tsx, empty-state.tsx, composer.tsx,
    suggestion-bar.tsx, message-list.tsx, message-user.tsx,
    message-assistant.tsx, message-activity.tsx, typing-indicator.tsx
  attachments/use-attachments-config.ts + attachment-preview.tsx
  tools/weather-card.tsx, stock-card.tsx, chart-card.tsx,
        generic-tool-card.tsx, highlight-note.tsx

Backend (src/agents/headless_complete.py): get_weather, get_stock_price,
get_revenue_chart tools; the chart tool replaces the previous Excalidraw
"Sketch a diagram" suggestion (MCP capability stays wired).

Bug fixes folded in:
- Tool-call cards stuck "running" forever — message-list.tsx indexes
  role:"tool" messages by toolCallId and passes the matching ToolMessage
  to renderToolCall so cards advance to "complete".
- Empty state was hugging the top of the viewport — Radix ScrollArea
  wraps content in a `display: table` div that breaks h-full propagation.
  Empty state now renders OUTSIDE the ScrollArea.
- SuggestionBar duplicated the empty-state prompts on first paint.
  Hidden until the conversation starts.
2026-05-06 23:15:26 -07:00
Tyler Slaton 83884b0be5 feat(showcase/langgraph-python): rebuild Headless UI: Simple
Minimum-viable headless chat that wires only `useAgent` + `useCopilotKit`,
dressed in shadcn/ui primitives. Five small single-purpose files so a
reader can grok the surface in under a minute:

  page.tsx        — provider + <Chat />, ~10 lines
  chat.tsx        — useAgent + useCopilotKit + send loop
  composer.tsx    — Textarea + send button
  empty-state.tsx — sparkles + sample prompts
  message-bubble.tsx + typing-indicator.tsx — render pieces

Wires runtimeUrl="/api/copilotkit" + agent="headless-simple" against
the neutral assistant graph registered in route.ts.

Also adds the v2 catch-all route at copilotkit-mcp-apps/[[...slug]]/route.ts
(used by Headless: Complete in the next commit). v2 hooks POST to subpaths
like /v2/agent/run; the previous flat route 404'd, leaving headless demos
stuck on "Thinking…".
2026-05-06 23:15:03 -07:00
Tyler Slaton 3ac7ac7977 chore(showcase/langgraph-python): scaffold showcase shell
Foundational layer that the per-demo work in subsequent commits builds on.

- Manifest-driven landing page (src/app/page.tsx) — auto-generated grid
  of demo cards from manifest.yaml, grouped by tag with explicit ordering
  for chat-ui / interactivity / generative-ui / agent-state / multi-agent
  / headless / platform.
- Per-demo titles (src/middleware.ts + src/app/demos/layout.tsx).
- Diagnostic console gated to NODE_ENV=production in app/layout.tsx so
  deployed showcases surface uncaught errors and iframe context.
- src/app/globals.css — Tailwind v4 @theme inline block that maps
  showcase CSS variables into Tailwind theme tokens (without it, shadcn
  utilities like bg-muted / text-foreground compile to nothing). body
  intentionally NOT given width:100% so <CopilotSidebar /> can shrink
  the document via marginInlineEnd.
- shadcn primitives under src/components/ui/ + src/lib/utils.ts.
- tsconfig.json — @/* path alias rooted at src/.

Removed:
- src/app/copilotkit-overrides.css (global override layer, superseded
  by per-demo theming)
- src/app/api/copilotkit-mcp-apps/route.ts (replaced with
  [[...slug]]/route.ts so v2 subpath POSTs like /v2/agent/run resolve)
2026-05-06 23:14:11 -07:00
Alem Tuzlak 48cd3b2c93 feat(showcase/voice): D5 mapping + sample-button bypasses /transcribe (#4674)
## Summary

- **Dashboard mapping fix.** `CATALOG_TO_D5_KEY` in
`showcase/shell-dashboard/src/lib/live-status.ts` was missing `voice →
["voice"]`, so `computeMaxPossible` capped the langgraph-python voice
cell at D4 even when the d5-voice probe row was green. The harness
`REGISTRY_TO_D5` already had the entry; only the dashboard mirror was
out of sync.
- **Sample-button decoupled from `/transcribe`.** The "Play sample"
button used to fetch `sample.wav` and POST it to the runtime's
transcription endpoint, which made the sample button and the mic
indistinguishable under aimock (both returned the canned transcription).
Reworked it into a synchronous static-text injector — sample button is
now a deterministic test/demo affordance, and the mic is the only path
that exercises real Whisper transcription. Synced across all 18
voice-enabled integrations. Phrase stays `"What is the weather in
Tokyo?"` so aimock's `weather in Tokyo` substring fixture still matches.
- **Probe-test parity.** Added the missing `d5-voice.test.ts` companion
(every other `d5-*.ts` script has one) — 9 tests covering registration,
`buildTurns`, `preFill` (sample-button click + textarea-poll path), and
the weather/Tokyo assertion.
- **QA + e2e cleanup** for langgraph-python: dropped the
no-longer-applicable "Transcribing…" mid-flight assertion and the `block
/demo-audio/sample.wav` error-state subsection. Other 16 integrations'
qa/e2e files follow in a parity sync PR.

## Test plan

- [x] `nx test @copilotkit/showcase-harness -- --run d5-voice` → 9/9
pass
- [x] `npm test` in `showcase/shell-dashboard` → 509/510 pass (1
skipped, 0 failed)
- [x] `nx build @copilotkit/showcase-harness` → clean
- [x] Local boot: `langgraph-cli dev` (port 8123) + `next dev` (port
3000) + dashboard (port 3002) — voice page at `/demos/voice` renders,
"Play sample" injects the canned phrase instantly, send → agent returns
weather, mic → real Whisper transcription with `OPENAI_API_KEY` set
- [ ] Reviewer: confirm the langgraph-python voice cell on the live
dashboard advances to D5 once the next d5-voice probe tick lands a green
row
2026-05-06 18:25:33 +02:00
Alem Tuzlak 728ed61ce8 feat(showcase/voice): D5 mapping + sample-button bypasses /transcribe
The langgraph-python voice cell sat at D4 even when its d5-voice probe
row was green. Root cause: the dashboard's CATALOG_TO_D5_KEY mirror in
showcase/shell-dashboard/src/lib/live-status.ts was missing voice ->
["voice"], so computeMaxPossible capped voice at D4 regardless of probe
state. The harness REGISTRY_TO_D5 already had the entry; only the
dashboard mirror was out of sync.

Separately, the "Play sample" button used to fetch sample.wav and POST
it to /transcribe. With aimock that meant both the sample button AND
the mic returned the same canned response, which made it impossible to
demo the mic path locally without conflating the two affordances.
Reworked the button into a synchronous static-text injector
(onTranscribed(sampleText)) so:

- Sample button = deterministic test/demo affordance, no runtime calls.
- Mic = real Whisper transcription via /transcribe.

Synced across all 18 voice-enabled integrations. Phrase stays "What is
the weather in Tokyo?" so aimock's "weather in Tokyo" substring fixture
still matches.

Also adds the missing d5-voice.test.ts companion (every other d5-* probe
script has one) and trims the langgraph-python qa/voice.md + e2e steps
that depended on the now-removed async behavior.
2026-05-06 18:11:24 +02:00
Alem Tuzlak 3eb53a8621 feat(showcase): D5 probe for headless-complete + extend headless-simple (langgraph-python)
Promotes /demos/headless-complete to its own D5 feature type so the
dashboard cell can reach D5 instead of riding on the headless-simple
probe (which was navigating to /demos/headless-simple regardless of
which catalog feature triggered it).

- New gen-ui-headless-complete D5 feature type + script that clicks
  each suggestion chip via preFill and asserts the right surface
  renders: WeatherCard (get_weather), StockCard (get_stock_price),
  HighlightNote (frontend useComponent), Excalidraw best-effort, and
  the canonical "Asia is the largest continent" text reply.
- Existing gen-ui-headless script now drives both turns by chip
  click (Profile card + Largest continent) instead of typing.
- Fixtures pin narration legs with both userMessage AND toolCallId
  and order them before the bare userMessage toolCall fixture —
  aimock's toolCallId matcher reads the LAST tool message in the
  request, but in a multi-turn probe that "last tool" stays on a
  previous turn's id until a new tool runs, which would otherwise
  hijack a later turn's prompt with a stale narration.
- headless-complete UserBubble + AssistantBubble now carry
  data-message-role so the harness conversation runner can detect
  message arrivals (mirrors the headless-simple convention).
- Mappings updated in lockstep:
    - REGISTRY_TO_D5:  headless-complete -> ["gen-ui-headless-complete"]
    - CATALOG_TO_D5_KEY (dashboard): same.
2026-05-06 16:21:43 +02:00
Alem Tuzlak 23d4770537 feat(showcase): hand-rolled headless-chat suggestion chips + parity across 17 integrations (#4669)
## Summary

Adds hand-rolled persistent suggestion chips to the `headless-simple`
and `headless-complete` demos in the langgraph-python north-star,
propagates the same surface to the other 17 showcase integrations, and
adds a deterministic aimock fixture so a new chip-click e2e test
("Largest continent") rounds-trips against a stable `Asia is the largest
continent…` response across all 18 demos.

## What changed

**Phase 0 — north-star (commit `7cbc5ea8`)**
- `showcase/aimock/feature-parity.json` — new fixture: `What is the
largest continent?` → `Asia is the largest continent — about 30% of
Earth's land area, home to over 4.6 billion people.`
- `langgraph-python/src/app/demos/headless-{simple,complete}/page.tsx` —
refactor `send` / `handleSubmit` to accept `(override?: string)` so chip
clicks dispatch synchronously without a `setInput` round-trip; render a
persistent `<div data-testid="headless-suggestions">` chip row above the
composer with 5 canonical entries; remove the dead
`useConfigureSuggestions` call from headless-complete (it was
registering suggestions nothing rendered).
- `langgraph-python/tests/e2e/headless-{simple,complete}.spec.ts` —
append one new test in each spec asserting chip click → user message →
`Asia` reply.

**Phase 1 — parity propagation across 17 integrations (commit
`4882c61f`)**
- Spec files `headless-simple.spec.ts` and `headless-complete.spec.ts`
are now byte-identical to the north-star in every integration (10 tests
each = 5 simple + 5 complete; verified via `cmp` for all 34 spec files).
- The 5-entry `suggestions` const is byte-identical between every
integration's simple and complete demos.
- All 17 integrations now expose the same selector surface (canonical
headings, empty-state text, `data-testid="headless-complete-messages"`,
dynamic placeholder, `rounded-br-sm` user bubble, no CopilotChat-default
testids).

**Glue preserved per integration** (verified by post-blitz code review):
- `built-in-agent`: `<CopilotKitProvider runtimeUrl="/api/copilotkit"
useSingleEndpoint>` + `agentId: "default"`
- `google-adk` / `llamaindex`: `agentId: "headless_simple"` /
`"headless_complete"` (Python-style underscores)
- `claude-sdk-typescript`: headless-complete
`runtimeUrl="/api/copilotkit-headless-complete"`
- `spring-ai`: 70-line `deduplicateMessages` adapter workaround +
`useMemo` import preserved verbatim
- All `@region[...]` markers preserved in place

**Adapter-specific decisions worth flagging in review:**
- `google-adk` headless-complete: rewrote `message-list.tsx` from
`msg-user`/`msg-assistant`/`agent-thinking` testid scheme to the
canonical `headless-complete-messages` wrapper; rewrote `input-bar.tsx`
placeholder to canonical dynamic; added the missing subtitle and
empty-state hint
- `ms-agent-dotnet`: extracted inline composer to a new `input-bar.tsx`
to match north-star structure
- `llamaindex`, `ms-agent-python`: added the canonical empty-state hint
(was missing entirely)
- `agno`, `built-in-agent`, `crewai-crews`, `mastra`, `ms-agent-dotnet`,
`pydantic-ai`: replaced per-integration empty-state hint with the
canonical Excalidraw line — chosen for parity over per-integration
accuracy (some demos don't actually wire an Excalidraw tool; alignment
was the explicit goal)

## Verification

- `validate-parity.ts`: 18/18 packages pass, 0 MUST failures
- `aimock-fixtures` test suite: 18/18 pass
- aimock fixture probed directly: `What is the largest continent?`
returns the canonical Asia response
- Each propagation slot reported `tsc --noEmit` clean (0 new errors) +
`playwright --list` shows all 10 expected tests
- Code review (`pr-review-toolkit:code-reviewer`) on the full diff: 0
Critical / Important / Minor findings, 1 stylistic nit (north-star
`input-bar.tsx` `onSubmit` type contravariant-loose, harmless)

## What was NOT done

Live per-integration Playwright runs against rebuilt Docker images. The
17 containers would each need a no-cache rebuild (~5-15 min each = hours
total) and the canonical local-test path is `showcase test <slug>` per
the existing CLI / CI pipeline. Static + structural verification covers
the propagation pattern.

## Test plan

- [ ] Run `showcase test <slug>` (or equivalent CI job) for at least one
drift-heavy integration: `google-adk` (testid scheme rewrite),
`built-in-agent` (provider glue), `spring-ai` (dedup workaround),
`llamaindex` (added testid + empty-state)
- [ ] Run the existing per-integration Playwright suites for at least
the north-star (`langgraph-python`) to confirm the new chip test passes
against a real backend + aimock
- [ ] Confirm aimock fixture validation still passes after deploy
2026-05-05 18:27:48 +02:00
Alem Tuzlak 602fb2d190 fix(showcase): trim headless-simple chips to in-surface set + add tool wildcard to google-adk/headless-complete
The validate-fixture-tool-surface check on PR #4669 flagged 18 drift
violations: every headless-simple demo carried 'Weather in Tokyo' /
'AAPL stock price' / 'Highlight a note' / 'Sketch a diagram' chips
that substring-match aimock fixtures returning tool calls
(get_weather / get_stock_price / highlight_note / etc.) — but
headless-simple demos only register 'show_card' via useComponent.
Tool-call dispatch had no matching renderer.

Trim the headless-simple chip list to two in-surface entries:
- 'Profile card' → 'Show me a profile card for Ada Lovelace' (existing
  show_card fixture; show_card is already registered by useComponent).
- 'Largest continent' → 'What is the largest continent?' (text-only
  fixture from Phase 0; no tool dependency).

The chip-click e2e test only asserts on the 'Largest continent' chip,
so the trim is test-compatible.

Headless-complete keeps the canonical 5-chip list (its tool surface
covers weather/stock/highlight/excalidraw via tool-renderers.tsx and
backend agents).

For google-adk/headless-complete: add a useDefaultRenderTool() wildcard
catch-all. The validator looks at page.tsx + hooks/* and a backend
agent file; google-adk's tool registrations live in tool-renderers.tsx
(unparsed) and there's no matching agents/headless_complete.py file,
so the validator saw an empty tool surface. The wildcard registers '*'
which matches every fixture tool — same pattern north-star already
uses in its own tool-renderers.tsx.
2026-05-05 18:03:03 +02:00
github-actions[bot] f6e184baa8 style: auto-fix formatting 2026-05-05 14:43:04 +00:00
Alem Tuzlak 8c7ea92bb1 fix(showcase/beautiful-chat): render A2UI surfaces (fixed + dynamic schema)
Search Flights and Sales Dashboard pills both produce visible surfaces
on the langgraph-python beautiful-chat demo. Three independent bugs were
masking each other:

- Flight TypedDict required `id` + `statusIcon`, which the aimock fixture
  doesn't supply. langchain rejected the call with `flights.0.id: Field
  required` and the agent surfaced the error string as the tool result.
  Made the type permissive (only the fields `_build_flight_components`
  reads need to be there).
- search_flights now expands flights into literal-children FlightCard
  components server-side instead of relying on the structural-children
  template form (the binder doesn't reliably expand it for our custom
  catalog — sibling demos avoid the form for the same reason).
- Sales Dashboard pill went into a tool-call loop because the
  userMessage+toolName fixtures matched both the initial call and the
  post-tool turn. Hoisted the toolCallId fixture above them so the
  follow-up turn returns content and breaks the loop.

Custom Row/Column reintroduced with `gap` support — the basic catalog's
versions ignore it, leaving cards squished. Children are array-of-strings
only (matches what the agent and fixture emit).

Two new e2e tests cover both pills end-to-end. 3s wait in beforeEach so
the v2 chat provider hydrates before the click dispatches. Full spec:
7/7 green.
2026-05-05 16:38:43 +02:00
Alem Tuzlak 7cbc5ea80e feat(showcase): hand-rolled suggestion chips + Largest-continent fixture in north-star headless demos 2026-05-05 14:22:13 +02:00
Alem Tuzlak 51db05f666 fix(showcase): emit reasoning events in langgraph-python and langgraph-fastapi (#4579)
## Summary

The `agentic-chat-reasoning` and `reasoning-default-render` cells in
`langgraph-python` and `langgraph-fastapi` never rendered any reasoning
content. Root cause: both agents were configured with `gpt-4o-mini` +
`use_responses_api=False`, so the underlying model produced no reasoning
content blocks and the Chat Completions API has no reasoning summary
surface in the first place. The frontend's `reasoningMessage` slot
stayed empty even though the cells are billed as reasoning demos.

This PR:

- Switches both agents (and their `tool_rendering_reasoning_chain`
siblings) to `gpt-5-mini` through the Responses API with
`reasoning={"effort":"medium","summary":"detailed"}`, mirroring the
`langgraph-typescript` and `pydantic-ai` agents that already worked.
Model is overridable via `OPENAI_REASONING_MODEL`.
- Updates the aimock `d5-all.json` fixture (and the matching harness
`reasoning-display.json`) to set the `reasoning` field on the `show your
reasoning step by step` match. Aimock now emits
`response.reasoning_summary_text.delta` events so the demo renders
deterministically without a real LLM call.
- Adds a `Show reasoning` `useConfigureSuggestions` pill on both
reasoning pages in both integrations so the demo is one click to
exercise.
- Tightens the `d5-reasoning-display` probe to also assert that a
reasoning-role message rendered (`[data-testid="reasoning-block"]` or
`[data-message-role="reasoning"]`), not just that the word "reasoning"
appears in the transcript.
- Un-skips the three streaming reasoning-block tests in
`agentic-chat-reasoning.spec.ts`, adds a suggestion-pill test, and
extends `reasoning-default-render.spec.ts` to cover the default
reasoning slot.
- Updates the `langgraph-python` QA doc to describe the new model +
Responses API setup and the pill flow.

Verified locally end-to-end: clicking the pill at
`/demos/agentic-chat-reasoning` renders the amber `ReasoningBlock` with
the fixture's reasoning text above the final answer bubble.

## Out of scope

Other integrations were audited and intentionally left alone:

- `langgraph-typescript`, `pydantic-ai` already use a reasoning model +
Responses API and work today.
- `agno`, `claude-sdk-python`, `ms-agent-python` use deliberate
workarounds (XML-tag reasoning + custom AGUI handler, Claude
extended-thinking deltas, `think` tool respectively) because their AG-UI
bridges either don't translate Responses-API reasoning items, run a
multi-call CoT loop incompatible with fixture replay, or don't emit
reasoning events at all.
- `llamaindex` uses `gpt-4.1` and surfaces reasoning inline as assistant
text. Its bridge (`llama-index-protocols-ag-ui`) does not translate
Responses-API reasoning items into AG-UI events; fixing that needs an
upstream patch and is out of scope here.

## Notes

Committed with `--no-verify` (explicit user request) — this worktree has
no `node_modules`, so the lefthook `test-and-check-packages` step
couldn't run locally. Changes are entirely under `showcase/` and CI runs
the same checks.

## Test plan

- [ ] CI fixture-validation passes on `showcase/aimock/d5-all.json`
- [ ] `showcase test langgraph-python --d5 --verbose` —
`reasoning-display` probe green (asserts `reasoning-block` selector +
keyword)
- [ ] `showcase test langgraph-fastapi --d5 --verbose` — same
- [ ] `nx run @copilotkit/showcase-langgraph-python:test:e2e -- --grep
reasoning` — un-skipped specs pass against the deployed Railway image
- [ ] Manual: visit `/demos/agentic-chat-reasoning` on a deployed
langgraph-python, click `Show reasoning`, confirm amber `REASONING —
Agent reasoning` block renders with italic step text above the final
answer bubble
- [ ] Manual: same on `/demos/reasoning-default-render`, confirm
CopilotKit's default `CopilotChatReasoningMessage` card renders
2026-05-01 13:40:10 +02:00
Ran Shemtov 41b7fa1cb3 Merge branch 'main' into chore/upgrade-langgraph-integration-demos 2026-05-01 13:35:10 +02:00
github-actions[bot] 6032b374c4 style: auto-fix formatting 2026-05-01 11:32:58 +00:00
Alem Tuzlak dca1b9894d fix(showcase): emit reasoning events in langgraph-python and langgraph-fastapi
The agentic-chat-reasoning and reasoning-default-render cells in
langgraph-python and langgraph-fastapi were configured with
gpt-4o-mini + use_responses_api=False, which never produces AG-UI
REASONING_MESSAGE_* events: gpt-4o-mini is not a reasoning model and
the Chat Completions API does not surface reasoning summary items at
all. The frontend's reasoningMessage slot was rendering nothing,
even though the cells were billed as "reasoning" demos.

- Switch both reasoning agents to gpt-5-mini (override via
  OPENAI_REASONING_MODEL) routed through the Responses API with
  reasoning={"effort":"medium","summary":"detailed"} so the model's
  chain of thought streams as content blocks that @ag-ui/langgraph
  translates into REASONING_MESSAGE_* events.
- Update the aimock d5-all.json and harness reasoning-display.json
  fixtures to include a "reasoning" field so aimock emits
  response.reasoning_summary_text.delta SSE events deterministically
  in CI without hitting a real LLM.
- Add a "Show reasoning" useConfigureSuggestions pill on both
  reasoning demo pages so the user can trigger the fixture-matched
  prompt with one click.
- Tighten the d5-reasoning-display probe: it now also asserts a
  reasoning-role message rendered via [data-testid="reasoning-block"]
  or [data-message-role="reasoning"], so a plain text response
  containing the word "reasoning" no longer falsely passes.
- Un-skip the three streaming reasoning-block tests in
  langgraph-python's agentic-chat-reasoning.spec.ts and add a
  suggestion-pill test; expand the reasoning-default-render spec to
  cover the default reasoning slot.
- Update the langgraph-python QA doc to describe the new model +
  Responses API setup and the suggestion-pill flow.
2026-05-01 13:30:28 +02:00
Alem Tuzlak f13c49f92f fix(showcase): drop hardcoded white chat background that broke dark mode (#4577)
## Summary

-
`showcase/integrations/{langgraph-python,langgraph-typescript,mastra,built-in-agent}/src/app/copilotkit-overrides.css`
(and the starter template that seeds new integrations) all forced
`.copilotKitChat { background-color: #fff !important; }`. The
`!important` won over the per-demo `ThemeProvider`, so the
`beautiful-chat` demo rendered a white chat panel in dark mode.
- `langgraph-fastapi` has no overrides file and was already correct —
this PR brings the other four to parity by deleting just the offending
rule (the `.copilotKitInput` border styles are kept).
- Updated `showcase/STYLING-GUIDE.md` with a warning so the example
block doesn't get pasted back in.

## Test plan

- [ ] Open `/demos/beautiful-chat` in `langgraph-python` with the OS in
dark mode — chat background follows the dark theme (no white panel).
- [ ] Same check for `langgraph-typescript`, `mastra`, and
`built-in-agent`.
- [ ] `langgraph-fastapi` unchanged (regression check on the working
baseline).
- [ ] Light mode in all four still renders correctly (chat picks up the
v2 light tokens).
2026-05-01 12:56:02 +02:00
Alem Tuzlak f396638c32 fix(aimock): add HITL 1:1-with-Alice fixture before broad Alice match (#4576)
## Summary

The hitl-in-chat demo's **"Schedule a 1:1 with Alice next week to review
Q2 goals."** suggestion was being intercepted by the broad `userMessage:
"Alice"` matcher used by the memory/context demo, which returns a
generic "Nice to meet you, Alice! I see you're in Tokyo — wonderful
city..." greeting. The HITL flow never fired and the user saw a
nonsensical reply.

Aimock's matcher uses `text.includes(match.userMessage)` (substring) +
first-fixture-wins by file order, so any message containing "Alice"
hijacked the suggestion before the HITL flow could trigger.

## Fix

Added a fixture pair earlier in `showcase/aimock/feature-parity.json`
with the **full suggestion sentence** as the matcher:

- `hasToolResult: false` → returns a `book_call` toolCall, letting the
frontend `useHumanInTheLoop` render the time-picker.
- `hasToolResult: true` → returns the booking confirmation message.

The substring-match-on-full-sentence is effectively exact — no other
realistic user message will contain that whole sentence — so the broad
`Alice` / `alice` fixtures stay scoped to the memory demo where the user
actually says "I'm Alice" or similar.

## Test plan

- [ ] Click "Schedule a 1:1 with Alice next week to review Q2 goals." in
the langgraph-python hitl-in-chat demo against an aimock-backed
deployment → expect the time-picker card to render and a booking
confirmation after picking a slot.
- [ ] The memory/context demo (where users type "I'm Alice") still gets
the Tokyo greeting — broad fixtures unchanged.
- [x] Pre-commit hooks pass (test, check-packages, commitlint).
2026-05-01 12:55:46 +02:00
Alem Tuzlak 25e03ef0e8 fix(showcase): drop hardcoded white chat background that broke dark mode
The shared `copilotkit-overrides.css` files in langgraph-python,
langgraph-typescript, mastra, built-in-agent, and the starter template
forced `.copilotKitChat { background-color: #fff !important; }`, which
won over the demo-level `ThemeProvider` and made the beautiful-chat
demo render a white panel in dark mode. langgraph-fastapi has no
overrides file and was unaffected — same fix gets the others to parity.

Also add a warning in showcase/STYLING-GUIDE.md so the example block
isn't pasted back in by the next contributor.
2026-05-01 12:43:15 +02:00
Alem Tuzlak 9845dadebb fix(aimock): re-key HITL confirmations on toolCallId so back-to-back flows work
Bug: in a single chat session, running both HITL booking flows
back-to-back (Alice 1:1 → then Sales call without refresh) used to
skip the time-picker on the second flow and jump straight to
"Booked ..." text.

Cause: confirmation fixtures were matched on `hasToolResult: true`,
which fires whenever the conversation has ANY tool message in
history. After the first flow finished, the second user message
short-circuited to a confirmation match before the second flow's
toolCall fixture (gated on `hasToolResult: false`) had a chance to
fire. The picker never rendered.

Fix: re-key the two confirmation fixtures on `toolCallId` (the
specific tool_call_id of the matching `book_call` invocation), which
only fires when the LAST conversation message is a tool result with
that id — exactly the moment we want the confirmation. Drop the
`hasToolResult: false` constraint on the toolCall fixtures so they
match a fresh user request regardless of prior tool history.

Add a back-to-back regression test to all 17 hitl-in-chat specs:
walk Alice flow to completion, then sales flow without refresh,
assert two `time-picker-card` elements rendered. If the multi-flow
regression returns, the second card never appears and the test
fails at `toHaveCount(2)`.
2026-05-01 12:42:53 +02:00
Ran Shem Tov 84af438694 chore: use latest cpk 2026-05-01 12:31:05 +02:00
Ran Shem Tov 8bae258b84 chore: fix peripherals for smoke tests and parity 2026-05-01 12:31:04 +02:00
Ran Shem Tov 0b41bebe23 chore: fix showcase drift 2026-05-01 12:31:04 +02:00
Alem Tuzlak 8cb84e88eb test(showcase): replicate hitl-in-chat regression spec across all 17 integrations
The hitl-in-chat demo ships in 17 integrations (langgraph-python plus
16 others — mastra, strands, ag2, agno, crewai-crews,
langgraph-typescript, langgraph-fastapi, pydantic-ai, llamaindex,
langroid, claude-sdk-python, claude-sdk-typescript, ms-agent-python,
ms-agent-dotnet, spring-ai, google-adk). All shipped placeholder e2e
specs that only checked the chat input was visible — none exercised
the actual booking flow.

Replace each with the full booking-flow spec written for
langgraph-python:
1. The "Schedule a 1:1 with Alice" suggestion renders the time-picker
   card AND the Tokyo greeting is absent (regression guard against
   the broad aimock `userMessage: "Alice"` matcher).
2. Picking a slot transitions to the picked-state card and produces
   a "Booked … Alice" assistant follow-up.
3. The "Book a call with sales" suggestion runs the same flow with
   the sales attendee.

Also add the matching aimock fixture pair for the sales suggestion
in feature-parity.json — without it, case 3 would only pass against
real OpenAI, not the aimock-backed CI deployments. The pair mirrors
the Alice fixture pair: `book_call` toolCall on first turn,
confirmation message after the picker resolves.

Per-integration coverage matters because each integration has its
own framework-specific HITL wiring (`useHumanInTheLoop` binding to
the agent, agent-side tool registration, run streaming protocol)
that can regress independently of the shared aimock fixture.
2026-05-01 12:25:36 +02:00
Alem Tuzlak 846a8a8938 test(showcase): add hitl-in-chat regression spec for Alice 1:1 suggestion
Pins the contract that the new full-sentence aimock fixture pair beats
the broad `userMessage: "Alice"` matcher:

1. Sending the suggestion `"Schedule a 1:1 with Alice next week to
   review Q2 goals."` renders `[data-testid="time-picker-card"]`,
   not the Tokyo greeting. The test explicitly asserts the Tokyo
   greeting is absent — `toHaveCount(0)` against
   `/Nice to meet you, Alice/i` — so any future broad-match
   regression fails here loudly.
2. Clicking a slot transitions to `[data-testid="time-picker-picked"]`
   and the assistant follow-up message contains "Booked ... Alice",
   verifying the `hasToolResult: true` branch of the fixture pair
   also wires through.
2026-05-01 12:17:24 +02:00
Alem Tuzlak e79cac1208 fix(showcase): unblock gen-ui-agent recursion + drop deepagents wrapper
Verified end-to-end against a local langgraph-python stack: agent now
walks plan → step1 in_progress → step1 completed → ... → final summary,
and the frontend renders a single inline progress card that updates in
place all the way to "All 3 steps complete".

Two real changes pulled out from the verification round:

1. agent.py: drop the `deepagents.create_deep_agent` wrapper for the
   plain `langchain.agents.create_agent` ReAct loop. The deepagents
   planner / sub-agent / write_todos middleware ate enough supersteps
   per turn that the run regularly tripped LangGraph's recursion
   limit before the agent could publish all three step transitions.
   The plain ReAct loop is one superstep per LLM/tool call, and
   `state_schema=GenUiAgentState` is supported directly so the
   middleware-only state-extension hack is gone.

2. route.ts: bake `recursion_limit: 100` into every LangGraphAgent
   via `assistantConfig`. `with_config({"recursion_limit": ...})` on
   the compiled Python graph does NOT propagate when the graph is
   served via the langgraph runs API — the wrapper is invisible to
   the assistant config the server hands to Pregel, which then falls
   through to langchain_core's hard-coded default of 25. Setting
   `assistantConfig.recursion_limit` on the JS side makes the limit
   travel with every run kicked off through this route, regardless
   of what the Python graph thinks its config is.
2026-05-01 11:47:40 +02:00