Commit Graph

51 Commits

Author SHA1 Message Date
Alem Tuzlak 9845dadebb fix(aimock): re-key HITL confirmations on toolCallId so back-to-back flows work
Bug: in a single chat session, running both HITL booking flows
back-to-back (Alice 1:1 → then Sales call without refresh) used to
skip the time-picker on the second flow and jump straight to
"Booked ..." text.

Cause: confirmation fixtures were matched on `hasToolResult: true`,
which fires whenever the conversation has ANY tool message in
history. After the first flow finished, the second user message
short-circuited to a confirmation match before the second flow's
toolCall fixture (gated on `hasToolResult: false`) had a chance to
fire. The picker never rendered.

Fix: re-key the two confirmation fixtures on `toolCallId` (the
specific tool_call_id of the matching `book_call` invocation), which
only fires when the LAST conversation message is a tool result with
that id — exactly the moment we want the confirmation. Drop the
`hasToolResult: false` constraint on the toolCall fixtures so they
match a fresh user request regardless of prior tool history.

Add a back-to-back regression test to all 17 hitl-in-chat specs:
walk Alice flow to completion, then sales flow without refresh,
assert two `time-picker-card` elements rendered. If the multi-flow
regression returns, the second card never appears and the test
fails at `toHaveCount(2)`.
2026-05-01 12:42:53 +02:00
Alem Tuzlak 8cb84e88eb test(showcase): replicate hitl-in-chat regression spec across all 17 integrations
The hitl-in-chat demo ships in 17 integrations (langgraph-python plus
16 others — mastra, strands, ag2, agno, crewai-crews,
langgraph-typescript, langgraph-fastapi, pydantic-ai, llamaindex,
langroid, claude-sdk-python, claude-sdk-typescript, ms-agent-python,
ms-agent-dotnet, spring-ai, google-adk). All shipped placeholder e2e
specs that only checked the chat input was visible — none exercised
the actual booking flow.

Replace each with the full booking-flow spec written for
langgraph-python:
1. The "Schedule a 1:1 with Alice" suggestion renders the time-picker
   card AND the Tokyo greeting is absent (regression guard against
   the broad aimock `userMessage: "Alice"` matcher).
2. Picking a slot transitions to the picked-state card and produces
   a "Booked … Alice" assistant follow-up.
3. The "Book a call with sales" suggestion runs the same flow with
   the sales attendee.

Also add the matching aimock fixture pair for the sales suggestion
in feature-parity.json — without it, case 3 would only pass against
real OpenAI, not the aimock-backed CI deployments. The pair mirrors
the Alice fixture pair: `book_call` toolCall on first turn,
confirmation message after the picker resolves.

Per-integration coverage matters because each integration has its
own framework-specific HITL wiring (`useHumanInTheLoop` binding to
the agent, agent-side tool registration, run streaming protocol)
that can regress independently of the shared aimock fixture.
2026-05-01 12:25:36 +02:00
github-actions[bot] 374b85bec4 style: auto-fix formatting 2026-05-01 08:06:25 +00:00
Jordan Ritter bba219102b style(showcase): format voice route files 2026-05-01 00:55:03 -07:00
Jordan Ritter e89107f8e3 fix(showcase): add voice agent backends and audio assets
- Add dedicated tool-free voice agents for strands, llamaindex,
  ms-agent-python (aimock returns tool calls when tools are registered,
  which the adapters don't loop on)
- Add sample_agent alias to langgraph-typescript langgraph.json
  (was only in dev-mode config)
- Add SampleAudioButton and voice route to google-adk
- Add sample.wav to agno, ms-agent-dotnet, ms-agent-python, google-adk
2026-05-01 00:52:30 -07:00
Jordan Ritter 6e35b71135 fix(showcase): port 9 voice routes from V1 to V2 multi-route handler
V1 CopilotRuntime in single-route mode rejects multipart/form-data
with 415 Unsupported Media Type. Port all 9 integrations to V2
createCopilotRuntimeHandler which handles the /voice sub-route
natively.

Integrations: claude-sdk-python, claude-sdk-typescript, crewai-crews,
llamaindex, ms-agent-dotnet, ms-agent-python, pydantic-ai, spring-ai,
strands
2026-05-01 00:52:19 -07:00
Jordan Ritter 3dc802a927 fix(showcase): remove from __future__ import annotations from 3 llamaindex agents
The `from __future__ import annotations` import causes all type
annotations to be stored as strings rather than evaluated at
definition time. When the LlamaIndex AGUIChatWorkflow validates
backend_tools via Pydantic, `Annotated[str, "..."]` parameters
fail with "class-not-fully-defined" because Pydantic cannot
resolve the deferred string annotations.

This caused RUN_ERROR on every request to these three agents,
which the D5 conversation runner cannot detect as a response,
leading to the 30s timeout.

Affected agents:
- tool_rendering_reasoning_chain_agent.py (4 backend tools)
- a2ui_fixed.py (display_flight backend tool)
- a2ui_dynamic.py (generate_a2ui backend tool)

Verified locally: all three endpoints now stream
TEXT_MESSAGE_CHUNK events and finish with RUN_FINISHED.
2026-04-30 17:24:59 -07:00
github-actions[bot] 751eb7d389 style: auto-fix formatting 2026-04-30 17:04:56 -07:00
Jordan Ritter faac42c313 fix(showcase): wire byoc-hashbrown backend agents correctly
- agno: add default agent alias + per-request runtime
- langgraph-fastapi: add default agent alias
- llamaindex: fix agent name mismatch (byoc_hashbrown → byoc-hashbrown-demo)
- mastra: create dedicated byocHashbrownAgent with hashbrown system prompt
  (was using weatherAgent which produced plain text instead of JSON)
- ms-agent-dotnet: upgrade byoc page to V2 CopilotKit import
2026-04-30 17:04:40 -07:00
Jordan Ritter d36660ba24 fix(showcase): add D5 probe testid to byoc-hashbrown across all integrations
The D5 conversation runner detects assistant responses via
data-testid="copilot-assistant-message". The byoc-hashbrown demo
overrides the assistantMessage slot with a custom HashBrown renderer,
which dropped that attribute. Without it the harness sees 0 messages
and times out.
2026-04-30 17:04:39 -07:00
Jordan Ritter d6b784ee9a feat(showcase): add interrupt demos to 12 integrations via Strategy B
Replace gen-ui-interrupt and interrupt-headless "not supported" stubs
with working demos using useFrontendTool + async Promise pattern.
Backend agents use system prompt + tools=[] — CopilotKit runtime
routes tool calls to the frontend handler. Pattern proven by
ms-agent-python/dotnet, now extended to ag2, agno, built-in-agent,
claude-sdk-python, claude-sdk-typescript, crewai-crews, google-adk,
langroid, llamaindex, mastra, pydantic-ai, strands.
2026-04-30 15:59:00 -07:00
Jordan Ritter 7db463a208 fix(showcase): wire all three slot overrides in llamaindex chat-slots demo
The llamaindex chat-slots demo only had the welcomeScreen slot,
missing the input.disclaimer and messageView.assistantMessage
overrides that langgraph-python (gold standard) provides. The D5
probe checks for data-testid="custom-assistant-message" after the
assistant responds, so the feature was stuck at D4.

Add CustomAssistantMessage and CustomDisclaimer components (matching
langgraph-python), wire them into CopilotChat props, update the e2e
spec to cover all three slots, and update manifest highlights.
2026-04-30 14:31:35 -07:00
Jordan Ritter 63878c3b12 fix(showcase): register request_user_approval stub in llamaindex hitl-in-app agent
AGUIChatWorkflow only emits TOOL_CALL_CHUNK events for tools registered
via the frontend_tools constructor argument. Without a backend stub for
request_user_approval the workflow silently dropped the tool call from
the aimock response, so CopilotKit never intercepted it and the approval
dialog never opened.

Add a FunctionTool stub (same pattern as book_call in hitl_in_chat_agent)
and pass it in frontend_tools so the AG-UI event chain fires correctly.

All 11 llamaindex D5 features now pass locally.
2026-04-30 06:41:34 -07:00
github-actions[bot] 2ed24a1cb5 style: auto-fix formatting 2026-04-30 08:49:30 +00:00
Jordan Ritter f0952e7656 fix(showcase): fix llamaindex D5 failures for tool-rendering, gen-ui-headless, hitl-steps
Three D5 probe failures fixed:

tool-rendering: get_weather was a backend tool so the workflow emitted
TOOL_CALL_CHUNK but no TOOL_CALL_RESULT, leaving useRenderTool stuck in
loading state forever. Move get_weather to frontend_tools so CopilotKit
manages the tool call lifecycle. Add ToolCallResultWorkflowEvent (which
subclasses ToolCallEndWorkflowEvent to pass the AG_UI_EVENTS isinstance
filter) and override aggregate_tool_calls to emit TOOL_CALL_RESULT for
render-only tools. The WeatherCard component now always renders with
data-testid="weather-card" even during loading, instead of switching to
a separate div that lacks the testid.

gen-ui-headless: the shared agent was missing a show_card frontend tool
stub, so the workflow never emitted TOOL_CALL_CHUNK for it. Add the
stub and register it in frontend_tools.

hitl-steps: human_in_the_loop was routed to the hitl-in-chat specialized
agent (which only has book_call), not the shared agent (which has
generate_task_steps). Move human_in_the_loop to sharedAgentNames so it
routes through the default agent. Introduce render_only_tool_names set
to distinguish render-only tools (get_weather) from interactive tools
(generate_task_steps, book_call, show_card) — only render-only tools
get premature TOOL_CALL_RESULT; interactive tools let CopilotKit manage
the result lifecycle to keep buttons enabled.
2026-04-30 01:47:37 -07:00
Jordan Ritter c1775f37ca fix(showcase): apply FixedAGUIChatWorkflow to all llamaindex HITL agents
PR #4464 only applied the upstream bug fixes to hitl_in_chat_agent.py.
The same three bugs (duplicate tool-call rendering, missing
parent_message_id, incorrect tool-result message roles) affect the
hitl_in_app_agent and the main agent router.

Switch both to use FixedAGUIChatWorkflow via workflow_factory so all
llamaindex demo features get the corrected AG-UI event stream.
2026-04-29 23:19:46 -07:00
Jordan Ritter 0d4cdc6625 fix(showcase): fix llamaindex hitl-text-input D5 test failure
The LlamaIndex AGUIChatWorkflow has three bugs that caused the
hitl-text-input test to fail:

1. ToolCallChunkWorkflowEvent emitted without parent_message_id,
   causing the AG-UI client to create a duplicate assistant message
   (one from MESSAGES_SNAPSHOT, one from dechunked TOOL_CALL_START).
   The duplicate produced 8 time-picker buttons instead of 4.

2. _snapshot_messages embeds toolCalls in the MESSAGES_SNAPSHOT AND
   emits separate TOOL_CALL_CHUNK events, so the client registers
   the same tool call twice on the same message.

3. ag_ui_message_to_llama_index_message converts AG-UI ToolMessages
   to ChatMessage(role='user') instead of role='tool', so the second-
   leg OpenAI API call has no tool-result message and aimock's
   hasToolResult matcher fails.

Fix: subclass AGUIChatWorkflow (FixedAGUIChatWorkflow) that:
- Assigns a stable ID to assistant messages and passes it as
  parent_message_id on ToolCallChunk events
- Strips toolCalls from MESSAGES_SNAPSHOT (letting TOOL_CALL_CHUNK
  be the sole mechanism)
- Corrects tool-result message role from 'user' to MessageRole.TOOL
  so the OpenAI API receives proper tool-result messages
2026-04-29 22:02:39 -07:00
Jordan Ritter 534cd1efa7 fix(showcase): D5 integration fixes across 12 frameworks
Per-framework fixes to pass D5 e2e-deep probes:
- agno: deduplicate agent_server routes
- claude-sdk-python: handle ParsedContentBlockStopEvent (SDK v0.97+)
- claude-sdk-typescript: remove orphan tool-rendering page
- crewai-crews: add backend tool_rendering agent + shared_state fix
- google-adk: add AGUIToolset to all ADK agents for frontend tools
- langgraph-typescript: remove stale import
- langroid: emit ToolCallResultEvent for backend tools + fix adapter
- llamaindex: v2 provider import, book_call stub, PYTHONPATH fix
- ms-agent-python: disable Responses API store for aimock compat
- pydantic-ai: simplify gen-ui page component
- spring-ai: raise tool iteration cap (1→5) + fix connection pooling
- strands: shared tools symlink + requirements update
2026-04-29 19:40:10 -07:00
Sam Julien 8ba692c426 fix(showcase): regenerate all 18 integration package-lock.json files
Recent feature commits added new dependencies to integration package.json
files (@copilotkit/voice, @hashbrownai/{core,react}, @json-render/{core,react})
and bumped Next.js from 15.4.10 to 15.5.15, but never regenerated the
corresponding package-lock.json. The Showcase Build & Deploy workflow runs
`npm ci --legacy-peer-deps` which strictly enforces lock sync, so every
deploy attempt has been failing at the install step. No new images have been
pushed to GHCR, so Railway services have stayed on stale code and any cell
added since each fw's last successful deploy iframes 404.

Regenerated all 18 lockfiles via `npm install --legacy-peer-deps
--package-lock-only --ignore-scripts` per integration. Verified each with
`npm ci --dry-run --legacy-peer-deps` — all clean.

Refs PDX-90.
2026-04-29 15:39:13 -07:00
github-actions[bot] be073359f9 style: auto-fix formatting 2026-04-29 21:51:40 +00:00
Sam Julien 815d5b377a fix(showcase): dedupe React imports in custom-bubbles sibling snippets
`custom-bubbles.snippet.tsx` in google-adk and llamaindex was built in
PR #4439 by concatenating two source files (user-bubble.tsx +
assistant-bubble.tsx). Each source had its own `import React from "react"`,
so the resulting sibling carried two copies and tripped
oxlint's no-redeclare rule.

Drops the second `import React from "react"` in each file. React is
already imported by the first one — the duplicate was always dead.
Resolves the 2 oxlint errors that were causing CI failure.
2026-04-29 14:49:39 -07:00
github-actions[bot] c3dbba44c8 style: auto-fix formatting 2026-04-29 14:49:39 -07:00
Sam Julien 3b45398251 fix(showcase): repair @endregion[sample-audio-button] placement broken by region-marker script
The marker-insertion script in ac3885fe0 used a brace counter that
counted opening braces from the destructured function parameters as
the start of the function body, then matched the destructuring's
closing `}` as the body's close. The result on every fw was an
`@endregion[sample-audio-button]` jammed onto the same line as the
destructuring's `}`, with the actual function body falling outside the
region — broken structure plus a format violation (`}// @endregion` on
one line).

Fixes both: strips the broken inline endregion and appends a proper
@endregion marker at end-of-file (which is where the function actually
ends, since these files contain only the single SampleAudioButton
function below the imports + interface). 17 files restored.
2026-04-29 14:49:39 -07:00
Sam Julien 9ac8e0644a docs(showcase): switch voice from siblings to region markers in actual demo source
Prior commit (878259e20) deployed sibling .snippet.* files for voice across
all 18 frameworks. That was the wrong call — siblings are a *fallback* for
demos that legitimately diverge from the canonical teaching shape. The
voice demos in 17 frameworks already match the canonical (V2 runtime +
TranscriptionService + sample-audio-button), so the right move is to tag
region markers on the real source.

Changes:
- 17 frameworks (everything except google-adk): add `@region[…]` markers
  to actual demo source for `voice-runtime`, `transcription-service-guard`,
  `voice-page`, `sample-audio-button`. 51 source files modified, no
  behavioral changes — just `// @region[name]` / `// @endregion[name]`
  comments wrapping existing code.
- crewai-crews/manifest.yaml: add `highlight:` block to the voice demo
  with the route file path so the bundler picks up the runtime regions.
  Every other framework already had this entry.
- 17 frameworks: delete the wrong sibling files (`voice-runtime.snippet.ts`
  and `voice-frontend.snippet.tsx`) that 878259e20 created.
- google-adk: KEEP the two siblings — google-adk genuinely diverges
  (uses the shared `/api/copilotkit` route rather than a dedicated
  `/api/copilotkit-voice`), which is exactly when the sibling fallback
  is the right answer.

Result: snippet audit B-docs-gap = 0; every framework's voice page
renders real demo code via `<Snippet>` refs. The 16 standard frameworks
pull from their actual route.ts / page.tsx / sample-audio-button.tsx;
google-adk pulls from its sibling.
2026-04-29 14:49:38 -07:00
Sam Julien 10cfd1009e docs(showcase): voice siblings + rewrite /voice.mdx to use <Snippet> refs
The first pass of /voice.mdx had inline code blocks. Rewrites the page
to use <Snippet> references against per-framework sibling files, matching
how the rest of shell-docs sources its code samples.

- Two siblings per framework (×18 fws = 36 files):
  - voice-runtime.snippet.ts: V2 CopilotRuntime + TranscriptionService
    setup, including the GuardedOpenAITranscriptionService wrapper that
    returns a clean 4xx when OPENAI_API_KEY is missing. Regions:
    `voice-runtime`, `transcription-service-guard`.
  - voice-frontend.snippet.tsx: chat surface with auto-mic-button, plus
    the SampleAudioButton that bypasses the mic for Playwright /
    screenshot flows. Regions: `voice-page`, `sample-audio-button`.
- /voice.mdx now uses 4 `<Snippet region="..." />` refs instead of
  inline code, so the docs reference real teaching code that lives next
  to each framework's actual demo (and stays in sync with the established
  per-framework sibling convention from PR #4439).
2026-04-29 14:49:38 -07:00
Sam Julien 933d37150b chore(showcase): introduce agent_config_pattern + auth_pattern manifest flags
Adds two new manifest pattern flags (matching the existing
`interrupt_pattern` / `a2ui_pattern` convention) so the canonical
`/agent-config` and `/auth` shell-docs pages can gate their per-pattern
sections via `<WhenFrameworkHas>` and only render the implementation that
applies to the framework the user has selected.

- `agent_config_pattern: shared-state | runtime-properties | null`
  - `runtime-properties` (1 fw): built-in-agent
  - `shared-state` (17 fws): everything else that wires agent-config

- `auth_pattern: langgraph | ag2-context-variables | microsoft-agent-framework | runtime-onrequest | null`
  - `langgraph` (3 fws): langgraph-python, langgraph-typescript, langgraph-fastapi
  - `ag2-context-variables` (1 fw): ag2
  - `microsoft-agent-framework` (2 fws): ms-agent-python, ms-agent-dotnet
  - `runtime-onrequest` (12 fws): everything else

Also fills in the previously-missing `a2ui_pattern` flag on 6 frameworks
that have wired demos but were rendering near-empty doc pages because
none of the existing `<WhenFrameworkHas>` gates matched. Audit-driven:
ag2/agno/claude-sdk-{python,typescript}/langroid use schema-loading;
built-in-agent uses schema-inline.
2026-04-29 13:25:02 -07:00
github-actions[bot] 7123b205f9 style: auto-fix formatting 2026-04-29 19:06:09 +00:00
Sam Julien 4bbb75d55f docs(showcase): clean noise from sibling snippets across all frameworks
Sweep across all `.snippet.*` files (existing + new in this branch) to
remove non-teaching content that distracts from the docs-page render.

Changes:
- 6 files (5 hitl + 1 tool-rendering): replace `(props: any)` +
  `eslint-disable-next-line @typescript-eslint/no-explicit-any` with
  proper structural prop types. Reads identical to the eye but no lint
  suppression in the rendered snippet.
- 1 file (state-streaming-middleware.snippet.py): drop 2
  `# type: ignore[name-defined]` markers. The stand-in identifiers
  (`write_document`, `AgentState`) already read as docs-only references.
- 1 file (delegation-log-frontend.snippet.tsx, BIA): rewrite the in-region
  JSDoc to be framework-agnostic. The file was ported from ag2 and still
  named `AG2 sub-agent` + referenced `ReplyResult` / `ContextVariables`
  in the BIA copy. Also drop a historical bug-fix note ("Per-status
  color map…") that is irrelevant outside ag2's commit history.
- 2 files (use-rendered-messages.snippet.tsx, google-adk + llamaindex):
  strip brittle internal-path references (`packages/react-core/src/v2/.../
  CopilotChatMessageView.tsx:542-612`, `react-core/v2/components/chat/
  CopilotChatToolCallsView.tsx`) that would rot within months. Replaced
  with conceptual references to the public component name only.

No region markers changed; audit still reports B-docs-gap: 0.
2026-04-29 11:42:13 -07:00
Sam Julien e857625c57 docs(showcase): headless-complete sibling snippets across google-adk and llamaindex
The shell-docs `/headless` page teaches the truly-headless composition
pattern via 5 regions:
- page-send-message (page-level wiring)
- use-rendered-messages-hook (composition hook)
- manual-activity-message-rendering (per-role dispatch)
- manual-tool-call-rendering (assistant + tool calls)
- custom-bubbles (pure-chrome user/assistant bubbles)

google-adk's headless-complete demo uses custom message-list/bubble
composition that diverges from this canonical pattern; llamaindex's
demo has its own use-rendered-messages.tsx but doesn't match the
canonical regions either. Per the established sibling convention, both
frameworks now ship docs-only siblings exposing all 5 regions adapted
from claude-sdk-typescript's canonical implementation.

Files per framework:
- use-rendered-messages.snippet.tsx (3 regions)
- custom-bubbles.snippet.tsx (combines user-bubble + assistant-bubble)
- page-send-message.snippet.tsx (1 region, minimal teaching variant)

Closes 10 B-docs-gap refs from PDX-83.
2026-04-29 11:07:50 -07:00
Sam Julien c94cac3d3c docs(showcase): hitl-in-chat sibling snippets (booking pattern) across 5 frameworks
The shell-docs `/human-in-the-loop` page teaches the booking pattern
(useHumanInTheLoop with a TimePickerCard rendering candidate slots)
via `<Snippet region="hitl-hook" />` and `<Snippet region="time-slots" />`.
agno, langroid, llamaindex, and spring-ai ship hitl-in-chat demos with
divergent (non-booking) hook wiring; built-in-agent's hitl-in-chat
cell maps to a generic approve/reject demo. Per the established sibling
convention, each framework now ships a docs-only
`hitl-hook-and-time-slots.snippet.tsx` exposing both regions with the
canonical booking shape.

Frameworks: agno, langroid, llamaindex, spring-ai (hitl-in-chat dir);
built-in-agent (hitl dir, where hitl-in-chat cell is routed).

Closes 9 B-docs-gap refs from PDX-83 (8 hitl-hook+time-slots across 4
fws + 1 time-slots for built-in-agent).
2026-04-29 11:07:36 -07:00
Sam Julien 72e9cf20c5 docs(showcase/llamaindex): add mcp-apps runtime-mcpapps-config region
Wraps the CopilotRuntime({ mcpApps: { servers: [...] } }) block in
src/app/api/copilotkit-mcp-apps/route.ts with
@region[runtime-mcpapps-config].
2026-04-29 08:51:19 -07:00
Sam Julien 7699e95166 chore(showcase): manifest a2ui_pattern + interrupt_pattern field values
Sets the per-framework values that drive the new <WhenFrameworkHas>
gating on /generative-ui/a2ui/fixed-schema and /human-in-the-loop/* docs
pages.

  a2ui_pattern values:
    schema-loading — backend loads schema from JSON at startup
                     (langgraph-python/typescript/fastapi, llamaindex,
                      crewai-crews, pydantic-ai, ms-agent-python,
                      google-adk)
    schema-inline  — backend defines schema inline in code
                     (spring-ai, ms-agent-dotnet)
    llm-driven     — backend generates schema dynamically per request
                     (mastra, strands)
    omit           — cell unshipped for the framework

  interrupt_pattern values:
    native        — framework has interrupt() primitive
                    (langgraph-python/typescript/fastapi)
    promise-based — demo uses useFrontendTool + Promise resolution
                    (ms-agent-python, ms-agent-dotnet)
    omit          — cells unshipped for the framework

Same commit also closes a presentation gap on the shell-dashboard
drilldown by adding the missing a2ui sibling files to highlight: lists:
- strands: catalog.ts, definitions.ts, renderers.tsx
- crewai-crews: same three
- google-adk: definitions.ts
2026-04-29 08:15:15 -07:00
Sam Julien e12b546b12 docs(showcase/llamaindex): add chat-slots disclaimer + assistant-message sibling
Reverses the PDX-76 deferral. llamaindex's chat-slots production demo
registers only the welcome slot, but the docs page also teaches the
disclaimer and assistant-message slot patterns — both framework-agnostic
CopilotKit primitives. Sibling teaching file gives those two regions
real teaching code (same pattern as google-adk's
chat-slots/disclaimer-and-assistant-message.snippet.tsx from PDX-71)
without changing the production demo's runtime behavior.
2026-04-29 08:14:26 -07:00
github-actions[bot] 7d131608c0 style: auto-fix formatting 2026-04-29 11:03:09 +00:00
Alem Tuzlak 1c7ca96514 feat(llamaindex): wire hitl-in-chat-booking and mark interrupt features unsupported
- Add hitl-in-chat-booking as a manifest alias to /demos/hitl-in-chat (frontend-tool pattern via useHumanInTheLoop — no backend interrupt required)

- Mark gen-ui-interrupt and interrupt-headless as not_supported_features: LlamaIndex FunctionAgent has no graph-interrupt API equivalent to LangGraph's interrupt()/Command(resume=...) primitive

- Stub page.tsx + README for each unsupported feature explaining the architectural gap and pointing to /demos/hitl-in-chat as the equivalent UX
2026-04-29 12:02:46 +02:00
Alem Tuzlak ba7a806a42 feat(llamaindex): wire mcp-apps README and refresh parity notes 2026-04-29 10:54:35 +02:00
Alem Tuzlak 3ffc66b587 feat(llamaindex): add beautiful-chat polished starter demo 2026-04-29 10:53:41 +02:00
Alem Tuzlak 1d699bb62c feat(llamaindex): add gen-ui-tool-based demo with chart components 2026-04-29 10:52:16 +02:00
Alem Tuzlak aacc37b184 feat(llamaindex): add cli-start and mcp-apps to manifest 2026-04-29 10:49:34 +02:00
Jordan Ritter 17e7e0a406 fix(showcase): add missing D5 demo entries and feature IDs to manifests
Add demo entries for hitl, hitl-in-app, hitl-in-chat, tool-rendering,
shared-state-read-write, and gen-ui-tool-based across 14 integrations.
Ensure every demo ID also appears in the features list so the showcase
matrix and D5 probes discover them correctly.
2026-04-28 22:20:58 -07:00
Jordan Ritter 2fc196aa8b fix(showcase): guard preferences-card.tsx against undefined interests
STATE_SNAPSHOT can deliver a Preferences object with interests undefined,
crashing .includes(), .filter(), and spread at 4 sites per file. Add
(value.interests ?? []) guards across all 17 integrations.
2026-04-28 22:20:53 -07:00
Jordan Ritter f1f3f07514 fix: resolve security vulnerabilities via dependency overrides (#3857)
## Summary

Comprehensive security vulnerability sweep via pnpm overrides and devDep
bumps. Reduces audit from **155+ to 3** unfixable vulnerabilities.

### Changes

**49 pnpm overrides** covering all resolvable transitive dependency
vulnerabilities:
- 12 initial overrides (phase 1)
- 7 upgraded to higher patched versions (phase 2)
- 30 new overrides added (phase 3)

**Direct dependency bumps:**
- storybook devDeps: ^10.1.10 → ^10.2.10 (root + react storybook
example)
- vitest in demo-agents: ^2.1.8 → ^4.1.3 (resolves vite 5.x vuln)
- next in chat-with-your-data: 15.6.0-canary.58 → 15.6.0-canary.61
- vite in react-router: ^6.0.0 → ~7.3.2

### Remaining 3 (truly unfixable)

| Package | Severity | Why |
|---------|----------|-----|
| parse-git-config | HIGH | No patch exists (patched: <0.0.0), dep of
danger |
| elliptic | LOW | No patch exists, deep in storybook crypto chain |
| next | MODERATE | Example on 15.x canary, advisory needs 16.x |

### Companion PR
ag-ui-protocol/ag-ui#1504

Part of CPK-7320
2026-04-28 13:42:41 -07:00
Sam Julien 09d205409a docs(showcase/llamaindex): region markers across 16 cells
Per-framework region pass on llamaindex. 16 (framework x cell) targets
advanced from B (regions missing) to A (ready). Same patterns as mastra
(#4326), smalls batch 1 (#4361), ms-agent batch 2 (#4363), and google-adk
(#4369).

Frontend regions (in-place markers):
- agentic-chat: provider-setup, configure-suggestions
- prebuilt-popup: popup-basic-setup
- prebuilt-sidebar: sidebar-basic-setup, sidebar-configuration
- frontend-tools: frontend-tool-registration + handler
- chat-customization-css: theme-css-import (page.tsx) + css-variables (theme.css)
- chat-slots: register-welcome-slot (lifted prop into named variable)
- tool-rendering: render-weather-tool
- tool-rendering-default-catchall: default-catchall-zero-config
- tool-rendering-custom-catchall: use-default-render-tool-wildcard
- headless-simple: use-agent-simple, message-list-simple
- readonly-state-agent-context: context-provider-sketch, use-agent-context-call
- open-gen-ui: minimal-runtime-flag (route.ts)
- open-gen-ui-advanced: advanced-runtime-config (same span as minimal)
- shared-state-read-write: notes-card-render, preferences-card-render
  (multi-file), use-agent-read, use-agent-write
- subagents: delegation-log-frontend

Sibling teaching files (mirror mastra/google-adk pattern):
- agentic-chat/chat-component.snippet.tsx (production Chat carries QA
  hooks - useFrontendTool, useRenderTool, useAgentContext - that aren't
  relevant to the prebuilt-chat docs page)
- tool-rendering/render-flight-tool.snippet.tsx (covers both
  render-flight-tool and catchall-renderer; production demo only
  registers a weather renderer)

Backend regions (Python files):
- tool-rendering: weather-tool-backend on src/agents/agent.py
  (manifest highlight extended to include this shared-agent file)
- a2ui-fixed-schema: backend-schema-json-load + backend-render-operations
  on src/agents/a2ui_fixed.py - wraps llamaindex's idiom (raw a2ui_operations
  dict ops returned from a backend tool) rather than langgraph-python's
  a2ui.render(...) helper

Manifest highlight added: tool-rendering picked up src/agents/agent.py so
the backend region is bundled with the cell. All other backend files were
already in highlight.

Deferred (documented divergence; will pick up via PDX-68 auto-config or
when showcase team aligns):

- chat-slots disclaimer + assistant-message slots - production demo only
  registers welcome; the other slots aren't implemented.
- declarative-gen-ui::runtime-inject-tool - cross-cutting (PDX-70).
- headless-complete (5 regions: custom-bubbles, manual-activity-message-
  rendering, manual-tool-call-rendering, page-send-message, use-rendered-
  messages-hook) - llamaindex's headless-complete is structurally simpler
  than the canonical pattern (single MessageList component, no separate
  bubble components, no useRenderedMessages hook). The regions don't fit
  cleanly. Defer until showcase team expands the demo or PDX-68 ships.
- shared-state-streaming - llamaindex demo is a TODO stub (page.tsx
  contains 'TODO: Implement State Streaming demo'); skip until the demo
  is implemented.
2026-04-28 13:21:02 -07:00
Jordan Ritter c272a795dc fix: remove stale starter: blocks from all 17 integration manifests
The packages/starters merge (PR #4351) eliminated starters as separate
deployable units. Remove the starter: block (path, name, description,
github_url, demo_url, clone_command) from all 17 integration manifests
to stop propagating stale showcase-starter-* Railway URLs through the
data pipeline.
2026-04-28 12:06:08 -07:00
Jordan Ritter c645e2e6aa feat(showcase): shared-state-read-write + subagents demos across 16 packages (#4359)
## Summary

Adds real working **Shared State (Read+Write)** and **Sub-Agents** demos
to 16 showcase packages, filling rows previously empty on the [coverage
dashboard](https://dashboard.showcase.copilotkit.ai/#coverage). Each
package mirrors the canonical `langgraph-python` and `google-adk`
reference implementations, adapted to the framework's native primitives.

**Packages affected (16):** ag2, agno, built-in-agent,
claude-sdk-python, claude-sdk-typescript, crewai-crews,
langgraph-fastapi, langgraph-typescript, langroid, llamaindex, mastra,
ms-agent-dotnet, ms-agent-python, pydantic-ai, spring-ai, strands

**Per-package deliverables:**
- Backend agent files (framework-native): preferences-injection
middleware/callback + `set_notes` tool; supervisor + 3 sub-agents
(research/writing/critique) wired as tools with running→completed/failed
delegation log
- Frontend `page.tsx` + `preferences-card.tsx` / `notes-card.tsx` for
SSRW; `delegation-log.tsx` for subagents — wired to `useAgent({ updates:
[OnStateChanged] })`
- Manifest entries (`features:` + `demos:` with `route` + `highlight`)
- Runtime route registration (`route.ts` and per-package agent server
config)
- QA scripts (real, replacing stubs)

## Approach

Built via parallel orchestration: 16 worktree-isolated agents
implemented one package each. Followed by a 7-agent code-review round
and a 13-package targeted fix wave (32 fix commits across 13 packages)
addressing the demo-breaking bugs the review surfaced.

## What was fixed during CR

Highlights from the 36 fix commits:
- **Sub-agent failure paths now correctly emit \`status: \"failed\"\`**
(was hardcoded \"completed\" or unreachable in
mastra/strands/langgraph-fastapi/langgraph-typescript/ag2)
- **Parallel-tool-call delegation race fixed** in langgraph-fastapi
(\`Annotated[list, add]\`) and langgraph-typescript (concat reducer) —
was last-write-wins
- **Silent data loss eliminated** in
claude-sdk-python/claude-sdk-typescript/crewai-crews — empty
\`JSON.parse\` catches now log + emit error events
- **\`ms-agent-dotnet\` \`set_notes\` writes to per-thread slot** (was
hardcoded \`thread: null\` → notes never reached UI)
- **\`mastra\` working-memory writes are deterministic** — new
\`tools/working-memory.ts\` helper writes directly via
\`memory.updateWorkingMemory\` (was LLM-prompted, non-deterministic)
- **\`built-in-agent\` e2e tests rewritten** to assert actual page UI
(specs were referencing recipe UI from a prior implementation)
- **\`spring-ai\` tool-call envelope IDs match supervisor\'s
\`tc.id()\`** (was random UUIDs that broke frontend correlation) + AG-UI
event ordering reordered + \`CopyOnWriteArrayList\` for parallel-call
safety
- **Stack trace + raw error message leaks scrubbed** across 8+ Next.js
routes — now log server-side with \`errorId\` + return \`{ error:
\"internal runtime error\", errorId }\` (mastra reference pattern
propagated)
- **Sub-agent calls no longer block event loops** in ag2
(\`asyncio.to_thread\`), langroid (\`llm_response_async\`), pydantic-ai
(async \`run\` + async tools)
- **\`langroid\` \`lru_cache\` cross-request contamination dropped** —
sub-agents rebuilt per call, no message-history leak between users
- **Numerous smaller items**: \`claude-sdk-python\` invalid model id
(\`claude-opus-4-5\` → dated id), \`Callable\` annotation, \`/health\`
endpoint exposed; \`built-in-agent\` floating \`latest\` deps pinned,
invalid \`X-Frame-Options\` removed, \`ignoreBuildErrors\` env-gated,
subagent role names aligned to canonical trio; \`crewai-crews\`
supervisor no longer resets delegations every turn; \`pydantic-ai\`
snapshot uses \`model_dump()\`

## Known follow-ups (deferred to follow-up PR)

These were classified as bucket (c)/(d) or Tier 2 during cr-loop and
intentionally deferred:
- **agno** sync \`sub_agent.run()\` blocks event loop (perf only — works
correctly)
- **ms-agent-python** \`asyncio.run\` thread fallback uses string-match
for runtime detection + \`worker.join()\` blocks; works but fragile
- **llamaindex** minor initial-state coercion when UI clears state via
\`agent.setState({})\`
- **Manifest highlight audit** (across packages):
\`langgraph-typescript\` \`headless-complete\` highlight points at
\`copilotkit-mcp-apps/route.ts\`; \`langgraph-fastapi\` \`byoc-*\`
missing route.ts highlights
- **\`agno\`** \`hitl-in-chat\` declared in demos but not features;
duplicate \`/demos/hitl-in-chat\` route across two demo entries
- **\`langgraph-typescript\` \`server.mjs\` \`graphSpec\`** only
registers 3 graphs while \`langgraph.json\` declares 23 — pre-existing
gap, this PR only added the 2 it needed
- **\`mastra\`** \`hitl\` legacy demo missing from features list
- **\`claude-sdk-python\` \`agents/agent.py\` line 474** also has the
legacy \`claude-opus-4-5\` default (out of CR scope)
- **PARITY_NOTES vs manifest mismatches** for \`hitl-in-app\` across
spring-ai, agno, ag2 — pre-existing
- **\`spring-ai\`** \`a2ui-fixed-schema\` missing from \`generative_ui\`
list; system-prompt dangling newline
- **\`built-in-agent\` zod v3↔v4 peer-dep mismatch** surfaces under
strict TS (\`ignoreBuildErrors\` env-gate now exposes them — was
previously hiding them)

## Build/test verification caveats

- **Windows MAX_PATH** prevented \`pnpm install\` at the worktree root
for several packages, so per-package \`tsc --noEmit\` was sometimes
deferred to CI. Verified pattern parity with reference implementations.
- **\`dotnet build\`** for \`ms-agent-dotnet\` not run locally — SDK
absent in worktree (only runtime). Code follows existing
\`SubagentsStore\`/\`AgentConfigAgent\` patterns; CI is the first
compile check.
- **\`mvn compile\`** for \`spring-ai\` not run — Maven absent locally.
Code uses only documented Spring AI 1.0.x + ag-ui-java APIs.
- **Lefthook \`test-and-check-packages\` hook bypassed** with
\`--no-verify\` on most fix commits — root \`node_modules\`/\`nx\`
absent in worktrees (Windows MAX_PATH/symlink issue). Failures unrelated
to changed files; rationale documented in commit bodies.

## Test plan

- [ ] CI runs \`tsc --noEmit\`, \`vitest\`, and per-package builds
across all 16 packages
- [ ] Manual QA against each package's \`qa/shared-state-read-write.md\`
and \`qa/subagents.md\` (deployed Railway services)
- [ ] Verify dashboard rows turn green for shared-state-read-write and
subagents on each integration column at
https://dashboard.showcase.copilotkit.ai/#coverage
- [ ] Spot-check spring-ai \`mvn compile\` and ms-agent-dotnet \`dotnet
build\` once SDK availability is sorted
- [ ] Confirm parallel-tool-call delegation race fix on
langgraph-fastapi/typescript by triggering parallel sub-agent calls
2026-04-28 11:52:49 -07:00
Jordan Ritter ee5952bd46 fix(showcase): fix tool-rendering D5 probes for llamaindex and ms-agent-dotnet
- Add missing `import os` to llamaindex agent.py (NameError crash on
  os.environ.get at module level prevented agent server startup)
- Update useRenderTool in both services to canonical v2 pattern:
  use `parameters` instead of `args`, add deps array, remove `: any`
  type cast
2026-04-28 10:37:11 -07:00
Jordan Ritter 6bc0db6a25 fix: harden showcase packages — dep pins + Docker image pins
Dependency version floors:
- next: ^15.0.0 → ^15.5.15 across all 19 showcase packages (CVE-2025-29927)
- express: ^4.21.0 → ^4.21.2 in claude-sdk-typescript (open redirect fix)
- hono: ^4.0.0 → ^4.6.0 in shell (path traversal fix)

Docker base image pins:
- node:20-slim → node:20.19-slim (18 Dockerfiles)
- python:3.12-slim → python:3.12.11-slim (12 Dockerfiles)
- aimock:latest → aimock:1.13.0 (1 Dockerfile)

Part of CPK-7320
2026-04-28 10:33:06 -07:00
github-actions[bot] 5fcb637bde style: auto-fix formatting 2026-04-28 16:37:55 +00:00
Alem Tuzlak 23a3b24a01 feat(showcase/integrations): shared-state-read-write + subagents demos across 15 packages
Adds real working Shared State (Read+Write) and Sub-Agents demos to 15
showcase integrations, mirroring the canonical langgraph-python and
google-adk reference implementations. Fills rows previously empty on
the showcase coverage dashboard.

Packages: ag2, agno, claude-sdk-python, claude-sdk-typescript,
crewai-crews, langgraph-fastapi, langgraph-typescript, langroid,
llamaindex, mastra, ms-agent-dotnet, ms-agent-python, pydantic-ai,
spring-ai, strands. (built-in-agent landed independently on main as
PR #4321 — its variant is canonical; this PR no longer touches it.)

Per-package deliverables: framework-native backend agents
(preferences-injection middleware/callback + set_notes tool;
supervisor + 3 sub-agents wired as tools with running -> completed
/failed delegation log); frontend page.tsx + preferences-card.tsx /
notes-card.tsx for SSRW and delegation-log.tsx for subagents — wired
to useAgent({ updates: [OnStateChanged] }); manifest entries; runtime
route registration + per-package agent server config; real QA
scripts.

Includes targeted hardening fixes from a 7-agent code-review loop:

- Sub-agent failure paths now correctly emit status: "failed"
  (previously hardcoded "completed" or unreachable in
  mastra/strands/langgraph-fastapi/langgraph-typescript/ag2)
- Parallel-tool-call delegation race fixed in langgraph-fastapi
  (Annotated[list, add]) and langgraph-typescript (concat reducer)
- Silent data loss eliminated in
  claude-sdk-python/claude-sdk-typescript/crewai-crews — empty
  JSON.parse catches now log + emit error events
- ms-agent-dotnet set_notes writes to per-thread slot via AsyncLocal
- mastra working-memory writes are deterministic via
  src/mastra/tools/working-memory.ts helper
- spring-ai tool-call envelope ids match supervisor's tc.id() and
  AG-UI event ordering reordered; CopyOnWriteArrayList for
  parallel-call safety
- Stack trace + raw error message leaks scrubbed across 8+ Next.js
  routes — log server-side with errorId + return generic envelope
- Sub-agent calls no longer block event loops in ag2
  (asyncio.to_thread), langroid (llm_response_async), pydantic-ai
  (async run + async tools)
- langroid lru_cache cross-request contamination dropped
- Numerous smaller items: claude-sdk-python invalid model id, Callable
  annotation, /health endpoint exposed; crewai-crews supervisor
  no longer resets delegations every turn; pydantic-ai snapshot uses
  model_dump()

CI fixes folded in:
- crewai-crews test_forwarded_props: extend the stubbed
  ag_ui_crewai.endpoint module to expose
  add_crewai_flow_fastapi_endpoint and add stub
  agents.shared_state_read_write / agents.subagents modules
- generate-catalog test: bump crewai-crews wired-cell expectation
  28 -> 30; replace hardcoded total-wired count with an invariant
  (wired + stub + unshipped = 737) plus a lower-bound floor
- oxfmt run on the qa/shared-state-read-write.md files in mastra +
  spring-ai

Rebased onto latest main (post showcase/packages -> showcase/integrations
rename + post built-in-agent landing). Original blitz history
preserved at the blitz-pre-rebase-snapshot tag.

Known follow-ups (deferred to follow-up PR):
- agno sync sub_agent.run() blocks event loop (perf only)
- ms-agent-python asyncio thread-fallback fragility
- llamaindex initial-state coercion when UI clears state
- Manifest highlight audit (langgraph-typescript headless-complete,
  langgraph-fastapi byoc-* missing route.ts highlights)
- agno hitl-in-chat declared in demos but not features; duplicate
  /demos/hitl-in-chat route
- langgraph-typescript server.mjs graphSpec only registers 3 graphs
  vs 23 in langgraph.json (pre-existing)
- mastra hitl legacy demo missing from features list
- claude-sdk-python agents/agent.py line 474 also has the legacy
  claude-opus-4-5 default
- PARITY_NOTES vs manifest mismatches for hitl-in-app across
  spring-ai/agno/ag2 (pre-existing)
- spring-ai a2ui-fixed-schema missing from generative_ui list
2026-04-28 18:36:13 +02:00
Jordan Ritter e9a2e143de fix(showcase): add shared-tools symlinks and refactor imports
Replace sys.path.insert hacks in Python agent files with direct
imports via symlinks to shared/{python,typescript}/tools.
Update Dockerfiles, entrypoints, and configs to support the new
symlink-based tool resolution. Add PARITY_NOTES for frameworks
that have known gaps.
2026-04-28 07:50:03 -07:00