mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
fix/aimock-starter-greeting
120 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
83404f8ca5 | style: auto-fix formatting | ||
|
|
efd35f0f8d | chore: align all a2ui instances with the latest implementations | ||
|
|
20aa327a8c |
fix(showcase): enable injectA2UITool for langgraph-python declarative-gen-ui
It was left at false while fastapi/typescript are true. Under the opt-in A2UI model false means no tool is injected, so the python demo rendered no surfaces (and the docs code-tab showed no injectA2UITool). Set true to match the other langgraph integrations. |
||
|
|
02b11c985d |
feat(showcase): drive dynamic A2UI via CopilotKitMiddleware (langgraph)
The declarative-gen-ui demo across the three langgraph integrations now relies on the middleware to inject and execute generate_a2ui — the agents collapse to create_agent + CopilotKitMiddleware with no hand-rolled tool. Adds render_a2ui fixtures for the new tool path and pins the integrations to the A2UI alpha SDKs (copilotkit 0.1.94a1, @copilotkit/sdk-js 1.59.3-alpha.1). |
||
|
|
3486944356 |
fix(showcase): controlled gen-UI — reliable 2nd-suggestion render + sidebar tag (OSS-137) (#5029)
## What & why [OSS-137](https://linear.app/copilotkit/issue/OSS-137/controlled-gen-ui-demo-optimize-2nd-suggestion-prompt-rename-sidebar) — the **Controlled Generative UI** demo (`gen-ui-tool-based`) had two issues, scoped here to **LangGraph-Python** and **Google ADK** (per the ticket; other 16 integrations roll out later). ### 1. 2nd suggestion didn't reliably render UI The "Traffic pie chart" chip (`"Show me a pie chart of website traffic by source."`) names a subject but supplies no numbers, so the agent **asked the user for data** instead of rendering a chart. **Fix:** a system-prompt directive (both LGP + ADK agents) instructing the agent to invent plausible illustrative sample values, call `render_*` immediately, and **never** reply with a clarifying question. The suggestion copy stays clean — behavior is carried by the system prompt, not by leaking "(use sample data)" hints into the UI. ### 2. Sidebar tag → product language Retagged the demo from `generative-ui` → `controlled-generative-ui` (LGP + ADK), so the dojo sidebar pill reads **"Controlled Generative UI"** — the established taxonomy already used in `shared/feature-registry.json` and the dashboard catalog. ## Tests Added D5 aimock fixture entries mirroring all three suggestion chips (bar / traffic-pie / market-share) so the suggestion-click path has deterministic coverage. The existing `"revenue by category"` probe message is **preserved**, so the [dashboard](https://dashboard.showcase.copilotkit.ai/#matrix:links,health) D5 row for the edited row stays green. ## Acceptance check - [x] 2nd suggestion renders UI without asking for data (system-prompt directive; verified locally against the live agent) - [x] Sidebar entry tagged "Controlled Generative UI" - [x] Sample/hallucinated data supplied via system prompt - [x] Scoped to LGP + ADK - [x] Tests augmented (D5 fixtures for every chip) - [x] D5 still shown for the edited row (probe message unchanged) ## Out of scope (left out deliberately) `package-lock.json` churn from a local reinstall (un-pins `latest`) was **not** committed. 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
0143b02ae3 |
WIP: D6 rollout — foundation + per-integration fixtures + conveyance fixes (#5109)
## Status
WIP / not ready to merge. Preserves in-flight D6 work so it isn't lost
mid-rollout. LGP is fixture-complete; other integrations are
mid-rollout.
### Latest banked work
- **langgraph-python — 185 / 0 / 2** (green). Achieved by narrowing a d4
chat matcher that was shadowing the d6 beautiful-chat search_flights
fixture (load order is shared -> d4 -> d6, first-match-wins) plus
refreshing the d6 tool-rendering and tool-rendering-custom-catchall AAPL
fixtures (turnIndex:0 -> hasToolResult:false so the first leg fires in
multi-pill threads). 2 skips are the by-design mcp-apps iframe gap.
- **langgraph-typescript — 185 / 0 / 2** (green). Mirrored the LGP d4
narrowing on the LGT side (3 matchers) and across the
d6/langgraph-typescript suite: replaced fragile turnIndex:0 gates with
hasToolResult:false, fixed em-dash escaping that broke literal matches
in multi-pill threads, and added jsFunctions payloads to the three
sandboxed-ui fixtures (`_from-feature-parity`, `headless-complete`,
`gen-ui-open-advanced`). The final fix removed the chain-tools
`hasToolResult` match gate (which checked the whole thread and made the
chain pill fall through to the broad weather matcher mid-thread),
mirroring LGP.
- **google-adk: 174/7/2 (was 134/52/5)** — conveyance + test-parity +
fixtures + pill-wiring rebuild; 7 residual (6 default-catchall framework
default-renderer testid version question, 1 beautiful-chat fixture).
- **Pill-parity staged across 13 integrations** (ag2, agno, mastra,
pydantic-ai, claude-sdk-python, claude-sdk-typescript, llamaindex,
langroid, strands, spring-ai, built-in-agent, crewai-crews,
langgraph-fastapi). The canonical LGP suggestion pill set is now
mirrored as `src/app/demos/*/suggestions.ts` files in each integration,
with targeted edits to existing `open-gen-ui-advanced` and
`byoc-hashbrown` files. **These new files are currently UNWIRED** — each
integration's `page.tsx` still defines its pill list inline via
`useConfigureSuggestions`. Banked so the canonical source survives; a
follow-up will rewire `page.tsx` to import from `suggestions.ts` and
drop the inline copies.
- **Fleet test-parity sweep**: 576 e2e specs across 15 integrations
aligned to LGP canonical (SHA-verified); 2 orphan specs removed.
- ms-agent-dotnet 177/5/7, ms-agent-python 174/11/2 (post
fixture-mirror); default-catchall green (page-level renderer, not
react-core-gated).
## Scope
### Conveyance (foundation)
Inbound `x-aimock-context` (and friends) must ride along on outbound LLM
HTTP calls so aimock fixture matching sees the inflight test's context.
Without this the call lands on the default project's aimock and silently
picks the wrong fixture. New per-integration
`_header_forwarding.{py,ts}` shim plus matching `agent_server` / route /
factory wiring covers: ag2, agno, built-in-agent, claude-sdk-python,
claude-sdk-typescript, crewai-crews, google-adk, langgraph-fastapi,
langgraph-python, langgraph-typescript, langroid, llamaindex, mastra,
ms-agent-python, pydantic-ai, strands.
For ADK/Gemini the global httpx hook is installed BEFORE any `agents.*`
import (google-genai constructs its client at module-import time).
### langgraph-python — 185 / 0 / 2
Fixture-complete via the conveyance shim + refreshed d6/langgraph-python
fixtures + copilotkit 0.1.93 bump + the latest d4-matcher-narrowing fix
(see banked work above).
### Per-integration fixtures
Mid-rollout snapshot of d6 fixtures across the cohort plus narrowing of
`aimock/shared/common.json`'s generic 'hello' fixture to 'hello world'
so it no longer shadows D6 pills whose prompts contain 'hello' as a
substring.
### Harness `--isolate` patch
`scripts/cli/_common.sh apply_isolation` now rewrites compose-file
relative paths to absolute (build/context/dockerfile/volumes/env_file),
enforces the docker compose `[a-z0-9_-]` project-name rule, and exports
`SHOWCASE_COMPOSE_FILE` / `SHOWCASE_INFRA_PORT_OFFSET` plus offset host
URLs. The TS harness CLI (`aimock-rebuild` / `config` / `doctor` /
`lifecycle`) honors the new env so concurrent isolated stacks stop
reporting each other's services as healthy.
## Lockfile decision flagged
`showcase/integrations/langgraph-python/pnpm-lock.yaml` was deleted in
this branch. Decision: keep the deletion. Rationale:
-
|
||
|
|
ceff27031c |
feat(showcase/langgraph-python): D6 conveyance shim + copilotkit 0.1.93 bump
- Add _header_forwarding_middleware.py paired with reasoning / subagent / tool-rendering-reasoning-chain agent edits so D6 fixture-matching sees the inflight x-aimock-context on outbound LLM calls - Bump copilotkit 0.1.92 → 0.1.93 in requirements.txt; regenerate package-lock.json against the post-npm-ci tree (sibling of |
||
|
|
d60285c337 |
fix(react-core): preserve generated thread tool followups (#5043)
## Summary - keep `CopilotChat` agents aligned to SDK-generated thread IDs even when `/connect` is intentionally skipped for non-explicit threads - stabilize `CopilotKitProvider` default object props so rerenders do not re-sync an empty local agent registry and replace the live remote/Intelligence agent mid-run - add regression coverage for SDK-generated thread frontend-tool follow-up runs and provider empty-agent rerender stability - add a focused langgraph-python showcase demo, aimock fixture, Playwright smoke, and QA checklist for ENT-658 - add a patch changeset for `@copilotkit/react-core` ## Testing - `npx nx run @copilotkit/react-core:test -- src/v2/components/chat/__tests__/CopilotChat.absentThreadConnect.test.tsx` - `npx nx run @copilotkit/react-core:test -- src/v2/providers/__tests__/CopilotKitProvider.stability.test.tsx` - Pre-commit hook passed: `pnpm run test` and `pnpm run check:packages` - Verified exact `CopilotKit/Intelligence` repro branch `mme/threadid-repro`: unchecked `Explicit threadId`, sent `invoke testFrontendToolCalling with label X`, confirmed user message/tool card/assistant reply remain visible - Verified the same Intelligence repro with `Explicit threadId` checked - `pnpm exec playwright test tests/e2e/threadid-frontend-tool-roundtrip.spec.ts --project=chromium --workers=1` from `showcase/integrations/langgraph-python` ## QA Checklist - [x] Reproduce the reset in `CopilotKit/Intelligence` branch `mme/threadid-repro` with `Explicit threadId` unchecked - [x] Confirm generated-thread frontend-tool round-trip preserves the user message, tool card, and assistant response - [x] Confirm explicit-thread frontend-tool round-trip still preserves the user message, tool card, and assistant response - [x] Open `/demos/threadid-frontend-tool-roundtrip` in the langgraph-python showcase demo - [x] Confirm `Explicit threadId` is unchecked and the chat starts in SDK-generated thread mode - [x] Send `invoke testFrontendToolCalling with label X` - [x] Confirm the user message remains visible - [x] Confirm the `testFrontendToolCalling` card remains visible and shows `label: X` plus `result: handled X` - [x] Confirm the assistant reply `Frontend tool finished for X.` appears - [x] Confirm the chat does not return to the empty state - [x] Repeat with `Explicit threadId` checked and confirm the explicit-thread path is unchanged ## Notes The visible reset had two frontend-side causes. First, the chat and agent could diverge when the SDK generated the thread ID. Second, in Intelligence mode, provider rerenders could re-sync an empty local agent registry and replace the live remote agent instance mid-run, dropping the in-memory chat stream. Both fixes live in `@copilotkit/react-core`. The Playwright file is intentionally a smoke test for the demo route/toggle. The source-level regressions live in `CopilotChat.absentThreadConnect.test.tsx` and `CopilotKitProvider.stability.test.tsx`. |
||
|
|
37db1c8e5b | Fix shell-docs setup packaging and framework nav | ||
|
|
daff12758a | fix(showcase): use uuid explicit thread demo id | ||
|
|
a879b8a062 | test(react-core): tighten thread roundtrip coverage | ||
|
|
24d93b52ad | fix(react-core): preserve generated thread tool followups | ||
|
|
478039786c |
fix(showcase): controlled gen-UI — reliable 2nd-suggestion render + sidebar tag (OSS-137)
The "Traffic pie chart" suggestion ("Show me a pie chart of website traffic
by source.") names a subject but supplies no numbers, so the agent asked the
user for data instead of rendering. Add a system-prompt directive (LGP + ADK)
telling the agent to invent illustrative sample values and render on the first
turn, never asking for data. The suggestion copy stays clean — the behavior is
carried by the system prompt, not parenthetical UI hints.
Also retag the gen-ui-tool-based demo (LGP + ADK) from `generative-ui` to
`controlled-generative-ui` so the dojo sidebar pill reads "Controlled
Generative UI" — the established product taxonomy (already a category in
shared/feature-registry.json and the dashboard catalog).
Scoped to LGP and ADK per the ticket; the other 16 integrations keep the old
tag until the taxonomy rolls out wider.
Tests: add D5 aimock fixture entries mirroring all three suggestion chips
(bar/traffic-pie/market-share) so the suggestion-click path has deterministic
coverage. The existing "revenue by category" probe message is preserved, so
the dashboard D5 row stays green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
4680eb9c16 |
fix(showcase/headless-complete): tag page-send-message region
The programmatic-control docs page renders a yellow "Missing snippet" box on the langgraph-python and google-adk variants because their headless-complete cells were never tagged with the page-send-message region the MDX requests. Add matching @region / @endregion markers around the useAgent / useCopilotKit / send / reset block in chat/chat.tsx so the Snippet component resolves on both integrations. |
||
|
|
64ffd19d8c |
fix(showcase): CR Round 2 cleanup — ADK reasoning graph name + dead CSS + comment + log tag
CR Round 2 confirmation surfaced one bucket (a) finding plus three
bucket (b) trivials worth rolling in together.
(a) `google-adk/src/app/demos/reasoning-{default,custom}/page.tsx`
comments said "Both demos share the same backend (`reasoning_agent`
graph)". That graph name is the langgraph-python convention —
`reasoning_agent.py` in LGP — but the ADK demo doesn't have a
graph by that name. `src/agents/registry.py:144-145` maps both
`reasoning-custom` and `reasoning-default` to
`AgentSpec(_thinking_chat)`, where `_thinking_chat` is built via
`build_thinking_chat_agent`. Round 1 fixed the same class of bug
in langgraph-typescript (which uses `agentic-chat-reasoning`) but
missed ADK; this is the matching fix.
(b1) `.../headless-simple/chat.tsx` (3 files) emitted
`console.error("[headless-simple] ...", err)` with no
integration-slug prefix. A user testing demos across frameworks
in the same browser session couldn't tell which integration's
runAgent failed. Tag with the framework slug:
`[google-adk:headless-simple]`, `[langgraph-python:headless-simple]`,
`[langgraph-typescript:headless-simple]`.
(b2) `globals.css` lines 133-137 — the `.shell-docs-sidebar
p[class*="sidebar-item-offset"] svg` rule (4×4 icons in accent
purple) was dead in fumadocs v16. The v16 sidebar emits separator
`<p>` elements with `inline-flex items-center gap-2` instead of
the v15 `sidebar-item-offset` class fragment; the live rule on
`p.inline-flex.gap-2 svg` (added earlier in this PR) already
handles the same styling at the correct 16×16 size. Drop the
dead rule.
(b3) `page-actions.tsx` — the regression-fix commit
(`0186ae9f2`) wedged `getClientBaseUrl()` between the cache-
describing block comment and the actual `cache = new Map(...)`
declaration. The comment now sits above its own subject again;
`getClientBaseUrl()` keeps its own JSDoc above its definition.
Call-site enumeration:
- ADK `_thinking_chat` reference — verified in
`showcase/integrations/google-adk/src/agents/registry.py` (line
144-145 + `build_thinking_chat_agent` import on line 23 + builder
invocation on line 108). Comment-only change; no symbol signatures
touched.
- Headless log tags — only the literal log string changes; no other
call site reads it.
- `globals.css` dead rule — verified no other selector in the file
depends on the removed lines (the section-header SVG color is set
by the surviving `p.inline-flex.gap-2 svg` rule).
- `page-actions.tsx` comment move — no functional change.
|
||
|
|
086e68c88b |
fix(showcase): log runAgent errors in headless-simple; correct LGT reasoning graph name
The Headless Simple demo's `chat.tsx` swallowed every `runAgent`
rejection with an empty arrow catch:
void copilotkit.runAgent({ agent }).catch(() => {});
This is the canonical "two hooks, your design system" example users
copy-paste as a starting point — silent swallow modeled broken practice
to every CopilotKit user, and the @region[use-agent-simple] block we
inline into `/<framework>/headless` docs surfaces the anti-pattern as
the recommended snippet. Replace the empty catch with a
`console.error("[headless-simple] runAgent failed", err)` so network
failures, transport disconnects, and runtime errors surface in the
developer's console. Applied across google-adk, langgraph-python, and
langgraph-typescript variants.
`langgraph-typescript/src/app/demos/reasoning-default/page.tsx` had a
comment claiming the demo backed onto the `reasoning_agent` graph, but
the LGT route map in `src/app/api/copilotkit/route.ts` actually points
both `reasoning-default` and `reasoning-custom` at the
`agentic-chat-reasoning` graph (the companion `reasoning-custom/page.tsx`
comment already gets this right). The `reasoning_agent` label is the
Python / ADK convention. Update the comment to match the TS route map.
Call-site enumeration:
- `copilotkit.runAgent` (in headless-simple/chat.tsx, 3 files) — the
return value is `Promise<void>`; existing callers don't await it, so
swapping the catch is non-breaking. The previous `void` operator
already discarded the promise value, so the runtime behavior of the
surrounding `send()` is unchanged.
- LGT `reasoning-default` page.tsx — comment-only change, no symbol
signatures touched.
|
||
|
|
5728611dfd |
feat(shell-docs): upgrade to fumadocs 16 / next 16, polish layout, add llms.txt + page actions
Stack upgrade - fumadocs-core/ui 15.8.5 → 16.8.12, next 15 → 16 (Turbopack), react 19 → 19.2 - Swap "next lint" → "oxlint ." to match the rest of the repo - New deps for the page-actions component: @radix-ui/react-popover, class-variance-authority, clsx, tailwind-merge Layout & brand polish - Sidebar floats as a rounded-2xl card with column-aligned padding; framework picker pill, accent-purple section icons (16px), accent active state, and a single divider line at the footer - New custom <ThemeSwitch> — single 50×28 neutral switch replaces the fumadocs sun/moon split (drops the vertical divider and purple tint) - Sidebar folder collapse state persists across navigations via SidebarFolderStatePreserver - BrandNav: wider top bar, lowercase "Talk to an engineer", BookIcon for Docs, GitHub/Discord icons rendered inline in our footer row - Mobile: nav clipping + content padding fixes, content grid-span-full - TOC-less pages: lift article max-width so content stretches into the empty TOC column on wide viewports New routes - /llms.txt — page index per fumadocs LLMs integration - /llms-full.txt — concatenated full text of every docs page - /<path>.md and /<path>.mdx — per-page raw markdown with <Snippet> regions inlined as fenced code blocks (resolver in lib/llm-text.ts reuses the same demo-content.json the <Snippet> runtime reads) - Page-actions bar: Copy Markdown + Open in Claude / Claude Code / Windsurf / Codex (Codex links to https://chatgpt.com/codex for universal coverage) Content fixes - Reasoning page (generative-ui/reasoning.mdx): rewrite to point at the real reasoning-default / reasoning-custom cells instead of the stale agentic-chat-reasoning / reasoning-default-render names - Strip <FeatureIntegrations /> chip list ("SUPPORTED BY ...") from 16 docs MDX files (component definition kept in mdx-registry) - Drop hideTOC: true from 11 pages so they pick up the lifted-cap rule - Default home (/) to the built-in-agent authored sidebar; fix active state matching on the home url - Restore default fumadocs Callout (drop the bespoke docs-callout) - OpsPlatformCTA redesign — light bordered card with accent stripe - FrameworkOverview redesign — drop atmospheric chrome, smaller hero - Homepage / docs-landing redesign Integrations (LGP / LGT / ADK) - Tag @region[default-reasoning-zero-config] in reasoning-default and @region[reasoning-block-render] in reasoning-custom for all three frameworks so the docs <Snippet> calls resolve - Tag @region[use-agent-simple] + @region[message-list-simple] in headless-simple and @region[use-rendered-messages-hook] + @region[manual-tool-call-rendering] + @region[manual-activity-message-rendering] + @region[custom-bubbles] across headless-complete Other - docs/components/layout/mobile-sidebar.tsx: lowercase "engineer" to match shell-docs - .claude/launch.json + .claude/preview/ — dev launch configs for the worktree so /preview brings up shell-docs on :3003 |
||
|
|
1a534ba9dd |
Merge remote-tracking branch 'origin/main' into tyler/laughing-burnell-67b26b
# Conflicts: # showcase/integrations/strands/package-lock.json |
||
|
|
7e1ec07b70 |
fix(shell-docs): CR Round 1 bucket-a fixes — content + nav + MDX overrides + script hardening
Six fixes from CR Round 1 partition, all bucket (a):
- frontend_tools.py: docstring claimed the file was "Chat Customization
(CSS) demo" but langgraph.json wires it as the Frontend Tools demo
graph, and the new MDX setup snippets cite this exact file via the
freshly-added `# region: middleware` markers. Users following the
langgraph-python copilot-middleware setup would see CSS-demo wording
on a Frontend Tools page. Rewrote the docstring to match what the
cell actually demonstrates (mirroring the sibling
frontend_tools_async.py phrasing).
- page.tsx mergeFrameworkNav: when introNode was non-null AND the root
nav had no "Get Started" section, introNode was prepended to rootNav
shifting every existing index +1. The adjustment block only added +1
when getStartedIdx !== -1, so the splice-back position for the
framework section was off-by-one in the no-Get-Started branch — the
framework header rendered one slot too early in the sidebar.
- docs-page-view.tsx h2/h3 overrides: `{...rest}` was spread AFTER
`id={id}`, so an MDX-supplied `<h2 id="custom">` would override the
slugified id and silently break the TOC anchor + any inbound deep-
links keyed on the slug. Reordered the spread so rest comes first
and the slug-id always wins.
- probe-shell-docs.ts: terminated with bare `main();` while every
sibling script (audit-docs-porting, verify-shell-docs) wraps in
`.catch(e => { console.error(e); process.exit(1); })`. A rejected
main() would surface as an unhandled rejection on older Node
runtimes and exit 0 in CI, masking failure. Aligned with the
established pattern.
- verify-shell-docs.ts: all four regex checks (InlineDemo refs,
Snippet regions, internal links, alias imports) scanned page.body
raw without first stripping fenced code blocks. Any docs page that
showed example code containing `<InlineDemo demo="x" />`,
`[link](/path)`, or `import x from "@/..."` triggered a false-
positive validator failure. Mirrors audit-docs-porting.ts's
FENCED_CODE_RE approach. Adds a regression test that fails without
the strip.
- 3 new MDX content fixes:
* mcp-apps.mdx + open-generative-ui.mdx: removed duplicate `<Callout>`
"Free course" blocks (the same Callout appeared twice on each
page, separated only by the Key Benefits list).
* subagents.mdx: changed `[OnStateChanged, OnRunStatusChanged]` to
`[UseAgentUpdate.OnStateChanged, UseAgentUpdate.OnRunStatusChanged]`
— the bare identifiers aren't exported (the reference doc
`useAgent.mdx` confirms the qualified form), so a user copying
the snippet would hit an import error.
Call-site enumeration:
- frontend_tools.py: only langgraph.json + the new setup MDX files
reference this file by name; both consume the region markers, not
the docstring. Docstring rewrite has zero call-site impact.
- mergeFrameworkNav: single caller (FrameworkScopedDocsPage at this
file's bottom). The new branch covers a strictly broader case;
the original splice/replace paths are unchanged.
- h2/h3: only used by the MDXRemote `components` map below. Spread
order is a local prop-precedence change; no upstream callers.
- probe-shell-docs main(): no external callers.
- verify-shell-docs check functions: 4 exported functions called
from runChecks() below + the test file. Strip is internal to each
function so signature is unchanged.
- UseAgentUpdate: confirmed exported from `@copilotkit/react-core/v2`
per reference doc useAgent.mdx; no implementation change needed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
662f757cd6 |
fix(showcase): un-skip gen-ui-interrupt tests — fix provisional agent race + resolve timing
Wait for CopilotKit runtime POST to complete before interacting so messages aren't silently dropped by the provisional agent stub. Defer resolve() via setTimeout so React commits the picked/cancelled badge before useInterrupt unmounts the card. Add candidateSlots() to the TS interrupt-agent to match the Python agent. Parse JSON-stringified interrupt values in interrupt-headless. Default playwright configs to local aimock. |
||
|
|
a805a8468f |
feat(shell-docs): framework-specific setup snippet system
Replace the LangGraph-flavoured <InstallSDKSnippet> / <InstallPythonSDK>
pattern with a package-owned setup mechanism:
- <FrameworkSetup concept="X" /> resolves
showcase/integrations/<framework>/docs/setup/X.mdx at render
time and returns null when the file is missing (silent absence).
- <DemoCode file="..." region="..." /> embedded in a concept file
pulls a live source excerpt from the same integration package, with
Shiki highlighting via the existing rehype-code pipeline (a static
source-rewrite pass expands the JSX into a fenced markdown block
before MDXRemote sees it).
- currentFramework is bound by DocsPageView's per-render override on
the components map - same pattern as MdxFrameworkOverview. Mirrored
in the framework-root after-features.mdx render.
- 6 agnostic root pages instrumented with one <FrameworkSetup> slot
each (frontend-tools, shared-state, human-in-the-loop, agent-config,
programmatic-control, multi-agent/subagents).
- LGP ships docs/setup/copilot-middleware.mdx as the proof-point with
a # region: middleware marker on src/agents/frontend_tools.py;
other frameworks ship nothing (slot renders silently).
Concept files resolve per package (not per docs folder) - LangGraph
variants share docs CONTENT under content/docs/integrations/langgraph/,
but each package owns its own source tree and therefore its own
docs/setup/ files. LGTS / Fastapi ship their own concept files when
their owners audit.
New Vitest setup in shell-docs covers extractRegion language dispatch,
duplicate-region handling, unterminated-region throws, resolveSetupConcept
path-traversal guards, and the rewriteDemoCode static-prop pre-expansion.
32 tests, all green.
The 18 legacy <InstallSDKSnippet> / <InstallPythonSDK> callers stay on
the old mechanism; the migration is a separate PR.
--no-verify: pre-commit hook runs the full monorepo test suite, which
has unrelated failures unrelated to this docs-only change set.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
3889d934e8 |
fix(showcase/langgraph-python): fix selector mismatches in e2e tests
tool-rendering-default-catchall: page.tsx had inline 3-pill config but
suggestions.ts exists with 4 pills (including "Chain tools"). Switched
page.tsx to import useSuggestions() from ./suggestions so all 4 pills
render, matching the test expectations.
frontend-tools: test used stale selectors ("background-container",
"var(--copilot-kit-background-color)", "Change background" pill) that
didn't match the actual demo code. Updated test to use the real
data-testid ("frontend-tools-background"), real default ("#4f46e5"),
and real pill names ("Sunset/Forest/Cosmic theme").
|
||
|
|
34b641874d |
fix(showcase): unified hoist across all integrations and sibling snippet files
Run the unified hoist codemod over showcase/integrations/* and adjacent source roots (src/lib, src/agent, src/mastra, src/main/java for Spring AI, agent/ for ms-agent-dotnet). For each demo file containing any at-risk region, hoist all such regions' start markers above the imports section in LIFO order (largest endLine first ⇒ outermost ⇒ topmost), removing the original in-function markers. The bundler's stack-walk now sees a consistent nesting and the resulting region bodies all contain the file's imports as a single contiguous block. Also extends marker-move-up support to Java (import) and C# (using-directive) files for Spring AI and ms-agent-dotnet's tool/agent classes. Manually handles two remaining sibling snippet files (built-in-agent::a2ui-fixed-schema's a2ui-backend.snippet.ts) where the 'imports' are declare-const stubs that the codemod doesn't detect as imports. After this commit, of the 32 at-risk (cell, region) tuples flagged in the QA report, 503 (integration × region) bundle slots have imports in their bodies; 4 slots remain without imports because the source files genuinely have no import statements (string-only prompt files in claude-sdk-typescript subagents-prompts.ts). Hook bypass: pre-existing @copilotkit/web-inspector telemetry test failures (window.localStorage + jsdom) are unrelated to this commit. |
||
|
|
e7cb02bfdd |
fix(showcase): include imports in demo region snippets across integrations
Apply marker-move-up across 260 demo files in 17 integrations. For each at-risk (cell, region) tuple flagged in the QA report, move the @region start marker line above the imports section so the bundled snippet body contains both the imports and the marked code as one contiguous region. End markers stay where they are. Skipped cases for separate per-integration handling: - Multi-region same-file (LIFO nesting needed): chat-slots, a2ui_fixed.py, tool-rendering/page.tsx, hitl-in-chat/page.tsx, subagents.py, voice route.ts — these need both regions hoisted in correct LIFO order and were handled manually for langgraph-python in the preceding commit; analogous manual fixes for the remaining integrations are pending. - Files where the target region is already wrapped by an outer region (e.g. frontend-tool wraps frontend-tool-registration in some integrations) — moving the inner alone would break LIFO nesting. Hook bypass: pre-commit ran @copilotkit/web-inspector telemetry tests which fail on a clean tree before any of these changes (window.localStorage not initialised under jsdom in some test cases). Pre-existing failure unrelated to this commit. |
||
|
|
240d881de8 |
fix(showcase/langgraph-python): include imports in demo region snippets
Move @region start markers above each demo file's imports so the bundled region body contains both the imports and the marked code as one contiguous block. Without this, snippets rendered in shell-docs were missing the imports they depended on (z, useState, tool, etc.), forcing readers to guess where each symbol came from. Where two regions share the same file and were sequential (not nested) in the original source, both start markers now sit at the top in proper LIFO nesting order, and the original in-function start markers are removed to avoid duplicate region slices being concatenated by the bundler. Affected regions in langgraph-python: - frontend-tool-registration (frontend-tools/page.tsx) - definitions-zod, create-catalog, provider-a2ui-prop (declarative-gen-ui) - definitions-types, catalog-creation, backend-schema-json-load, backend-render-operations (a2ui-fixed-schema + a2ui_fixed.py) - sandbox-function-registration (open-gen-ui-advanced) - bar-chart-renderer (gen-ui-tool-based) - render-weather-tool, render-flight-tool, weather-tool-backend (tool-rendering + tool_rendering_agent.py) - headless-useinterrupt-primitives (interrupt-headless) - hitl-hook, time-slots (hitl-in-chat) - backend-interrupt-tool, frontend-useinterrupt-render (gen-ui-interrupt + interrupt_agent.py) - subagent-setup, supervisor-delegation-tools (subagents.py) - context-provider-sketch (readonly-state-agent-context) - state-streaming-middleware (shared_state_streaming.py) - transcription-service-guard, voice-runtime (voice route.ts) Hook bypass: pre-commit ran @copilotkit/web-inspector telemetry tests which fail on a clean tree before any of these changes (window.localStorage not initialised under jsdom in some test cases). Pre-existing failure unrelated to this commit. |
||
|
|
fcc2cef9b2 |
fix(showcase): simplify health endpoints to local-only (no agent proxy)
All 18 integration health endpoints previously proxied to the backend agent /health with a 3s timeout, causing false reds when agents were slow but functional. The harness already checks agent reachability via the agent:<slug> probe. Health endpoints now return a simple 200 confirming the Next.js process is alive. |
||
|
|
2482317ccc |
style: apply ruff format to Python codebase
320 files reformatted. One-time alignment to match the ruff format check added to CI in #4812. |
||
|
|
c41c2dec71 |
fix(showcase/beautiful-chat): pin canvas to "beautiful-chat" agent id so shared state renders
`<CopilotKit agent="beautiful-chat">` routes the chat to agent id
"beautiful-chat", but ExampleCanvas called `useAgent()` with no args and
fell back to DEFAULT_AGENT_ID ("default"). The frontend's agent registry
tracks state per id, so `manage_todos` state-deltas from the chat run
landed on "beautiful-chat" and never reached the canvas's "default"
subscription — the Task Manager pill auto-flipped the panel to App mode
but the To Do column stayed empty. Drop the unused "default" alias from
the runtime route and pin the canvas to `useAgent({ agentId:
"beautiful-chat" })` so both halves share one ProxiedCopilotRuntimeAgent
instance. Adds a Playwright regression test asserting the 3 verbatim
todo titles render after the pill click, plus 3 aimock fixtures for the
multi-turn flow (enableAppMode -> manage_todos -> confirmation).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
cdb5f44fb5 |
test(reasoning-chain): add regression coverage at runtime, harness, and e2e layers
Three layers of regression guards for the runtime reasoning-role filter and
the chained demo behavior:
1. Runtime unit test — packages/runtime/.../run-message-filtering.test.ts:
- Verifies `LangGraphAgent.run` strips `role:"reasoning"` from
`input.messages` before delegating to super.run.
- Verifies user/assistant/system/tool messages pass through in order.
- Verifies empty + missing messages arrays are tolerated.
- Verifies pre-existing forwardedProps.streamSubgraphs default + override
behavior is preserved.
- 6/6 tests pass against the runtime package's vitest config.
2. D5 harness probe — showcase/harness/.../d5-tool-rendering-reasoning-chain.ts:
- Expanded from one chained turn (flights→weather) to all three chained
pills in a single thread (stocks AAPL→MSFT, dice d20→d6, flights→weather).
- This is the canonical multi-pill regression at the harness layer:
without the runtime reasoning-role filter, the second pill would crash
before the model was called.
- Each turn asserts the per-turn delta of reasoning-block mounts (idx+1),
the minimum card count for each tool group, and unique transcript
substrings that scope to that turn.
3. Playwright e2e spec — showcase/integrations/langgraph-python/tests/e2e/
tool-rendering-reasoning-chain.spec.ts:
- Mirrors the pattern of the sibling tool-rendering-default-catchall spec
(notably its multi-pill regression at lines 162-212).
- Page-loads test verifies the 3 pills mount and no cards leak from a
prior session.
- One test per chained pill (stocks, dice, flights+weather) asserts the
full chain renders with reasoning-block + correct per-tool cards +
narration matching the aimock fixture text.
- Sequential-pills regression test clicks all 3 pills in one thread,
asserts each chain renders independently AND the reasoning-block count
increases monotonically across turns.
Agent: extend `get_stock_price` to accept optional `price_usd` and
`change_pct` arguments (mirrors the basic tool-rendering agent's signature
introduced in #4770). The aimock fixtures script the chained AAPL/MSFT
comparison by passing deterministic prices via these args; without the
wider signature, pydantic rejects the tool call and the card never mounts.
The runtime unit test is the strongest guard — it would catch any
regression on the role-filter logic without depending on the full Docker
stack. The harness probe and Playwright spec catch end-to-end regressions
in the canonical CI environment.
|
||
|
|
bb6554c433 |
fix(showcase/langgraph-python): chain reasoning-chain demo pills end-to-end
The tool-rendering-reasoning-chain demo previously promised chained tool
calls in its pill titles but the agent and fixtures only delivered single
tools — clicking "Weather + flights to Tokyo" produced just a WeatherCard,
"Compare two stocks" only fetched AAPL, "Find flights from SFO to JFK"
showed flights but no destination weather. Three changes close the gap.
Agent: replace the soft "call 2+ tools when relevant" system prompt with
concrete per-pill chain examples mirroring the pattern already used by the
langgraph-typescript `tool-rendering` agent (weather→flights, ticker→peer,
roll→contrast die, flights→destination weather).
Pills: drop the redundant Tokyo pill (it was the SFO/JFK chain in reverse)
and reword each remaining pill message to PRE-DISCLOSE the chain so the
model commits to the follow-up call:
- "Compare AAPL and MSFT stocks for me."
- "Roll a 20-sided die for me and compare it to a smaller one."
- "Find flights from SFO to JFK and show me the weather there."
Fixtures: 9 fixtures (3 per pill: final-content → second-leg → first-leg,
ordered by toolCallId specificity for first-match-wins). Each fixture is
scoped by a langgraph-python-UNIQUE userMessage tail ("Compare AAPL and
MSFT stocks", "compare it to a smaller one", "show me the weather there").
Those substrings appear nowhere else across the 14+ integrations sharing
showcase-aimock on Railway, so the new fixtures cannot cross-contaminate
the other reasoning-chain demos that still ship the older prompt set.
A toolName-based gate was considered and rejected because most fleet
agents register `roll_dice` and aimock's `toolName` matcher is a tool-LIST
gate, not a tool-CALL gate — it would NOT have isolated this demo.
Probe: collapse the two-turn flow (Tokyo + SFO/JFK) into one chained turn
(SFO→JFK + JFK weather) that asserts BOTH per-tool renderers
(FlightListCard + WeatherCard) mount in a single response. Same coverage
at half the wall-clock and exercises the actual chained-tool path.
|
||
|
|
04d8008ea7 |
fix(showcase/langgraph-python): unbreak shared-state pills, auth sign-out, gen-ui-agent progression, multimodal D5
Four independent showcase production bugs Alem reported, plus the
D5 multimodal harness regression they unblocked.
Shared-state-read-write: "Greet me" ("Say hi and introduce yourself.")
and "Plan a weekend" ("Suggest a weekend plan based on my interests.")
were matching the bare `hi` and `plan` catch-alls in feature-parity.json
and returning the generic showcase-assistant blurb / 5-step content plan
instead of shared-state-aware responses. Added pill-specific fixtures in
shared-state.json (mirrored into d5-all.json) so the longer userMessage
substrings win first-match-wins ahead of feature-parity.
Auth sign-out: signing out unmounted CopilotKit entirely and bounced
the user back to the SignInCard, so the demo never showcased the
runtime returning 401 — its whole point. The QA contract in
qa/auth.md spelled out the intended UX. Restored it: CopilotKit stays
mounted after the first sign-in, the AuthBanner flips to an amber
"Signed out — the agent will reject your messages" state with a
re-Sign-in button, and CopilotKit's `onError` callback drives a
`data-testid="auth-demo-error"` surface that displays the runtime's
401 the moment the user sends an unauthenticated message. Updated the
e2e spec to match (the old "SignInCard re-mounts after sign-out" test
pinned the regression).
Gen-ui-agent: the aimock fixture short-circuited the 7-step
progression spelled out in `gen_ui_agent.py`'s SYSTEM_PROMPT to a
single set_steps call with all three steps already `completed`, so
the InlineAgentStateCard rendered the final 3/3 state instantly with
no sequential pending → in_progress → completed animation.
Regenerated as a 7-leg toolCallId chain per pill (8 fixtures × 3
pills): seed leg keyed on userMessage with NO `hasToolResult` gate
(matching PR #4770's pattern — `hasToolResult: false` would block the
seed from firing on the second pill in a multi-pill session), then
six toolCallId-keyed transitions, then a final narration. Fixture
order: toolCallId legs FIRST so the most specific match wins.
Multimodal D5: the sample-attachment buttons auto-send via
`agent.addMessage + copilotkit.runAgent` (restored in PR #4761), but
the D5 harness still typed `input` + pressed Enter via the runner
after `preFill`, sending a second user message that competed with the
in-flight image upload — the v1 LangGraph runtime SSE stream got
tangled (browser DevTools showed `statusCode: pending` indefinitely)
and the assistant message never rendered. Added `skipSend?: boolean`
to ConversationTurn (distinct from `skipFill`, which still presses
Enter once the textarea has content) and switched d5-multimodal.ts to
`skipSend: true` with `responseTimeoutMs: 60_000` so the runner waits
on the assistant response without poking the chat further. Bumped the
PDF auto-prompt fixture in feature-parity.json to include the word
"document" so the existing `buildModalityAssertion("document")` check
still lands.
D5 result: 37 → 39 of 40 features passing. Only
`tool-rendering-reasoning-chain` remains and is a separate
agent/runtime bug (Tokyo Responses-API `reasoning` message survives
into the next turn's conversation history, runtime returns
`RUN_ERROR: "message role is not supported"`).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
792ae78799 |
fix(showcase/langgraph-python): remove dead "Show reasoning" pill from chat-slots
The chat-slots cell is wired to the neutral sample_agent graph (plain ChatOpenAI, no Responses API, no reasoning config), so it never emits AG-UI REASONING_MESSAGE_* events. The pill could never light up the wrapped messageView.reasoningMessage slot, and its prompt didn't match any fixture in showcase/aimock/d5-all.json — aimock-backed runs hit "No fixture matched". Drop the pill (the QA doc and the e2e spec already only expect "Write a sonnet" and "Tell me a joke") and leave a note pointing reasoning demos at /demos/reasoning-default and /demos/reasoning-custom where the dedicated reasoning_agent graph lives. |
||
|
|
9bbcf876d8 |
fix(showcase/langgraph-python): drop nested-flex-gap arbitrary variant
The `[&_div[style*='flex-direction:_row']]:gap-4` arbitrary variant (quotes inside doubly-nested brackets) is the most exotic Tailwind syntax in this PR and lines up exactly with when the Vercel form-filling deploy started failing. Tailwind v4's content scanner is likely choking on the apostrophes in the nested attribute selector. The Metric `flex-1 min-w-[120px]` and the chart `flex-1 min-w-0` already give us even distribution inside the basic catalog's gap-less Row; the auto-injected nested gap was nice-to-have, not load-bearing. |
||
|
|
cc6dd5c4a9 | Merge branch 'main' into fix/showcase-declarative-gen-ui-card-width | ||
|
|
ba5369c930 |
refactor(showcase/langgraph-python): make A2UI renderers fill their slot instead of widening the chat
Drop the inline <style> override that widened the chat's `cpk:max-w-3xl` column and the outer max-w-6xl bump. The chat keeps its normal width; the real bug was that the basic catalog's Row/Column primitives are bare `display: flex` divs with no gap and no min-width control on children, so when the agent dropped multiple Metrics or charts into a Row they collapsed to content width and looked glued together. Four targeted fixes inside our renderers: - Metric gains `flex-1 min-w-[120px]` so a row of KPI tiles distributes the available width evenly inside the gap-less basic Row, instead of shrinking to content. - PieChart and BarChart switch from a hardcoded max-w to `flex-1 min-w-0` so two charts side-by-side each take half the card column with Recharts' ResponsiveContainer doing the rest, instead of one chart insisting on 640px and overflowing. - Card's CardContent picks up a Tailwind arbitrary variant `[&_div[style*='flex-direction:_row']]:gap-4` (plus the column equivalent) that injects a gap into any nested basic Row/Column the agent drops in. Underscores in the arbitrary value compile to literal spaces, matching React's serialized inline `flex-direction: row`. - Card itself drops the old `min-w-[260px]` floor in favour of `min-w-0` so it cooperates if the agent ever stacks Cards horizontally inside a Row. |
||
|
|
bedc53a652 |
fix(showcase/langgraph-python): widen declarative-gen-ui surface and polish InfoRow
The chat shell caps its scroll column at cpk:max-w-3xl (~768px), which left A2UI-generated cards (KPI dashboards, charts, status reports) feeling pinched on the declarative-gen-ui demo. Locally widen that wrapper to 64rem via a scoped attribute selector on the demo and bump the outer page wrapper from max-w-4xl to max-w-6xl so the card column actually has room to grow. While here, fix the InfoRow trailing-separator artifact: each row now draws its own border-bottom with last:border-b-0 so the final row in a Card (e.g. the Status Report demo) no longer leaves a dangling line, regardless of whether the agent wraps the rows in a Column or drops them directly into the Card child slot. Right-align the value with tabular-nums for cleaner stacks. Card itself gains w-full overflow-hidden so it stretches into the now-wider column instead of sitting at its min-width. |
||
|
|
25f1f921a8 |
fix(showcase): unbreak multimodal demo end-to-end (auto-send, dedupe, proxy) (#4761)
## Summary Re-lands the multimodal-attachments fix from #4584 (May 1, never merged) onto current `main`, ported to the post-refactor file layout where `page.tsx` was split into `legacy-converter-shim.tsx`, `multimodal-chat.tsx`, and `file-to-data-attachment.ts`. Auto-send was the visible regression: clicking **Try with sample image / Try with sample PDF** only queued the attachment chip instead of sending the canned prompt. This PR restores the full end-to-end behavior plus five regression tests so it can't silently break again. ## What was broken and what changed 1. **Random uploads crashed with `Failed to fetch`.** aimock returned HTTP 404 on no-match, the LangGraph SDK surfaced `NotFoundError`, the AG-UI stream surfaced a `RUN_ERROR`, the demo crashed. → Added `--proxy-only` + `--provider-openai https://api.openai.com` to the local aimock command so unmatched user prompts fall through to real OpenAI (mirrors Railway). 2. **Bundled-sample fixtures keyed on user-visible canned prompts.** Auto-prompts are now natural and specific ("can you tell me what is in this demo image/pdf I just attached") so they render cleanly as the user message bubble AND can't collide with arbitrary user prompts — random uploads phrase questions differently and fall through to the proxy. 3. **Sample buttons now auto-send via `useAgent`.** The previous DataTransfer path queued the attachment via the chat's hidden file input but required clicking send while the attachment was still uploading — `CopilotChat.onSubmitInput` rejects submits during upload AND clears the input regardless, so the canned prompt was eaten. Rewrite calls `agent.addMessage(...)` + `copilotkit.runAgent({ agent })` directly with the base64'd content part. 4. **PDF flattened text bled into the rendered user message.** `_PdfFlattenMiddleware` ran in `before_model` and persisted the rewrite to agent state. Switched to `wrap_model_call` so the PDF→text rewrite is scoped to the model request only. 5. **Attachments doubled (and PDFs rendered as broken `<img>`).** The `@ag-ui/langgraph` round-trip mis-tags PDFs as `image` and re-injects the user's original modern part, doubling chips. Added `dedupeUserMessageMedia` subscriber on `onMessagesSnapshotEvent` + `onRunFinalized` to dedupe by `source.value` and re-key type from mimeType. Also flipped `onRunInitialized` from REPLACE to APPEND so the modern part stays for the UI alongside a legacy `binary` sibling for the converter. 6. **Regression suite (`tests/e2e/multimodal.spec.ts`).** Five focused tests, all pass against live local stack (15.4s): - page loads with all expected affordances - sample image: auto-sends, EXACTLY ONE `<img>`, assistant references the logo - sample PDF: auto-sends, EXACTLY ONE `DocumentAttachment` chip ("PDF" label), NO `<img>`, no `[Attached document]` text bleed - image then PDF in the same session: each message keeps its own single chip - PDF then image in the same session: symmetric ## Test plan - [x] `showcase up langgraph-python` — both sample buttons auto-send; image renders as `<img>`, PDF renders as PDF chip; random paperclip uploads go through proxy - [x] `BASE_URL=http://localhost:3100 CI=1 npx playwright test multimodal.spec.ts` — **5 / 5 passing** - [ ] Post-merge: e2e-deep cycle for langgraph-python multimodal cell stays green ## Closes Closes #4584. |
||
|
|
7c3edca2b7 |
fix(showcase): unbreak multimodal demo end-to-end (sample buttons auto-send, dedupe, proxy)
The langgraph-python multimodal-attachments demo had a stack of bugs that compounded each other. Fixing them required touching the local docker-compose, the aimock fixtures, the LangChain middleware, the client-side AG-UI shim, and the sample-attachment buttons. This commit lands the full set together because they only make sense as a unit — verified end-to-end against `showcase up langgraph-python` in a headed browser. New e2e suite pins each regression. Supersedes #4584 (the original fix from May 1 that never landed — this is a fresh port onto the post-refactor file layout where page.tsx is split into legacy-converter-shim.tsx, multimodal-chat.tsx, file-to-data-attachment.ts). What was broken and what changed: 1. Random uploads crashed with `Failed to fetch`. aimock returned HTTP 404 on no-match, the LangGraph SDK surfaced `NotFoundError`, the AG-UI stream surfaced a `RUN_ERROR`, the demo crashed. Added `--proxy-only` + `--provider-openai https://api.openai.com` to the local aimock command so unmatched user prompts fall through to real OpenAI (mirrors the Railway aimock setup). 2. Bundled-sample fixtures keyed on user-visible canned prompts. The auto-prompts are deliberately long, specific, and natural- reading ("can you tell me what is in this demo image/pdf I just attached") so they (a) render cleanly as the user message bubble, and (b) can't collide with arbitrary user prompts — random uploads phrase questions differently and fall through to the proxy. 3. Sample buttons now auto-send via `useAgent`. The previous DataTransfer-based path queued the attachment via the chat's hidden file input, then required clicking send while the attachment was still uploading — `CopilotChat.onSubmitInput` rejects submits during upload AND clears the input regardless, so the canned prompt was eaten. Rewrite to call `agent.addMessage(...)` + `copilotkit.runAgent({ agent })` directly with the base64'd content part, sidestepping the upload race entirely. 4. PDF flattened text bled into the rendered user message. `_PdfFlattenMiddleware` ran in `before_model` and returned `{"messages": rewritten}`, which persisted to agent state. The chat UI then rendered the `[Attached document]\n<pdf body>` text part inline with the user prompt. Switched to `wrap_model_call` so the PDF→text rewrite is scoped to the outgoing model request only and never pollutes state. 5. Attachments doubled (and PDFs rendered as broken `<img>`). The `@ag-ui/langgraph` round-trip translates outgoing `binary` parts to LangChain `image_url` and incoming `image_url` back to `image` AG-UI parts — regardless of mimeType, so PDFs came back as `type: "image"` with `mimeType: "application/pdf"` and were forced into `ImageAttachment`, where the load failed and the chat showed two "Failed to load image" boxes. Plus the user's original modern part survived alongside the round-tripped one, doubling visible chips. Added a `dedupeUserMessageMedia` subscriber on both `onMessagesSnapshotEvent` and `onRunFinalized` to: - dedupe media parts by `source.value` so the local + round- tripped copy collapse to one chip - re-key part `type` from `mimeType` so PDFs route to `DocumentAttachment` (icon + filename) and images to `ImageAttachment`. Also flipped the `onRunInitialized` shim from REPLACE to APPEND — keep the modern part for the UI AND emit a legacy `binary` sibling for the converter. 6. Regression suite (`tests/e2e/multimodal.spec.ts`). Replaces the pre-rewrite suite with five focused tests: - page loads with all expected affordances - sample image: auto-sends, EXACTLY ONE `<img>`, assistant references the logo - sample PDF: auto-sends, EXACTLY ONE `DocumentAttachment` chip ("PDF" label), NO `<img>`, no `[Attached document]` text bleed - image then PDF in the same session: each message keeps its own single chip, no cross-contamination - PDF then image in the same session: symmetric All 5 pass against the live local stack (15.4s). |
||
|
|
985bebf39c |
fix(showcase/voice): use real Whisper for mic transcription
The voice route's OpenAI client previously fell through to OPENAI_BASE_URL, which docker-compose.local.yml sets to http://aimock:4010/v1. Aimock has a catchall transcription fixture that returns "What is the weather in Tokyo?" for every audio file, so the mic button always produced that phrase no matter what the user actually said. Pin baseURL to real OpenAI (overridable via OPENAI_TRANSCRIPTION_BASE_URL). The sample-audio button stays as synchronous text injection — that's the documented design, and what the e2e + d5 probe rely on. Also: - Tidy the sample button label ("Try a sample question" -> "Try a sample audio") so the affordance matches what it does. - Realign tests/e2e/voice.spec.ts with the shipped component (the voice-sample-audio container testid and Sample: "..." caption it asserted on never existed on HEAD) and add cold-start timeout headroom for the mic-button render and the agent-flow test. - Add "env": ".env" to langgraph.json so langgraph_cli dev picks up OPENAI_API_KEY locally. Docker/Railway paths inject env vars directly so this is a no-op there. |
||
|
|
32d0237cb5 |
fix(shell-docs): cutover-blocker fixes from Phase 4 validation (#4741)
## Summary
Phase 4 validation (run today against `docs.showcase.copilotkit.ai`)
surfaced four cutover-blocking issues. This PR fixes all of them in four
focused commits.
## Commits
1. **chore(showcase): close last 2 yellow Missing snippet boxes for
cutover** — adds in-place region markers to
`langgraph-python::frontend-tools` and a sibling
`slot-overrides.snippet.tsx` teaching file for
`langgraph-python::chat-slots`. Both pages now render zero `Missing
snippet` warnings.
2. **fix(shell-docs): correct feature-viewer slug + demo-id translation
for code tab** — adds `getFeatureViewerSlug()` and
`getFeatureViewerDemoId()` helpers in `registry.ts`, with explicit
override maps for the 8 framework-name mismatches (built-in-agent,
google-adk, claude-sdk-{python,typescript}, ms-agent-{python,dotnet},
crewai-crews, llamaindex) and 5 demo-id mismatches (gen_ui_tool_based,
gen_ui_agent, shared_state_streaming, shared_state_read_write,
hitl_in_chat). When the framework or demo has no feature-viewer
equivalent, the helper returns null and the Code tab is hidden on that
page (graceful degradation; Demo tab still renders).
3. **fix(shell-docs): redirect catalog hygiene (self-loops +
framework-scoped gap)** — removes 31 self-redirect entries from
`seo-redirects.ts` (where source === destination caused infinite 301
loops on canonical URLs like /frontend-tools, /faq, /human-in-the-loop).
Reorders middleware logic so the redirect catalog is consulted before
the framework-scoped short-circuit fires, fixing 97 framework-scoped
slug-rename URLs that were soft-404ing instead of redirecting. Adds
defense-in-depth skip-when-equal guard in middleware.
4. **fix(shell-docs): resolve 53 sitemap 500s before cutover** — three
independent root causes in MDX rendering:
- 4 missing \`<Component />\` registrations in \`mdx-registry.tsx\`
(CopilotCloudConfigureCopilotKit,
SelfHostingCopilotRuntimeConfigureCopilotKit, CloudCopilotKit, Content)
— affected ~37 pages.
- 3 langgraph tutorial pages had markdown lists immediately preceding
JSX closing tags, causing remark to bail. Fix = blank line between
bullet and close tag (~9 pages).
- 5 MDX files used escaped JSX comments that Acorn cannot parse —
replaced with proper unescaped form (~7 pages).
## Verification
- Production build clean: 27 static pages generated, 0 errors
- Local probes: all 53 sitemap-500 URLs return 200; all 31 former
self-loop URLs return 200; all 85 framework-scoped non-loop catalog
entries 301 to expected destinations
- Code tab: 47/60 (framework × demo) URLs resolve to real code panels
post-fix; remaining 13 are feature-viewer per-framework deploy-coverage
gaps in the \`ag-ui-protocol/ag-ui\` repo, not this one
- Both langgraph-python pages render 0 \`Missing snippet\` boxes after
re-bundling demo-content
## Test plan
- [ ] CI passes
- [ ] Sitemap probe: every URL in \`/sitemap.xml\` returns 200
- [ ] Redirect probe: spot-check 5 framework-scoped slug renames (e.g.
\`/agno/frontend-actions\`, \`/pydantic-ai/use-agent-hook\`)
- [ ] Self-loop probe: \`/frontend-tools\` returns 200, not infinite 301
- [ ] Code tab: visit \`/built-in-agent/frontend-tools\` (tab should be
hidden), \`/langgraph-python/agentic-chat\` (tab should render code),
\`/llamaindex/agentic-chat\` (tab should render after llama-index slug
rename)
- [ ] Yellow boxes: visit \`/langgraph-python/frontend-tools\` and
\`/langgraph-python/custom-look-and-feel/slots\` — zero Missing snippet
warnings
|
||
|
|
5c0a87e18b | Merge branch 'main' into test/lgp-a2ui-regression-coverage | ||
|
|
e831c72a8f | style: auto-fix formatting | ||
|
|
70e2fb13c8 |
refactor(showcase): rename byoc-* slugs to declarative-* + sort index by manifest features
User-facing renames so the showcase reads the way a cold visitor would
expect:
- `byoc-hashbrown` → `declarative-hashbrown` (and `byoc-json-render` →
`declarative-json-render`). The display titles already said
"Declarative UI: …"; only the URL slugs and folder paths still
leaked the internal BYOC ("Bring Your Own Components") jargon.
Renamed:
/demos/byoc-hashbrown → /demos/declarative-hashbrown
/demos/byoc-json-render → /demos/declarative-json-render
/api/copilotkit-byoc-* → /api/copilotkit-declarative-*
src/app/demos/byoc-* → src/app/demos/declarative-*
qa/byoc-*.md → qa/declarative-*.md
tests/e2e/byoc-*.spec.ts → tests/e2e/declarative-*.spec.ts
Internal Python module names + langgraph graph IDs stay legacy
(`byoc_hashbrown_agent.py`, `byoc_hashbrown`) — those are not
user-facing and renaming them is a separate cross-codebase pass.
- `a2ui-fixed-schema` slug intentionally unchanged.
- Tool Rendering trio parenthetical rename (Default → Catch-all →
Custom progression reads clearly as "how much do I customize?"):
Tool Rendering (Default) — unchanged
Tool Rendering (Custom default) → Tool Rendering (Catch-all)
Tool Rendering (Specific) → Tool Rendering (Custom)
- `tool-rendering-reasoning-chain` cell renamed from
"Generative UI: Rendering multiple tools" to
"Generative UI: Tool calls + reasoning" (the demo is about combining
reasoning + tool rendering, not about quantity of tools).
- `Open Generative UI: Default` / `Open Generative UI: Custom`
descriptions expanded so a visitor understands how Open Generative UI
differs from Tool Rendering (agent composes UI from a registered
library vs. attaching a renderer to a *named* backend tool).
- Showcase index now sorts demos within each tag by `manifest.features`
order. Previously demos appeared in manifest declaration order, which
ignored the team's curated "polished flagship → simplest start →
variants" arc.
Cross-cutting registry / harness / dashboard updates that fall out of
the rename:
- `shared/feature-registry.json` adds the two new IDs alongside the
legacy `byoc-*` (so the catalog stays valid; the other 17
integrations still declare `byoc-*` in their manifests).
- `shared/constraints.yaml` adds the new IDs to the
generative-ui-approach allow-list.
- `scripts/__tests__/generate-catalog.test.ts` updates the cell-count
expectations (45 features × 18 integrations = 810; 792 after docs-
only exclusion; 45 LGP cells = 38 wired + 1 stub + 6 unshipped).
- Harness probe `d5-byoc.ts` + `d5-byoc.test.ts` now route both slug
families through `preNavigateRoute` and exercise the new branches.
- `d5-feature-mapping.ts` and `shell-dashboard/live-status.ts` mirror
the dual-ID mapping so both legacy and renamed slugs roll up under
the same `byoc` D5 featureType.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
08ba26590e |
refactor(showcase/langgraph-python): clean up agents + demos for didactic clarity
Extracts duplicated logic into shared modules so each demo file reads
as the feature it teaches, not the boilerplate around it.
Python agents:
- `_a2ui_utils.py` (new) — `sanitize_a2ui_components` and
`has_root_component`, consumed by both `a2ui_dynamic.py` and
`beautiful_chat.py` (these previously duplicated the same defensive
validator inline).
- `byoc_hashbrown_prompt.py` (new) — the 56-line system prompt extracted
from `byoc_hashbrown_agent.py` so the agent file stays focused on the
`create_agent(...)` wiring.
- `beautiful_chat.py` secondary `ChatOpenAI` now passes
`streaming=True` (matching `a2ui_dynamic.py`) so aimock's SSE-only
fixture matcher sees the call in replay mode — without it the demo
surfaced "An internal error occurred" on every load.
- `multimodal_agent.py` — top-level `from pypdf import PdfReader` (was
lazy with three layers of `# pragma: no cover` exception handling
for stages that never failed independently); kept a single log line
at the outer except so Railway logs stay triageable.
- `gen_ui_agent.py` — switched from `deepagents.create_deep_agent` (whose
planner+sub-agent middleware ate enough supersteps to trip LangGraph's
default recursion limit on this single-tool ReAct loop) to plain
`langchain.agents.create_agent`. Comment explains the math.
- `tool_rendering_agent.py` — docstring no longer claims to back the
`tool-rendering-reasoning-chain` cell (it has its own agent file).
TypeScript / TSX demos:
- `_shared/parse-json-result.ts` (new) — extracted from
`tool-rendering/parse-json-result.ts`; now consumed by three demos.
- `_shared/slot-override.ts` (new) — `makeSlotOverride<T>` centralizes
the 11 `as unknown as` casts that `chat-slots/page.tsx` previously
scattered across the slot-override block.
- `shared-state-read` — extracted `RecipeCard` component + `types.ts`,
dropped the dual-state-sync pattern that had the read-only demo
locally mutating recipe state. Page is now a thin shell that publishes
edits via `agent.setState` and reads back via `agent.state.recipe`.
Also surfaces `runAgent` rejections via `console.error` instead of
the previous silent `.catch(() => {})`.
- `headless-complete/hooks/use-auto-scroll.ts`,
`headless-complete/hooks/use-typing-indicator.ts` (new) — extracted
from `chat.tsx`. The chat file shrinks by ~40 lines. Same silent-
rejection fix on `runAgent` as shared-state-read.
- `frontend-tools-async/fake-notes-db.ts` (new) — extracted from
`page.tsx`; the demo file no longer leads with 60 lines of fake-DB
scaffolding before `useFrontendTool` (the actual feature) appears.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
a3d93266e9 |
fix(showcase/langgraph-python): correctness fixes for demos
Real demo-time bugs in the langgraph-python integration:
- `interrupt-headless` was rendering hardcoded stale slot dates from a
`DEFAULT_SLOTS` constant instead of reading `payload.slots` that the
backend `interrupt(...)` already supplies. Sibling `gen-ui-interrupt`
does this correctly. Headless variant now matches.
- `_shared/interrupt-fallback-slots.ts` (new) — the JS fallback for
when the backend returns no `slots` array. Generates relative to
Date.now() so the picker never shows past dates. Used by both
`interrupt-headless` and `gen-ui-interrupt`. Deletes the old stale-
literal `gen-ui-interrupt/fallback-slots.ts` (dates from April that
had already decayed).
- `interrupt_agent.py` — replaced hardcoded `timezone(timedelta(-7))`
PDT with `zoneinfo.ZoneInfo("America/Los_Angeles")` so the demo
doesn't lie about offsets in winter. Also fixed a Sunday edge case
where `next_monday` collapsed to the same date as `tomorrow` (both
Python and the JS fallback) — added a `<= 1` skip-a-week guard.
- `package.json#dev` ran `uvicorn agent_server:app --port 8000` but
`src/agent_server.py` was a 3-line stub with no `app` symbol AND the
API route pointed at port 8123. Replaced with the canonical
`langgraph_cli dev --port 8123` invocation; deleted the dead stub.
- `entrypoint.sh:47` smoke-checked `src/agents/tools.py` (a file that
never existed) so every Railway boot logged a phantom ERROR. Dropped.
- `reasoning_agent.py` and `tool_rendering_reasoning_chain_agent.py`
intentionally omit `CopilotKitMiddleware` (they exercise only
reasoning-token streaming, no frontend tools / app context). Added a
one-line comment so a future maintainer doesn't cargo-cult it back in.
- `subagents._invoke_sub_agent` text-block walker had a regression
where a `{"type":"text","text":null}` payload (a known provider
quirk) would crash `"".join(parts)` with `TypeError: sequence item
N: expected str instance, NoneType found`. Restored the
`isinstance(block.get("text"), str)` guard.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
2f6816fe3f |
feat(showcase/langgraph-python): align chat surfaces + shadcn recipe overhaul (#4753)
## Summary
Four LangGraph Python showcase demos updated:
- **HITL In-app**: `CopilotChat` → `CopilotPopup`
(`defaultOpen={true}`). Tickets panel fills the viewport; chat is a
popup in the corner. Approval dialog still portaled to `<body>`.
- **Shared State: Streaming**: `CopilotChat` (custom aside) →
`CopilotSidebar`. Document view fills the page; sidebar opens by
default.
- **Shared State: Read + Write**: `CopilotPopup` → `CopilotSidebar`.
Card grid breakpoint bumped from `lg:` to `xl:` so cards stack instead
of pinching when the sidebar consumes ~480px on mid-size laptops;
`overflow-y-auto` on the wrapper keeps the page scrollable.
- **Shared State: Read (Recipe)**: visual overhaul with shadcn
primitives (`Card`, `Input`, `Select`, `Textarea`, `Badge`, `Button`,
`Spinner`, `Separator`). All `data-testid` attributes and
QA-doc-asserted strings preserved verbatim ("Make Your Recipe", "AI
Recipe Assistant", "+ Add Ingredient", "+ Add Step", "Improve with AI",
default ingredients/instructions, all 7 dietary preference labels, all 5
cooking-time labels, all 3 suggestions, etc.). React state-sync logic
preserved verbatim (no behavior change).
## Why
The chat surfaces were inconsistent across demos — HITL used a full-pane
chat where the demo's premise is "approval modals appear *outside* the
chat surface", and the shared-state demos mixed `CopilotChat` and
`CopilotPopup` instead of using the prebuilt `CopilotSidebar`. The
Recipe demo was raw Tailwind while the rest of the suite uses shadcn
primitives.
## Plumbing fixes that fell out of the work
- **`clsx` + `tailwind-merge` added to `package.json`** as direct deps.
They were imported by `src/lib/utils.ts` (the `cn` helper used by all
shadcn primitives) but only transitively resolvable, which broke `next
dev --turbopack`'s stricter module resolution.
- **`globals.css` rule**: `body[data-scroll-locked] { padding-right: 0
!important }`. Radix overlays (Select, Dialog, Popover) use
`react-remove-scroll-bar`, which measures the gap between viewport and
body inner width to detect scrollbar size — and our body uses
`margin-inline-end: 480px` to make room for `<CopilotSidebar />`. The
library mis-reads that 480px as scrollbar width and "compensates" with a
matching `padding-right`, which shrinks the content area and shifts the
centered card every time a dropdown opens. We have no real scrollbar
(`body { overflow: hidden }`), so the compensation is unnecessary —
neutralize it. Comment in the diff explains the rationale.
## Test plan
- [x] `next build` clean across all 54 routes
- [x] `oxlint` reports zero new warnings (only pre-existing patterns the
diff preserved verbatim)
- [x] Manually verified `/demos/hitl-in-app`,
`/demos/shared-state-streaming`, `/demos/shared-state-read-write`,
`/demos/shared-state-read` in browser (Chrome, dev server)
- [x] Verified the Recipe page's Select dropdowns no longer shift the
card when opened (the original motivation for the `globals.css` rule)
- [ ] HITL e2e suite (`tests/e2e/hitl-in-app.spec.ts`) — selectors
preserved (`getByPlaceholder("Type a message")`, suggestion-pill
testids, approval-dialog-* testids); CopilotPopup's `defaultOpen={true}`
mounts the chat input on first paint
- [ ] Streaming and Read e2e suites are stale pre-PR (reference content
that doesn't exist in source — "AI Document Editor", "Sales Pipeline")
and remain in the same state
## CR loop
Ran `cr-loop` (Round 1, 7 agents). Reclassified one finding to bucket
(d) per user direction (npm/pnpm lockfile-convention concern is a
different PR's subject — pre-existing across all
`showcase/integrations/*` and out of scope here). Procedure 3 audit
promoted no items to bucket (a); 24 bucket (c) items confirmed as
pre-existing and routed to follow-up backlog.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
|
||
|
|
3d47397957 |
feat(showcase/langgraph-python): align chat surfaces + shadcn recipe overhaul
Switch HITL In-app to CopilotPopup; Shared State Streaming and Read+Write to CopilotSidebar; widen card-grid breakpoint and add overflow scroll on Read+Write so content stays usable when the sidebar consumes ~480px of viewport. Re-skin the read-only Recipe demo with shadcn primitives, preserving every data-testid and QA-doc-asserted string. Add clsx and tailwind-merge as direct deps (previously only transitively resolvable, breaking turbopack dev) and a globals.css rule to neutralize the Radix scroll-lock padding-right that react-remove-scroll-bar injects when it mis-detects body's margin-inline-end (CopilotSidebar) as scrollbar width. QA docs updated to reference the new chat surfaces. |
||
|
|
46c5b56881 |
fix(showcase/langgraph-python): force reasoning emission in reasoning-chain agent
The tool-rendering-reasoning-chain D5 probe has been red since the probe was added (PR #4743) — the reasoningMessage slot's <ReasoningBlock> never mounts because the agent doesn't emit reasoning summaries. Compared to the working reasoning_agent.py (which powers reasoning-default + reasoning-custom and is green), the only meaningful difference was the reasoning config: reasoning_agent.py (works): reasoning={"effort": "medium", "summary": "detailed"} tools=[] tool_rendering_reasoning_chain_agent.py (was broken): reasoning={"effort": "low", "summary": "auto"} tools=[get_weather, search_flights, ...] `summary: "auto"` lets the model decide whether to emit a reasoning summary. With tools present the model often skips it (chain-of-thought goes straight into the tool call without a summary). `summary: "detailed"` forces emission on every response, which lights up the ReasoningBlock slot. Bumping effort medium so the surface visible to the user reads as a real chain-of-thought, matching reasoning-display. Probe diagnostics from prod confirmed the toolCalls were firing fine (WeatherCard mounted, assistant text landed) — only the reasoning emission was missing. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
89e4c4e52c |
chore(showcase/langgraph-python): bump model to gpt-5.4 across all agents
Swap every ChatOpenAI / init_chat_model call across the LGP agents from the prior mix (gpt-4o-mini, gpt-4o, gpt-4.1, gpt-5-mini) to the unified `gpt-5.4` model. Also updates the OPENAI_REASONING_MODEL env default in reasoning_agent.py and tool_rendering_reasoning_chain_agent.py so reasoning demos pick up the new model unless explicitly overridden. Verified locally with the user's gpt-5.4 OpenAI access — all 39 LGP demos render and respond correctly. The probe sweep on prod will confirm the model name resolves end-to-end. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
3d4912f323 |
fix(showcase/langgraph-python): close remaining D5 cells (framework + fixtures)
Lands the Bucket A framework fix and the fixture-correctness changes needed to flip the remaining D4/D2 cells in the langgraph-python column to D5. Full local D5 sweep is green (1 passed, 0 failed). Framework: every Python `ToolMessage` constructed via `Command(update=...)` now sets `name=` and `id=str(uuid.uuid4())`. Without these, @ag-ui/langgraph synthesises TOOL_CALL_START events with `toolCallName: null` and `parentMessageId: null`, which @ag-ui/client@0.0.53's Zod schema rejects; the rejection is silently swallowed by `withAbortErrorHandling -> EMPTY`, completing the SSE observable mid-stream so post-tool state never reaches the consumer. Fix is applied across shared_state_streaming, shared_state_read_write, gen_ui_agent, beautiful_chat, subagents (2 sites). Single-flag change in a2ui_dynamic flips the secondary `_design_a2ui_surface` LLM call to `streaming=True` so aimock's record/replay (SSE-only) sees it. Fixtures (d5-all.json): - toolCallId follow-ups for set_steps (3), display_flight, generate_a2ui (4), schedule_meeting (2), generateSandboxedUi (7), and revenue chart so multi-turn probes don't recurse into recursion-limit loops - four hand-crafted secondary `_design_a2ui_surface` fixtures so A2UI dynamic renders without a real LLM - mcp-apps fixture rewritten to emit `create_view` tool call with a minimal Excalidraw element payload; runtime middleware fetches the UI resource and the iframe mounts - AAPL and revenue `hasToolResult: true` follow-ups tightened to `toolCallId` so they don't match cross-turn after prior turns' tool results - voice fast-path content-only fixture - beautiful-chat-schedule-meeting first-turn fixture gains content so the conversation runner sees an assistant message before the picker click assertion Probes: bumped per-card waitForSelector in d5-gen-ui-headless-complete from 15s to 60s — recharts ResponsiveContainer can be slow under 4 sequential turns. Shell-dojo: hide CLI Start Command from the dojo navigation via `HIDDEN_DOJO_FEATURE_IDS`. Registry/manifests untouched so harness/parity/dashboard still see it. |