Commit Graph

12110 Commits

Author SHA1 Message Date
Tyler Slaton bd2db6009a Add dark theme support to default tool renderer 2026-06-15 11:09:12 -07:00
Jordan Ritter 5e641d88c9 fix(showcase/harness): override LOCAL_SERVICES_JSON in --isolate generator to target the requested slug
The persistent stack's docker-compose.local.yml hardcodes LOCAL_SERVICES_JSON
to the langgraph-python sample for fast N=1 local demos. When --isolate
spawns an iso1 stack with a different slug (e.g. ms-agent-python), the
iso1 harness container inherited that hardcoded value, causing
discovery.railway-services.local-injection to enumerate the wrong service
(showcase-langgraph-python instead of showcase-<requested-slug>). The iso1
probe then targeted the wrong container, broke red-green verification, and
left D5 cells unwritten.

Inject a per-slug LOCAL_SERVICES_JSON override into the iso1 compose
generator so iso1 always probes the slug passed via --isolate.
2026-06-15 10:38:06 -07:00
GeneralJerel 4dfa6cbe90 refactor(showcase): render the workflow recorder inside the chat card
The recorder was a floating HUD overlay, which read as a different UI element
from the in-chat cards (e.g. "Save this workflow?"). Fold the live step feed
into the demonstration card itself: the `awaitDashboardDemonstration` card is
now titled "Recording your workflow" (with a REC pulse), keeps the
non-directional copy, and embeds a reactive `<RecordingSteps/>` child that
narrates each captured action inside the card. Removes the floating panel and
its CSS. Same card chrome as the other cards, so it reads consistently in the
conversation.

`RecordingSteps` subscribes to the recording context itself (not via the host
card's render closure), so it updates live as each step is logged without a
stale-closure dep.
2026-06-15 10:29:15 -07:00
GeneralJerel 3107d3c7c5 feat(showcase): live recorder HUD + non-directional teach copy; fix save-card stall
Address review feedback on the banking self-learning "teach a workflow" loop
(PR #5266):

- Recorder HUD: a floating "Recording your workflow" panel narrates each officer
  action live (Opened Dashboard -> Opened Transactions -> Opened Pending approval
  -> Opened the exception form -> Filed the policy exception -> Approved the
  charge), driven by logStep on the nav / tab / file-exception / approve call
  sites. New recording-feed.tsx; steps + logStep added to recording-context.

- Non-directional demonstration: the await card is retitled "Show me how" and no
  longer lists the steps ("go ahead and do it yourself now and I'll watch and
  learn"); the agent's spoken handoff is likewise generic.

- Fix the "nothing happens after I'm done" stall: the model sometimes asked
  "should I save this?" in prose instead of calling saveLearnedWorkflow, leaving
  no Save card to click. The await tool-result is now directive (call
  saveLearnedWorkflow; the card is how you ask), reinforced in the prompt.

- Harden the ending: after saving, the agent treats the demonstrated charge as
  already cleared and waits, instead of re-running the freshly-saved procedure
  on it.

Verified end-to-end in OSS dev (taught Google Ads, recalled AWS); lint + build
green.
2026-06-15 10:11:03 -07:00
Benjamin Taylor 8de9ac5f5b ci(e2e-dojo): bump dojo suites to 4-vCPU runner (experiment)
Measure wall-clock impact of a larger Depot runner on the e2e suites.
The long pole (langgraph-python, ~10.4min) spends ~50% on build/prep
(CPU-bound) and ~43% on the Playwright run. 2->4 vCPU + NX_PARALLEL 4
should speed build/prep; the test-phase gain reveals whether it is
CPU-bound (big win) or LLM-latency-bound (then sharding is the lever).
2026-06-15 12:07:29 -05:00
Jordan Ritter 87b369fdb4 feat(showcase): declarative gen-UI demo as a sales-analyst dashboard (OSS-136) (#5396)
## Summary
- Reworks the `declarative-gen-ui` demo (LangGraph Python + Google ADK)
to feel like Beautiful Chat's sales dashboard, per
[OSS-136](https://linear.app/copilotkit/issue/OSS-136/declarative-gen-ui-demo-rework-to-feel-like-beautiful-chats-sales)
- Suggestion pills are natural business questions; chart-type steering
moved from user prompts into the agent system prompt + frontend context
(`sales-context.ts`, shared verbatim by both integrations)
- Every pill renders a dashboard-grade surface:
- **Show my sales dashboard** — bare KPI strip + regional revenue donut
+ 6-month revenue bars (no surrounding card)
  - **Team performance** — rep table + quota-attainment bar chart
- **Anything at risk?** — risk KPI strip over three severity cards (icon
badges, reason + next action)
  - **Top account details** — account fact card + product-line donut
- Renderers ported to beautiful-chat's visual language: card chrome,
metric typography, recharts donut/bars, shared palette, lucide severity
icons
- Test triad synced: D5 probe (per-pill newly-mounted testid
assertions), Playwright e2e (incl. click-dispatch guard), aimock
fixtures **captured from live model responses**, QA docs

## Verification
- Fixture validation: 738/738 · harness probe unit tests: 17/17
- Container e2e (fixture replay, both integrations rebuilt): LGP 6/6 ·
ADK 6/6
- Live-LLM iteration on LGP: all four pills produce the steered
composition consistently (11/11 captured runs + repeated interactive
verification)

## Test plan
- [ ] `pnpm exec nx run @copilotkit/showcase-harness:test --
d5-gen-ui-declarative`
- [ ] `pnpm --filter @copilotkit/showcase-scripts test aimock-fixtures`
- [ ] `BASE_URL=<container> npx playwright test
declarative-gen-ui.spec.ts` per integration
- [ ] Post-deploy:
https://dashboard.showcase.copilotkit.ai/#matrix:links,health row
"Declarative UI: Dynamic A2UI" still reaches D5 (CV badge) for
langgraph-python and google-adk
2026-06-15 09:42:45 -07:00
Jordan Ritter d5152eaa83 fix(showcase/e2e+qa): composition exclusions + KPI=4 contract alignment
- Scope clickPill locator to data-message-role='user' bubble so the pill
  button itself can no longer satisfy the dispatch guard
- Dedup clickPill retry: skip click if the user bubble already exists
- Hero pill: assert declarative-card count=0 (OSS-136 no-Card rule),
  metric count >=4 (was >=3 — KPI strip is 4 tiles per composition rule)
- At-risk pill: assert no chart and no table testids (composition rule)
- Top-account pill: assert no data-table and no status-badge testids
- Rename hero test title to 'KPI strip + pie + bar (no surrounding card)'
  so the title no longer falsifies the body
- QA docs: replace 'card + metrics + pie + bar' Expected Results with
  '4 KPI metrics + 1 PieChart + 1 BarChart, no surrounding Card per OSS-136'
- Probe responseTimeoutMs derived from FIRST_SIGNAL_TIMEOUT_MS so it
  matches the e2e 90s budget
2026-06-15 09:35:46 -07:00
Jordan Ritter 3ec2432964 fix(showcase/harness): newly-mounted gate + minCounts delta + per-pill chart asserts
- Migrate readDeclarativeTestIds from booleans to counts so leftover vs
  newly-mounted is distinguishable
- everyNewlyMounted gate uses current[k] > baseline[k] (was boolean
  !baseline[k] against a count, which falsely blocked at non-zero baseline)
- minCounts enforce newly-mounted delta, not raw current count
  (fixes the cross-pill bleed: at-risk metric:3 floor used to pass on
  hero's 3 leftover metrics with no fresh mount)
- Per-pill minCounts add the chart sibling asserts D5 was missing:
  hero=4 metric+1 pie+1 bar, team=1 table+1 bar, top-account=1 info-row+1 pie
- 31/31 harness tests green
2026-06-15 09:35:45 -07:00
Jordan Ritter 0f58f04e50 fix(showcase/renderers): stable row keys + per-card id + no-silent-zero charts
- DataTable rowKey uses first-column value + index instead of bare index,
  with JSON.stringify(row) fallback (stops re-mount on dynamic A2UI re-emits)
- Card emits data-card-id={props.title} so multi-card pills no longer
  collide on a single declarative-card testid
- PieChart/BarChart value coercion replaced 'Number(x) || 0' with
  finite-number check + console.warn on drift (no longer masks legitimate 0)
2026-06-15 09:35:45 -07:00
Jordan Ritter 8711326b5f fix(showcase/a2ui): tighten Zod schemas
- PrimaryButton.action: z.any() -> z.unknown() (forces caller narrowing)
- Row.justify/align + Column.align: z.string() -> z.enum() matching the
  renderer's CSS map
- DataTable rows accept numeric cells (z.union([string, number]))
- DataTable column-key refine documented in description (host
  CatalogComponentDefinition requires ZodObject, blocks .refine)
2026-06-15 09:35:45 -07:00
Jordan Ritter 0964823f3c fix(showcase/sales-context): honest duplication notice + extract TODO
Replace the misleading 'single source of truth' claim with an explicit
DUPLICATION NOTICE describing the per-integration parity convention and
a TODO(OSS-136) for the future shared-module extraction. Both copies
remain byte-identical.
2026-06-15 09:35:44 -07:00
Jordan Ritter df48df3587 fix(showcase/google-adk): align aimock fixture + suggestions comment
- Correct suggestions.ts file-path reference (was pointing at LP a2ui_dynamic.py)
- Tighten userMessage matchers to full pill prompts
- Strip unschema'd weight/variant fields from Metric/Card/Chart/Text payloads
2026-06-15 09:35:44 -07:00
Jordan Ritter e57ea6b864 fix(showcase/langgraph-python): align agent + aimock with ADK parity
- Replace fake gpt-5.4 with env-overridable real model (default gpt-4o)
- Register generate_a2ui tool matching SYSTEM_PROMPT + ADK structure
- Stub tool raises RuntimeError if middleware bypassed (fail-loud)
- Reorder LP fixture entries: inner render_a2ui before outer generate_a2ui
  to match ADK first-match-wins ordering
- Tighten userMessage matchers to full pill prompts (no substring hijack)
- Drop dead _design_a2ui_surface mirrors; strip unschema'd weight/variant fields
- Honest SYSTEM_PROMPT comment cross-referencing ADK _INSTRUCTION
2026-06-15 09:35:43 -07:00
Alem Tuzlak be26c104f8 fix(examples): enable Messages tab in slack app manifest (YAML)
The YAML manifest was missing the app_home block the JSON variant has, so
apps created from it showed "Sending messages to this app has been turned
off". Add app_home with the Messages tab enabled (not read-only), matching
slack-app-manifest.json.
2026-06-15 18:34:59 +02:00
Alem Tuzlak 9e25746557 chore(deps): force workspace linking for bot packages via root pnpm overrides
examples/slack depends on @copilotkit/bot* at "~0.0.1" so the example stays
deployable (mirrors a real npm install). Under pnpm 10 (link-workspace-packages
defaults off) that resolved the PUBLISHED 0.0.1 from npm instead of the local
workspace packages, so the example couldn't exercise local changes. Add root
pnpm.overrides mapping the three @copilotkit/bot* packages to workspace:*, which
forces local installs to link the workspace copies while leaving the example's
published version range intact.
2026-06-15 18:32:18 +02:00
Jordan Ritter 388c69e684 fix(showcase): align D6 tool-rendering catchall userMessage to D5 probe input
The D5 e2e-deep probe for tool-rendering-{default,custom}-catchall sends
"weather in Tokyo" as the test input (harness/src/probes/scripts/
d5-tool-rendering-{default,custom}-catchall.ts). The fixture userMessage
matcher uses substring match. Main's catchall fixtures had a stale
"check Tokyo weather forecast" string that could not substring-match
the probe input, causing the fixture to miss and the probe to fall
through to the live LLM — surfacing as CV x red D4 on the dashboard.

Rename to the canonical "weather in Tokyo" string across affected
integrations' tool-rendering-{default,custom}-catchall.json files.

No agent or page.tsx changes; fixture content otherwise unchanged.
2026-06-15 09:17:42 -07:00
GeneralJerel dac136d4f6 feat(showcase): demonstrate the self-learning workflow on the real dashboard
Rework the banking demo's teach-a-workflow loop so the officer demonstrates the
over-limit unlock on the actual dashboard instead of an inline chat card, and so
the first over-limit approve request no longer shows an approval card that fails.

When asked to approve an over-limit charge it has no saved procedure for, the
agent now declines and offers to record (no approval card). The officer opens the
new Dashboard -> Transactions -> Pending approval view, files a policy exception
and approves the charge there; a waiting card holds the chat until they click
"I'm done". The agent then saves the procedure and applies it itself to a
different over-limit charge.

Move the teach/recall HITL tools (offerWorkflowRecording,
awaitDashboardDemonstration, saveLearnedWorkflow, openPolicyException,
finalizePolicyException, approveTransaction) and the agent data/permission
readables into the global CopilotContext. A route-scoped registration unmounts
when the officer navigates to the dashboard, which drops the in-progress card and
prevents the followUp from firing; global registration survives navigation and
renders on every route.

The demonstrated exception code is captured via the recording context and handed
to the Save step, so the saved procedure names the exact code used. The agent
prompt is updated for the decline+offer and dashboard handoff, and still never
spells out the unlock or names a justifying code.

Verified end-to-end in OSS dev (lint + build green, gate smoke 3/3): Beat 1 shows
no card, the dashboard demonstration clears the Google Ads charge, and recall
clears the AWS charge via the learned procedure.
2026-06-15 08:43:38 -07:00
github-actions[bot] a574bbad90 style: auto-fix formatting 2026-06-15 15:13:03 +00:00
Alem Tuzlak 21e4b87a43 docs(shell-docs): add WhatsApp bot guide 2026-06-15 17:08:27 +02:00
Alem Tuzlak 210d5ded66 docs(bot-whatsapp): package README/ARCHITECTURE + example setup guide 2026-06-15 17:08:06 +02:00
Alem Tuzlak 4952181379 feat(whatsapp-example): runnable WhatsApp bot demo (MCP-wired) 2026-06-15 16:40:20 +02:00
Alem Tuzlak 865d5f812d fix(bot-whatsapp): address code review — command history, id round-trip, hmac, deps 2026-06-15 16:02:11 +02:00
Alem Tuzlak a925e33be9 docs(bot-slack): document assistant pane + native streaming; enable in slack example
Update package READMEs + ARCHITECTURE for the assistant pane, native streaming,
and the new onThreadStarted / setSuggestedPrompts / setTitle surface. Reverse
the slack.mdx callout that told users to delete the assistant scopes (now
required), and enable the pane in the examples/slack manifest (assistant_view +
assistant:write + assistant_thread_* events) with a dev-ex onThreadStarted
greeting and the assistant config.
2026-06-15 15:47:46 +02:00
Alem Tuzlak 779baef81e feat(bot-slack): agent-native Slack APIs — assistant pane + native streaming
Activate Slack's assistant pane and native streaming as the default experience,
with zero config and safe degradation. The portable surface lands in
@copilotkit/bot (onThreadStarted lifecycle, capability-gated
thread.setSuggestedPrompts/setTitle, two SurfaceCapabilities flags + optional
adapter methods); @copilotkit/bot-slack adds the Bolt Assistant middleware
(assistant.ts), the chat.startStream transport with automatic legacy fallback
(native-stream.ts), a per-thread no-double-delivery listener guard, and
renderer pane-status mode. Everything degrades, never throws.
2026-06-15 15:42:03 +02:00
Alem Tuzlak d2c6375e80 chore(hooks): make lint-fix script double-quote-free
lefthook invokes a multi-line `run` as `sh -c "<script>"` and (in lefthook
2.1.1 / Windows git-sh) does not escape the script's embedded double quotes,
so `"$@"` / `[ "$#" ]` / `x=""` close the `-c` string early and the shell
aborts with "unexpected EOF", failing every commit. Rewrite the lint-fix
script double-quote-free (unquoted $@/$# — staged paths have no spaces; JSON
is already glob-excluded so oxfmt never sees package.json; ruff scoped to .py
via case).
2026-06-15 15:37:03 +02:00
Alem Tuzlak c60674bfb5 feat(bot-whatsapp): public exports 2026-06-15 15:24:16 +02:00
GeneralJerel 654c96f7d0 fix(showcase): keep best-effort recording from raising a dev error overlay
The self-learning recorder POSTs to the annotate endpoint, which only exists with an
Intelligence backend; in OSS mode it returns 422. Call sites logged that rejection with
console.error, which Next.js 16 surfaces as a full-screen dev error overlay mid-demo even
though the failure is expected and harmless.

Swallow the failure in the recorder seam (useRecordUserActionInCurrentThread): catch it and
log quietly via console.debug instead of letting it reject. Recording stays best-effort — a
no-op without an Intelligence backend, and unchanged (records normally) with one.
2026-06-15 06:13:54 -07:00
GeneralJerel 8e7c6e2fb6 feat(showcase): narrate the self-learning teach-a-workflow loop
Drive the FOR-137 self-learning story as an agent-orchestrated, narrated loop. When an
over-limit approval is rejected, the agent offers to record a workflow; the officer
demonstrates by filing a policy exception; the agent summarizes and saves the procedure;
then it applies that procedure itself to a different over-limit charge. Same-session recall
works by echoing the learned procedure back into the thread.

page.tsx: three new human-in-the-loop tools (offerWorkflowRecording,
recordExceptionDemonstration, saveLearnedWorkflow) plus a transactions agent-readable so the
agent resolves a charge the user names to the right id instead of guessing.
openPolicyException now returns the new exception id, and the agent-driven exception tools
are followUp:true so the recall chain (open then finalize then approve) does not stall.

route.ts: TEACH & RECALL prompt rules and an ACTION DISCIPLINE carve-out. The prompt does
not restate the unlock procedure, preserving the learning invariant.

policy-exception-inline.tsx: surface the demonstrated exception code via onFiled(code).

scripts/over-limit-gate-smoke.mjs: regression guard proving only a finalized
justifying-code exception lifts the policy-limit gate.

Verified end-to-end in OSS dev mode (lint and build green): the demonstration clears the
Google Ads charge and recall clears the AWS charge with a single successful approve.
2026-06-15 06:05:27 -07:00
Alem Tuzlak c1ee1a818f feat(bot-whatsapp): WhatsAppAdapter implementing PlatformAdapter 2026-06-15 15:03:20 +02:00
Alem Tuzlak 9e168bd937 feat(bot-whatsapp): default context (formatting + no-streaming guidance) 2026-06-15 14:50:44 +02:00
Alem Tuzlak fa239e99fc feat(bot-whatsapp): webhook ingress — listener + signed HTTP server 2026-06-15 14:46:50 +02:00
Alem Tuzlak 2dc28f4901 feat(bot-whatsapp): buffered run renderer (no streaming) 2026-06-15 14:40:05 +02:00
Alem Tuzlak b61cdc0e94 feat(bot-whatsapp): pluggable HistoryStore + per-turn history replay 2026-06-15 14:35:46 +02:00
Alem Tuzlak d261e70ebe feat(bot-whatsapp): interaction decode, Cloud API client, inbound media parts 2026-06-15 14:30:54 +02:00
Alem Tuzlak 4cbc178f6d feat(bot-whatsapp): map IR to Cloud API text/button/list payloads 2026-06-15 14:12:16 +02:00
Alem Tuzlak 185814ba61 feat(bot-whatsapp): add types, render limits, and markdown->WhatsApp transform 2026-06-15 14:03:56 +02:00
Alem Tuzlak 61c13f86bb chore(bot-whatsapp): scaffold package 2026-06-15 13:51:01 +02:00
Jordan Ritter 51d2775940 fix(showcase/aimock): BIA tool-rendering cross-pill shadow on hasToolResult narration (#5446)
## Summary

The `built-in-agent:tool-rendering-custom-catchall` probe sends
`"weather in Tokyo"` then `"What's the current price of AAPL?"` in one
session. After the weather tool runs in turn 1, `hasToolResult` is true
across the rest of the thread — which fires `tool-rendering.json`'s AAPL
`hasToolResult:true` narration prematurely on turn 2 iteration 1,
returning prose without ever emitting `get_stock_price`. The
custom-catchall assertion (both tools rendered through the wildcard
testid) then fails with missing `get_stock_price`.

The bug was structural: the (hasToolResult:true narration +
hasToolResult:false/turnIndex:0 emitter) layered fallbacks were authored
as if hasToolResult tracked the CURRENT pill's tool, but the matcher
checks for ANY tool result in history.

## Fix

Replace the layered hasToolResult fallbacks with a `sequenceIndex:0`
emitter ordered BEFORE a bare `userMessage+context` narration. The
per-test fixture-match counter resets each run, so the emitter fires
exactly once on iteration 1 regardless of prior pills' tool history,
then falls through to the narration on iteration 2. The toolCallId-keyed
narration above each block is retained for the non-BIA fast path.

Applied symmetrically to the two AAPL blocks in `tool-rendering.json`
(the `"What's the current price of AAPL?"` block at the top and the
legacy `"current price of AAPL"` alias block lower down).

The earlier partial fix to `tool-rendering-custom-catchall.json` is kept
— those fixtures never match real probe traffic (they use unique `"check
Tokyo weather forecast"` substring) but the reordering is consistent
with the cross-file pattern.

## Verification

- `./bin/showcase test built-in-agent:tool-rendering --d6 --direct` →
green
- `./bin/showcase test built-in-agent:tool-rendering-custom-catchall
--d6 --direct` → green (turn 2 emits get_stock_price; cross-tool
signature pass)
- `./bin/showcase test built-in-agent:tool-rendering-default-catchall
--d6 --direct` → green
- `pnpm vitest run __tests__/aimock-fixtures.test.ts` (showcase/scripts)
→ 737 pass; collision/shadow ceilings unchanged.

## Test plan

- [x] tool-rendering local green
- [x] tool-rendering-custom-catchall local green
- [x] tool-rendering-default-catchall local green
- [x] aimock-fixtures.test.ts collision/shadow ceilings unchanged
- [ ] CI green
2026-06-15 00:25:19 -07:00
Jordan Ritter 97031e54e1 fix(showcase/aimock): break BIA tool-rendering cross-pill shadow on hasToolResult fallback
The custom-catchall probe sends 'weather in Tokyo' then 'AAPL'. After the
weather tool runs in turn 1, hasToolResult is true across the rest of the
thread — which fires tool-rendering.json's AAPL 'hasToolResult:true' narration
prematurely on turn 2 iteration 1, returning prose without ever emitting the
get_stock_price tool. The custom-catchall assertion (both tools rendered
through the wildcard testid) then fails with missing get_stock_price.

Replace the (hasToolResult:true narration + hasToolResult:false/turnIndex:0
emitter) layered fallbacks with a sequenceIndex:0 emitter ordered before a
bare userMessage+context narration. The per-test fixture-match counter resets
each run, so the emitter fires exactly once on iteration 1 regardless of
prior pills' tool history, then falls through to the narration on
iteration 2. Applied symmetrically to the two AAPL blocks in
tool-rendering.json (the 'What\'s the current price of AAPL?' block at the
top and the legacy 'current price of AAPL' alias block lower down). The
toolCallId-keyed narration above each block is retained for the non-BIA
fast path.

The earlier partial fix to tool-rendering-custom-catchall.json is kept (it
adds toolCallId-scoped narration legs ordered before the existing
hasToolResult:true narrations); those fixtures never match real probe
traffic (the probe sends 'weather in Tokyo' / 'current price of AAPL', not
the unique 'check Tokyo weather forecast' substring in this file) but the
reordering is consistent with the cross-file pattern and harmless.

Verified locally: built-in-agent:tool-rendering and
built-in-agent:tool-rendering-custom-catchall both green via
`./bin/showcase test ... --d6 --direct`; built-in-agent:tool-rendering-default-catchall
also green; aimock-fixtures.test.ts (collision/shadow ceilings) unchanged.
2026-06-15 00:18:53 -07:00
wuyangfan 650b283016 docs: fix LangGraph documentation links 2026-06-15 12:25:27 +08:00
Jordan Ritter cb9fbe4c55 feat(showcase/built-in-agent): D6 BIA component port — tool-rendering + headless-complete + LGP-canonical naming (#5427)
## Summary

BIA D6 component port (#4 from PR #5413 followup list). Companion to PR
#5407 (claude-sdk-python), PR #5413 (initial BIA D6), PR #5421 (BIA D6
small follow-ups). Brings BIA closer to LGP gold-standard parity.

### What's in

- **tool-rendering**: 5 LGP-mirrored companion components
(`weather-card`, `flight-list-card`, `stock-card`, `d20-card`,
`custom-catchall-renderer`) + extracted `tool-renderers.tsx` wiring.
**D6 GREEN.**
- **headless-complete**: 2 new components (`stock-card`, `chart-card`) +
`get_revenue_chart` server tool + LGP-aligned 4 pill suggestions +
`data-message-role` on bubble cascade + `get_weather` tool name fix in
tool-renderers. Turns 1 (weather) + 2 (stock) GREEN.
- **roll_dice → roll_d20 rename** across BIA backend + reasoning-chain
references (LGP canonical naming).
- **aimock fixtures**: BIA-namespaced `tool-rendering.json` updated +
`tool-rendering-reasoning-chain.json` realigned to the rename.
- **PARITY_NOTES**: documents the server-tool reprompt loop
architectural gap blocking turns 3+4 of headless-complete.

### Known issue (documented in PARITY_NOTES.md, out-of-PR-scope)

headless-complete turns 3 (highlight_note) and 4 (revenue_chart) RED due
to BIA's TanStack multi-turn server-tool reprompt cycle + aimock
userMessage-keyed fixtures looping until timeout. Three remediation
options identified — all require changes outside this PR's scope (BIA
agent architecture / aimock matcher precedence / fixture matcher
gating).

### Test plan

- [x] Local `--direct` D6 on tool-rendering: GREEN (2.8s)
- [x] Local `--direct` D6 on headless-complete: turns 1+2 GREEN, turns
3+4 RED (documented)
- [x] CI green on PR
2026-06-14 20:15:30 -07:00
Tyler Slaton f655013dd2 fix(showcase): harden built-in-agent root rollout 2026-06-12 17:17:18 -07:00
Tyler Slaton 88de23324a Merge remote-tracking branch 'origin/main' into codex/bia-root-review-fix 2026-06-12 16:47:25 -07:00
Tyler Slaton 1adaf1b839 fix(shell-docs): harden built-in-agent root migration 2026-06-12 16:41:13 -07:00
Tyler Slaton 6d6a65c531 docs: testing that previews are up
Signed-off-by: Tyler Slaton <tyler@copilotkit.ai>
2026-06-12 16:27:57 -07:00
Tyler Slaton 5c8c56071b feat(shell-docs): serve Built-in Agent docs at the docs root
The default framework's docs now live at bare root URLs (/quickstart,
/server-tools, ...) instead of under /built-in-agent/. The root
catch-all resolves BIA-authored pages first, /built-in-agent/:path*
permanently redirects to /:path*, and sidebar/landing/selector hrefs
are root-relative.

The whole root surface shares ONE sidebar: the Built-in Agent IA with
the agnostic root sections (Concepts, Runtime, Deploy, Platforms, Other)
folded in via buildRootSurfaceNav. Empty ---Section--- placeholders in
the BIA meta.json position each folded-in section; appendSharedRootSections
fills them, dropEmptySections clears any that stay empty, and route-group
nodes (e.g. the (other) tree) are excluded so the fold never emits a bogus
/(other)/... href. Without this, navigating from a BIA page to an agnostic
page (/concepts/*, /backend/*) swapped the sidebar between two overlapping
IAs. The fold is scoped to the root surface only — deepagents keeps
Platforms-only and generated frameworks are untouched.

Duplicate pages that the fold would otherwise double up are consolidated
onto their canonical root homes:
- The three BIA backend wrappers (copilot-runtime, custom-agent, ag-ui)
  are retired; the folded-in Runtime section is their single home. This
  resolves the /backend/ag-ui collision (the wrapper had shadowed the
  real root page). The bare /ag-ui segment belongs to the AG-UI protocol
  docs, so old /built-in-agent/ag-ui links redirect to /backend/ag-ui.
- The BIA troubleshooting wrappers are retired in favor of the canonical
  /troubleshooting/* pages surfaced by the folded-in Other section, so
  Troubleshooting appears once.

Stale redirect R25 (/runtime-server-adapter -> /backend/copilot-runtime)
is removed: runtime-server-adapter is a distinct, current 'Deploy to any
runtime' page linked from the sidebar, and the redirect had made it
unreachable at its own URL.

Middleware rules that would shadow or loop against the new surface
(M2 /quickstart, BIA_DEFAULT_ROOT_REDIRECTS, MV-telemetry) are retired,
and remaining live destinations move off the old prefix. The sitemap,
llms.txt, per-page .md/.mdx, and OG image routes resolve the same
content the pages serve. The client-side RouterPivot bounce is removed
since root URLs now render real content in place.
2026-06-12 16:08:35 -07:00
Maxim 954e3b613d feat(showcase): align card internals and add severity icons to StatusBadge
Override the basic catalog's Text (its built-in 8px margin misaligned
card rows), keep badges content-sized instead of stretched by flex
parents, and prefix each badge with a hardcoded lucide icon per variant
(error/warning/success/info). Renderer-only — payloads and fixtures are
unaffected.
2026-06-13 00:16:54 +02:00
Maxim 1e0d200f53 feat(showcase): dashboard-grade surfaces on every declarative-gen-ui pill
Hero loses its surrounding card (bare KPI strip over the chart cards,
pinned to all six months); team performance pairs the rep table with a
quota-attainment bar chart; top account pairs the fact card with a
product-line pie (new dataset entry); at-risk becomes a risk panel — KPI
strip (ARR at risk / accounts / biggest exposure) over three side-by-side
severity cards with reason + next action. Fixtures re-captured from live
responses; D5 probe drops declarative-card from the hero set; e2e asserts
the accompanying charts and the risk panel; QA docs updated.
2026-06-13 00:16:53 +02:00
Maxim 06c819d0fb feat(showcase): match declarative-gen-ui renderers to beautiful-chat's sales dashboard
Ports beautiful-chat's exact visual language into the catalog renderers:
DashboardCard chrome (12px radius, 20px padding, soft shadow) for Card and
chart wrappers, its Metric typography with colored trend deltas, a recharts
donut (innerRadius 40, paddingAngle 2, tooltip, no legend) replacing the
custom SVG donut, and uniform blue bars on a dashed grid. E2E pie
fingerprints move from circle/legend assertions to recharts sectors; the
hero surface-count guard allows the two ResponsiveContainers (pie + bar)
one composed dashboard now produces.
2026-06-13 00:16:53 +02:00
Maxim 667114cfa4 test(showcase): assert pill clicks dispatched in declarative-gen-ui e2e
Click a pill, then require the user-message bubble before asserting on
the surface; retry the click if it was swallowed. On slow dev-server
hydration the first click can land before the chat send pipeline is
wired, which previously burned the full surface-assertion budget and
masked the real failure point.
2026-06-13 00:14:49 +02:00