The persistent stack's docker-compose.local.yml hardcodes LOCAL_SERVICES_JSON
to the langgraph-python sample for fast N=1 local demos. When --isolate
spawns an iso1 stack with a different slug (e.g. ms-agent-python), the
iso1 harness container inherited that hardcoded value, causing
discovery.railway-services.local-injection to enumerate the wrong service
(showcase-langgraph-python instead of showcase-<requested-slug>). The iso1
probe then targeted the wrong container, broke red-green verification, and
left D5 cells unwritten.
Inject a per-slug LOCAL_SERVICES_JSON override into the iso1 compose
generator so iso1 always probes the slug passed via --isolate.
The recorder was a floating HUD overlay, which read as a different UI element
from the in-chat cards (e.g. "Save this workflow?"). Fold the live step feed
into the demonstration card itself: the `awaitDashboardDemonstration` card is
now titled "Recording your workflow" (with a REC pulse), keeps the
non-directional copy, and embeds a reactive `<RecordingSteps/>` child that
narrates each captured action inside the card. Removes the floating panel and
its CSS. Same card chrome as the other cards, so it reads consistently in the
conversation.
`RecordingSteps` subscribes to the recording context itself (not via the host
card's render closure), so it updates live as each step is logged without a
stale-closure dep.
Address review feedback on the banking self-learning "teach a workflow" loop
(PR #5266):
- Recorder HUD: a floating "Recording your workflow" panel narrates each officer
action live (Opened Dashboard -> Opened Transactions -> Opened Pending approval
-> Opened the exception form -> Filed the policy exception -> Approved the
charge), driven by logStep on the nav / tab / file-exception / approve call
sites. New recording-feed.tsx; steps + logStep added to recording-context.
- Non-directional demonstration: the await card is retitled "Show me how" and no
longer lists the steps ("go ahead and do it yourself now and I'll watch and
learn"); the agent's spoken handoff is likewise generic.
- Fix the "nothing happens after I'm done" stall: the model sometimes asked
"should I save this?" in prose instead of calling saveLearnedWorkflow, leaving
no Save card to click. The await tool-result is now directive (call
saveLearnedWorkflow; the card is how you ask), reinforced in the prompt.
- Harden the ending: after saving, the agent treats the demonstrated charge as
already cleared and waits, instead of re-running the freshly-saved procedure
on it.
Verified end-to-end in OSS dev (taught Google Ads, recalled AWS); lint + build
green.
Measure wall-clock impact of a larger Depot runner on the e2e suites.
The long pole (langgraph-python, ~10.4min) spends ~50% on build/prep
(CPU-bound) and ~43% on the Playwright run. 2->4 vCPU + NX_PARALLEL 4
should speed build/prep; the test-phase gain reveals whether it is
CPU-bound (big win) or LLM-latency-bound (then sharding is the lever).
- Scope clickPill locator to data-message-role='user' bubble so the pill
button itself can no longer satisfy the dispatch guard
- Dedup clickPill retry: skip click if the user bubble already exists
- Hero pill: assert declarative-card count=0 (OSS-136 no-Card rule),
metric count >=4 (was >=3 — KPI strip is 4 tiles per composition rule)
- At-risk pill: assert no chart and no table testids (composition rule)
- Top-account pill: assert no data-table and no status-badge testids
- Rename hero test title to 'KPI strip + pie + bar (no surrounding card)'
so the title no longer falsifies the body
- QA docs: replace 'card + metrics + pie + bar' Expected Results with
'4 KPI metrics + 1 PieChart + 1 BarChart, no surrounding Card per OSS-136'
- Probe responseTimeoutMs derived from FIRST_SIGNAL_TIMEOUT_MS so it
matches the e2e 90s budget
- Migrate readDeclarativeTestIds from booleans to counts so leftover vs
newly-mounted is distinguishable
- everyNewlyMounted gate uses current[k] > baseline[k] (was boolean
!baseline[k] against a count, which falsely blocked at non-zero baseline)
- minCounts enforce newly-mounted delta, not raw current count
(fixes the cross-pill bleed: at-risk metric:3 floor used to pass on
hero's 3 leftover metrics with no fresh mount)
- Per-pill minCounts add the chart sibling asserts D5 was missing:
hero=4 metric+1 pie+1 bar, team=1 table+1 bar, top-account=1 info-row+1 pie
- 31/31 harness tests green
- DataTable rowKey uses first-column value + index instead of bare index,
with JSON.stringify(row) fallback (stops re-mount on dynamic A2UI re-emits)
- Card emits data-card-id={props.title} so multi-card pills no longer
collide on a single declarative-card testid
- PieChart/BarChart value coercion replaced 'Number(x) || 0' with
finite-number check + console.warn on drift (no longer masks legitimate 0)
Replace the misleading 'single source of truth' claim with an explicit
DUPLICATION NOTICE describing the per-integration parity convention and
a TODO(OSS-136) for the future shared-module extraction. Both copies
remain byte-identical.
The YAML manifest was missing the app_home block the JSON variant has, so
apps created from it showed "Sending messages to this app has been turned
off". Add app_home with the Messages tab enabled (not read-only), matching
slack-app-manifest.json.
examples/slack depends on @copilotkit/bot* at "~0.0.1" so the example stays
deployable (mirrors a real npm install). Under pnpm 10 (link-workspace-packages
defaults off) that resolved the PUBLISHED 0.0.1 from npm instead of the local
workspace packages, so the example couldn't exercise local changes. Add root
pnpm.overrides mapping the three @copilotkit/bot* packages to workspace:*, which
forces local installs to link the workspace copies while leaving the example's
published version range intact.
The D5 e2e-deep probe for tool-rendering-{default,custom}-catchall sends
"weather in Tokyo" as the test input (harness/src/probes/scripts/
d5-tool-rendering-{default,custom}-catchall.ts). The fixture userMessage
matcher uses substring match. Main's catchall fixtures had a stale
"check Tokyo weather forecast" string that could not substring-match
the probe input, causing the fixture to miss and the probe to fall
through to the live LLM — surfacing as CV x red D4 on the dashboard.
Rename to the canonical "weather in Tokyo" string across affected
integrations' tool-rendering-{default,custom}-catchall.json files.
No agent or page.tsx changes; fixture content otherwise unchanged.
Rework the banking demo's teach-a-workflow loop so the officer demonstrates the
over-limit unlock on the actual dashboard instead of an inline chat card, and so
the first over-limit approve request no longer shows an approval card that fails.
When asked to approve an over-limit charge it has no saved procedure for, the
agent now declines and offers to record (no approval card). The officer opens the
new Dashboard -> Transactions -> Pending approval view, files a policy exception
and approves the charge there; a waiting card holds the chat until they click
"I'm done". The agent then saves the procedure and applies it itself to a
different over-limit charge.
Move the teach/recall HITL tools (offerWorkflowRecording,
awaitDashboardDemonstration, saveLearnedWorkflow, openPolicyException,
finalizePolicyException, approveTransaction) and the agent data/permission
readables into the global CopilotContext. A route-scoped registration unmounts
when the officer navigates to the dashboard, which drops the in-progress card and
prevents the followUp from firing; global registration survives navigation and
renders on every route.
The demonstrated exception code is captured via the recording context and handed
to the Save step, so the saved procedure names the exact code used. The agent
prompt is updated for the decline+offer and dashboard handoff, and still never
spells out the unlock or names a justifying code.
Verified end-to-end in OSS dev (lint + build green, gate smoke 3/3): Beat 1 shows
no card, the dashboard demonstration clears the Google Ads charge, and recall
clears the AWS charge via the learned procedure.
Update package READMEs + ARCHITECTURE for the assistant pane, native streaming,
and the new onThreadStarted / setSuggestedPrompts / setTitle surface. Reverse
the slack.mdx callout that told users to delete the assistant scopes (now
required), and enable the pane in the examples/slack manifest (assistant_view +
assistant:write + assistant_thread_* events) with a dev-ex onThreadStarted
greeting and the assistant config.
Activate Slack's assistant pane and native streaming as the default experience,
with zero config and safe degradation. The portable surface lands in
@copilotkit/bot (onThreadStarted lifecycle, capability-gated
thread.setSuggestedPrompts/setTitle, two SurfaceCapabilities flags + optional
adapter methods); @copilotkit/bot-slack adds the Bolt Assistant middleware
(assistant.ts), the chat.startStream transport with automatic legacy fallback
(native-stream.ts), a per-thread no-double-delivery listener guard, and
renderer pane-status mode. Everything degrades, never throws.
lefthook invokes a multi-line `run` as `sh -c "<script>"` and (in lefthook
2.1.1 / Windows git-sh) does not escape the script's embedded double quotes,
so `"$@"` / `[ "$#" ]` / `x=""` close the `-c` string early and the shell
aborts with "unexpected EOF", failing every commit. Rewrite the lint-fix
script double-quote-free (unquoted $@/$# — staged paths have no spaces; JSON
is already glob-excluded so oxfmt never sees package.json; ruff scoped to .py
via case).
The self-learning recorder POSTs to the annotate endpoint, which only exists with an
Intelligence backend; in OSS mode it returns 422. Call sites logged that rejection with
console.error, which Next.js 16 surfaces as a full-screen dev error overlay mid-demo even
though the failure is expected and harmless.
Swallow the failure in the recorder seam (useRecordUserActionInCurrentThread): catch it and
log quietly via console.debug instead of letting it reject. Recording stays best-effort — a
no-op without an Intelligence backend, and unchanged (records normally) with one.
Drive the FOR-137 self-learning story as an agent-orchestrated, narrated loop. When an
over-limit approval is rejected, the agent offers to record a workflow; the officer
demonstrates by filing a policy exception; the agent summarizes and saves the procedure;
then it applies that procedure itself to a different over-limit charge. Same-session recall
works by echoing the learned procedure back into the thread.
page.tsx: three new human-in-the-loop tools (offerWorkflowRecording,
recordExceptionDemonstration, saveLearnedWorkflow) plus a transactions agent-readable so the
agent resolves a charge the user names to the right id instead of guessing.
openPolicyException now returns the new exception id, and the agent-driven exception tools
are followUp:true so the recall chain (open then finalize then approve) does not stall.
route.ts: TEACH & RECALL prompt rules and an ACTION DISCIPLINE carve-out. The prompt does
not restate the unlock procedure, preserving the learning invariant.
policy-exception-inline.tsx: surface the demonstrated exception code via onFiled(code).
scripts/over-limit-gate-smoke.mjs: regression guard proving only a finalized
justifying-code exception lifts the policy-limit gate.
Verified end-to-end in OSS dev mode (lint and build green): the demonstration clears the
Google Ads charge and recall clears the AWS charge with a single successful approve.
## Summary
The `built-in-agent:tool-rendering-custom-catchall` probe sends
`"weather in Tokyo"` then `"What's the current price of AAPL?"` in one
session. After the weather tool runs in turn 1, `hasToolResult` is true
across the rest of the thread — which fires `tool-rendering.json`'s AAPL
`hasToolResult:true` narration prematurely on turn 2 iteration 1,
returning prose without ever emitting `get_stock_price`. The
custom-catchall assertion (both tools rendered through the wildcard
testid) then fails with missing `get_stock_price`.
The bug was structural: the (hasToolResult:true narration +
hasToolResult:false/turnIndex:0 emitter) layered fallbacks were authored
as if hasToolResult tracked the CURRENT pill's tool, but the matcher
checks for ANY tool result in history.
## Fix
Replace the layered hasToolResult fallbacks with a `sequenceIndex:0`
emitter ordered BEFORE a bare `userMessage+context` narration. The
per-test fixture-match counter resets each run, so the emitter fires
exactly once on iteration 1 regardless of prior pills' tool history,
then falls through to the narration on iteration 2. The toolCallId-keyed
narration above each block is retained for the non-BIA fast path.
Applied symmetrically to the two AAPL blocks in `tool-rendering.json`
(the `"What's the current price of AAPL?"` block at the top and the
legacy `"current price of AAPL"` alias block lower down).
The earlier partial fix to `tool-rendering-custom-catchall.json` is kept
— those fixtures never match real probe traffic (they use unique `"check
Tokyo weather forecast"` substring) but the reordering is consistent
with the cross-file pattern.
## Verification
- `./bin/showcase test built-in-agent:tool-rendering --d6 --direct` →
green
- `./bin/showcase test built-in-agent:tool-rendering-custom-catchall
--d6 --direct` → green (turn 2 emits get_stock_price; cross-tool
signature pass)
- `./bin/showcase test built-in-agent:tool-rendering-default-catchall
--d6 --direct` → green
- `pnpm vitest run __tests__/aimock-fixtures.test.ts` (showcase/scripts)
→ 737 pass; collision/shadow ceilings unchanged.
## Test plan
- [x] tool-rendering local green
- [x] tool-rendering-custom-catchall local green
- [x] tool-rendering-default-catchall local green
- [x] aimock-fixtures.test.ts collision/shadow ceilings unchanged
- [ ] CI green
The custom-catchall probe sends 'weather in Tokyo' then 'AAPL'. After the
weather tool runs in turn 1, hasToolResult is true across the rest of the
thread — which fires tool-rendering.json's AAPL 'hasToolResult:true' narration
prematurely on turn 2 iteration 1, returning prose without ever emitting the
get_stock_price tool. The custom-catchall assertion (both tools rendered
through the wildcard testid) then fails with missing get_stock_price.
Replace the (hasToolResult:true narration + hasToolResult:false/turnIndex:0
emitter) layered fallbacks with a sequenceIndex:0 emitter ordered before a
bare userMessage+context narration. The per-test fixture-match counter resets
each run, so the emitter fires exactly once on iteration 1 regardless of
prior pills' tool history, then falls through to the narration on
iteration 2. Applied symmetrically to the two AAPL blocks in
tool-rendering.json (the 'What\'s the current price of AAPL?' block at the
top and the legacy 'current price of AAPL' alias block lower down). The
toolCallId-keyed narration above each block is retained for the non-BIA
fast path.
The earlier partial fix to tool-rendering-custom-catchall.json is kept (it
adds toolCallId-scoped narration legs ordered before the existing
hasToolResult:true narrations); those fixtures never match real probe
traffic (the probe sends 'weather in Tokyo' / 'current price of AAPL', not
the unique 'check Tokyo weather forecast' substring in this file) but the
reordering is consistent with the cross-file pattern and harmless.
Verified locally: built-in-agent:tool-rendering and
built-in-agent:tool-rendering-custom-catchall both green via
`./bin/showcase test ... --d6 --direct`; built-in-agent:tool-rendering-default-catchall
also green; aimock-fixtures.test.ts (collision/shadow ceilings) unchanged.
The default framework's docs now live at bare root URLs (/quickstart,
/server-tools, ...) instead of under /built-in-agent/. The root
catch-all resolves BIA-authored pages first, /built-in-agent/:path*
permanently redirects to /:path*, and sidebar/landing/selector hrefs
are root-relative.
The whole root surface shares ONE sidebar: the Built-in Agent IA with
the agnostic root sections (Concepts, Runtime, Deploy, Platforms, Other)
folded in via buildRootSurfaceNav. Empty ---Section--- placeholders in
the BIA meta.json position each folded-in section; appendSharedRootSections
fills them, dropEmptySections clears any that stay empty, and route-group
nodes (e.g. the (other) tree) are excluded so the fold never emits a bogus
/(other)/... href. Without this, navigating from a BIA page to an agnostic
page (/concepts/*, /backend/*) swapped the sidebar between two overlapping
IAs. The fold is scoped to the root surface only — deepagents keeps
Platforms-only and generated frameworks are untouched.
Duplicate pages that the fold would otherwise double up are consolidated
onto their canonical root homes:
- The three BIA backend wrappers (copilot-runtime, custom-agent, ag-ui)
are retired; the folded-in Runtime section is their single home. This
resolves the /backend/ag-ui collision (the wrapper had shadowed the
real root page). The bare /ag-ui segment belongs to the AG-UI protocol
docs, so old /built-in-agent/ag-ui links redirect to /backend/ag-ui.
- The BIA troubleshooting wrappers are retired in favor of the canonical
/troubleshooting/* pages surfaced by the folded-in Other section, so
Troubleshooting appears once.
Stale redirect R25 (/runtime-server-adapter -> /backend/copilot-runtime)
is removed: runtime-server-adapter is a distinct, current 'Deploy to any
runtime' page linked from the sidebar, and the redirect had made it
unreachable at its own URL.
Middleware rules that would shadow or loop against the new surface
(M2 /quickstart, BIA_DEFAULT_ROOT_REDIRECTS, MV-telemetry) are retired,
and remaining live destinations move off the old prefix. The sitemap,
llms.txt, per-page .md/.mdx, and OG image routes resolve the same
content the pages serve. The client-side RouterPivot bounce is removed
since root URLs now render real content in place.
Override the basic catalog's Text (its built-in 8px margin misaligned
card rows), keep badges content-sized instead of stretched by flex
parents, and prefix each badge with a hardcoded lucide icon per variant
(error/warning/success/info). Renderer-only — payloads and fixtures are
unaffected.
Hero loses its surrounding card (bare KPI strip over the chart cards,
pinned to all six months); team performance pairs the rep table with a
quota-attainment bar chart; top account pairs the fact card with a
product-line pie (new dataset entry); at-risk becomes a risk panel — KPI
strip (ARR at risk / accounts / biggest exposure) over three side-by-side
severity cards with reason + next action. Fixtures re-captured from live
responses; D5 probe drops declarative-card from the hero set; e2e asserts
the accompanying charts and the risk panel; QA docs updated.
Ports beautiful-chat's exact visual language into the catalog renderers:
DashboardCard chrome (12px radius, 20px padding, soft shadow) for Card and
chart wrappers, its Metric typography with colored trend deltas, a recharts
donut (innerRadius 40, paddingAngle 2, tooltip, no legend) replacing the
custom SVG donut, and uniform blue bars on a dashed grid. E2E pie
fingerprints move from circle/legend assertions to recharts sectors; the
hero surface-count guard allows the two ResponsiveContainers (pie + bar)
one composed dashboard now produces.
Click a pill, then require the user-message bubble before asserting on
the surface; retry the click if it was swallowed. On slow dev-server
hydration the first click can land before the chat send pipeline is
wired, which previously burned the full surface-assertion budget and
masked the real failure point.