## Summary
- Adds Deep Agents as a top-level integration in the docs (selector,
grid, quickstart dropdown, feature matrix)
- Creates `docs/integrations/deepagents/` with landing page and
quickstart (content ported from `langgraph/deep-agents`)
- Adds `deepagents` rewrite rule in `next.config.mjs` so `/deepagents/*`
routes correctly
- Adds callout on `langgraph/deep-agents` page pointing to the new
integration for discoverability
Closes CPK-7187
## Test plan
- [ ] `/deepagents` loads the landing page
- [ ] `/deepagents/quickstart` loads the quickstart guide
- [ ] Deep Agents appears in the integration selector dropdown
- [ ] Deep Agents appears in the integration grid
- [ ] Deep Agents appears in the quickstart dropdown
- [ ] `/langgraph/deep-agents` shows the callout linking to
`/deepagents`
🤖 Generated with [Claude Code](https://claude.com/claude-code)
- thread-store-registry: makeStore now attaches a __testId via
intersection so callers can distinguish stubs at a glance during
debugging instead of relying on identity-by-allocation alone.
- thread-store-registry subscriber-isolation test: assert the
diagnostic content ("Subscriber onThreadStoreRegistered error") and
Error argument, not just that some error was logged.
- handle-threads identifyUser-throws test: assert
"Error identifying intelligence user" with an Error argument, since
the throw originates inside resolveIntelligenceUser which logs and
returns 500 before subscribeToThreads is reached.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add the missing protected connect() override returning EMPTY so the
mock matches TestAgent and ThrowingAgent. Without it, a clone() that
ever exercised connect() would fall through to AbstractAgent.connect()
and may try to open a real transport in tests.
- clone() now forwards this.agentId directly instead of coercing
undefined to "". The constructor accepts string | undefined to make
this type-safe — coercion would silently turn "no agent id" into
"empty agent id", a different state per AgentConfig.agentId?: string.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- MockSocket.disconnect() now flips connected to false to match real
Phoenix sockets, so reconnect-cycle assertions are not vacuous.
- MockChannel.off(event, ref) guards against the case where a prior
off(event) without a ref already deleted the entry, preventing a
TypeError from filter() on undefined.
- Use vi.stubGlobal("fetch", ...) + afterAll(unstubAllGlobals) so the
fetch mock no longer leaks into sibling test files in the same worker.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
## Summary
Bumps `@ag-ui/client`, `@ag-ui/core`, and `@ag-ui/encoder` from `0.0.52`
→ `0.0.53` to pick up the ESM interop fix in
[ag-ui-protocol/ag-ui#1578](https://github.com/ag-ui-protocol/ag-ui/pull/1578).
## Why
`@ag-ui/client@0.0.52`'s ESM bundle compiled `import jsonpatch from
"fast-json-patch"` down to `import * as a from "fast-json-patch"`. Under
Node's native ESM loader (and webpack/turbopack in Next.js), that
namespace import is empty for `fast-json-patch@3.x` because the module
populates its exports via `Object.assign(exports, ...)` — the CJS→ESM
named-export detector cannot see them.
Result: `a.applyPatch is not a function` thrown on every `STATE_DELTA`
and `ACTIVITY_DELTA` event. The most visible symptom is open generative
UI with LangGraph agents, where the console floods with:
```
[ui] Failed to apply activity patch for 'call_xxx-activity': a.applyPatch is not a function
```
`0.0.53` switches the source back to a default import so the emitted
bundle works under both ESM and CJS consumers.
## Test plan
- [x] `pnpm install` resolves cleanly to
`@ag-ui/{client,core,encoder}@0.0.53` in the lockfile
- [x] Pre-commit hooks pass (lint, tests, package checks)
- [ ] Verify open generative UI with a LangGraph agent no longer logs
`applyPatch is not a function`
🤖 Generated with [Claude Code](https://claude.com/claude-code)
handleClearThreads is intentionally synchronous, but neither test
asserted the return type. Add `expect(response).not.toBeInstanceOf(Promise)`
on both branches so a regression that starts awaiting I/O updates the
synchronous call sites.
The handleGetThreadEvents/handleGetThreadState 501 tests asserted the
status code but not that intelligence stayed untouched. Stub spies for
both `listThreads` and a hypothetical `getThreadEvents`/`getThreadState`,
then assert neither was called — so a regression that drops the early
return and falls through to platform calls fails this test even after
the response code changes.
The handleSubscribeToThreads 500 test created an `errorSpy` to silence
console output but never asserted on it. Add `expect(errorSpy).toHaveBeenCalled()`
so a regression that quietly drops the diagnostic log is caught.
The InMemoryAgentRunner getThreadEvents test claimed the synthetic
terminal event was present but only asserted on TEXT_MESSAGE_*. Add an
explicit assertion on `RUN_ERROR` with `code: "INCOMPLETE_STREAM"` so
finalizeRunEvents' contract is locked in — a regression that stops
appending the synthetic event would leave the inspector showing an
in-progress thread forever.
Declare `onNewMessage` on `MessagePopulatingTestAgent.runAgent`'s options
type to match the runner's call site and `TestAgent` above. Without it,
a regression that starts depending on `onNewMessage` here would compile
cleanly even though the mock would silently drop the call.
Drop the dead `MockChannel.channels` field — never read or populated.
Stop auto-firing `onOpen` from `MockSocket.connect()`. Real Phoenix sockets
fire `onOpen` once per upgrade, so tests should drive the transition
explicitly via `triggerOpen()`. The auto-fire would either double-fire
when a test also called `triggerOpen()` or hide cases where production
code forgets to await the open before joining a channel. No tests in this
file relied on the auto-fire.
Convert the archive/delete fetch assertions to filter by URL+method, the
same way the rename test was already written. Hardcoded `mock.calls[2]`
and `[3]` indices broke the moment any startup fetch was added or
reordered; the filter-based form survives that without losing
specificity.
Reset `mockUseCopilotKit` at the start of `beforeEach` before re-priming
via `setupCopilotKit()`. `mockReturnValue` is stable across calls, but a
future test using `mockReturnValueOnce` would otherwise leak un-consumed
queued returns into the next test.
R1's notify-ordering fix only made the sync subscriber path safe; async
subscribers still race against the next register(). When notifySubscribers
awaits inside Promise.all, control returns to register() which assigns the
new store to the same agentId before the async handler resumes — at which
point registry.get(agentId) returns the new store, not the unregistered one.
Carry the previous store on the onThreadStoreUnregistered payload so
subscribers tearing down state don't need to consult the registry. Same
treatment for unregister(). Doc the "do not call registry.get(agentId) in
this callback" contract on the subscriber type itself.
Also tighten getAll(): cache the snapshot, freeze it (so the Readonly<>
claim is honest at runtime), and return the same reference between
mutations. Stable identity matters for useSyncExternalStore consumers that
compare snapshots to skip re-renders. Tests updated to assert prevStore
delivery, frozen-snapshot mutation throws, and reference stability across
calls. Mock notifySubscribers now uses Promise.all to mirror production
parallel dispatch.
- Add buffer length check before timingSafeEqual to prevent RangeError
on missing/malformed X-Hub-Signature-256 headers
- Switch Dockerfile from pnpm to npm with package-lock.json (pnpm
lockfile lives at monorepo root, not in the package directory)
<!--
Thank you for sending the PR! We appreciate you spending the time to
work on these changes.
Help us understand your motivation by explaining why you decided to make
this change.
**Please PLEASE reach out to us first before starting any significant
work on new or existing features.**
By the time you've gotten here, you're looking at creating a pull
request so hopefully we're not too late.
We love community contributions! That said, we want to make sure we're
all on the same page before you start.
Investing a lot of time and effort just to find out it doesn't align
with the upstream project feels awful, and we don't want that to happen.
It also helps to make sure the work you're planning isn't already in
progress.
As described in our contributing guide, please file an issue first:
https://github.com/ag-ui-protocol/ag-ui/issues
Or, reach out to us on Discord: https://discord.com/invite/6dffbvGU3D
You can learn more about contributing to copilotkit here:
https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md
Happy contributing!
-->
## What does this PR do?
(Describe the changes introduced in this PR)
## Related PRs and Issues
- (Direct link to related PR or issue, if relevant)
## Checklist
- [ ] I have read the [Contribution
Guide](https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md)
- [ ] If the PR changes or adds functionality, I have updated the
relevant documentation
- [ ] "Allow edits by maintainers" is checked (lets us help iterate on
your PR directly — faster turnaround for everyone)
Add --ci flag to eval orchestrator that skips Docker lifecycle and
assumes services are already running. Add ci-native-eval.sh helper
that installs deps, starts next dev + agent servers natively, health-
waits, then runs showcase eval --ci. Fix on-demand E2E workflow with
langgraph-python support and agent-type detection.
New showcase_eval_check.yml creates a Check Run with "Run Showcase
Eval" action button on every PR. Modified showcase_eval.yml adds
workflow_dispatch trigger with dispatch-gate job, Check Run update
in post-result, and devops bot token for Checks API calls.
Hono web server that receives check_run.requested_action webhooks from
GitHub, authenticates as the devops bot, updates the Check Run to
in_progress, and dispatches showcase_eval.yml via workflow_dispatch.
Includes GHCR build workflow and pnpm workspace registration.
## Summary
- Add D2/API badge to per-cell status row in shell-dashboard (cells now
show `API ✓ RT ✓ CV ✓` instead of just `RT ✓ CV ✓`)
- Wire `d2` field in `CellState` to the existing `agent:<slug>` PB rows
from the agent-check probe
- Add API (Agent) dimension to the CellDrilldown popover panel
- Fix stale FP/D6 test references left over from #4513
## Test plan
- [x] `live-status.test.ts` — 3 new tests: d2
present/absent/rollup-exclusion (42 total, all pass)
- [x] `cell-drilldown.test.tsx` — updated "renders all 5 badge
dimensions" to check API (Agent) instead of FP
- [x] `overlay-selector-integration.test.tsx` — FP → API in health-only,
health+docs, testing-kind assertions
- [x] `cell-pieces.test.tsx` — CP8 tests updated from FP to API
- [x] `compute-tally-detail.test.tsx` — added `d2` to `makeCellState` to
satisfy TS
- [x] TypeScript type-check clean (no new errors)
The cell status row rendered RT (D4/e2e) and CV (D5) badges but was
missing the D2/API badge despite the legend already documenting it.
- Add `d2` field to `CellState` sourced from `agent:<slug>` rows
- Render API LiveBadge before RT in CellStatus (D2 < D4 ordering)
- Add API (Agent) to CellDrilldown DIMENSIONS array
- Fix stale FP/D6 references in tests left over from #4513
## Summary
`available: "always"` suggestion configs (static and dynamic) didn't
render on the welcome screen when the chat used `runtimeUrl` to fetch
agents instead of registering them locally with
`agents__unsafe_dev_only`.
The bug has been latent since v2 first landed (Dec 2025); a per-thread
cloning change masked it for some flows from Mar 31 → Apr 23, and the
Apr 23 revert (#3525 backout) re-exposed it.
This PR is welcome-screen only — the `!isConnecting && !isRunning` UI
gate is unchanged, so suggestions still hide during connect/replay and
during runs as before.
## What was broken
With `runtimeUrl`, the agents registry is empty during the initial
`/info` fetch. Two compounding issues meant `available: "always"`
configs never got off the ground:
1. **`SuggestionEngine.reloadSuggestions`** bailed early when the
consumer agent wasn't in the registry yet. The first reload fires from
`useConfigureSuggestions` on mount — at that moment the registry is
empty, so every config got skipped. Static pills never appeared, and
dynamic pills never even started generating.
2. **`useConfigureSuggestions`'s global-config path** (no
`consumerAgentId` or `"*"`) only iterated the current agents map. Empty
map → zero reload calls → suggestions stuck empty until something else
triggered a reload.
## Fix
Three small changes, scoped to the welcome screen path:
1. `SuggestionEngine.reloadSuggestions` no longer bails when the agent's
missing — defaults `messageCount` to 0 and processes static configs
anyway. Dynamic configs still skip until a real agent arrives.
2. `useConfigureSuggestions`'s global path also calls
`reloadSuggestions(targetAgentId)` directly (covers the empty-map case
where the agent the chat is bound to isn't yet in the registry).
3. `useConfigureSuggestions` subscribes to `onAgentsChanged` *only* for
dynamic configs, *only* when the target agent isn't yet present, and
*unsubscribes after firing once*. Dynamic pills catch up after the
runtime fetch completes, without piling up overlapping generations as
multiple hooks mount.
## What's preserved
- `hasSuggestions` keeps `!isConnecting && !isRunning` — bootstrap
replay and run-in-flight both still hide pills (no mid-replay layout
jump, no stale-context flash mid-run).
- `available: "always"` is the *eligibility window* (welcome screen vs
after first message), not a "render through transitions" override.
- Threading behavior (thread switch, explicit `threadId`, multi-chat) is
unchanged. None of the new code paths fire on thread connects.
## Test plan
- [x] `packages/core` — engine unit tests for `reloadSuggestions` when
no agent is present + when only static "always" config exists. All 59
core-suggestions tests pass.
- [x] `packages/react-core` — 4 integration tests in
`CopilotChat.suggestionsAlways.test.tsx`:
- shows on welcome screen with explicit `consumerAgentId`
- shows on welcome screen with global config
- hides during a run, reappears after (default scroll)
- hides during a run, reappears after (pin-to-send)
- [x] All 1157 react-core tests pass.
- [ ] Manual smoke in `examples/v2/react/demo` with both static and
dynamic `available: "always"` configs — welcome screen pills, hide
during run, regenerate after.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
vitest types invocationCallOrder as `number[]`, but TS narrows
indexed access to `number | undefined` under noUncheckedIndexedAccess.
Capture the entries first, assert each is defined, then compare —
keeps the ordering check intact while satisfying tsc.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Realtime metadata-deletion test now identity-checks the surviving
thread (id=t-1) so a regression that drops the wrong thread surfaces.
- Rename test finds the PATCH call by URL+method instead of indexing
fetchMock.mock.calls[2], which was brittle against any change in
startup fetch order.
- Register/unregister test uses mockReturnValue (not mockReturnValueOnce)
so the same spies are returned across all renders, and the test
explicitly sets runtimeConnectionStatus=Connected to exercise the
fully-wired flow.
- Connecting-gate test replaces the 20ms wall-clock setTimeout with
chained microtask flushes inside act(), making the "no fetch while
Connecting" assertion deterministic on slow runners.
- Socket-teardown test sources the threshold from production
(ɵMAX_SOCKET_RETRIES) and asserts both the pre-threshold (no
premature teardown) and post-threshold (teardown fires) states.
- MockSocket.connect() now fires registered onOpen handlers
synchronously, mirroring real Phoenix sockets so production code
awaiting onOpen is exercised by the same lifecycle.
- MockChannel.join() now returns a fresh MockPush per call so stale
ok/error callbacks from a prior join cannot fire against a new
join's listeners.
- getMockSockets is typed as MockSocketLike[] so socket-API typos
surface at compile time instead of only at runtime.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Tests in @copilotkit/react-core need the WebSocket retry budget so they
can validate teardown semantics without hardcoding the threshold
separately from production. Exposing it under the ɵ-prefixed internal
namespace keeps it out of the public API surface while letting the
test bench import a single source of truth.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two-part welcome-screen regression for `available: "always"` suggestion configs
when the chat connects to agents via `runtimeUrl` instead of registering them
locally:
1. SuggestionEngine.reloadSuggestions bailed early when the agent wasn't yet
in the registry. With runtimeUrl, the registry is empty during the initial
/info fetch, so the very first reload (fired by useConfigureSuggestions on
mount) skipped every config — static pills never appeared on the welcome
screen, dynamic pills never started generating. Now: don't bail, default
`messageCount` to 0, run static configs anyway. Dynamic configs still need
a real agent and skip until one arrives.
2. useConfigureSuggestions's global-config path only iterated the current
agents map, which compounded the problem above — the empty map meant zero
reloads. Now: also reload for the chat's resolved consumer agent (covers
the empty-map case), and subscribe to onAgentsChanged for dynamic configs
only, firing exactly once when the target agent first appears (so dynamic
pills catch up after the runtime fetch completes, without piling up
overlapping generations as multiple useConfigureSuggestions hooks mount).
`hasSuggestions` keeps the `!isConnecting && !isRunning` UI gate. `available:
"always"` controls eligibility windows (welcome screen vs. after first
message), not whether to render through connect/replay or through a run —
those still hide and the end-of-run reload regenerates against the new
context.
Tests:
- Engine: added unit coverage for reloadSuggestions when no agent is present.
- React: 4 integration tests covering welcome screen (specific + global
consumerAgentId) and the run lifecycle (hide during run, reappear after) in
default and pin-to-send modes.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
handle-threads.test.ts:
- handleGetThreadMessages intelligence-path test now asserts the response
body verbatim, so a regression that swaps in a stub body is caught.
- 422-no-intelligence test issues a real DELETE request for the delete
path instead of cloning a POST request.
- handleClearThreads block carries a comment explaining why the handler
is intentionally synchronous (no I/O on either branch).
- The identifyUser-throws test now silences console.error for the
duration of the assertion, matching the pattern already used by the
subscribe-throws test.
in-memory-runner.test.ts:
- ThrowingAgent test asserts RUN_ERROR is the last emitted event and
that no RUN_FINISHED is emitted, locking in terminal-event semantics.
- getThreadEvents test asserts the full TEXT_MESSAGE triple is present
in the persisted event log, and the comment now reflects the real
finalizeRunEvents behaviour (it appends a synthetic terminal event
when the agent does not emit one, so terminal events ARE persisted).
- getThreadState multi-run test gains a cross-thread isolation
assertion: a snapshot on a different thread must not bleed into the
original thread's state.
- Bumped the inter-thread sort delay from 5ms to 20ms to absorb timer
jitter on slow CI runners.
- Removed four redundant `agent.agentId = "test-agent"` reassignments
(the constructor already sets it via super({ agentId })).
- Aligned MessagePopulatingTestAgent.runAgent with TestAgent: `onEvent`
is now required so the runner contract is exercised consistently.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- register() now snapshots and clears the previous slot before queuing the
unregister notification, then assigns the new store before queuing the
register notification. The unregister microtask is queued first so
subscribers tear down stale subscriptions before receiving the
replacement.
- All notifySubscribers calls now attach .catch handlers so notification
failures surface via console.error instead of being silently swallowed
by `void`.
- getAll() returns a shallow copy so callers cannot mutate the registry's
internal state through the returned reference.
- Test mock now mirrors the real notifySubscribers signature (handler +
errorMessage) and wraps each subscriber call in try/catch, exercising
the same error-isolation behaviour as production.
- Replaced the `as unknown as ɵThreadStore` cast with a satisfies-typed
minimal stub matching the real interface.
- Added tests for ordering (unregister before second register), throwing
subscriber isolation, and getAll() snapshot isolation.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
## Summary
Critical-path content fixes from the shell-docs QA triage. Fixes
everything that breaks copy-paste or click-through on existing pages.
**Items addressed:**
- 2.2, 8.1, 13.1, 13.2 — code-block import hygiene
- 12.1, 14.1, 15.1, 15.2, 22.1 — cross-page link sweep
- 9.1, 9.2 — `gpt-5.2*` → `gpt-5.4*` sweep + extend model-name CI
validator
- 7.1 — `migrations.enabled` premium walkthrough corrections
## Visual inspection
1. Start the dev server:
```bash
nx run shell-docs:dev
```
Open http://localhost:3003.
2. **Imports — copy-paste check.** Visit each page below; copy each
`route.ts` / `app/layout.tsx` / `app/page.tsx` block into a scratch
TypeScript file and verify it has all needed imports:
- `/built-in-agent/quickstart` (BIA, default integration — should now
import `CopilotKit` from `@copilotkit/react-core/v2`)
- `/agent-spec/quickstart`
- `/microsoft-agent-framework/quickstart`
- Switch framework picker to each of: `langgraph`, `mastra`,
`pydantic-ai`, `adk`, `agno`, `aws-strands`, `llamaindex`. Re-check the
same blocks.
- `/auth`, `/agentic-protocols/mcp`, `/backend/copilot-runtime`,
`/multimodal-attachments` — verify code blocks now show import lines.
3. **Links — click-through check.** On each page below, click every
external link in the prose:
- `/agentic-protocols` (the index — 7 link rewrites; all should resolve)
- `/multimodal-attachments` (the migration Callout should be GONE)
- `/faq` ("What's available?" table — the rows should be plain bold
labels, not links)
- `/langgraph/agent-app-context` (mid-prose link should now point at
langgraph fw page, not BIA)
- `/inspector` (the agent-app-context relative link should resolve to
langgraph fw page)
- `/troubleshooting/common-issues` (the `copilot-runtime` and
`model-selection` links should now resolve)
4. **Model names — search check.** In the dev server (or via grep),
confirm `gpt-5.2` and `gpt-5.2-mini` no longer appear anywhere in
`showcase/shell-docs/src/content/`. All references should now be
`gpt-5.4` / `gpt-5.4-mini`.
5. **Premium walkthrough — read-through.** Walk `/premium/self-hosting`
end to end as if installing fresh. The walkthrough should now tell you
to set `migrations.enabled: true` BEFORE the install command. The "Job
will appear as Completed" prose should only fire if migrations were
enabled.
6. **Anti-checks (these should NOT have changed):**
- The legitimate `gpt-4o`, `gpt-4.1`, `gpt-5.4` references in JSDoc /
source code.
- Code blocks where imports were intentionally omitted because a prior
block on the same page established them.
- Reference pages under `/reference/v1/` (those are owned by a separate
PR).
## Summary
Sidebar / IA / HITL cleanup from the shell-docs QA triage. Adds fallback
callouts to HITL pages for non-native frameworks, fixes gate scoping
bugs, wires orphan pages into nav, adds missing `meta.json`.
**Items addressed:**
- 3.1 — `useInterrupt` and `headless` fallback callout for the 13
frameworks without `interrupt_pattern` (via new `absent` mode on
`<WhenFrameworkHas>`)
- 3.3 — `headless.mdx` `useHeadlessInterrupt` symbol scope fix (wrap
"Driving it from plain UI" inside the native gate)
- 3.5 — `multi-agent/meta.json` added so breadcrumbs use the explicit
"Multi-Agent" label
- 17.1 — `agent-config.mdx` shared lead-in between `<InlineDemo>` and
the gated branches
- 18.1 — `ag-ui-middleware.mdx` moved under `agentic-protocols/`,
registered in `meta.json`, links to upstream AG-UI guide, 302 from the
old `/ag-ui-middleware` path
## Visual inspection
1. Start the dev server:
```bash
nx run shell-docs:dev
```
Open http://localhost:3003.
2. **HITL fallback callout.** Switch the framework picker to each of:
`built-in-agent`, `mastra`, `crewai-crews`, `ag2`, `google-adk`,
`pydantic-ai`, `agno`, `claude-sdk-python`, `claude-sdk-typescript`,
`llamaindex`, `langroid`, `spring-ai`, `strands`. For each, navigate to:
- `/human-in-the-loop/useInterrupt`
- `/human-in-the-loop/headless`
The page should render a callout pointing readers at
`/human-in-the-loop` (`useHumanInTheLoop`), NOT collapse to an empty
middle.
3. **HITL native framework check (no regression).** Switch to
`langgraph-fastapi`, `langgraph-typescript`, `langgraph-python`,
`ms-agent-dotnet`, `ms-agent-python`. Same two pages should still render
their full native/promise-based content. Especially check
`/human-in-the-loop/headless` for proper `useHeadlessInterrupt` usage
now within its gate.
4. **Multi-agent sidebar.** With any framework selected, confirm
"Multi-Agent" appears in the sidebar with `subagents` as a sub-entry.
5. **agent-config lead-in.** Visit `/agent-config`. Verify there's a
smooth prose transition between the InlineDemo and the "How it works"
section.
6. **AG-UI middleware sidebar.** Confirm `ag-ui-middleware` appears in
the sidebar under the AG-UI / agentic-protocols section. Click through
and verify the page renders correctly. Verify the 302 from the old
`/ag-ui-middleware` URL works.
7. **Anti-checks (should NOT have changed):**
- Per-framework HITL nav for the 5 native fws unchanged in shape.
- Other sidebar sections unaffected.
- Existing `<WhenFrameworkHas>` gates on other pages still work.
## Summary
- Reorder dashboard legend so all depth/level explanations (D0-D4, L1-L4
Strip, D4 RT, D5 CV, D6 FP) are grouped together on the left
- Regression indicator, color chips, and status symbols shift right
- Moved `▼ depth regression` from DepthLegend into HealthLegend to keep
it after the FP explanation
## Test plan
- [ ] Visual check: legend items render in new order with depth
explanations grouped left
Reorder legend items so depth-related explanations (D0-D4, L1-L4 Strip,
D4 RT, D5 CV, D6 FP) are grouped together on the left, followed by the
regression indicator, color chips, and status symbols.
## Summary
- The DepthChip (colored badge) correctly shows D4/D5/D6 depth numbers
-- unchanged.
- The descriptive text row below each cell's chip was duplicating the
D4/D5/D6 labels instead of showing the human-readable RT (Round Trip) /
CV (Conversation) / FP (Feature Parity) abbreviations.
- Updates LiveBadge name props in CellStatus, DIMENSIONS labels in
CellDrilldown, and all related tests.
## Test plan
- [x] All overlay-selector-integration tests pass (15 tests)
- [x] All cell-drilldown tests pass (7 tests)
- [x] cell-pieces tests pass (pre-existing timestamp format failures
only, confirmed identical on main)
The DepthChip (colored badge) correctly shows D4/D5/D6 depth numbers.
The descriptive text row below each cell's chip was also showing D4/D5/D6
instead of the human-readable RT (Round Trip) / CV (Conversation) /
FP (Feature Parity) labels. Fix the LiveBadge name props in CellStatus,
the DIMENSIONS labels in CellDrilldown, and all related tests.
## Summary
Partially reverts #4506: badge chips now display depth-layer numbers
(D4, D5, D6) instead of abbreviations (RT, CV, FP). The descriptive
labels appear in parentheses in the cell drilldown and as legend
explanations only.
- **Chips**: D4 / D5 / D6 (the depth numbers operators are familiar
with)
- **Cell drilldown dimensions**: "D4 (Round Trip)", "D5 (Conversation)",
"D6 (Feature Parity)"
- **Legend**: Chip color indicators use D4/D5/D6; descriptive section
explains what each depth means with RT/CV/FP in prose
- **Tooltip**: "D4 per feature" in column tally
## Test plan
- [x] All 23 tests in cell-pieces.test.tsx and cell-drilldown.test.tsx
pass
- [x] All 15 overlay-selector-integration tests pass
- [x] 372/372 tests passing (same 3 pre-existing failures as before,
unrelated)
Partially reverts #4506: chips display depth-layer numbers (D4, D5, D6)
instead of abbreviations (RT, CV, FP). The descriptive labels Round Trip,
Conversation, and Feature Parity now appear in parentheses in the cell
drilldown dimensions and as legend explanations, keeping the abbreviations
in prose only.
- cell-pieces: chip names D4/D5/D6 (were RT/CV/FP)
- cell-drilldown: dimension labels "D4 (Round Trip)" etc.
- adaptive-legend: chip refs use D4/D5/D6, descriptive text uses RT/CV/FP
- feature-grid tooltip: D4 per feature
- tests updated to match
## Summary
Renames per-cell badge chip labels in the showcase dashboard for
clarity:
- **E2E** -> **RT** (Round Trip): single message, full-stack response
verification
- **D5** -> **CV** (Conversation): multi-turn scripted dialogue with
tool calls and content assertions
- **D6** -> **FP** (Feature Parity): cross-framework behavioral
consistency check
Reorganizes the adaptive legend to group chip color indicators first,
then adds descriptive labels explaining what RT, CV, and FP mean.
Updates cell-drilldown dimension labels, feature-grid tooltip, and all
affected tests.
## Test plan
- [x] All shell-dashboard vitest tests pass (372/372 passing; 3
pre-existing failures unrelated to this change)
- [x] No remaining references to old E2E/D5/D6 chip labels in source or
test files (D5/D6 in chips-explainer are depth layer IDs D0-D6,
correctly unchanged)
Rename per-cell badge chip labels for clarity:
- E2E -> RT (Round Trip): single message, full-stack response verification
- D5 -> CV (Conversation): multi-turn scripted dialogue with tool calls
- D6 -> FP (Feature Parity): cross-framework behavioral consistency check
Reorganize the adaptive legend to group chip color indicators first,
followed by descriptive labels explaining what RT, CV, and FP mean.
Update cell-drilldown dimension labels to match the new naming.
Update all affected tests to use the new label strings.