1198 Commits

Author SHA1 Message Date
Sam Julien 10b87af572 docs: organize Channels guides by provider and framework (#6193)
## Summary

- Make Slack and Microsoft Teams the production-ready Channels choices
in the top provider picker, with framework-aware routes under
`/slack/...` and `/teams/...`.
- Put ten task-oriented guides inside every provider/framework journey
and remove the standalone Channels overview from navigation.
- Restore the global Channels SDK reference at `/reference/channels`,
including 34 current core, UI, state, transcript, and direct-adapter
entries.
- Add the complete CopilotKit Intelligence setup walkthrough, product
screenshot, and a stable architecture-diagram slot that Mike's final
artwork can replace in place.
- Qualify Discord and WhatsApp correctly: their direct adapters already
ship, while managed Intelligence support is coming soon.

## Why

Developers should choose their chat provider and agent framework first,
then stay in that context while they build and operate the integration.
The previous structure mixed provider guides, an extra overview layer,
and stale provider-specific reference pages, making it hard to find the
supported path or understand which behavior was managed versus
developer-operated.

This update keeps the guide journey provider-specific while returning
API material to the normal global Reference surface. It also documents
operational boundaries that matter in production instead of adding pages
for their own sake.

## How

- Reuse provider-aware MDX across Slack, Teams, and all 19 public
agent-framework integrations; the built-in agent keeps the shorter route
without a framework segment.
- Organize the sidebar into Getting started, Build, Production, and API
reference with guides for Intelligence, tools, rich and interactive
messages, commands and reactions, files, state, persistence,
transcripts, and operations.
- Pin the verified `@copilotkit/channels@0.4.0` and
`@copilotkit/runtime@1.64.1` pair and align the copy with current SDK
source plus live Intelligence behavior.
- Document managed capabilities and provider-specific realities,
including active/standby runtimes, optional hosted endpoint defaults,
output-free turn finalization, Slack manifest scopes, Teams attachment
shapes and consent, delivery acceptance semantics, and durable state
requirements.
- Preserve useful direct-adapter discoverability for Slack, Teams,
Discord, Telegram, and WhatsApp without restoring obsolete symbol pages.
- Add one-hop redirects for retired routes and cover navigation,
framework selection, raw-doc URLs, search, reference discovery, and
sitemap output.

Validation:

- Full docs suite: 51 files / 348 tests
- Typecheck
- Lint with no errors (existing baseline warnings only)
- Production build: 222 static pages
- Live HTTP checks: Slack Intelligence, Teams + Mastra files, Slack rich
messages, and direct-adapter reference all return 200
- Independent read-only correctness passes against the current Channels
SDK, Runtime, Intelligence, and every public framework setup

Linear:
https://linear.app/copilotkit/issue/OSS-615/channels-sdk-documentation-audit
2026-07-29 08:50:10 -07:00
Alem Tuzlak 5c89395dc8 fix(showcase): unbreak gen-ui-agent and declarative-json-render on a real LLM
Two built-in-agent demos were broken against a real model while their D5/D6
rows stayed green, because aimock exercises neither failure. Both root causes
were confirmed against the live OpenAI API.

gen-ui-agent stopped mid-plan on every run. `@tanstack/ai`'s `chat()` applies
`maxIterations(5)` when no `agentLoopStrategy` is passed, and nothing errors
when the budget runs out — the run just ends. GEN_UI_AGENT_PROMPT scripts 7
`set_steps` calls (1 initial + in_progress/completed per step x 3) plus a
closing message, so the walk died two calls short with the last step pinned at
`pending` and no narration. Reproduced with the real model and the real prompt:

  default budget -> 5 calls, "completed, completed, pending", no message
  maxIterations(25) -> 7 calls, all completed, 333-char summary

Every demo factory now passes the shared DEMO_AGENT_LOOP_STRATEGY (25 —
several times the longest scripted walk, still bounded). The two non-streaming
tool-free `chat()` calls keep the default: they have no loop to exhaust.

declarative-json-render rendered nothing at all: RUN_STARTED -> RUN_FINISHED,
no events, no console error, no banner. `text.format: { type: "json_object" }`
has a server-side precondition that the word "json" appear in the request
`input`, but the adapter sends `systemPrompts` as `instructions` and only
`messages` as `input` — so with the JSON directive living solely in
SYSTEM_PROMPT the API rejected every run with

  400 Response input messages must contain the word 'json' in some form to
      use 'text.format' of type 'json_object'.   (param: input)

The directive now rides in `messages` (as `user`, since TanStackChatMessage
admits no `system` role and the runtime hoists system messages into
systemPrompts — the half that is not input). json_object enforcement is kept;
verified live that the run then streams a complete, JSON.parse-able spec.

That 400 was invisible because the hand-rolled converters forward a whitelist
of chunk types and dropped RUN_ERROR — every one except reasoning-factory. The
`type: "tanstack"` factories were fine (the runtime's converter rethrows), so
the four `type: "custom"` converters now call the shared throwOnRunError.

Also: the gen-ui-agent progress card announced "All N steps complete" whenever
the RUN ended, ignoring the step data, so a truncated run read as a UI glitch
instead of an agent that stopped early. The wording is now derived from the
steps (`describeProgress`, extracted pure so it is testable without a DOM) and
a stalled run says so.

Both gotchas recorded in showcase/GOTCHAS.md.
2026-07-29 15:30:37 +02:00
Lukas Moschitz 3f0a7bda05 fix(showcase/lgt): copy manifest.yaml into runtime image
The x-pathname middleware added in this branch activates layout.tsx's
request-time readFileSync(cwd/manifest.yaml) in generateMetadata. Without
the manifest in the runner stage, every /demos page 500s with ENOENT once
middleware sets the header. Mirrors langgraph-python's Dockerfile and
satisfies scripts/__tests__/runtime-manifest-copy.test.ts.
2026-07-29 14:46:50 +02:00
Lukas Moschitz f1513e3554 docs(showcase/lgt): fix reasoning-agent fallback-model comment (gpt-5-mini) 2026-07-29 14:34:19 +02:00
Lukas Moschitz 3100b3ea80 docs(showcase/lgt): fix recovery-agent ALS-store comment (wrapModelCall, not beforeAgent) 2026-07-29 14:34:19 +02:00
Lukas Moschitz aa685b3957 fix(showcase/lgt): stop mcp-apps route leaking error message/stack to client
The mcp-apps POST catch block returned the raw error message and stack
to the client and logged nothing server-side, leaking internal paths and
library versions to anyone who could hit /api/copilotkit-mcp-apps. Mirror
the sibling copilotkit/route.ts hardening: log full details server-side
under a randomUUID correlation id and return only the id plus a generic
message with status 500.
2026-07-29 14:34:19 +02:00
Lukas Moschitz b1838b9de6 fix(showcase/lgt): add x-pathname middleware for per-demo titles (LGP parity) 2026-07-29 14:34:19 +02:00
github-actions[bot] 12ff61f74b style: auto-fix formatting 2026-07-29 11:40:41 +00:00
Lukas Moschitz 8f6123a5ec Merge remote-tracking branch 'origin/main' into lukas/oss-583-align-langgraph-typescript-showcase-demos-and-code
# Conflicts:
#	showcase/GOTCHAS.md
2026-07-29 13:38:42 +02:00
Lukas Moschitz 72ad5396c7 fix(showcase): langgraph-typescript multimodal handles PDF attachments (no 400)
The @ag-ui/langgraph converter collapses EVERY attachment (image and document)
into a LangChain image_url data-URL, so a PDF reached rewritePart as an
image_url part and hit neither the image nor document branch -> passed through
unchanged -> sent to gpt-4o as an image -> OpenAI 400 'Invalid MIME type. Only
image types are supported.' (The existing document-branch pdf-parse code was
dead — the part never arrives as 'document'.)

Add an image_url branch to rewritePart that routes on the data-URL MIME (mirrors
langgraph-python's multimodal_agent): image/* passes through unchanged; any
non-image (application/pdf) is flattened to text via the existing pdf-parse path
instead of being forwarded as an image. Source-only, no dep/model change.

multimodal D6 still green (PNG path unchanged); PDF upload verified live (text
extracted, no 400).
2026-07-29 13:30:15 +02:00
Alem Tuzlak f57151990f fix(showcase/ms-agent-dotnet): wire system prompts to Instructions + demo repairs
Root cause across nearly every ChatClientAgent: system prompts were passed as
`description:` (agent metadata) instead of `instructions:` (the actual system
message). ChatClientAgent(instructions, name, description, …) therefore ran
with null instructions, so models ignored tool guidance and BYOC JSON demos
emitted prose.

Fixes reported staging failures:
- shared-state-read-write pills: instructions now reach the model + stronger set_notes guidance
- declarative-hashbrown / declarative-json-render: instructions + ChatResponseFormat.Json (LGP parity)
- declarative-gen-ui: catalog-specific design prompt (no DashboardCard), stronger outer agent
- shared-state-streaming: feature was demo-only in the manifest → shell "Backend fixture unavailable"; added to features

Verified: unit suite 79/79 in Docker SDK 9.
2026-07-29 13:00:08 +02:00
Alem Tuzlak f58d24109a fix(showcase): repair ms-agent .NET real-LLM demo defects
Staging click-through on ms-agent-dotnet / ms-agent-harness-dotnet hit
several GOTCHAS #8 defects: aimock D6 was green while live LLMs failed.

A2UI (beautiful-chat sales dashboard, declarative-gen-ui pills):
- Force the page-registered catalogId (models invent "sales_dashboard").
- Sanitize/normalize flat components; salvage type-as-key nests; drop
  entries missing id/component (SummaryCard without id, charts without type).
- Strengthen secondary design prompts with the flat catalog contract.

Shared state + subagents side panels:
- Tool invocation drops AsyncLocal set by SetActiveThread, so writes landed
  in the global slot while snapshots keyed by AgentSession/AgentThread.
  Mirror writes and fall back on read (same pattern as D5ParityAgents).
- Wire the dead TryBuildDeterministicReply path for the "Remember something"
  pill so notes update without relying on the model calling set_notes.

Open generative UI advanced + beautiful-chat calculator:
- Prompt for clickable keypad (not form/submit) and notifyHost ping wiring.

Verified: ms-agent-dotnet unit suite 79/79 green in Docker SDK 9; harness
agent builds clean and cvdiag tests 5/5.
2026-07-29 11:42:32 +02:00
Tyler Slaton 968f7bf497 Merge branch 'main' into agent/oss-615-channel-docs 2026-07-28 17:17:59 -07:00
copilotkit-qa-bot 3efed933e6 docs(voice): address FAC-61 review feedback 2026-07-28 16:10:07 -07:00
copilotkit-qa-bot 0db84764dd docs(voice): clarify Google ADK voice route setup 2026-07-28 15:29:51 -07:00
Tyler Slaton 8a2d385a94 chore(docs): remove merge-only lint drift 2026-07-28 16:02:26 -04:00
Tyler Slaton af13ab447e Merge origin/main into agent/oss-615-channel-docs 2026-07-28 15:31:13 -04:00
Lukas Moschitz c70eff232c fix(showcase): fix langgraph-typescript a2ui-recovery (dom-missing / flaky heal)
Two langgraph-typescript-scope causes:
1) Fixture prompt drift: the a2ui-recovery.json userMessage keys + suggestions.ts
   pills still carried langgraph-python's copied prompt while the shared probe
   sends a langgraph-typescript-UNIQUE prompt -> aimock 404 -> heal surface never
   mounts. Retarget the 4 keys + 2 pills to the probe's unique prompt (kept
   distinct: a2ui-recovery fixtures have no x-aimock-context, so per-slug-unique
   prompts are load-bearing to avoid cross-framework collisions).
2) Flaky heal (green-then-red): TS @ag-ui/langgraph getA2UITools invokes its inner
   render_a2ui sub-agent via a config-less model.stream(), so config-based header
   forwarding never reaches it; the inner aimock call carries no x-test-id, its
   sequenceIndex falls into the never-reset DEFAULT_TEST_ID bucket, and the
   seq0->seq1 heal staging only works on the first run. Add wrapModelCall/
   wrapToolCall middleware + AsyncLocalStorage + a custom OpenAI fetch that
   forwards inbound x-* headers onto every outbound call (outer emit AND inner
   render) — mirroring the mechanism the green TS sibling mastra uses.

Not the shared Python recovery-loop defect (mastra, also TS getA2UITools, is
green — the TS path is fixable). D6 a2ui-recovery now green + STABLE (6 real ~7.7s
runs); aimock fixtures test 844 passed. Divergence documented in PARITY_NOTES.md.
2026-07-28 18:29:43 +02:00
Alem Tuzlak 9a501d5e0d fix(showcase/built-in-agent): repair six real-LLM defects D6 was masking (#6200)
Six defects reported on staging. All six reproduce; five share one
cause.

## The shared cause

`built-in-agent/src/app/api/copilotkit/route.ts` maps **~20 demos to the
same prompt-less `createBuiltInAgent()`**. The reference wires each to
its own graph **and its own system prompt** (28 graphs in
`langgraph-python/langgraph.json`).

aimock replays a scripted tool-call sequence keyed on `userMessage` +
`context`, so D6 is green whether or not the model could have reasoned
its way there. This is `showcase/GOTCHAS.md` #8 verbatim — including its
stated "structural tell": *grep the integration's `route.ts` for demos
left in the generic fallthrough that LGP wires to a dedicated graph.*

## Per-defect

| Demo | Symptom | Root cause |
|---|---|---|
| `gen-ui-tool-based` | charts plot zeros | No prompt. The assistant
says so itself: *"I used placeholder values since no sales figures were
provided."* Nondeterministic — one run gave a real 0..220 axis, the next
`0 1 2 3 4` with no bars, which is how it passed review. |
| `gen-ui-agent` | wall of text, steps frozen at "step 1 of 4" |
`set_steps` + its `/steps` STATE_DELTA already worked; nothing told the
model to *walk* pending→in_progress→completed. |
| `subagents` | left panel always empty | `grep -rn delegations src/ \|
grep -v demos/subagents` → **no output**. Frontend reads
`agent.state.delegations`; the converter handles `/steps` and `/notes`
and has no `delegations` case. Chat works because the tools do run. |
| `declarative-json-render` | raw JSON for pills 2 & 3 | Intercepted the
SSE: **the wire is already one closing brace short** — nothing is
dropped in transport, and the client parser correctly requires a
balanced object. Pill 1 only works because the model copies the prompt's
worked example verbatim (`\$1.24M`, `+18% vs Q2`). Model-level JSON
enforcement had been removed with the note *"the system prompt already
enforces JSON-only output"* — it does not. |
| `a2ui-recovery` | 5 identical cards | Reuses declarative-gen-ui's
prompt with no "call once" constraint, while the pills literally ask it
to *"self-correct a malformed first attempt"*. The retry loop is
**inside** the tool and returns on first valid pass, so every supervisor
retry is pure duplication. |
| `declarative-gen-ui` | D4 though it works live | **The inverse case.**
Badges: `UI ✓ BE ✓ 1P ✗ D6 —` (*"gated — blocked by a lower rung"*). Not
a live bug — a fixture bug. See below. |

## The D4 cap (`declarative-gen-ui`, `a2ui-recovery`)

Both demos' secondary design-call fixtures gated on
`match.responseFormat: \"json_object\"`. That matcher **can never match
this backend**:

- The agent talks to the OpenAI *Responses* API, where JSON mode is
`text.format`.
- aimock has **no `text.format` handling at all** —
`responsesToCompletionRequest` forwards only a top-level
`response_format` (absent in 1.19.1, forwarded in 1.37.4), a key the
Responses API doesn't accept and this client doesn't send.
- So `effective.response_format?.type` is always `undefined` and
`router.ts` skips the fixture → the design call never matched → surface
never painted → D5 red, D6 never ran.

Re-keyed onto `match.toolName`, which aimock *does* normalize out of a
Responses request (`responsesToolsToCompletionsTools`): the outer call
declares `generate_a2ui`, the in-tool design call declares nothing.
Tool-less secondary fixtures moved last, because several pills' brief is
a substring of the pill text (`\"Build my Q2 revenue summary …\"`
contains `\"Q2 revenue summary\"`) and would otherwise win the outer
request.

## Verification

**Reproduced live** (staging, real LLM): the zero-axis chart plus its
self-incriminating message, the unbalanced JSON on the wire, and the
reference rendering the *same pill* correctly for contrast.

**Two mutation-verified test suites** — both fail with the fix reverted:

- `aimock-a2ui-routing.test.ts` drives aimock's **real `matchFixture`**.
On the old fixtures 14/16 cases fail with `no fixture matched the
secondary design call for brief \"…\"` — the D5 red, reproduced as a
unit test.
- `tanstack-factory.test.ts` covers the `/delegations` and `/steps`
deltas (3 delegation tests fail with the branch disabled; the 2 controls
still pass).

**No regressions:** `tsc --noEmit` is 61 errors before *and* after (all
pre-existing — `gpt-5.4` missing from the adapter's model union, zod
v3/v4 skew); oxlint clean on changed files; the repo-wide fixture
validator passes (858 assertions).

`text.format` was verified by reading the request mapping in the
**pinned** `@tanstack/ai-openai@0.15.6` → `@tanstack/openai-base@0.9.2`:
`modelOptions` is spread straight into `responses.create()`,
`validateTextProviderOptions` only inspects
`metadata`/`conversation`/`previous_response_id`, and the adapter sets
`text.format` itself only when an `outputSchema` is passed (none here,
so nothing is clobbered).

## What is NOT verified

**The four prompt/`text.format` changes have not had a real-LLM
click-through** — that needs an OpenAI key this environment doesn't
have, and D6/aimock cannot verify them by construction (that's the whole
point of gotcha #8). Reasoning is documented inline at each site. This
area has burned the repo before: a previous `response_format` attempt
made the call return an empty string "verified against real OpenAI",
which is why `text.format` — the Responses API's own param — is used
instead. **Worth a live click-through on the staging deploy before this
is considered closed.**

## Deliberate non-goals

- **`a2ui-recovery` only demonstrates recovery under aimock.** The
heal/exhaust branches need a designer LLM that emits invalid surfaces on
demand; a real one succeeds on attempt 1. The single-call constraint
fixes the 5-card bug and makes it honest; making recovery visible live
needs deliberate fault injection — a product decision, recorded in
`PARITY_NOTES.md`.
- **The reference pins `openai:gpt-4o-mini`** (`gen_ui_agent.py:92`).
Flagged, not touched.

## Docs

`GOTCHAS.md` gains built-in-agent as a second worked instance of #8
(including that the masking runs *both* ways, and that `1P ✗ D6 —` means
gated, not failing), plus the `text.format` and dead-`responseFormat`
traps. `PARITY_NOTES.md` records that the named-agent registry is
**not** prompt-neutral, and that any new `state.<slot>` needs a
converter branch.
2026-07-28 17:32:20 +02:00
Lukas Moschitz 557f6f7318 fix(showcase): fix langgraph-typescript tool-rendering (get_weather card missing)
tool-rendering red (expected weather-card, 0 elements): the 'weather in Tokyo'
fixture returns assistant content + a get_weather tool call together; langchain-js
streamed-chunk reassembly left the tool call only under
additional_kwargs.tool_calls (surfacing intermediate chunks as invalid_tool_calls)
and produced a bare 'generic' message, so top-level tool_calls stayed empty.
shouldContinue then routed to __end__, the tool never ran, no TOOL_CALL_* events,
no weather card. LGP's Python create_agent parses the identical fixture correctly.

Add normalizeAssistantMessage() in tool-rendering.ts (called in chatNode): when
the response isn't a well-formed AIMessage with populated top-level tool_calls,
reconstruct a clean AIMessage, promoting additional_kwargs.tool_calls into parsed
tool_calls (no-op passthrough when already correct). Restores TS<->Python parity.

D6 tool-rendering now green (two real ~5s runs); fixture + frontend untouched
(byte-identical to LGP). Only the backend agent changed.
2026-07-28 16:50:35 +02:00
Alem Tuzlak 1d5cdb35cf fix(showcase/built-in-agent): repair six real-LLM defects D6 was masking
Every affected cell was D6-green (or D4-gated) while the deployed demo was
visibly broken. The common cause is GOTCHAS #8: aimock replays a scripted
tool-call sequence keyed on userMessage + context, so the fixture answers a
question the model was never asked.

Prompt gaps — ~20 demos shared one prompt-less `createBuiltInAgent()` where the
reference wires each to its own graph AND its own system prompt. Adds
`createBuiltInAgent({ systemPrompt })` + `demo-prompts.ts`, ported from the
reference graphs:

- gen-ui-tool-based plotted zeros; the assistant said "I used placeholder values
  since no sales figures were provided". Nondeterministic — some runs invented
  real values, which is how it passed review.
- gen-ui-agent published its plan once then narrated, freezing the progress card
  on step 1 of 4.
- a2ui-recovery painted five identical cards: nothing constrained the supervisor
  to one `generate_a2ui` call, and the pills literally ask it to "self-correct".
  The retry loop lives inside the tool, so supervisor retries are duplication.

subagents' delegation panel was permanently empty — the frontend reads
`agent.state.delegations` and no code emitted that slot. The converter now emits
a `/delegations` delta per sub-agent result (whole-array `add`, since initial
state is `{}` and strict fast-json-patch rejects unresolvable paths while
@ag-ui/client swallows the throw), and ports the reference's
`_MAX_CRITIQUE_ITERATIONS = 1` cap.

declarative-json-render dumped raw JSON for any prompt the model couldn't crib
from the worked example. Captured the SSE: the wire is already one closing brace
short, so nothing is dropped in transport — model-level enforcement had been
removed on the grounds that "the system prompt already enforces JSON-only
output". It does not. Restores it via the Responses API's `text.format`, the
param the removed `response_format` maps to, verified against the pinned
@tanstack/ai-openai@0.15.6 -> @tanstack/openai-base@0.9.2 request mapping.

declarative-gen-ui and a2ui-recovery were the inverse: capped at D4 (UI ✓ BE ✓
1P ✗ D6 gated) while working live. Their secondary design-call fixtures gated on
`match.responseFormat`, which aimock can never satisfy here — it has no
`text.format` handling and only forwards a top-level `response_format` the
Responses API doesn't accept. Re-keyed onto `match.toolName` (the outer call
declares `generate_a2ui`, the in-tool design call declares nothing) with the
tool-less fixtures moved last, since several pills' brief is a substring of the
pill text.

Tests (both mutation-verified — they fail with the fix reverted):
- aimock-a2ui-routing.test.ts drives aimock's real `matchFixture`; 14/16 cases
  fail on the old fixtures with "no fixture matched the secondary design call".
- tanstack-factory.test.ts covers the `/delegations` and `/steps` deltas.

Typecheck unchanged at 61 pre-existing errors; oxlint clean on changed files.

NOT verified live: the four prompt/`text.format` changes need a real-LLM
click-through, which needs a key this environment doesn't have. Reasoning and
the exact request mapping are documented inline.
2026-07-28 16:49:43 +02:00
Lukas Moschitz 629a70d918 fix(showcase): fix langgraph-typescript reasoning-display (text-unstable)
reasoning-display red with text-unstable: @langchain/openai@1.4.4's streaming
Responses converter pushes the reasoning-summary delta and the answer
output_text delta to the same content-block index (both 0), so the streaming
reducer collapses them into one reasoning block that swallows the answer.
@ag-ui/langgraph then routes the whole turn to REASONING_MESSAGE_* events and no
assistant TEXT_MESSAGE renders — the probe never settles.

Set disableStreaming:true in reasoning-agent.ts to force the non-streaming
Responses converter (processes final output items, yields separate reasoning +
text blocks), so reasoning and answer both render, matching langgraph-python's
content. Document the one behavioral divergence in PARITY_NOTES.md: the summary
no longer token-streams (arrives at once) until the upstream langchain-openai
index-collision bug is fixed.

D6 reasoning-display now green (real runs, reasoning-custom + reasoning-default);
fixture untouched (LGP-identical shape). Only the backend agent changed.
2026-07-28 16:30:49 +02:00
Alem Tuzlak 734294155a fix(showcase/ms-agent-dotnet): copy manifest.yaml into the runtime image
PR #6130 added `src/middleware.ts`, which sets the `x-pathname` header that
`src/app/demos/layout.tsx`'s `generateMetadata()` reads. That header is what
makes the layout actually call `loadDemoIndex()` — a request-time
`readFileSync(process.cwd()/manifest.yaml)`. The Dockerfile's runner stage
never copied `manifest.yaml`, so every `/demos/*` route now 500s with "An
error occurred in the Server Components render" (ENOENT), while the
statically-prerendered home page keeps returning 200.

On the dashboard that reads as "service is up, every cell pinned at D3": D4
fails on `page.type` waiting for `textarea`, D5 times out, D6 fails on
`waitForSelector('[role="textbox"]')` — the chat input never mounts because
the page is an error boundary.

langgraph-python hit this exact bug and fixed it with the same one-line COPY;
ms-agent-dotnet is the only integration shipping `middleware.ts` without it.
Adds a ratchet test over that invariant (mutation-verified: fails with the
COPY removed).
2026-07-28 15:10:44 +02:00
Lukas Moschitz f6c768344c fix(showcase): register gen_ui_tool_based in langgraph-typescript server.mjs
The langgraph-typescript agent runs a custom server.mjs whose hardcoded
graphSpec — NOT langgraph.json — is the runtime graph registry. gen_ui_tool_based
was missing from it, so the langgraph server returned 404 for that graph and the
Tool-Based Generative UI cell fell through. Add gen_ui_tool_based and drop the
three redundant graph ids (reasoning-default-render, tool-rendering-{default,
custom}-catchall) that route.ts now consolidates onto shared graphs, keeping
server.mjs in sync with langgraph.json.
2026-07-28 14:38:13 +02:00
github-actions[bot] cc06f4430d style: auto-fix formatting 2026-07-28 12:19:24 +00:00
Lukas Moschitz 267f98ddda fix(showcase): align langgraph-typescript to langgraph-python north star
Bring the langgraph-typescript Showcase integration to parity with the
reference langgraph-python: byte-identical demo frontends, the same demo set
and manifest overview, framework-native TypeScript backend wiring, and
canonical mirrored aimock D6 fixtures. Scope is limited to
integrations/langgraph-typescript and its aimock/d6 fixtures.

Frontend: realign ~25 drifted demo files to the north star; add missing
src/components/ui primitives, READMEs, declarative-gen-ui/sales-context, and
the threadid-frontend-tool-roundtrip page; remove LGT-only drift; replace the
hand-rolled landing page with the manifest-driven one. Only the sanctioned
"LangGraph (TypeScript)" identity strings differ.

Manifest: rebuild to match langgraph-python's feature order, demo order,
names, tags, descriptions, and routes; drop two bogus NSF entries that were
never real features (reasoning-default-render, agentic-chat-reasoning); restore
shared-state-streaming and tool-rendering-reasoning-chain as supported cells.
not_supported_features now matches the north star (gen-ui-interrupt /
interrupt-headless — a shared upstream react-core resume-path limitation).

Backend: add the dedicated gen-ui-tool-based data-viz graph; consolidate the
tool-rendering variants and reasoning cells onto their shared graphs (matching
langgraph-python) and drop redundant duplicate graph ids; register
threadid-frontend-tool-roundtrip; move the MCP Apps runtime to the
[[...slug]] catch-all route.

Fixtures: re-mirror all 43 aimock/d6 fixtures from the canonical langgraph-python
set with match.context re-keyed, replacing previously drifted per-integration
fixtures.
2026-07-28 14:17:11 +02:00
Alem Tuzlak 809bb72e72 feat(showcase): bring ms-agent-dotnet to D6 (frontend parity + shared-state-streaming, a2ui-recovery, threadid) (#6130)
## What

Brings the **ms-agent-dotnet** (Microsoft Agent Framework .NET) showcase
integration from D5 to **D6**, using **langgraph-python** as the
north-star reference.

### 1. Frontend parity with langgraph-python
Restores near-identical frontends where ms-agent-dotnet had drifted,
while **preserving the load-bearing .NET adaptations** (per the showcase
iron rules — differences belong in fixtures/minimal backend, not the
shared frontend):
- Root shell: `globals.css` (Tailwind `@theme` block + brand green),
manifest-driven index `page.tsx`, `layout.tsx`, new `middleware.ts`
(`x-pathname`), `tsconfig` include.
- `declarative-gen-ui` subtree restored (fixes divergent pill testids
the shared probe asserts).
- Doc-snippet `@region` markers, import-style normalization, `subagents`
revert, stale-file cleanup, `auth` inspector flag.
- **Kept** (load-bearing, not reverted): `parse-json-result` 3-layer
unwrap, multimodal legacy-shim, tool-based `hitl` (MAF has no
`interrupt()`), `agent-config` `properties=`.

### 2. shared-state-streaming → per-token (removed from
`not_supported_features`)
`write_document`'s `document` arg now streams into `state.document`
per-token via a `createSharedStateStreamingAgent` route shim (mirrors
the proven `createGenUiAgent` bridge, with a partial-JSON string
decoder), since the .NET AG-UI host has no `predict_state_config`.

### 3. a2ui-recovery cell (new)
First MS-Agent-Framework implementation of the A2UI
validate→retry→`a2ui_recovery_exhausted` recovery loop. Because the MAF
AG-UI adapter can't emit the custom `ACTIVITY_SNAPSHOT{status:"failed"}`
the exhausted card needs, it's a **raw-SSE `MapPost` endpoint**
(`RecoveryAgent.cs`) — the same adapter-bypass pattern already shipped
for `/multimodal`. Adds the demo frontend, API route, deterministic
aimock fixture (heal seq0-invalid→seq1-valid; exhaust always-invalid),
and a unique per-slug `PROMPTS` entry in the shared probe.

### 4. threadid-frontend-tool-roundtrip demo (parity)
Added for demo-set parity (reuses the `frontend_tools` passthrough; not
a D6-scored feature, mirroring the reference).

`gen-ui-interrupt` / `interrupt-headless` remain honestly quarantined
(upstream `@copilotkit/react-core` `useInterrupt` resume-path bug — not
a backend gap).

## Verification
- Code was authored in parallel worktree-isolated slots, each
cross-verified against the reference + the shared probe contracts; the
a2ui-recovery fixture was cross-checked against
`RecoveryAgent.ValidateComponents`.
- Local D6 harness: the image builds and the stack + probes run, but
**full local green was blocked by Windows-only harness friction**
(`core.symlinks=false` breaks `stage_shared`'s `[ -L ]` materialization;
`--direct` doesn't context-scope the `x-aimock-context` header so
context-keyed a2ui fixtures miss). These are environmental, not code
issues. **Relying on CI's Linux harness (real symlinks + fleet worker)
for authoritative D6.**

## Follow-up (not in this PR)
- `stage_shared()` should also materialize Windows symlink-as-file
entries (detect a regular file whose content is a relative path), so
forced local rebuilds work on `core.symlinks=false` checkouts.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-28 13:09:06 +02:00
lukasmoschitz eca1d741c8 fix(showcase): align langgraph-fastapi to north-star langgraph-python (OSS-582) (#6184)
## Summary

Aligns the **langgraph-fastapi** showcase integration to the north-star
**langgraph-python** (OSS-582 — "Align LangGraph (FastAPI) Showcase
demos and code"). Moves the integration onto the product-centric demo
set: backend agent graphs, runtime wiring, frontend chrome, the
demo-browser overview (`manifest.yaml`), and aimock fixtures.

Linear: OSS-582.

## What changed

**Backend agents (`src/agents/src/`)**
- Ported the missing dedicated graphs: `agentic_chat`, `gen_ui_agent`,
`gen_ui_tool_based`, `shared_state_streaming`.
- v2 `create_agent` ports + config alignment for `agent_config_agent`,
`headless_complete`, `reasoning_agent`, `tool_rendering_agent`,
`tool_rendering_reasoning_chain_agent`, `a2ui_dynamic`.

**Runtime wiring**
- `route.ts`: wired the canonical demos to their dedicated graphs and
removed them from the generic `sample_agent` fallthrough loop (this is
what made several demos behave correctly against a real LLM — see
below); added `recursion_limit`.
- `langgraph.json`: registered the 3 new graphs.
- Renamed the dedicated API routes
`copilotkit-byoc-{hashbrown,json-render}` →
`copilotkit-declarative-{hashbrown,json-render}`.

**Overview / content**
- `manifest.yaml`: demos + features aligned to LGP (names, descriptions,
tags, order, canonical ids; deprecated aliases migrated; phantom
`hitl-in-chat-booking` removed; `shared-state-streaming` un-quarantined
now that it works). Integration identity (name/slug/logo) preserved. The
demo-browser overview is now card-for-card identical to
langgraph-python.
- Restored two missing `@region` markers (factory-automation snippet
extraction).

**Frontend chrome**
- `globals.css`, `layout.tsx`, `middleware.ts` (was missing),
`declarative-gen-ui/*` + `sales-context.ts`, `beautiful-chat`;
`Dockerfile` now copies `manifest.yaml` (fixed an RSC crash on every
`/demos/*`).

**aimock fixtures (`aimock/d4|d6/langgraph-fastapi/`)**
- D6 fixture fixes for the previously-red cells (stale `turnIndex`
gates, cross-file substring shadowing, missing `chunkSize`, prompt
narrowing).

**Showcase tooling / docs**
- Fixed harness services racing on a shared image tag
(`docker-compose.local.yml`).
- `GOTCHAS.md` #8: documents that aimock D6 can be green while a demo is
broken against a real LLM (fixtures replay scripted tool calls
regardless of which graph ran), plus how to catch it.
- `PARITY_NOTES.md`: sanctioned divergences (a2ui-recovery per-slug
prompt isolation; declarative-json-render scoped-test divergence).

## Verification

- **D6 sweep: 38 green / 1 red.** The single red is **a2ui-recovery**,
which reproduces **identically on the north star** (shared Python
`ag_ui_langgraph` recovery loop; mastra's TS impl is green). It is not a
fastapi defect and is tracked as a separate PR against the north star.
Documented in `PARITY_NOTES.md`.
- `validate-manifests` (manifest → registry, all 20 integrations) —
green.
- `validate-routes --all` (runtime-route wiring) — green.
- Build + TypeScript typecheck (`next build` via the Docker image build)
— green.
- `@region` audit — all region pairs balanced and LGP-consistent (0
orphans/typos/mismatches).
- Live manual QA against real OpenAI (`:3102` vs `:3100`) confirmed the
key demos (reasoning-custom, shared-state-streaming, gen-ui-tool-based,
agentic-chat).

## Notable finding

Several demos were D6-green but broke against a real model because they
fell back to the generic `sample_agent` instead of their dedicated graph
— the aimock fixture masked the wiring bug by replaying scripted tool
calls. This PR fixes the wiring and documents the gap (GOTCHAS #8).
D6-green is necessary but not sufficient for graph/tool-dependent demos;
a real-LLM click-through is required.

## Deferred / follow-ups

- **a2ui-recovery** — shared north-star defect in the Python
`ag_ui_langgraph` recovery loop; separate PR against the north star.
- **d5-byoc probe** — always sends the hashbrown pill even on the
json-render page (fleet-wide harness limitation); tracked as a
follow-up. The gating `byoc` grid cell is green.

## Scope note

Only `langgraph-fastapi` integration code, its aimock fixtures, and
showcase tooling/docs changed. **langgraph-python (the north star) was
not touched.**
2026-07-28 12:49:57 +02:00
Tyler Slaton 6c15645b6b docs: organize Channels guides by provider and framework 2026-07-27 23:50:31 -04:00
Mike Ryan 9988c23890 docs: render shared concepts for Angular 2026-07-27 09:38:00 -07:00
github-actions[bot] b4cff204d1 style: auto-fix formatting 2026-07-27 15:10:04 +00:00
Lukas Moschitz c9d8a9ac14 fix(showcase): restore missing @region markers on 2 langgraph-fastapi agents
The factory automation extracts curated snippets via @region markers.
readonly_state_agent_context.py and shared_state_read_write.py were byte-aligned
to langgraph-python except their @region markers had been stripped, so the
factory would extract nothing for those two regions. Restore them to match LGP:
- agent-context-setup around create_agent in readonly_state_agent_context.py
- shared-state-setup around create_agent in shared_state_read_write.py
Comment-only; no behavior change. Audit confirms all 49 fastapi region pairs are
now balanced and LGP-consistent (0 orphans/typos/mismatches).
2026-07-27 17:03:08 +02:00
Lukas Moschitz efe51e7deb chore(showcase): remove fe-parity tool
fe-parity.ts was a scratch parity checker for this alignment work; it is no
longer needed in the worktree. Drop it and de-reference it from
langgraph-fastapi/PARITY_NOTES.md (reword the intro, convert the
machine-readable fe-parity-allow allowlist into a plain human-readable
'sanctioned file-level divergences' list). No other references existed. D6
behavior remains the parity judge.
2026-07-27 16:55:23 +02:00
Lukas Moschitz b8f37c0c85 docs(showcase): document declarative-json-render scoped-D6 divergence as sanctioned
The byoc D6 featureType covers both declarative demos; the grid cell is green
(hashbrown route+pill). A scoped declarative-json-render --d6 run reds for two
non-fastapi reasons: (1) the shared d5-byoc probe always sends the hashbrown
pill even on the json-render page (harness limitation — build context lacks
demos[]; tracked as a separate follow-up), and (2) fastapi deliberately keeps
hashbrown (@hashbrownai/react) and json-render (@json-render/react) as separate
library integrations with their own pills/fixtures, vs LGP's unified
render_dashboard contract. Sanctioned; do not collapse fastapi's json-render to
the unified contract (content downgrade) — the real fix is probe-side.
2026-07-27 16:49:05 +02:00
Lukas Moschitz bf2de51705 fix(showcase): wire langgraph-fastapi demos to dedicated graphs (drop sample_agent fallbacks)
Live QA surfaced demos that behaved wrong on fastapi because they fell back to
the generic sample_agent (or were unregistered) instead of the dedicated graph
LGP uses. aimock D6 masked these (fixtures script the tool calls), so they were
green in the grid but broken against a real model. Audit of fastapi's
agentNames fallthrough vs LGP's neutralAssistantCells found exactly these:

- gen-ui-tool-based: port gen_ui_tool_based graph (tools=[], frontend supplies
  render_bar/pie_chart via useComponent). Was sample_agent, whose query_data
  tool + prompt made the model loop on data queries instead of rendering.
- shared-state-streaming: port shared_state_streaming graph (StateStreaming
  middleware + write_document tool + document state). Was sample_agent, which
  never emits state.document, so it only wrote to chat. Remove from
  not_supported_features (now works, D6 green).
- agentic-chat: port agentic_chat graph (tools=[]). Was sample_agent (7+ tools).
- threadid-frontend-tool-roundtrip: wire to frontend_tools (was unregistered).
- reasoning-custom: align reasoning_agent config to LGP (gpt-5.4 / effort
  medium / summary detailed; were gpt-5-mini / low / auto).

Register the 3 new graphs in langgraph.json; remove the 3 names from the
sample_agent fallthrough loop in route.ts. D6 green x2 for reasoning-display,
gen-ui-custom, shared-state-streaming, agentic-chat; agent-config +
tool-rendering-reasoning-chain re-verified (no regression).
2026-07-27 16:24:50 +02:00
Lukas Moschitz 7bd1ca6e70 docs(showcase): record a2ui-recovery D6 red as a shared north-star defect
Soften the stale 'D6 green'/'under investigation' wording. A live trace
confirmed the a2ui-recovery surface-missing D6 red reproduces identically on
the langgraph-python north star (shared Python ag_ui_langgraph recovery loop;
mastra's TS impl is green), so fastapi is already behaviorally aligned and
there is nothing to fix under integrations/langgraph-fastapi/. Fixing the
shared loop is a separate PR against the north star / the package.
2026-07-27 14:30:21 +02:00
Lukas Moschitz 6af08217b2 fix(showcase): align langgraph-fastapi demo overview + canonical cell wiring to parity
The demo-browser overview (page.tsx, identical to LGP) groups cards by
demo.tags[0] and orders within a group by manifest.features[] index, so the
card arrangement is fully driven by manifest.yaml. fastapi's manifest was an
older curation: a different naming scheme (~26 demos), deprecated feature/demo
ids (agentic-chat-reasoning, reasoning-default-render, byoc-hashbrown,
byoc-json-render, hitl, phantom hitl-in-chat-booking), divergent tags/order,
and a missing shared-state-streaming/shared-state-read card.

Align manifest.yaml demos+features to LGP (names, descriptions, tags, order,
canonical ids), preserving fastapi identity (name/slug/logo/description) and
the sanctioned a2ui-recovery prompt divergence. Overview is now card-for-card
identical to langgraph-python.

Canonicalizing the ids pulled two cells into the D6 tested set that were
mis-wired to the old names; wire them to LGP's shape:
- reasoning-custom/reasoning-default -> reasoning_agent in copilotkit/route.ts
  (were wired under agentic-chat-reasoning/reasoning-default-render).
- rename api routes copilotkit-byoc-{hashbrown,json-render} ->
  copilotkit-declarative-{hashbrown,json-render}; align endpoints + the
  declarative-hashbrown-demo agent id to what the demo pages request.

Full D6 sweep: 37 green / 1 red; reasoning-display and byoc now green (two
real runs each); the lone red is a2ui-recovery (known shared north-star
defect, identical on LGP).
2026-07-27 14:30:21 +02:00
Lukas Moschitz a19ed403cf fix(showcase): port langgraph-fastapi tool-rendering-reasoning-chain to parity
Two independent fastapi divergences from the green north-star kept this red:

1. Missing chunkSize:9999 on the 8 tool-call fixtures. Under aimock's global
   8-byte chunking the large tool-call args failed to JSON-parse in one piece,
   so the AAPL->MSFT reasoning chain stopped at leg 1 (turns 1 & 2). Byte-align
   to LGP (and the green mastra/langgraph-typescript siblings carry the same
   pattern).

2. Fixture-pool shadowing broke turn 3. aimock pools all d4+d6 fixtures for a
   context and matches by userMessage substring, first-match-wins in load order
   (d4 before d6). fastapi had broad keys LGP doesn't:
   - d4/chat.json: 'weather' -> 'weather in San Francisco', 'flights from SFO
     to JFK' -> 'Find flights from SFO to JFK.' (period-terminated so it stops
     being a substring of turn 3's 'Find flights from SFO to JFK and show me
     the weather there.').
   - tool-rendering-{custom,default}-catchall.json: bare 'Find flights' ->
     'Find flights from SFO to JFK.' (now matches LGP's exact key).

Also byte-align the backend agent to LGP: scriptable get_stock_price signature,
the detailed chaining system prompt, model gpt-5.4, reasoning summary detailed.

D6 tool-rendering-reasoning-chain green (two real ~21s runs). Shared-fixture
regression all green: agentic-chat, tool-rendering, tool-rendering-custom-catchall,
tool-rendering-default-catchall, headless-complete (Tokyo-weather consumer).
2026-07-27 11:01:18 +02:00
Mark a3ad8960c9 Merge branch 'main' into mark/ms-agent-dotnet-apikeyresolver 2026-07-26 21:44:10 -07:00
Jordan Ritter 5a09e0e787 fix(showcase): pin ag2<1.0.0 to unblock CI after upstream module rename
ag2 1.0.0 (published 2026-07-27T00:58:37Z) renames its top-level module
from `autogen` to `ag2` and ships no `autogen/` package at all. The ag2
integration floats on `ag2[openai,ag-ui]>=0.9.0`, so every fresh CI
resolution now picks up 1.0.0 and collection fails at:

    src/agents/_multimodal_normalize.py:75
        from autogen.ag_ui import AGUIStream, RunAgentInput
    E   ModuleNotFoundError: No module named 'autogen'

This breaks `Python unit tests (3.10)` and `(3.12)` in showcase_validate
on main and on every open PR. It is time-based rather than commit-based:
main only looks green because it has not re-run since the release.

Pin rather than migrate. 1.0.0 is not a rename -- it deletes the stable
`autogen/ag_ui/adapter.py` and promotes the former `autogen.beta.ag_ui`
stream into its place, dropping `AGUIStream(event_interceptors=...)` and
replacing `dispatch(context=...)` with `dispatch(variables=...)`. Since
`_multimodal_normalize.py` overrides `dispatch` and forwards `context=`,
rewriting the import path alone would swap an import error for a runtime
TypeError. A real migration needs to port ~20 `autogen` import sites and
re-verify every ag2 demo cell, which is not an urgent-unblock change.
2026-07-26 21:03:03 -07:00
Mark a9ce37b75c fix(showcase/ms-agent-dotnet): resolve OPENAI_API_KEY-first via ApiKeyResolver
Port ms-agent-harness-dotnet's ApiKeyResolver into ms-agent-dotnet so the 15
agent factories resolve the OpenAI credential as OPENAI_API_KEY (env) ->
config[OPENAI_API_KEY] -> GitHubToken, and the endpoint via OPENAI_BASE_URL ->
default, instead of hardcoding configuration["GitHubToken"] per agent.

Previously ms-agent-dotnet's main chat clients authenticated only with the
GitHub token. On a fixture-miss fall-through, aimock proxies to real
api.openai.com, which rejects the GitHub token (invalid_api_key, surfaced as
502). ms-agent-harness-dotnet already works because it resolves OPENAI_API_KEY
first; this brings ms-agent-dotnet to parity so its fall-through returns 200.

Interim showcase-side mitigation while the aimock cross-provider guard
(PNI-108, CopilotKit/aimock#340) is deferred.

- Copy agent/ApiKeyResolver.cs verbatim from ms-agent-harness-dotnet (code
  copy, no new dependency)
- Rewire 15 factories + A2uiSecondaryToolCaller to ResolveApiKey/ResolveEndpoint;
  drop the dead per-file DefaultOpenAiEndpoint const
- Add tests/ApiKeyResolverTests.cs (precedence, mock-endpoint fallback,
  non-mock fail-fast)

Verified: dotnet build 0/0, dotnet test 76/76, whitespace format clean, and a
live OpenAI call through the resolver returned 200.
2026-07-27 03:56:33 +00:00
Jordan Ritter bde2d99a5b fix(showcase/ms-agent-python): keep the user's prompt on the multimodal PDF turn (#6159)
`d6:ms-agent-python/multimodal` has been red in staging and prod since
2026-05-30. Turn 1 (image) passes; turn 2 (PDF) fails. This fixes it —
**without touching the fixture**, because the fixture was never the
problem.

## The verbatim turn-2 error

Backend (`showcase-ms-agent-python`), and reproduced locally:

```
[/multimodal] Streaming failed
openai.InternalServerError: Error code: 503 - {'error': {'message': 'Strict mode: no fixture matched',
  'type': 'invalid_request_error', 'param': None, 'code': 'no_fixture_match'}}
The above exception was the direct cause of the following exception:
agent_framework.exceptions.ChatClientException: ("<class
  'agent_framework_openai._chat_completion_client.OpenAIChatCompletionClient'> service failed to
  complete the prompt: Error code: 503 - {'error': {'message': 'Strict mode: no fixture matched', …
```

Surfaced in the browser as `An internal error has occurred while
streaming events.`, with the probe reporting `failure_turn: 2`,
`turns_completed: 1`.

## Request-shape diagnosis

This reads like a fixture gap and is not one. I pulled the **actual
outbound request** off the local aimock's `GET /__aimock/journal` during
a failing run. Turn 2, verbatim (bodies elided):

```
[0] role=system  "You are a helpful assistant. The user may attach images or documents…"
[1] role=user    "can you tell me what is in this demo image I just attached"
[2] role=user    [image_url <data:image/png;base64,iVBORw0K…>]
[3] role=user    [image_url <data:image/png;base64,iVBORw0K…>]
[4] role=assistant "The attached image is the CopilotKit logo — a clean, geometric mark…"
[5] role=user    "can you tell me what is in this demo pdf I just attached"
[6] role=user    "[Attached document]\nCopilotKit Quickstart\nAdd AI copilots to your React…"
[7] role=user    "[Attached document]\nCopilotKit Quickstart\nAdd AI copilots to your React…"
```

One logical user turn arrived as **three separate user messages**, and
the *last* one carries only the flattened document — the question is
nowhere in it. That is why aimock's strict mode refused it:
`userMessage` is a substring match against the last user turn, and the
last user turn was a PDF dump.

**Root cause:** `agent_framework_openai` emits **one OpenAI message per
`Content`**. `_chat_completion_client._prepare_message_for_openai`
builds a fresh `args` dict on every iteration of its content loop, so a
user `Message` carrying `[prompt_text, flattened_doc_text]` serialises
to two consecutive user messages — prompt-only, then document-only.
`_PdfFlattenChatMiddleware` was appending the flattened `[Attached
document]` text as a *second* text `Content` beside the prompt, which is
exactly the shape that gets split.

Two corroborating details that make the mechanism airtight:

- **Why turn 1 (image) passes.** aimock already skips *text-less*
trailing user messages (`getLastUserText` in `router.ts`, whose comment
documents this exact MS Agent Framework behavior). The image turn's
split-off trailing message has no text at all, so aimock falls back to
the prompt message and matches. The PDF turn's trailing message *does*
have text — the document — so there is nothing to skip past.
- **Why `langgraph-python` is green** doing the identical `[Attached
document]` flattening: LangChain keeps multiple text parts *inside one
message* rather than splitting them into separate messages.

This is a product bug, not a mock artefact. Against a real LLM it would
not 503 — the model would just answer the wrong thing, because the
question is buried behind a document dump instead of being the current
turn.

## The fix

`showcase/integrations/ms-agent-python/src/agents/multimodal_agent.py`

1. **Merge** the flattened document *into* the message's existing prompt
text content instead of appending it as a second content. The turn stays
a single text content and serialises to a single user message:
`"<prompt>\n[Attached document]\n<body>"`.
2. The merge **copies** the prompt `Content` rather than mutating it.
This is load-bearing: the middleware restores the original `contents`
list after `call_next`, and that restore only undoes the *list* swap —
an in-place mutation would leak the raw PDF body into the AG-UI
`MESSAGES_SNAPSHOT` and render a wall of PDF text in the user's chat
bubble. There is a test for this.
3. **Attachment-only turns** (a PDF with no question) still work: with
no text content to merge into, the flattened document stands alone as
the message body.
4. **Dedupe identical flattened blocks.** The page's
`LegacyConverterShim` appends a legacy `binary` mirror alongside every
modern attachment part, so the same PDF reached the middleware twice and
its body was being sent to the model twice (visible as the duplicated
`[6]`/`[7]` above). Now emitted once.

Post-fix outbound turn 2, same journal endpoint:

```
[5] role=user "can you tell me what is in this demo pdf I just attached\n[Attached document]\nCopilotKit Quickstart\nAdd AI copilots to your React application with CopilotKit…"
matched fixture userMessage: "can you tell me what is in this demo pdf I just attached"
```

One user message, prompt intact, document intact, emitted once.

## The fixture is untouched

```
$ git diff --stat origin/main -- showcase/aimock/
(empty)
```

The existing `userMessage` match key was always correct; the corrected
request shape is what satisfies it. Relaxing or re-recording the fixture
to match the broken request was an explicit non-goal — it would have
made the cell actively certify a model that never sees the user's
question.

## Same-pattern audit

- `_PdfFlattenChatMiddleware` is the **only** `ChatMiddleware` in
`ms-agent-python`, and the only place in the integration that constructs
`Content` or reassigns `message.contents` (`grep` for `ChatMiddleware` /
`Content.from_text` / `.contents =` across `src/` returns hits in this
one file only). No second instance of the pattern to fix.
- `ms-agent-python` is the only MS-Agent-Framework Python integration
doing PDF flattening — `ms-agent-dotnet` has a multimodal e2e spec but
no Python agent. The other `[Attached document]` implementations
(`langgraph-python`, `langgraph-fastapi`, `agno`, `claude-sdk-python`,
`langroid`, `pydantic-ai`, `langgraph-typescript`, `built-in-agent`) run
on frameworks that do not split a message's contents into separate wire
messages, so they are not exposed to this. The upstream
one-message-per-`Content` behavior is pinned by a dedicated test, so if
it ever changes we find out by that test failing rather than by a silent
regression.
- The file is a regular per-integration file, not a `shared/` symlink
(`git ls-files -s` → `100644`). No shared code touched;
`validate-shared-symlinks.ts` confirms no new erosion.

## Red / green / control

All three on the real probe surface, from a clean worktree at
`origin/main` `38613623f4`.

### RED — before the change

```
$ bin/showcase test ms-agent-python:multimodal --d6 --direct --verbose --cycle --isolate

[conversation-runner] turn 1/2 — assistant settled { bubbleIndex: 0, textLength: 100, hasAssertions: true }
[conversation-runner] turn 1/2 — assertions passed
[conversation-runner] turn 2/2 — sending message { inputLength: 29, timeoutMs: 60000 }
[conversation-runner] turn 2/2 — FAILED {
  errorCategory: 'assertion-failed',
  turnsCompleted: 1,
  elapsedMs: 1577,
  bodyTextLength: 421,
  hasTextarea: true,
  hasErrorBoundary: false
}
[warn] CVDIAG component=harness-d6 boundary=fixture-match … status=miss … error=chat errored: copilot-error-banner visible — An internal error has occurred while streaming events.
[info] probe.e2e-full.service-complete {"slug":"ms-agent-python","passed":0,"failed":1,"skipped":0,"incapable":0,"total":1,"state":"red","durationMs":9384}
  ✗ d6:ms-agent-python red (9.5s)
    multimodal: chat errored: copilot-error-banner visible — An internal error has occurred while streaming events.

  0 passed, 1 failed (9.5s)
⚠ Tests failed for ms-agent-python:multimodal (exit 1)
```

Evidence the outbound request lacked the prompt — aimock journal from
that run, 8 entries, `200,503,503,503,200,503,503,503` (2 attempts × 3
retries on turn 2):

```
[5] role=user STRING "can you tell me what is in this demo pdf I just attached"
[6] role=user STRING "[Attached document]\nCopilotKit Quickstart\nAdd AI copilots to…"
[7] role=user STRING "[Attached document]\nCopilotKit Quickstart\nAdd AI copilots to…"
status: 503
```

### GREEN — after the change, fixture unchanged

```
$ bin/showcase test ms-agent-python:multimodal --d6 --direct --verbose --rebuild --keep --isolate

[conversation-runner] turn 1/2 — assistant settled { bubbleIndex: 0, textLength: 100, hasAssertions: true }
[conversation-runner] turn 1/2 — assertions passed
[conversation-runner] turn 2/2 — assistant settled { bubbleIndex: 1, textLength: 233, hasAssertions: true }
[conversation-runner] turn 2/2 — assertions passed
[conversation-runner] conversation completed successfully { turnsCompleted: 2, totalDurationMs: 8279 }
[info] probe.e2e-full.feature-complete {"slug":"ms-agent-python","featureType":"multimodal","pass":true,"durationMs":8788}
[info] probe.e2e-full.service-complete {"slug":"ms-agent-python","passed":1,"failed":0,"skipped":0,"incapable":0,"total":1,"state":"green","durationMs":10187}
  ✓ d6:ms-agent-python green (10.5s)

  1 passed (10.5s)
✓ Tests passed for ms-agent-python:multimodal
```

Both turns pass. aimock journal for that run: **2 entries, statuses
`200,200`** (down from 8 entries with six 503s — no retries needed).
**The fixture was not modified**; `git diff origin/main --
showcase/aimock/` is empty and the diff is two files, both under
`showcase/integrations/ms-agent-python/`.

### CONTROL — an already-green integration, same command, same stack

```
$ bin/showcase test langgraph-python:multimodal --d6 --direct --isolate

[conversation-runner] turn 2/2 — assistant settled { bubbleIndex: 1, textLength: 233, hasAssertions: true }
[conversation-runner] turn 2/2 — assertions passed
[conversation-runner] conversation completed successfully { turnsCompleted: 2, totalDurationMs: 8395 }
  ✓ d6:langgraph-python green (9.1s)

  1 passed (9.1s)
✓ Tests passed for langgraph-python:multimodal
```

Local harness, shared probe, shared frontend and fixtures are all sound
— the red was specific to this integration.

## Covering test


`showcase/integrations/ms-agent-python/tests/python/test_multimodal_pdf_prompt.py`
— 7 tests. Not fakes: each one drives the real
`_PdfFlattenChatMiddleware` and then the real
`OpenAIChatCompletionClient._prepare_message_for_openai`, and asserts
against the actual OpenAI wire payload. The PDF is the bundled
`public/demo-files/sample.pdf` through real `pypdf`, and the prompt
asserted on is **read out of the real aimock fixture** rather than
hardcoded, so the test fails if either side drifts.

Test-level red→green (stash the source change, keep the tests):

```
# pre-fix
FAILED test_multimodal_pdf_prompt.py::test_pdf_turn_last_user_message_contains_the_prompt
FAILED test_multimodal_pdf_prompt.py::test_pdf_turn_serialises_to_a_single_user_message
FAILED test_multimodal_pdf_prompt.py::test_duplicate_pdf_parts_are_flattened_once
3 failed, 4 passed in 2.37s
```

with the primary failure reading:

```
AssertionError: expected the PDF turn to serialise to 1 user message, got 2:
  ['can you tell me what is in this demo pdf I just attached',
   '[Attached document]\nCopilotKit Quickstart\nAdd AI copilots to']
```

```
# post-fix — full integration suite (6 pre-existing CVDIAG + 7 new), CI's exact invocation
$ PYTHONPATH=".:src" python -m pytest tests/python/ -q
13 passed in 2.40s
```

Coverage: prompt survives to the final user turn; the turn stays one
user message; the upstream one-message-per-`Content` split is pinned;
original `contents` restored and the prompt `Content` not mutated;
duplicate mirror parts flattened once; attachment-only turn still
flattens; image turn left byte-identical.

## Pre-push

`validate-parity.ts` 20/20 pass · `validate-shared-symlinks.ts` no new
erosion · `aimock-fixtures.test.ts` 842 pass · full `tests/python/`
suite 13 pass · lefthook `lint-fix` + `commitlint` clean · Python lines
≤88 cols matching the file's existing style · no lockfile churn, two
files in the diff.

## Scope

One cell, one middleware, one integration. The other five red
`multimodal` cells from the same sweep have five different root causes
and are not addressed here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01PYdjeveT8Xof9TyHWMLoJr
2026-07-26 00:11:39 -07:00
Jordan Ritter db75a04837 chore(showcase): mark multimodal unsupported for llamaindex and crewai-crews (#6158)
## What this is

`multimodal` (Attachments) has never worked on **`llamaindex`** or
**`crewai-crews`**, but both manifests
listed it under `features`, so the fleet probed it and reported **red**.
A red chip says "this regressed".
The truth is "this was never built". This PR marks both cells
**unsupported** instead. It does **not**
implement the feature, and it does **not** suppress the cell — the cell
still exists, the demo stays wired,
and the chip renders the 🚫 unsupported glyph.

## Mechanism used (existing, not invented)

The repo already has exactly one way to declare a feature unsupported
for an integration:
**`not_supported_features` in the integration's `manifest.yaml`**. The
full derivation chain:

| Step | Location |
|---|---|
| Declaration | `showcase/integrations/<slug>/manifest.yaml` ->
`not_supported_features:` |
| Schema | `showcase/shared/manifest.schema.json` — *"feature IDs that
this integration's framework cannot architecturally support … excluded
from parity computation"* |
| Status fold |
`showcase/harness/src/shared/catalog/catalog-flatten.ts:239` —
`determineCellStatus()` checks `not_supported_features` **first**,
returns `status: "unsupported"` |
| Input mapping |
`showcase/harness/src/shared/cell-model/catalog-input.ts:53` —
`isSupported: cell.status !== "unsupported"` |
| Model | `showcase/harness/src/shared/cell-model/cell-model.ts:847` —
`if (!isSupported) return UNSUPPORTED;` (the frozen singleton at `:551`:
`supported: false`, `chipColor: "gray"`, `isRegression: false`) |
| `/api/matrix` | `showcase/harness/src/http/matrix.ts:206` ->
`matrix-compute.ts:53` — projects that same model, so the API value
**is** the rendered chip by construction |
| Render |
`showcase/shell-dashboard/src/components/unified-cell.tsx:308` — `if
(!model.supported)` renders `data-testid="unified-cell-unsupported"`
with 🚫 and `title="Not supported by this framework"` |

The mechanical guard at `catalog-flatten.ts:169` rejects a feature that
appears in **both** `features` and
`not_supported_features`, so each entry was **moved**, not duplicated.

Note this mechanism is strictly stronger than a probe-side skip:
`buildCellModel` returns `UNSUPPORTED`
regardless of what the live PocketBase row says. Verified against the
existing not-supported cells on these
same two integrations, which carry **green** PB rows and still render 🚫:

```
llamaindex/gen-ui-interrupt        matrix: chip=gray supported=False | PB rows: d5=green d6=green e2e=green
llamaindex/shared-state-streaming  matrix: chip=gray supported=False | PB rows: d5=green d6=green e2e=green
crewai-crews/mcp-apps              matrix: chip=gray supported=False | PB rows: d5=green d6=green e2e=green
```

So the cell can never read green *or* red once declared here — which is
the property we want.

## Per-integration reason (recorded inline in each manifest)

**`llamaindex` — upstream gap.** The pinned
`llama-index-protocols-ag-ui==0.2.2`
(`llama_index/protocols/ag_ui/utils.py:82-85`) passes an AG-UI
`UserMessage`'s `content` straight into
`ChatMessage(...)`. A text-only turn passes a plain string (fine — every
other llamaindex cell is green);
an attachment turn passes a **list** of AG-UI content-part models, which
pydantic routes into
`ChatMessage.blocks`, a union discriminated on `block_type` — a field
AG-UI's `TextInputContent` /
`ImageInputContent` / `BinaryInputContent` do not carry. Live backend
error:

```
pydantic_core._pydantic_core.ValidationError: 3 validation errors for ChatMessage
blocks.0
  Unable to extract tag using discriminator 'block_type' [type=union_tag_not_found,
    input_value=TextInputContent(type='te... image I just attached'), input_type=TextInputContent]
```

Fixing it needs a content-part ->
`TextBlock`/`ImageBlock`/`DocumentBlock` conversion, upstream or on our
side of `get_ag_ui_workflow_router`. This PR also corrects that module's
docstring, which asserted the
router *"normalizes them via the OpenAI `input_file` path"* — it does
not.

**`crewai-crews` — never implemented, ours.** `src/agent_server.py`
registers a dedicated AG-UI endpoint for
every other demo but **no `/multimodal` route** (the block ends at the
catch-all
`add_crewai_crew_fastapi_endpoint(app, LatestAiDevelopment(), "/")`),
and there is no
`src/agents/multimodal_agent.py`. So
`src/app/api/copilotkit-multimodal/route.ts` aliases the generic shared
crew, which has no vision handling and dies on a content-part message:

```
[CopilotKit] Error (agent_run_error_event): Error: thread=… run=…: CrewAI flow failed; see server logs
```

That route file's own header comment already concedes the gap ("A
dedicated per-demo crew with vision-tuned
agent prompts is tracked as follow-up work").

## Proof

Method: the **real** `GET /api/matrix` handler (`registerMatrixRoute`)
driven over the **real** production
PocketBase `status` collection (all 3082 rows, fetched verbatim from
`showcase-pocketbase-production.up.railway.app`) and the **real**
on-disk manifests (default `loadCells` =
`buildCatalogCells`, the single flattening authority). Same rows, same
fixed clock
(`now = max(observed_at) + 60s = 1784932566438`) for both runs — the
only variable is the manifest diff.

### BEFORE — real `GET /api/matrix`, real prod PocketBase rows,
manifests at `origin/main` (38613623f4)

```
llamaindex/multimodal
  {"chipColor": "red", "supported": true, "achievedDepth": 4, "ceilingDepth": 6, "isRegression": true, "surfaceState": "red", "isStaleCell": false}
crewai-crews/multimodal
  {"chipColor": "red", "supported": true, "achievedDepth": 4, "ceilingDepth": 6, "isRegression": true, "surfaceState": "red", "isStaleCell": false}
```

### AFTER — same route, same rows, same clock, manifests with this PR

```
llamaindex/multimodal
  {"chipColor": "gray", "supported": false, "achievedDepth": 0, "ceilingDepth": 0, "isRegression": false, "surfaceState": "gray", "isStaleCell": false}
crewai-crews/multimodal
  {"chipColor": "gray", "supported": false, "achievedDepth": 0, "ceilingDepth": 0, "isRegression": false, "surfaceState": "gray", "isStaleCell": false}
```

### Full multimodal column, AFTER (control)

```
ag2                      chip=green  supported=True  depth=6/6  [unchanged]
agno                     chip=red    supported=True  depth=4/6  [unchanged]
built-in-agent           chip=red    supported=True  depth=4/6  [unchanged]
claude-sdk-python        chip=green  supported=True  depth=6/6  [unchanged]
claude-sdk-typescript    chip=green  supported=True  depth=6/6  [unchanged]
crewai-crews             chip=gray   supported=False depth=0/0  [CHANGED]
google-adk               chip=green  supported=True  depth=6/6  [unchanged]
langgraph-fastapi        chip=green  supported=True  depth=6/6  [unchanged]
langgraph-python         chip=green  supported=True  depth=6/6  [unchanged]
langgraph-typescript     chip=green  supported=True  depth=6/6  [unchanged]
langroid                 chip=green  supported=True  depth=6/6  [unchanged]
llamaindex               chip=gray   supported=False depth=0/0  [CHANGED]
mastra                   chip=red    supported=True  depth=4/6  [unchanged]
ms-agent-dotnet          chip=green  supported=True  depth=6/6  [unchanged]
ms-agent-harness-dotnet  chip=green  supported=True  depth=6/6  [unchanged]
ms-agent-python          chip=red    supported=True  depth=4/6  [unchanged]
pydantic-ai              chip=green  supported=True  depth=6/6  [unchanged]
spring-ai                chip=green  supported=True  depth=6/6  [unchanged]
strands                  chip=green  supported=True  depth=6/6  [unchanged]
strands-typescript       chip=green  supported=True  depth=6/6  [unchanged]
```

### CONTROL — an already-supported multimodal cell is untouched

`langgraph-python/multimodal` (the reference integration) is `chip=green
supported=true depth=6/6` **before
and after**, byte-identical. So is every other integration's multimodal
cell, including the four that are
red for unrelated reasons (`agno`, `mastra`, `ms-agent-python`,
`built-in-agent`) — those stay **red**, they
were not swept up.

### Complete set of cells whose state changed — exactly 2 of 1000

Diffing every field of every cell in the `/api/matrix` body, before vs
after (identical 1000-cell keyset):

```
crewai-crews/multimodal   red -> unsupported
llamaindex/multimodal     red -> unsupported
```

Nothing else. Aggregate confirms no green was manufactured:

```
                BEFORE                                   AFTER
chipColor       green=632  gray=317  red=51              green=632  gray=319  red=49
supported=false 94                                       96
```

`green` is **unchanged at 632** — this PR turned two reds into
unsupported and created zero greens.

Second, independent derivation (the dashboard's own generated
`catalog.json`, via
`showcase/scripts/generate-registry.ts`, which runs full AJV validation
first) agrees, and also changes
exactly 2 cells:

```
crewai-crews/multimodal (integrated): status wired -> unsupported, max_depth 4 -> 0
llamaindex/multimodal   (integrated): status wired -> unsupported, max_depth 4 -> 0

metadata BEFORE: total_cells 980, wired 688, stub 0, unshipped 198, unsupported 94, docs_only 20
metadata AFTER : total_cells 980, wired 686, stub 0, unshipped 198, unsupported 96, docs_only 20
```

`parity_tier` is unchanged on both columns (already `partial`), and no
other integration's cells moved.

### Live dashboard render

`shell-dashboard` run locally against the prod PocketBase, Playwright
over the real DOM. On the
`feature-row-multimodal` ("Attachments") row, exactly two of twenty
columns carry
`data-testid="unified-cell-unsupported"`:

```
CrewAI (Crews)     unsupportedGlyph=TRUE   text="🚫"
LlamaIndex         unsupportedGlyph=TRUE   text="🚫"
LangGraph (Python) unsupportedGlyph=false  text="Demo ↗ Code </> D6 UI ✓ BE ✓ 1P ✓ D6 ✓"
Agno               unsupportedGlyph=false  text="Demo ↗ Code </> D4 UI ✓ BE ✓ 1P ✗ D6 —"
… 16 more, all unsupportedGlyph=false
```

Visually: a grey outlined 🚫 badge in those two columns — plainly not a
green `D6` pill, and not a red `✗`.
Control row `feature-row-agentic-chat` has `unsupportedGlyph=false` in
all twenty columns.

## Quality gates

- `oxfmt --check` on all three changed files — clean
- `oxlint` — 0 warnings, 0 errors
- `generate-registry.ts --validate-only` (AJV against
`manifest.schema.json`) — passes; confirms no
  `features` / `not_supported_features` overlap
- `shell-dashboard`: production build ✓, **68 test files / 1333 tests
passed**, 1 skipped
- `harness`: `tsc --noEmit` clean; **171 test files / 3625 tests
passed**. 3 pre-existing failures
(`d5-mapping-drift`, `frontend-matrix`, `d5-representatives` — the last
complains about
`browser-use-smoke`) fail **identically on clean `origin/main`**,
verified by stashing this diff and
  re-running. Unrelated to this change.
- Diff is 3 files, no lockfile churn, no generated artifacts, no stray
worktree files.

## Deliberately not done

- Not implementing the feature for either integration (the upstream
conversion for llamaindex and the
  vision crew for crewai-crews remain open work).
- Not touching `agno` / `mastra` / `ms-agent-python` / `built-in-agent`,
whose multimodal cells are red for
  four unrelated reasons and stay red here.
- Not loosening the D5 assertion, not dropping `skipSend`, not deleting
the cell, not skipping the probe,
and not special-casing the probe to pass — every one of those would
green a broken cell.

## Follow-up commit: CI-caught wired-count bound

The first push failed **Showcase: Validate** -> "Run build pipeline
tests" (the `showcase/scripts` vitest,
which I had not run locally — my mistake; the harness and dashboard
suites both passed):

```
FAIL __tests__/generate-catalog.test.ts > parity tier: crewai-crews wired cells render
     at_parity or partial against the elected reference
AssertionError: expected 29 to be greater than or equal to 30
 ❯ __tests__/generate-catalog.test.ts:257:32
```

Real and caused by this PR: reclassifying `crewai-crews/multimodal` from
`wired` to `unsupported` drops that
integration's wired-cell count 30 -> 29.

That `30` is a **snapshot floor, not an invariant** — the assertion's
own comment states the partial parity
tier requires only `intersection >= 3` with the reference's wired set.
29 clears that by a wide margin, and
the two tier assertions immediately below it (`parity_tier` in
`["at_parity","partial"]`, uniform across the
column) still pass untouched. So the floor was updated to 29 with a
comment recording exactly which cell
moved and why, rather than being deleted or loosened to a no-op.

Nothing else in that suite moved: `metadata.total_cells` is still 980
and the
`wired + stub + unshipped + unsupported == total_cells` sum invariant
still holds (686 + 0 + 198 + 96 = 980).

Re-run locally after the fix: **72 test files / 2323 tests passed, 0
failed.**
`oxfmt --check` and `oxlint` clean on the changed test file.
2026-07-25 23:18:35 -07:00
Jordan Ritter 59f275eedc fix(showcase/mastra): restore shared-tools symlink, shrink erosion ratchet (#6161)
## What

`showcase/integrations/mastra/shared-tools` was a **real committed
directory** where a **symlink into the single source of truth** belongs.
This restores the symlink (`-> ../../shared/typescript/tools`) and
shrinks the erosion ratchet accordingly.

This is the iron-rule violation described in `showcase/AGENTS.md` → "The
single-source symlink mechanism": `*/shared-tools`, `*/tools`, and
`*/_shared` are meant to be symlinks into `showcase/shared/...`, so
edits to the shared source silently do NOT reach an integration whose
symlink has been clobbered by a real directory.

## Root cause

| | |
|---|---|
| Symlink originally added | `93d5815cdb` — *"refactor: add shared tools
symlink to mastra, update path alias"* (target
`../../shared/typescript/tools`) |
| Symlink clobbered | `534cd1efa7` — *"fix(showcase): D5 integration
fixes across 12 frameworks"* — deleted the symlink (`shared-tools \| 1
-`) and committed **14 real files** in its place |

That commit clobbered `shared-tools` for **three** TS integrations at
once: `mastra`, `langgraph-typescript`, and `claude-sdk-typescript` —
all three grandfathered in
`showcase/scripts/validate-shared-symlinks.baseline.json`.

mastra's Dockerfile already documented the symlink as the expected
state, so the tree had drifted from its own documented contract:

```
# shared-tools/ is a symlink to ../../shared/typescript/tools — resolved at
# build time by CI which copies the target into the build context.
```

## Divergence inventory — **NONE**

Diffed the real directory against `showcase/shared/typescript/tools`
exhaustively **before** changing anything. **Zero divergence**:
identical file set, identical git blob hashes, identical sha256s across
all 14 files.

| File | git blob (main) | classification |
|---|---|---|
| `__tests__/generate-a2ui.test.ts` | `b3c8223e` | identical copy |
| `__tests__/get-weather.test.ts` | `57dd5b94` | identical copy |
| `__tests__/query-data.test.ts` | `ef284b0b` | identical copy |
| `__tests__/sales-todos.test.ts` | `8c5c0d6e` | identical copy |
| `__tests__/schedule-meeting.test.ts` | `ab470b3b` | identical copy |
| `__tests__/search-flights.test.ts` | `f64b8782` | identical copy |
| `generate-a2ui.ts` | `7d76c00c` | identical copy |
| `get-weather.ts` | `0b444f7a` | identical copy |
| `index.ts` | `3b0e2a45` | identical copy |
| `query-data.ts` | `48a2cbdc` | identical copy |
| `sales-todos.ts` | `96b1874f` | identical copy |
| `schedule-meeting.ts` | `d5e626e8` | identical copy |
| `search-flights.ts` | `7a943023` | identical copy |
| `types.ts` | `f87f26d1` | identical copy |

Classification totals:
- **stale** (shared moved on, mastra behind): none
- **mastra-specific** (intentional or accidental local edit): **none**
- **additive** (file only in mastra's copy): none
- files present in one and not the other: **none**

```
$ git diff --no-index --stat integrations/mastra/shared-tools shared/typescript/tools
(no output — identical)
```

Because the divergence set is empty, restoring the symlink is a **pure
structural fix with zero content change** — nothing mastra-specific
existed in the copy, so nothing can be silently lost. The restored link
blob is `ddc634b6` — **the same git object** the pre-erosion symlink
had.

## AFTER — structural proof

```
$ ls -la showcase/integrations/mastra/shared-tools
lrwxr-xr-x  shared-tools -> ../../shared/typescript/tools

$ git ls-files -s showcase/integrations/mastra/shared-tools
120000 ddc634b6d9 0	showcase/integrations/mastra/shared-tools
```

Verified on a **fresh checkout of this commit** (not merely the working
tree): mode `120000`, link resolves, and the content mastra resolves to
is byte-identical to `showcase/shared/typescript/tools`.

Validator on a fresh checkout — what CI sees:

```
ℹ 2 single-source slot(s) are ERODED
  • [known] showcase/integrations/claude-sdk-typescript/shared-tools  (real-dir)
  • [known] showcase/integrations/langgraph-typescript/shared-tools  (real-dir)

✔ single-source symlinks OK — no NEW erosion (2/2 baselined slot(s) still eroded)
```

`mastra/shared-tools` removed from
`validate-shared-symlinks.baseline.json` per the shrink-only ratchet —
the validator reports stale entries specifically so this step is
mechanical. Baseline: **3 → 2**.

## Docker-image symlink verification

A `COPY` can dereference or break these symlinks, and there is prior art
in this repo of `docker build` eroding these exact links — so this was
tested directly rather than assumed.

**1. A naked `docker build` with the symlink FAILS** — expected, and
exactly why `stage_shared()` exists (the target is outside the build
context):

```
ERROR: failed to compute cache key: "/shared-tools": not found
```

**2. After the `stage_shared()` dereference the build succeeds and REAL
FILES land in the image** — no dangling symlink:

```
=== /app/shared-tools inode type ===
drwxr-xr-x  shared-tools        <- real dir, not a link
=== file count === 8
```

**3. Verified inside the actual running showcase container** built from
this branch:

```
$ docker exec showcase-iso2-mastra sh -c 'test -L /app/shared-tools && echo IS_SYMLINK || echo IS_REAL_DIR'
IS_REAL_DIR
```

All 8 in-container `sha256sum` values match
`showcase/shared/typescript/tools` byte-for-byte.

**4. Every build path stages first** — this is not a new mechanism, it
is the one 13 other integrations already rely on:

- `stage_shared()` in `showcase/scripts/cli/_common.sh` (called by
`cmd-recreate.sh` under `trap restore_symlinks EXIT`)
- `stageSharedModules()` in `showcase/harness/src/cli/lifecycle.ts` (4
call sites, each paired with `restoreSymlinks()`)
- `.github/workflows/test_showcase-frontend-matrix.yml` runs
`stage_shared` in the step immediately before its raw `docker build`

Both stagers iterate `["tools", "shared-tools", "_shared"]` generically,
so `mastra/shared-tools` is handled with **no code change**. 13 other
integrations already `COPY` a symlinked shared dir in their Dockerfile
(`ag2`, `agno`, `claude-sdk-python`, `crewai-crews`, `google-adk`,
`langgraph-fastapi`, `langgraph-python`, `langroid`, `llamaindex`,
`ms-agent-dotnet`, `ms-agent-harness-dotnet`, `ms-agent-python`,
`pydantic-ai`, `strands`) and depend on this same staging.

Confirmed the full `stage_shared` → build → `restore_symlinks`
round-trip: after a complete probe run the symlink is intact and the
tree is clean.

**Consequence worth stating plainly:** because the staged bytes are
identical before and after this change, the Docker build context is
**bit-identical** either way — the COPY layer even cache-hits. The built
image cannot differ, which is the structural reason no cell can flip.

## Behavior proof — probe verdicts

Ran the **real D5 probes** (`bin/showcase test <slug>:<feature> --d5
--direct`, real Docker containers + Playwright), on two features, from
two separate worktrees: a pristine `origin/main` checkout (BEFORE) and
this branch (AFTER). Not unit tests against fakes.

| feature | BEFORE (`origin/main`, real dir) | AFTER (symlink) | delta |
|---|---|---|---|
| `mastra:tool-rendering` | **green** — 1 passed, 0 failed | **green** —
1 passed, 0 failed | **unchanged** |
| `mastra:multimodal` | **red** — dom-missing | **red** — dom-missing
(identical signature) | **unchanged** |

```
BEFORE tool-rendering: {"slug":"mastra","passed":1,"failed":0,...,"state":"green","durationMs":6443}
AFTER  tool-rendering: {"slug":"mastra","passed":1,"failed":0,...,"state":"green","durationMs":4697}

BEFORE multimodal:     {"slug":"mastra","passed":0,"failed":1,...,"state":"red","durationMs":122339}
AFTER  multimodal:     {"slug":"mastra","passed":0,"failed":1,...,"state":"red","durationMs":121494}
```

**No cell changed state.**

`tool-rendering` is the load-bearing one — it is precisely the feature
that consumes `shared-tools`. Its manifest highlights
`shared-tools/{get-weather,query-data,schedule-meeting,search-flights}.ts`,
and `src/mastra/tools/index.ts` imports the impls through the alias that
resolves into the converted directory:

```ts
import {
  getWeatherImpl, queryDataImpl, scheduleMeetingImpl,
  searchFlightsImpl, generateA2uiImpl, buildA2uiOperationsFromToolCall,
} from "@copilotkit/showcase-shared-tools";   // tsconfig alias -> ./shared-tools (now the symlink)
```

Its staying green is direct evidence that the symlink resolves, the
shared tool impls are found and executed, and Next.js bundles them
correctly.

### On the multimodal RED

`mastra:multimodal` is **already red on `origin/main`** — I reproduced
it on the pristine BEFORE worktree with the byte-identical failure
signature:

```
waitForTurnComplete: turn 1 did not complete within 60000ms
  (reason=dom-missing, runsFinished=0, count=0, attrPresent=true, runningNow=false, runStartCount=0)
```

Its cause is unrelated to `shared-tools`: `mastra` ships **no
`public/demo-files/` directory at all**, so the probe's `sample.png` /
`sample.pdf` attachment 404s. Verified independently — `mastra/public/`
contains only `.gitkeep`, two logo SVGs, `demo-audio/` and the `angular`
symlink. `agno` has the same gap. This is a separate pre-existing
packaging bug and is explicitly **not** addressed here.

Also confirmed the symlink survives a full probe run: after
`stageSharedModules()` dereferenced it tree-wide and `restoreSymlinks()`
ran, the link is intact and the tree is clean.

## Other slots / other integrations

`mastra` has **no** `tools/` or `_shared/` slot at all, so
`shared-tools` was its only exposure.

Full audit of all three slots across all 17 integrations on a clean
checkout — the only remaining real-dirs are the two still-baselined:

| integration | slot | state |
|---|---|---|
| `mastra` | `shared-tools` | **symlink (fixed here)** |
| `langgraph-typescript` | `shared-tools` | real dir — still baselined |
| `claude-sdk-typescript` | `shared-tools` | real dir — still baselined
|
| all 14 others | `tools` / `_shared` | healthy symlinks |

`langgraph-typescript/shared-tools` and
`claude-sdk-typescript/shared-tools` are eroded the same way, by the
same commit, and are **also byte-identical** to the shared source — so
both are trivially convertible. They are deliberately **left for
follow-up** rather than folded in here so that each conversion carries
its own behavior proof; `langgraph-typescript` in particular imports via
a relative `from "../../shared-tools"` rather than mastra's tsconfig
path alias, which warrants its own verification. Healing those two takes
the baseline to `[]`, at which point the guard becomes **fully
enforcing**.

## Control

`langgraph-python` (an integration that already had correct symlinks) is
untouched by this diff — its `tools`/`_shared` symlinks remain intact
and resolve to content identical to the shared source. This commit
touches only `showcase/integrations/mastra/shared-tools` and
`showcase/scripts/validate-shared-symlinks.baseline.json`.

I also verified my own worktree checkout had not itself mangled symlinks
(a real risk with `git worktree`) **before** drawing any conclusion
about mastra — every other integration's symlinks were intact in the
same checkout, so the mastra real-dir was genuine and not an artifact.

## Tests

- `showcase/scripts/__tests__/validate-shared-symlinks.test.ts` —
**11/11 pass** on a clean checkout, including the case asserting zero
fresh erosion and zero stale baseline entries against the real repo
- `showcase/scripts/__tests__/extract-starter.test.ts` — pass
- `npx tsx showcase/scripts/extract-starter.ts mastra` — correctly
dereferences the symlink into real files in the extracted starter (its
docstring names `mastra` as the example; that path is symlink-aware by
design)
- `prettier --check` clean on the changed baseline JSON

> Operational note: running the validator test **while** a `showcase
test` run is in flight reports 27 spurious erosions, because
`stageSharedModules()` has temporarily dereferenced every symlink
tree-wide mid-run. Not a defect — but the validator test is not safe to
run concurrently with a probe run.
2026-07-25 23:18:18 -07:00
Jordan Ritter c2e9264dde Merge branch 'main' into fix/ms-agent-python-multimodal-prompt 2026-07-25 23:04:20 -07:00
Jordan Ritter 4a303f8bed Merge branch 'main' into chore/multimodal-unsupported-llamaindex-crewai 2026-07-25 23:04:14 -07:00
Jordan Ritter 2825b10ba1 Merge branch 'main' into fix/mastra-shared-tools-symlink 2026-07-25 23:04:09 -07:00
Jordan Ritter b907a2660a Merge branch 'main' into fix/showcase-integration-page-titles 2026-07-25 23:04:04 -07:00
Jordan Ritter 1c24e0644b chore(showcase): store demo assets uniformly as Git LFS pointers (#6163)
## What

Demo assets under
`showcase/integrations/*/public/{demo-files,demo-audio}/` were stored
two different ways. This makes all of them LFS pointers — the root
`.gitattributes` convention that the majority already follow.

**Before this PR, on `main`:**

| form | integrations |
|---|---|
| **LFS pointers** (10) | `ag2`, `built-in-agent`, `claude-sdk-python`,
`claude-sdk-typescript`, `crewai-crews`, `langroid`, `llamaindex`,
`ms-agent-harness-dotnet`, `pydantic-ai`, `spring-ai` |
| **Raw blobs + per-integration `.gitattributes` override** (8) |
`google-adk`, `langgraph-fastapi`, `langgraph-python`,
`langgraph-typescript`, `ms-agent-dotnet`, `ms-agent-python`, `strands`,
`strands-typescript` |
| **Missing entirely** (2) | `agno`, `mastra` — being added as pointers
on #6157 |

Committed sizes: pointer-mode `sample.png` 130 B / `sample.pdf` 129 B /
`sample.wav` 130 B; raw-mode `sample.png` 10083 B / `sample.pdf` 2486 B
/ `sample.wav` 87078 B.

## Why the overrides existed, and why they no longer need to

Each override re-declared the same paths `-filter -diff -merge` so the
files would be committed as real binaries. The stated reason (in their
own comments) was that the image build didn't run `git lfs pull`, so a
pointer stub shipped into the image and the multimodal sample-attachment
magic-bytes guard rejected it.

That premise no longer holds. The deploy build's Checkout step in
`showcase_build.yml` hardcodes `lfs: true` (added in `7bde1eef3aa`), so
every integration image gets real binaries regardless of storage form.
The 10 pointer-mode integrations already ship through that exact
pipeline today. The overrides are now dead weight that only buys
divergence.

## What changed

- Deleted all 8 override `.gitattributes`. Each contained **nothing
but** demo-asset exemptions (the `langgraph-*`/`strands*` ones also
exempted `public/demo-audio/sample.wav`), so each is removed in full.
- Renormalized the 21 affected assets through the LFS clean filter.
- Root `.gitattributes` untouched. `.github/` untouched. `agno`/`mastra`
untouched.

Deleting an override necessarily un-exempts that integration's
`demo-audio/sample.wav` too, so the 5 raw `.wav` files are converted in
the same pass — leaving them raw under an inherited `filter=lfs` would
produce a permanently dirty checkout.

## Byte-identity

Storage form changes, content does not. Every asset's sha256 **already
equals** the LFS OID the pointer-mode integrations reference, so each
renormalized blob is bit-for-bit the pointer blob already committed on
`main`:

```
sample.png  10083 B  oid 01aa5681de99461247543e9215c1e4da3242e26b2bee11593fcdbe209672d973
sample.pdf   2486 B  oid 3da2afae36a1a81fd2c02f15e54bfc38b6c22e41655c31a5b54ff1e0e3daab41
sample.wav  87078 B  oid bd4aa7b049f1c3e324dfd15af4068d7f8fbf2eae1dd044df270dddc5f38a5c57
```

Verified per file (21/21): worktree sha256 before == worktree sha256
after == the pointer's OID, and each staged blob hash equals the
canonical pointer blob (`a780acbd` png / `6446acbb` pdf / `f156ab54`
wav) already on `main`.

**No new LFS objects are introduced at all**, so a dangling pointer is
impossible by construction. Confirmed anyway:

- LFS batch API returns `download` actions for all three OIDs.
- All three objects downloaded and confirmed to hash to their OID.
- `git lfs push --dry-run origin <branch>` reports nothing to upload.
- `git lfs fsck --pointers HEAD` reports no new violations (only the two
pre-existing `examples/teams/appPackage/*.png` ones, also present on
`main`).

## On the CI signal, and the real proof

Being precise about what CI here does and does not prove:

- **`showcase_build.yml`** (deploy, `lfs: true`) is the job that
determines whether the deployed image gets real assets. It runs **only
on push to `main`**, not on PRs. It is unchanged by this PR.
- **`showcase_build_check.yml`** (the PR job) checks out with `lfs: ${{
matrix.service.lfs }}`, which is `false` for all 8 converted
integrations. Its images therefore contain pointer stubs. It is
build-only (`push: false`, no smoke test, no asset assertion), so it
will pass — but **a green `build_check` here does not prove real assets
ship.** I am not resting the case on it.

I did **not** rely on a local Docker build either — it smudges LFS from
the local checkout and would prove nothing.

### The real proof: prod already serves pointer-mode assets correctly

Instead of a weak build signal, here is what the **live deployed prod
images** serve today. Pointer-mode integrations are unchanged by this PR
and demonstrate that the pointer storage form already ships real bytes
through the `lfs: true` deploy checkout — for all three asset types:

```
--- .png : POINTER-mode on main (unchanged) ---
ag2                   200  10083  01aa5681de  REAL-exact-match
spring-ai             200  10083  01aa5681de  REAL-exact-match
llamaindex            200  10083  01aa5681de  REAL-exact-match
pydantic-ai           200  10083  01aa5681de  REAL-exact-match
langroid              200  10083  01aa5681de  REAL-exact-match
--- .png : RAW-mode on main (converted by this PR) ---
langgraph-python      200  10083  01aa5681de  REAL-exact-match
strands               200  10083  01aa5681de  REAL-exact-match
google-adk            200  10083  01aa5681de  REAL-exact-match
ms-agent-python       200  10083  01aa5681de  REAL-exact-match
langgraph-typescript  200  10083  01aa5681de  REAL-exact-match

--- .wav : ALREADY pointer-mode on main (precedent for the 5 wav conversions) ---
google-adk            200  87078  bd4aa7b049  REAL-exact-match
ms-agent-python       200  87078  bd4aa7b049  REAL-exact-match
ms-agent-dotnet       200  87078  bd4aa7b049  REAL-exact-match
ag2                   200  87078  bd4aa7b049  REAL-exact-match
spring-ai             200  87078  bd4aa7b049  REAL-exact-match
--- .wav : RAW on main (converted by this PR) ---
langgraph-python      200  87078  bd4aa7b049  REAL-exact-match
strands               200  87078  bd4aa7b049  REAL-exact-match
langgraph-typescript  200  87078  bd4aa7b049  REAL-exact-match
langgraph-fastapi     200  87078  bd4aa7b049  REAL-exact-match
strands-typescript    200  87078  bd4aa7b049  REAL-exact-match

--- .pdf : pointer-mode vs raw-mode ---
ag2            (ptr)  200   2486  3da2afae36  REAL-exact-match
spring-ai      (ptr)  200   2486  3da2afae36  REAL-exact-match
langgraph-python(raw) 200   2486  3da2afae36  REAL-exact-match
strands        (raw)  200   2486  3da2afae36  REAL-exact-match
```

Both storage forms already serve **byte-identical** content in prod
(sha256 equals the LFS OID in every case), including `.wav` where 3
integrations are already pointer-mode. Converting raw to pointer
therefore cannot change what prod serves: both forms resolve to the same
bytes through the same deploy checkout. That is the safety argument —
verified, not asserted.

## Sequencing

- **#6160** (`ci(showcase): fetch Git LFS in every workflow that serves
integration assets`) should merge **first**. It makes
`showcase_build_check.yml` fetch LFS and removes the dead per-slot `lfs`
flag; after it lands, `build_check` on this kind of change becomes a
real signal instead of a weak one. No file overlap with this PR (it
touches only `.github/workflows/`), so no merge conflict either way.
- **#6157** (agno/mastra assets) is independent — disjoint paths, any
order, no conflict.

This PR assumes the deploy build's hardcoded `lfs: true` already on
`main` and complements #6160's non-deploy workflow coverage.

## Cell impact

**Changed-cell set: empty (expected and predicted).** This changes
storage form only. The prod probes above show both storage forms already
serve byte-identical content, so the multimodal cells for the 8
converted integrations have no mechanism by which to change state. The
cells cannot be re-measured against this branch pre-merge (the deploy
build runs only on push to `main`), so this is a prediction grounded in
the byte-identity and prod-parity evidence rather than a post-deploy
dashboard reading — worth a dashboard confirmation after merge.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01PYdjeveT8Xof9TyHWMLoJr
2026-07-25 09:59:39 -07:00
Jordan Ritter 51cc63d974 chore(showcase): store demo assets uniformly as Git LFS pointers
Demo assets under showcase/integrations/*/public/{demo-files,demo-audio}/
were stored two different ways. Ten integrations committed them as LFS
pointers (the root .gitattributes convention); eight carved themselves out
with a per-integration .gitattributes that re-declared the same paths
`-filter -diff -merge`, committing raw binaries instead.

Those carve-outs were added when the image build did not fetch LFS, so a
pointer stub shipped into the image and the multimodal sample-attachment
magic-bytes guard rejected it. That premise no longer holds: the deploy
build's Checkout step hardcodes `lfs: true` (7bde1eef3a), so every
integration image now gets real binaries regardless of storage form. The
overrides are dead weight that only buys divergence.

Delete all eight override files and renormalize the 21 affected assets
through the LFS clean filter. Each override contained nothing but demo-asset
exemptions, so each is removed in full; the root .gitattributes is untouched.

Storage form changes, content does not. Every asset's sha256 already equals
the LFS OID the pointer-mode integrations reference, so each renormalized
blob is bit-for-bit the pointer blob already committed on main -- no new LFS
objects are introduced and no pointer can dangle:

  sample.png  10083 B  oid 01aa5681de99461247543e9215c1e4da3242e26b2bee11593fcdbe209672d973
  sample.pdf   2486 B  oid 3da2afae36a1a81fd2c02f15e54bfc38b6c22e41655c31a5b54ff1e0e3daab41
  sample.wav  87078 B  oid bd4aa7b049f1c3e324dfd15af4068d7f8fbf2eae1dd044df270dddc5f38a5c57

All three OIDs return download actions from the LFS batch API and were
downloaded and confirmed to hash to their OID.
2026-07-24 16:33:14 -07:00