Commit Graph

31 Commits

Author SHA1 Message Date
Jordan Ritter ca2bc2c152 fix(showcase): specialize strands + spring-ai declarative-hashbrown/json-render backends (+ fixtures, manifests, controller guards) 2026-06-07 13:53:28 -07:00
Jordan Ritter 182e3e2fd9 fix(showcase): add follow-up-turn aimock fixtures for multi-pill D6 sessions 2026-06-06 02:30:17 -07:00
Ran Shem Tov 02b11c985d feat(showcase): drive dynamic A2UI via CopilotKitMiddleware (langgraph)
The declarative-gen-ui demo across the three langgraph integrations now
relies on the middleware to inject and execute generate_a2ui — the agents
collapse to create_agent + CopilotKitMiddleware with no hand-rolled tool.
Adds render_a2ui fixtures for the new tool path and pins the integrations
to the A2UI alpha SDKs (copilotkit 0.1.94a1, @copilotkit/sdk-js 1.59.3-alpha.1).
2026-06-04 18:36:12 +02:00
Jordan Ritter fe95444314 fix(showcase/aimock): reorder gen-ui-interrupt d6 fixture legs fleet-wide (toolName before toolCallId)
Audit follow-up to #5232, which fixed the gen-ui-interrupt d6 fixture
leg mis-order for langgraph-python + langgraph-typescript. aimock's
matchFixture is first-match-wins in array order and the loader dedups
on userMessage, so a toolCallId resume-leg placed BEFORE its
toolName:schedule_meeting first-leg shadows the tool-emitting leg on
turn-2 — the interrupt never fires and time-picker-card never mounts.

Fleet audit of every showcase/aimock/d6/*/gen-ui-interrupt.json found
ag2 as the only remaining mis-ordered fixture (both toolCallId
resume-legs preceded their toolName first-legs). Reorder ag2 so each
pill's toolName first-leg precedes its toolCallId resume-leg, matching
the proven-correct mastra/pydantic-ai/langgraph pattern. Reorder only;
response payloads and match keys are unchanged.

All other gen-ui-interrupt-supporting integrations (built-in-agent,
claude-sdk-*, langgraph-fastapi, langroid, ms-agent-dotnet/python,
pydantic-ai, spring-ai, strands) were already correctly ordered.
2026-06-04 07:36:43 -07:00
Jordan Ritter 798ed56b03 fix(showcase/aimock): reorder gen-ui-interrupt d6 fixture legs so turn-2 emits the tool call (langgraph-python+typescript)
The gen-ui-interrupt d6 probe drives two interrupt turns in one thread.
Turn-2 failed because the langgraph-python/typescript fixtures ordered each
pill's toolCallId resume-leg BEFORE its toolName first-leg. aimock dedups on
userMessage (first-match-wins), so the text-only resume-leg shadowed the
tool-emitting first-leg on the duplicate userMessage — the agent emitted a
plain text reply with no schedule_meeting tool call, the interrupt never
fired, and the time-picker-card never mounted.

Reorder each pill so the toolName:schedule_meeting first-leg precedes its
toolCallId resume-leg, matching the passing mastra/pydantic-ai ordering. The
resume-leg still only matches when the last message is its tool result
(toolCallId guard), so confirmation still works.
2026-06-04 01:53:16 -07:00
Jordan Ritter b6b35bd406 fix(showcase): restore sandboxed-UI jsFunctions in gen-ui-open-advanced fixtures 2026-06-02 17:57:47 -07:00
Jordan Ritter 3ca0a26f11 fix(showcase): gate LGP tool-rendering AAPL + Find-flights fixtures on toolName 2026-06-02 17:57:47 -07:00
Jordan Ritter 1a7c466e27 fix(showcase/aimock): add missing gen-ui-headless-complete D6 fixtures
The gen-ui-headless-complete probe
(showcase/harness/src/probes/scripts/d5-gen-ui-headless-complete.ts)
references fixtureFile "gen-ui-headless-complete.json", but that file
existed in no D6 slug, so the probe's first leg 503'd under strict.

Add gen-ui-headless-complete.json for all 18 D6 slugs, modeled on the
green langgraph-python headless-complete.json pattern: the four gen-UI
pills (weather/stock/highlight/revenue) with narration (toolCallId)
fixtures FIRST and toolcall (userMessage+context) fixtures AFTER, no
turnIndex gate, so every one of the probe's four sequential turns in one
chat thread matches regardless of prior assistant/tool history. Interrupt
fixtures are intentionally omitted (interrupt-headless is a separate
cluster item).

These 8 fixtures per slug share match keys with the pre-existing
headless-complete.json for the same context (the demos share pills and
are disambiguated at runtime by probe path), which raises the
aimock-fixtures collision-detection exact-duplicate count by 46. Bump
KNOWN_DUPLICATE_CEILING 230 -> 276 to match, consistent with how prior
per-integration feature fixtures bumped the baseline; the substring-shadow
ceiling is unchanged (no new shadows).

Validated: all 732 aimock-fixtures schema/collision tests pass, and a
local aimock --strict run returns 200 for each of the four pills across
turns (langgraph-python + pydantic-ai spearheads, plus spring-ai).
2026-06-01 00:58:12 -07:00
Jordan Ritter 5d0f9ecc20 fix(showcase/aimock): drop stale turnIndex:0 gate from multi-turn D6 pill fixtures
D6 drives each feature's pills as sequential turns in one chat thread.
aimock matches turnIndex against the count of assistant messages in
history, so a "turnIndex": 0 gate on a per-pill tool-call leg only
matches the FIRST pill — pills 2+ (turnIndex 1/2/3) match no fixture,
503 under strict, and the harness reports "timeout: assistant did not
respond".

Remove the turnIndex:0 gate from the multi-turn tool-call legs of
frontend-tools (sunset/forest/cosmic) and tool-rendering-reasoning-chain
(the three chained pills: Compare AAPL/MSFT, compare-to-smaller dice,
weather-there flights) so they match on their distinct userMessage
substrings + context, mirroring the already-green langgraph-python
fixtures. Each pill's substring is unambiguous and its toolCallId-keyed
follow-up fixture is ordered before the de-gated first leg (first-match
wins), so follow-up narration still wins and there is no cross-matching.

The reasoning-chain "Find flights from SFO to JFK." pill keeps its
turnIndex:0 gate to match the green langgraph-python reference exactly.

Proven locally with aimock --strict: multi-turn requests for these pills
returned 503 with the gate (negative control) and now return 200 across
all turns with the correct fixture/narration.
2026-06-01 00:57:59 -07:00
Jordan Ritter dfeb06121b fix(showcase/d6): drop turnIndex:0 from multi-turn goldfish agentic-chat fixtures
The aimock matcher gates turnIndex against the request's assistant-message
count (assistantCount !== turnIndex → skip). The agno, spring-ai, and
langgraph-fastapi agentic-chat fixtures baked turnIndex:0 on all three
goldfish conversation turns. Turn 2+ carries >=1 assistant message, so
turnIndex:0 could never match those turns — aimock returned no-match and the
request fell through to proxy (503), failing the d6 cell.

Mirror the canonical clean pattern used by the other 15 frameworks
(langgraph-python et al.): omit turnIndex on the multi-turn conversation
turns and disambiguate purely by userMessage. Single-turn fixtures in the
same files keep turnIndex:0 (unchanged). Verified locally against the aimock
matcher: turn-2 goes 404→200 for all three frameworks, turn-1 unregressed.
2026-05-31 20:45:13 -07:00
Jordan Ritter bd4424bdf6 Merge branch 'blitz/d6-parity/maf-dotnet' into blitz/d6-parity/integration 2026-05-30 20:17:57 -07:00
Jordan Ritter 811afa9025 fix(showcase): ms-agent-dotnet D6 residuals (toolCallId strip, fixtures, hitl pages)
Address the four residual D6 failures on the ms-agent-dotnet integration
(baseline was 177/5).

route.ts toolCallId strip completeness
- Extend stripReplaySafeToolCallIdsFromMessage to also clean the
  snake_case `tool_calls[].id` array and any nested OpenAI-style
  `function.tool_call_id`. AG-UI canonical uses `toolCalls`, but some
  runtime / message-converter paths emit the OpenAI shape and the
  replay-safe `__ck_run_<uuid>` suffix was leaking through to aimock on
  those paths. Apply the same coverage in applyToolResultDecisionSuffix
  so decision-suffix routing (`__approved` / `__rejected` /
  `__cancelled`) lands on every tool-call shape.
- Add a universal strip middleware in createAgent itself so EVERY
  registered agent (not only the replay-safe ones) clears the suffix
  before the request reaches the backend / aimock. Decision suffixing
  remains scoped to createReplaySafeAgent because the suffix is
  non-idempotent and the inner middleware re-runs the same logic.
- Drop the now-redundant strip calls inside createGenUiAgent,
  createReadonlyContextAgent, createSharedStateReadWriteAgent, and
  createReasoningAgent — the outer createAgent middleware already
  canonicalised inbound messages by the time these run.

chat-slots fixture
- Mirror LGP's turnIndex:1 'Give me a fun fact' entry so the
  chat-slots e2e's second-turn assistant slot has a deterministic
  reply (the bare 'Give me a fun fact' fixture in headless-simple.json
  is turnIndex:0-gated and was never matching chat-slots' turn 1).

HITL reject + interrupt-headless cancel branches
- Add a reject-branch fixture in render-a2ui.json keyed on
  toolCallId `call_d5_generate_steps_001__rejected` so the hitl.spec
  reject-flow's 'will not execute the Mars trip plan' assertion lands
  the right narration. Keep the legacy hasToolResult:true fixture as a
  fallback for paths that don't apply a decision suffix. Add an
  `__approved` variant for symmetry.
- Add cancel-branch fixtures in interrupt-headless.json for both the
  sales-intro and 1:1-with-alice pills, keyed on `__cancelled`
  toolCallIds, returning the Denied/not-booked narration the
  interrupt-headless cancel spec expects.

hitl-in-app + hitl-in-chat demo pages already mirror LGP (only an extra
README.md per directory in this integration), so no page changes were
needed.
2026-05-30 20:17:08 -07:00
Jordan Ritter 4edcd2f129 fix(showcase): ms-agent-python D6 residuals (toolCallId strip, fixtures, hitl pages, multimodal)
Port the replay-safe toolCallId stripping middleware from the ms-agent-dotnet
sibling route.ts so HITL/interrupt demos match toolCallId-keyed aimock
fixtures across the 2nd-turn request. The strip walks every inbound message
(role=tool, role=assistant.toolCalls[].id, toolCallId, tool_call_id) before
the AG-UI HttpAgent forwards them onto the FastAPI backend, and the outbound
event stream rewrites toolCallId on TOOL_CALL_* events to embed the
deterministic per-run suffix so the next turn's fixture matcher still
keys on the original (suffix-stripped) id.

Wraps the human_in_the_loop, interrupt-adapted, hitl-in-app, and
hitl-in-chat agents with the new replay-safe middleware. Adds gen-ui-agent
set_steps state-snapshot synthesis, readonly-state-agent-context system
message injection, and shared-state-read-write preference-as-system
injection — all ported verbatim from the dotnet sibling so the per-agent
shaping behavior is identical across the MAF runtimes.

chat-slots.json: add the turnIndex:1 "Give me a fun fact" fixture mirrored
from langgraph-python/chat-slots.json so the chat-slots.spec.ts second-
turn assertion ("second assistant turn is also wrapped in the custom
slot") gets a deterministic reply instead of falling through to headless-
simple's fun-fact fixture or the live proxy.

hitl-in-app and hitl-in-chat demo pages were already mirrored from LGP
(only a minor consumerAgentId addition on the in-app suggestions module).
No frontend page changes needed.

The B2 header-forwarding conveyance in src/agents/_header_forwarding.py
is untouched and verified intact (httpx hook + Starlette HTTP middleware
plus ContextVar bridge).

Multimodal diagnosis: the 5-test multimodal.spec.ts timeout is NOT a
routing/wiring gap. The agent_server.py mounts /multimodal FIRST (before
the catch-all "/"), the dedicated Next.js route /api/copilotkit-multimodal
registers the HttpAgent under "multimodal-demo", and the page wires
runtimeUrl + agent correctly. The multimodal fixture's userMessage
"describe the sample image" is the same stale phrasing LGP uses (which
passes at 185/0/2 — the test asserts a /image/i regex on the assistant
transcript, so fallthrough to the proxy still satisfies it). The
remaining suspects are (a) agent_framework_ag_ui's AG-UI -> AF adapter
mishandling inbound `binary` content parts, (b) the dual chat_client
init pattern (multimodal_chat_client built after the global httpx hook
is installed — the hook is idempotent so this should be safe), or
(c) >30s real-OpenAI vision latency under D6 record-replay. Capture
agent_server stderr during a single multimodal run to confirm.
2026-05-30 20:14:36 -07:00
Jordan Ritter 0143b02ae3 WIP: D6 rollout — foundation + per-integration fixtures + conveyance fixes (#5109)
## Status

WIP / not ready to merge. Preserves in-flight D6 work so it isn't lost
mid-rollout. LGP is fixture-complete; other integrations are
mid-rollout.

### Latest banked work

- **langgraph-python — 185 / 0 / 2** (green). Achieved by narrowing a d4
chat matcher that was shadowing the d6 beautiful-chat search_flights
fixture (load order is shared -> d4 -> d6, first-match-wins) plus
refreshing the d6 tool-rendering and tool-rendering-custom-catchall AAPL
fixtures (turnIndex:0 -> hasToolResult:false so the first leg fires in
multi-pill threads). 2 skips are the by-design mcp-apps iframe gap.
- **langgraph-typescript — 185 / 0 / 2** (green). Mirrored the LGP d4
narrowing on the LGT side (3 matchers) and across the
d6/langgraph-typescript suite: replaced fragile turnIndex:0 gates with
hasToolResult:false, fixed em-dash escaping that broke literal matches
in multi-pill threads, and added jsFunctions payloads to the three
sandboxed-ui fixtures (`_from-feature-parity`, `headless-complete`,
`gen-ui-open-advanced`). The final fix removed the chain-tools
`hasToolResult` match gate (which checked the whole thread and made the
chain pill fall through to the broad weather matcher mid-thread),
mirroring LGP.
- **google-adk: 174/7/2 (was 134/52/5)** — conveyance + test-parity +
fixtures + pill-wiring rebuild; 7 residual (6 default-catchall framework
default-renderer testid version question, 1 beautiful-chat fixture).
- **Pill-parity staged across 13 integrations** (ag2, agno, mastra,
pydantic-ai, claude-sdk-python, claude-sdk-typescript, llamaindex,
langroid, strands, spring-ai, built-in-agent, crewai-crews,
langgraph-fastapi). The canonical LGP suggestion pill set is now
mirrored as `src/app/demos/*/suggestions.ts` files in each integration,
with targeted edits to existing `open-gen-ui-advanced` and
`byoc-hashbrown` files. **These new files are currently UNWIRED** — each
integration's `page.tsx` still defines its pill list inline via
`useConfigureSuggestions`. Banked so the canonical source survives; a
follow-up will rewire `page.tsx` to import from `suggestions.ts` and
drop the inline copies.
- **Fleet test-parity sweep**: 576 e2e specs across 15 integrations
aligned to LGP canonical (SHA-verified); 2 orphan specs removed.
- ms-agent-dotnet 177/5/7, ms-agent-python 174/11/2 (post
fixture-mirror); default-catchall green (page-level renderer, not
react-core-gated).

## Scope

### Conveyance (foundation)

Inbound `x-aimock-context` (and friends) must ride along on outbound LLM
HTTP calls so aimock fixture matching sees the inflight test's context.
Without this the call lands on the default project's aimock and silently
picks the wrong fixture. New per-integration
`_header_forwarding.{py,ts}` shim plus matching `agent_server` / route /
factory wiring covers: ag2, agno, built-in-agent, claude-sdk-python,
claude-sdk-typescript, crewai-crews, google-adk, langgraph-fastapi,
langgraph-python, langgraph-typescript, langroid, llamaindex, mastra,
ms-agent-python, pydantic-ai, strands.

For ADK/Gemini the global httpx hook is installed BEFORE any `agents.*`
import (google-genai constructs its client at module-import time).

### langgraph-python — 185 / 0 / 2

Fixture-complete via the conveyance shim + refreshed d6/langgraph-python
fixtures + copilotkit 0.1.93 bump + the latest d4-matcher-narrowing fix
(see banked work above).

### Per-integration fixtures

Mid-rollout snapshot of d6 fixtures across the cohort plus narrowing of
`aimock/shared/common.json`'s generic 'hello' fixture to 'hello world'
so it no longer shadows D6 pills whose prompts contain 'hello' as a
substring.

### Harness `--isolate` patch

`scripts/cli/_common.sh apply_isolation` now rewrites compose-file
relative paths to absolute (build/context/dockerfile/volumes/env_file),
enforces the docker compose `[a-z0-9_-]` project-name rule, and exports
`SHOWCASE_COMPOSE_FILE` / `SHOWCASE_INFRA_PORT_OFFSET` plus offset host
URLs. The TS harness CLI (`aimock-rebuild` / `config` / `doctor` /
`lifecycle`) honors the new env so concurrent isolated stacks stop
reporting each other's services as healthy.

## Lockfile decision flagged

`showcase/integrations/langgraph-python/pnpm-lock.yaml` was deleted in
this branch. Decision: keep the deletion. Rationale:

- 03bed3b76 (fix(showcase): regenerate 18 lockfiles in isolation; switch
to npm ci) migrated all showcase integrations off pnpm onto npm ci.
- The integration's Dockerfile uses `npm ci --legacy-peer-deps`.
- Every sibling integration committed only `package-lock.json` after
03bed3b76.
- The orphan pnpm-lock.yaml only risks tooling drift.

If anyone wants it restored: `git checkout origin/main --
showcase/integrations/langgraph-python/pnpm-lock.yaml`.

## Commits

- feat(showcase): D6 conveyance — forward x-aimock-context headers to
LLM clients
- feat(showcase/langgraph-python): D6 conveyance shim + copilotkit
0.1.93 bump
- feat(showcase/langgraph-typescript): D6 conveyance — propagate request
headers into ChatOpenAI
- feat(showcase/built-in-agent): D6 conveyance — header-forwarding shim
+ factory wiring
- feat(showcase/harness): support concurrent --isolate runs
- test(showcase): D6 langgraph-python fixtures — drive to 180/5/2
- test(showcase): D6 per-integration aimock fixtures + shared narrowing
- docs(showcase): GOTCHAS entry for D6 conveyance + --isolate notes
- feat(showcase): D6 conveyance — wire header-forwarding shims into
remaining entrypoints
- fix(showcase): unblock LGP D6 beautiful-chat + custom-catchall via d4
matcher narrowing
- fix(showcase): narrow LGT D6 d4 shadows + wire sandboxed-ui
jsFunctions
- chore(showcase): copy LGP canonical suggestion pills into 13
integrations
- fix(showcase): forward x-aimock-context per-request in google-adk
routes
- test(showcase): align google-adk e2e specs to langgraph-python
canonical
- fix(showcase): align google-adk D6 fixtures to LGP contract
- fix(showcase): wire google-adk default-catchall to shared 4-pill
suggestions
2026-05-30 16:32:19 -07:00
Jordan Ritter 430dbecc09 feat(showcase): apply D6 parity sweep to pydantic-ai (mirror LGP demos + specs + fixtures)
pydantic-ai had never received the fleet D6 parity sweep — its e2e specs and demo pages
were a pre-sweep, integration-specific set (only 3/26 suggestion files; missing canonical
demos; non-canonical byoc-*/agentic-chat-reasoning/reasoning-default-render variants).
Sitting at 69/109/2.

This change mirrors langgraph-python's canonical frontend (demos + specs + aimock
fixtures) into pydantic-ai, preserving pydantic-ai's Python backend untouched. The
per-demo agent.py files that pydantic-ai carries inside demo directories are preserved.

Changes:
- tests/e2e/: rsync LGP canonical 37-spec set over pydantic-ai (byte-identical). Removes
  non-canonical byoc-hashbrown.spec.ts, byoc-json-render.spec.ts, shared-state-write.spec.ts.
  Adds canonical declarative-hashbrown.spec.ts, declarative-json-render.spec.ts,
  reasoning-custom.spec.ts, reasoning-default.spec.ts.
- src/app/demos/: rsync LGP demos over pydantic-ai. Removes non-canonical demos
  (byoc-hashbrown, byoc-json-render, agentic-chat-reasoning, reasoning-default-render,
  shared-state-write). Adds canonical demos (declarative-hashbrown, declarative-json-render,
  reasoning-default, reasoning-custom) and the _shared/ helpers + demos/layout.tsx
  pydantic-ai was missing. Restores pydantic-ai-specific agent.py files into the 9 demo
  dirs that survived the mirror.
- src/app/demos/frontend-tools/page.tsx: patched agent slug from "frontend_tools" (LGP)
  to "frontend-tools" (matches pydantic-ai's main route.ts registry).
- src/app/api/copilotkit-byoc-{hashbrown,json-render}/ renamed to copilotkit-declarative-*
  to match the canonical frontend wiring. Internals still use HttpAgent against the
  pydantic backend's /byoc_hashbrown/ + /byoc_json_render/ mounts (Python backend
  untouched per scope). copilotkit-declarative-hashbrown/route.ts updates the registered
  agent slug from "byoc-hashbrown-demo" to "declarative-hashbrown-demo" to match the
  canonical demo. copilotkit-declarative-json-render/route.ts updates only the endpoint
  path string (the agent slug "byoc_json_render" is the canonical LGP convention).
- src/app/api/copilotkit/route.ts: renamed reasoning agent registrations from
  agentic-chat-reasoning + reasoning-default-render to reasoning-custom + reasoning-default
  to match canonical demo slugs. Both still proxy to the same /reasoning/ backend mount.
- manifest.yaml: features[] + demos[] updated to reflect the canonical demo set
  (byoc-* + agentic-chat-reasoning + reasoning-default-render removed; declarative-* +
  reasoning-default + reasoning-custom added).
- aimock/d6/pydantic-ai/: added gen-ui-custom.json (mirrored from LGP with
  context-swap + copiedFrom marker, per established fixture convention). Removed
  orphan gen-ui-open-advanced.json (no LGP counterpart in the canonical set).

The Python backend (agent.py / src/agent_server.py / src/agents/) is unchanged.
Some pydantic-ai backend mounts continue to exist that the mirrored frontend no longer
references (e.g. /reasoning/ remains, the deleted demos' agent slugs are still
registered in route.ts but harmlessly orphaned) — these are intentional carry-overs
to avoid touching Python backend code per scope.
2026-05-30 16:05:40 -07:00
Jordan Ritter 47d3b594ce fix(showcase): align ms-agent-python D6 fixtures to LGP (152/33 → 174/11)
Same pattern as ms-agent-dotnet: full d4 chat.json rewrite + d6 mirrors. Dropped stale
turnIndex gates and broad shadow matchers. Default-catchall now green via page-level
shadcn-catchall-renderer. 11 residual: declarative-gen-ui charts, multimodal conveyance,
tool-rendering, reasoning-chain.
2026-05-30 09:53:35 -07:00
Jordan Ritter afa4cf0cf2 fix(showcase): align ms-agent-dotnet D6 fixtures to LGP (155/27 → 177/5)
Full d4 chat.json rewrite + d6 mirrors. Dropped stale turnIndex gates and broad shadow
matchers that no longer reflect LGP-canonical conveyance. Default-catchall now green via
page-level shadcn-catchall-renderer (not react-core-gated). 5 residual: chat-slots, hitl,
interrupt-headless, readonly-state.
2026-05-30 09:53:25 -07:00
Jordan Ritter a49ba1ef74 fix(showcase): align google-adk D6 fixtures to LGP contract
Narrowed d4 shadow matchers, removed whole-thread gates, render_a2ui inner tool
name, root chunkSize for emoji surrogate pairs, sentinel-disabled the duplicate
open-gen-ui-advanced matcher.
2026-05-29 23:59:41 -07:00
Jordan Ritter 9baafac952 fix(showcase): drop chain-tools hasToolResult gate to reach LGT D6 185/0/2
The whole-thread `hasToolResult: false` gate on the Chain-tools first-turn fixture caused
the chain pill to fall through to the broad "weather in Tokyo" matcher mid-thread.
Removing the gate (mirroring LGP) yields LGT 185/0/2.
2026-05-29 21:22:56 -07:00
Jordan Ritter e8f38c6129 fix(showcase): narrow LGT D6 d4 shadows + wire sandboxed-ui jsFunctions
Mirrors the LGP first-match-wins fix on the langgraph-typescript side. Three broad
matchers in d4/langgraph-typescript/chat.json were shadowing d6 fixtures; narrowed
them so the d6 pills win.

Across the d6/langgraph-typescript suite: dropped fragile turnIndex:0 gates in favour
of hasToolResult:false for first-leg tool emissions, fixed em-dash escaping that broke
literal string matches in multi-pill threads, and added jsFunctions payloads to the
three sandboxed-ui fixtures (_from-feature-parity, headless-complete,
gen-ui-open-advanced) so the sandbox renderer has executable handlers.

LGT D6 now passes 184/1/2 locally. Residual 1 fail is the custom-catchall multi-pill
follow-up case; tracked separately.
2026-05-29 21:08:05 -07:00
Jordan Ritter 508bea0e88 fix(showcase): unblock LGP D6 beautiful-chat + custom-catchall via d4 matcher narrowing
Aimock fixture load order is shared -> d4 -> d6, and matching is first-match-wins. A
broad d4 langgraph-python chat fixture for search_flights was shadowing the d6
beautiful-chat fixture and preventing the intended response from firing. Neutered the
d4 matcher to a non-matching sentinel so the d6 fixture wins.

Also refreshed d6/tool-rendering-custom-catchall.json and d6/tool-rendering.json:
replaced fragile turnIndex:0 gates with hasToolResult:false so the AAPL first-leg
fixture fires correctly when D5 probes run a prior 'weather in Tokyo' pill in the
same thread (multi-pill turnIndex>=2).

LGP D6 now passes 185/0/2 locally (2 skips are the by-design mcp-apps iframe gap).
2026-05-29 21:08:05 -07:00
Jordan Ritter c414fbad23 test(showcase): D6 per-integration aimock fixtures + shared narrowing
Refresh d6 fixtures across the rollout cohort: ag2, built-in-agent,
claude-sdk-{python,typescript}, crewai-crews, google-adk, langgraph-fastapi,
langgraph-typescript, langroid, llamaindex, mastra, ms-agent-{dotnet,python},
pydantic-ai, strands. Companion d4/{langgraph-typescript,mastra,ms-agent-dotnet}
chat.json refreshes. Add the missing ms-agent-python/gen-ui-custom.json
to bring the integration up to the standard pill set.

Also narrow aimock/shared/common.json's generic 'hello' fixture to
'hello world' so it no longer shadows D6 pills whose prompts contain
'hello' as a substring (e.g. langgraph-python headless-simple sends
'Say hello in one short sentence.'). 'hello world' is unused by any
current demo pill, so the fixture remains a manual-typing fallback
without poisoning fixture matching.

This is a mid-rollout snapshot — fixture coverage is uneven across
integrations and rides alongside the conveyance shims landed earlier
in this branch.
2026-05-29 16:16:12 -07:00
Jordan Ritter 5f21584b3a test(showcase): D6 langgraph-python fixtures — drive to 180/5/2
Refresh d6/langgraph-python fixtures (beautiful-chat, chat-slots,
frontend-tools-async, gen-ui-{agent,custom,declarative,interrupt,open},
headless-complete, hitl-in-app, interrupt-headless, shared-state-streaming,
tool-rendering, tool-rendering-reasoning-chain, _from-feature-parity) plus
d4/langgraph-python/chat.json. Combined with the conveyance shim landed
in this branch, LGP is fixture-complete at 180 pass / 5 fail / 2 skip.

The residual 5 fails trace to react-core/v2's custom-wildcard-renderer
bug surfaced through gen-ui-interrupt (shared useInterrupt hook misbehaves
on the second interrupt; cross-integration D6 blocker, not a fixture or
conveyance defect). The 2 skips are the by-design mcp-apps iframe gap.
2026-05-29 16:15:57 -07:00
Martha Schumann 24d93b52ad fix(react-core): preserve generated thread tool followups 2026-05-27 10:27:29 -07:00
Jordan Ritter 43ddf0372c chore(showcase-agno): mark gen-ui-interrupt + interrupt-headless as not_supported
agno uses useFrontendTool (Strategy B) — async Promise handler in the
frontend — rather than LangGraph's native interrupt() primitive. The D5
probe asserts via useInterrupt hook which routes through the LangGraph
interrupt event path. agno doesn't emit those events; the D5 probe
fundamentally cannot pass against agno's HITL architecture without
per-integration probe-code divergence (which would violate the D6
apples-to-apples invariant).

Excluding these two features at the manifest level (matching google-adk
precedent) lets the D6 probe skip them cleanly. Removes the
corresponding D6 aimock fixtures since they're no longer reachable.

If agno gains LangGraph-style interrupt() support, or if aimock gains
AG-UI-event fixture authoring, this exclusion can be reverted.
2026-05-26 15:20:58 -07:00
Jordan Ritter 03ce685ae6 feat(showcase-aimock): D6 fixtures for ag2, agno, crewai-crews, langroid,
llamaindex, mastra, pydantic-ai, spring-ai, strands, built-in-agent

Adds per-integration D6 fixtures across 10 additional showcase
integrations, completing coverage for all 18 integrations.
2026-05-26 14:02:18 -07:00
Jordan Ritter 89d8440e7d feat(showcase-aimock): D6 fixtures for google-adk
Adds D6 fixtures for google-adk; skips gen-ui-interrupt and
interrupt-headless per the integration's not_supported_features.
2026-05-26 14:02:17 -07:00
Jordan Ritter abe69c3d56 feat(showcase-aimock): D6 fixtures for Microsoft Agent Framework integrations
Adds per-integration D6 fixtures for ms-agent-python and ms-agent-dotnet.
2026-05-26 14:02:17 -07:00
Jordan Ritter 4df07816b8 feat(showcase-aimock): D6 fixtures for Claude SDK integrations
Adds per-integration D6 fixtures for claude-sdk-python and
claude-sdk-typescript with Anthropic-style request/response shapes.
2026-05-26 14:02:17 -07:00
Jordan Ritter 69c964aca2 feat(showcase-aimock): D6 fixtures for LangGraph-based integrations
Adds per-integration D6 fixtures for langgraph-python, langgraph-typescript,
and langgraph-fastapi. Each fixture is keyed by match.context for cross-
integration isolation and uses hasToolResult:false on toolCall responses
to prevent re-match loops.
2026-05-26 14:02:17 -07:00
Jordan Ritter 1e66a5f8d2 feat(showcase-aimock): per-framework D4/D6 fixture reorg into d4/d6/shared dirs
Move monolithic d5-all.json + feature-parity.json + smoke.json into
per-integration directories under d4/<slug>/, d6/<slug>/, and shared/.
Every fixture file is now context-scoped to enable server-side aimock
routing via match.context. Migrated 12 HITL fixtures from main's
d5-all.json additions into shared/_migrated-from-d5-all-hitl.json
for follow-up distribution into per-integration files.
2026-05-26 11:25:30 -07:00