Commit Graph

42 Commits

Author SHA1 Message Date
Alem Tuzlak a3586fb62a fix(showcase): revert reasoning field on aimock d5 reasoning fixture
PR #4579 added a `reasoning` field to the "show your reasoning step by
step" fixture so aimock would emit response.reasoning_summary_* deltas
for the OpenAI Responses API path. Side effect: aimock's Chat
Completions handler also emits non-standard `reasoning_content` deltas
(DeepSeek/Qwen-style) ahead of the role/content chunks. Many
integrations' OpenAI client adapters don't expect those deltas and
either hang or fail to parse the stream — manifesting as "assistant
did not respond within 30000ms" across most reasoning cells in
production.

Restore the original content-only fixture. The langgraph-python /
langgraph-fastapi agent fixes from #4579 still work against real
OpenAI (gpt-5-mini + Responses API streams real reasoning summaries),
but the aimock-driven path no longer exercises the role-reasoning
render — keyword-only assertion in the d5 probe handles that.
2026-05-01 14:07:21 +02:00
Alem Tuzlak 51db05f666 fix(showcase): emit reasoning events in langgraph-python and langgraph-fastapi (#4579)
## Summary

The `agentic-chat-reasoning` and `reasoning-default-render` cells in
`langgraph-python` and `langgraph-fastapi` never rendered any reasoning
content. Root cause: both agents were configured with `gpt-4o-mini` +
`use_responses_api=False`, so the underlying model produced no reasoning
content blocks and the Chat Completions API has no reasoning summary
surface in the first place. The frontend's `reasoningMessage` slot
stayed empty even though the cells are billed as reasoning demos.

This PR:

- Switches both agents (and their `tool_rendering_reasoning_chain`
siblings) to `gpt-5-mini` through the Responses API with
`reasoning={"effort":"medium","summary":"detailed"}`, mirroring the
`langgraph-typescript` and `pydantic-ai` agents that already worked.
Model is overridable via `OPENAI_REASONING_MODEL`.
- Updates the aimock `d5-all.json` fixture (and the matching harness
`reasoning-display.json`) to set the `reasoning` field on the `show your
reasoning step by step` match. Aimock now emits
`response.reasoning_summary_text.delta` events so the demo renders
deterministically without a real LLM call.
- Adds a `Show reasoning` `useConfigureSuggestions` pill on both
reasoning pages in both integrations so the demo is one click to
exercise.
- Tightens the `d5-reasoning-display` probe to also assert that a
reasoning-role message rendered (`[data-testid="reasoning-block"]` or
`[data-message-role="reasoning"]`), not just that the word "reasoning"
appears in the transcript.
- Un-skips the three streaming reasoning-block tests in
`agentic-chat-reasoning.spec.ts`, adds a suggestion-pill test, and
extends `reasoning-default-render.spec.ts` to cover the default
reasoning slot.
- Updates the `langgraph-python` QA doc to describe the new model +
Responses API setup and the pill flow.

Verified locally end-to-end: clicking the pill at
`/demos/agentic-chat-reasoning` renders the amber `ReasoningBlock` with
the fixture's reasoning text above the final answer bubble.

## Out of scope

Other integrations were audited and intentionally left alone:

- `langgraph-typescript`, `pydantic-ai` already use a reasoning model +
Responses API and work today.
- `agno`, `claude-sdk-python`, `ms-agent-python` use deliberate
workarounds (XML-tag reasoning + custom AGUI handler, Claude
extended-thinking deltas, `think` tool respectively) because their AG-UI
bridges either don't translate Responses-API reasoning items, run a
multi-call CoT loop incompatible with fixture replay, or don't emit
reasoning events at all.
- `llamaindex` uses `gpt-4.1` and surfaces reasoning inline as assistant
text. Its bridge (`llama-index-protocols-ag-ui`) does not translate
Responses-API reasoning items into AG-UI events; fixing that needs an
upstream patch and is out of scope here.

## Notes

Committed with `--no-verify` (explicit user request) — this worktree has
no `node_modules`, so the lefthook `test-and-check-packages` step
couldn't run locally. Changes are entirely under `showcase/` and CI runs
the same checks.

## Test plan

- [ ] CI fixture-validation passes on `showcase/aimock/d5-all.json`
- [ ] `showcase test langgraph-python --d5 --verbose` —
`reasoning-display` probe green (asserts `reasoning-block` selector +
keyword)
- [ ] `showcase test langgraph-fastapi --d5 --verbose` — same
- [ ] `nx run @copilotkit/showcase-langgraph-python:test:e2e -- --grep
reasoning` — un-skipped specs pass against the deployed Railway image
- [ ] Manual: visit `/demos/agentic-chat-reasoning` on a deployed
langgraph-python, click `Show reasoning`, confirm amber `REASONING —
Agent reasoning` block renders with italic step text above the final
answer bubble
- [ ] Manual: same on `/demos/reasoning-default-render`, confirm
CopilotKit's default `CopilotChatReasoningMessage` card renders
2026-05-01 13:40:10 +02:00
Alem Tuzlak dca1b9894d fix(showcase): emit reasoning events in langgraph-python and langgraph-fastapi
The agentic-chat-reasoning and reasoning-default-render cells in
langgraph-python and langgraph-fastapi were configured with
gpt-4o-mini + use_responses_api=False, which never produces AG-UI
REASONING_MESSAGE_* events: gpt-4o-mini is not a reasoning model and
the Chat Completions API does not surface reasoning summary items at
all. The frontend's reasoningMessage slot was rendering nothing,
even though the cells were billed as "reasoning" demos.

- Switch both reasoning agents to gpt-5-mini (override via
  OPENAI_REASONING_MODEL) routed through the Responses API with
  reasoning={"effort":"medium","summary":"detailed"} so the model's
  chain of thought streams as content blocks that @ag-ui/langgraph
  translates into REASONING_MESSAGE_* events.
- Update the aimock d5-all.json and harness reasoning-display.json
  fixtures to include a "reasoning" field so aimock emits
  response.reasoning_summary_text.delta SSE events deterministically
  in CI without hitting a real LLM.
- Add a "Show reasoning" useConfigureSuggestions pill on both
  reasoning demo pages so the user can trigger the fixture-matched
  prompt with one click.
- Tighten the d5-reasoning-display probe: it now also asserts a
  reasoning-role message rendered via [data-testid="reasoning-block"]
  or [data-message-role="reasoning"], so a plain text response
  containing the word "reasoning" no longer falsely passes.
- Un-skip the three streaming reasoning-block tests in
  langgraph-python's agentic-chat-reasoning.spec.ts and add a
  suggestion-pill test; expand the reasoning-default-render spec to
  cover the default reasoning slot.
- Update the langgraph-python QA doc to describe the new model +
  Responses API setup and the suggestion-pill flow.
2026-05-01 13:30:28 +02:00
Alem Tuzlak a549de3a41 fix(aimock): add fixtures for beautiful-chat suggestions + e2e regression
The 6 beautiful-chat demos (spring-ai, strands, langroid, agno,
claude-sdk-typescript, claude-sdk-python) ship three identical
suggestion chips: "Plan a 3-day Tokyo trip", "Explain RAG like I'm
12", and "Draft a launch email". Against the deployed aimock-backed
showcase, all three were broken:

- Tokyo trip: hijacked by the broad `userMessage: "hi"` fixture,
  because the substring "hi" appears inside "arc**hi**tecture" in
  the prompt. Returned a generic "Hi there! I'm your showcase
  assistant..." greeting with nothing about Tokyo.
- RAG explain: no fixture matched, aimock returned an error.
- Launch email: same — no fixture, error.

Add three on-topic fixtures with the full suggestion sentence as
`userMessage` (effectively-exact substring match). Place them
before the broad "hi" fixture in the file so first-match-wins
routes each suggestion to the right response.

Add a `beautiful-chat.spec.ts` regression suite to all 6
integrations: send each suggestion, assert the right keywords
appear in the assistant reply ("Day 1/2/3" for Tokyo,
"open-book/RAG" for RAG, "Subject:/co-pilot" for email), AND
assert the hijacked greeting is absent. If the broad "hi" fixture
re-broadens or the new fixtures are reordered/removed, these
tests fail loudly.
2026-05-01 13:17:44 +02:00
Jordan Ritter c30e40f371 fix(showcase): add mastra get-weather hyphen variant to aimock fixtures
Mastra registers the weather tool as get-weather (with hyphen) instead
of get_weather. Add a duplicate fixture entry so aimock matches correctly.
2026-05-01 00:52:43 -07:00
Jordan Ritter 463b0b7d0b feat(showcase): D5 voice test for langgraph-python
Add D5 voice test that exercises sample-audio transcription via aimock.
Infrastructure: voice in D5 feature type registry + mapping, skipFill
support in conversation runner (9 new tests), inputValue forwarding
in e2e-deep Page wrappers, aimock transcription fixture, tool-free
weather fallback fixture for agents without tools. Verified locally:
D5 suite passes green on langgraph-python (60.4s).
2026-04-30 22:15:45 -07:00
Alem Tuzlak 0eb81b1406 fix(showcase): unbreak agent-config + byoc D5, reaching 31/31 green
agent-config: drop the AgentConfigLangGraphAgent subclass and use plain
LangGraphAgent. The subclass repacked CopilotKit provider properties
into forwardedProps.config.configurable.properties so the Python graph
could read them via RunnableConfig.configurable.properties — but
@ag-ui/langgraph@0.0.31 builds the LangGraph SDK request as
{ ..., config, context: { ...input.context, ...config.configurable } }
which merges configurable INTO context. LangGraph 0.6.0+ then rejects
with HTTP 400 'Cannot specify both configurable and context' on every
chat round-trip. Net effect: chat sent the user message, runtime 400'd,
no assistant response ever rendered. Removing the subclass unbreaks
the round-trip; the Python agent falls back to its DEFAULT_* constants
so the demo's frontend toggles no longer steer the system prompt
(known regression, tracked separately pending @ag-ui/langgraph fix
that decouples context from configurable).

byoc:
- D5 probe now sends the 'Sales dashboard' pill prompt (matches the
  fixtures added in main:f0a89b843 in feature-parity.json) instead of
  the previous generic 'render a byoc hashbrown' prompt that had no
  matching JSON-shaped fixture. Removed the now-obsolete byoc.json D5
  fixture file and regenerated the d5-all.json bundle (52 -> 50
  fixtures).
- Added data-testid='copilot-assistant-message' + data-message-role=
  'assistant' to the byoc-hashbrown and byoc-json-render renderer
  wrapper divs. The CopilotChat default assistantMessage slot includes
  these markers; overriding the slot with a custom JSON-rendering
  component dropped them, so the e2e-deep conversation runner's
  settle-detection cascade (which counts these selectors) never saw
  the response and timed out at 30s. Re-attaching the markers is a
  purely additive change that doesn't affect the renderers'
  behavior.
- D5 byoc assertion now waits for [data-testid='metric-card'] AND a
  chart (bar-chart or pie-chart) to render — a structural check on
  the BYOC contract output, not a transcript-keyword check that the
  custom renderer would never produce.

E2E status: 31/31 passing locally against
./bin/showcase up langgraph-python aimock with this branch's bundle.
2026-04-30 15:11:19 +02:00
Alem Tuzlak a0abd81c6b Merge main into blitz/lgp-d5-coverage-design/integration 2026-04-30 14:09:58 +02:00
Alem Tuzlak 8e67d42a66 Merge branch 'main' into fix/showcase-byoc-aimock-fixtures 2026-04-30 14:07:33 +02:00
Alem Tuzlak f0a89b8430 fix(showcase): add aimock fixtures for byoc-json-render + byoc-hashbrown pills
Both BYOC demo pills were getting eaten by older generic substring
fixtures ("dashboard", "revenue by category as a pie chart", "monthly
expenses as a bar chart") — the json-render pills came back as foreign
tool calls (render_pie_chart) and the hashbrown sales pill came back as
the "Working on your request..." placeholder, leaving both renderers
empty.

Add a dedicated fixture per pill keyed on the full pill prompt sentence,
placed before the conflicting generic rules so first-fixture-wins lands
on the right one. Each fixture returns the exact content shape the
demo's renderer expects:

  byoc-json-render -> {root, elements} flat spec for <Renderer />
  byoc-hashbrown   -> {ui:[{<tag>:{props:{...}}}]} for useJsonParser

Removes two pre-existing duplicates that had been added at the bottom of
the file but never matched because the conflicting generic rules came
first.
2026-04-30 13:54:29 +02:00
Alem Tuzlak 99868c7c46 feat(showcase): regenerate d5-all.json bundle and fix DOM typing in chat-css probe (F) 2026-04-30 13:25:27 +02:00
Alem Tuzlak f1b02a4616 fix(showcase): stop infinite tool-call loop in beautiful-chat + restore brand styling
Beautiful Chat suggestion clicks looped forever because feature-parity.json
tool-calling fixtures lacked an `id` and a paired `toolCallId` followup.
After the agent ran the tool and re-prompted aimock, the same userMessage
substring matched again and the same toolCall was returned indefinitely.
Added explicit ids to 10 broken fixtures (pieChart, barChart, render_*_chart,
scheduleTime, search_flights, toggleTheme) and 11 paired toolCallId
followups returning content summaries — same convention the file already
uses for show_card, weather, etc.

Beautiful Chat layout also showed a black/white split and a broken logo on
the 8 integrations using the full ExampleLayout pattern (crewai-crews,
langgraph-fastapi, langgraph-python, langgraph-typescript, mastra,
ms-agent-dotnet, ms-agent-python, pydantic-ai). Two issues:

1. globals.css hardcoded `body { background: #fafaf9 }` and never defined
   the brand tokens (--background, --foreground, --card, --primary, …) that
   the layout, mode-toggle, todo card/column, and chart components reference
   via Tailwind 4 arbitrary values. ThemeProvider was also adding `dark` to
   <html> from system preference, so CopilotKit's chat went dark while body
   stayed cream.
2. example-layout/index.tsx renders <img src="/copilotkit-logo.svg" /> but
   the file did not exist in any integration's public/.

Added the full token set (light + dark) under :root and :root.dark/.dark,
registered the Tailwind 4 dark variant, switched body to var(--background)
/var(--foreground), and copied copilotkit-logo.svg + copilotkit-logo-mark.svg
into each integration's public/ from examples/integrations/langgraph-python.
2026-04-30 12:42:26 +02:00
Jordan Ritter e4a332e095 fix: add showcase-aimock to CI build & deploy pipeline
Production showcase-aimock was running a week-old image because fixture
file changes in showcase/aimock/ did not trigger a CI rebuild. This adds
showcase-aimock to the Build & Deploy workflow matrix so it auto-deploys
on merge, creates a thin Dockerfile that bakes fixture files into the
image, documents the local-vs-production parity requirement in
docker-compose.local.yml, and adds an aimock fixture deployment section
to the RUNBOOK.
2026-04-29 20:23:41 -07:00
Jordan Ritter 8d0bc3cd93 fix(showcase): migrate D5 fixtures from turnIndex to hasToolResult
hasToolResult checks whether tool-role messages exist in the request,
correctly disambiguating multi-turn fixtures regardless of the backend
AG-UI implementation's re-invocation behavior.
2026-04-29 19:40:10 -07:00
Jordan Ritter 6a4c25a24d fix: add nested sub-agent fixtures for D5 mcp-subagents demo
Add 3 fixtures for the independent LLM calls made by each sub-agent
tool (research_agent, writing_agent, critique_agent) so the demo works
without --proxy-only mode. Interleave them with supervisor fixtures for
readability. Rebuild d5-all.json bundle.
2026-04-28 20:45:00 -07:00
Jordan Ritter a2bc64692f fix(showcase): migrate D5 fixtures from toolCallId to turnIndex
aimock v1.16.0 ships native turnIndex matching — counts assistant
messages in the request's message array instead of relying on
toolCallId in the last tool message. Stateless, concurrent-safe,
fixture-ordering-insensitive.

Removes all toolCallId match criteria from 9 D5 fixture files (8
source + 1 bundle), adds turnIndex values, and reorders fixtures
to natural 0→N reading order.
2026-04-28 18:10:25 -07:00
Jordan Ritter bd32bcff78 fix(showcase): replace sequenceIndex with toolCallId in D5 fixtures
aimock sequenceIndex counter is server-global, causing fixture match
failures when multiple integrations run concurrently. Replace with
stateless toolCallId matching.
2026-04-28 16:55:46 -07:00
Jordan Ritter 83b077db87 refactor(showcase): update scripts, docs, and aimock references for harness rename
Update comment references in showcase scripts (create-integration,
redirect-decommission, validate-pins, verify-railway-image-refs),
RAILWAY.md docs, aimock/d5-all.json fixture comment, and
built-in-agent route comment.
2026-04-28 13:48:30 -07:00
Alem Tuzlak a16eaa1b74 fix(showcase): repair agent-config (LangGraph 1.x context) + multimodal fixtures
agent-config — empty stream
The /demos/agent-config runtime returned RUN_ERROR on every send. The
existing route wrapper repacked the provider's properties into
forwardedProps.config.configurable.properties so the agent could read
them via RunnableConfig["configurable"]["properties"]. LangGraph 1.x
(deployed langgraph 1.1.x / langgraph-api 0.7.x) now rejects any run
that sends configurable with "Cannot specify both configurable and
context. Prefer setting context alone." because it auto-injects an
empty context and refuses both channels at once.

Switch to LangGraph 1.x's context API: route wrapper repacks user props
onto forwardedProps.context instead, the Python graph defines
context_schema=AgentConfigContext and reads via the
Runtime[AgentConfigContext] parameter, and RunnableConfig is no longer
consulted. Unit tests updated to pass the flat context dict directly.
Add context to RESERVED_FORWARDED_PROPS_KEYS so a caller that
explicitly sets forwardedProps.context is preserved (and merged with
provider-supplied properties) instead of being treated as a user prop
and double-wrapped.

multimodal — PDF "doesn't work"
The multimodal pipeline (legacy-shape rewrite shim → ag-ui-langgraph
converter → _PdfFlattenMiddleware running pypdf) is fine; sample.pdf
extraction shows up correctly in the run state snapshot. The visible
breakage is the aimock fixture file: aimock matches userMessage as a
substring, and the generic "hi" fixture lives early enough to swallow
any prompt containing "this" — including "What is in this PDF?" and
"What is in this image?" — short-circuiting before the request reaches
real OpenAI and returning a generic "I'm your showcase assistant"
greeting that never references the attachment.

Add two more-specific fixtures ("this PDF", "this image") ahead of the
"hi" entry so multimodal prompts return responses grounded in the
bundled samples. The "hi" fixture stays in place so the several E2E
specs that fill literal "hi" still match it.
2026-04-28 17:44:07 +02:00
Jordan Ritter c70d86f6f9 fix: convert all D5 fixtures from toolCallId to sequenceIndex matching
Aimock's Gemini adapter synthesizes tool_call_ids (call_gemini_<name>_N)
that don't match the D5 fixture IDs (call_d5_<name>_001). This breaks
ALL two-leg and multi-leg fixtures for google-adk and any future Gemini
backend because Gemini's functionResponse format has no id field.

sequenceIndex counts per-userMessage matches regardless of provider,
making fixtures provider-agnostic.

Also adds missing generate_haiku fixtures to d5-all.json (were in
gen-ui-custom.json but never bundled) and fixes hitl-text-input
fixture arg name→attendee to match frontend schema.
2026-04-28 03:31:55 -07:00
Jordan Ritter 5c31e171cd docs: add showcase-ops memory budget note to RAILWAY.md
Documents the D5 e2e-deep Chromium memory envelope (4 services x 2
features = 8 concurrent contexts, ~2.4GB peak) and the two tuning
knobs operators should adjust if OOM occurs.
2026-04-27 15:02:07 -07:00
Jordan Ritter 3a5d29f0c4 fix(showcase): correct doc counts, fix unfalsifiable registry test, update comments 2026-04-27 11:42:57 -07:00
Jordan Ritter 63f36e4a24 fix(showcase): stringify d5 fixture arguments for aimock validation 2026-04-27 11:24:52 -07:00
Jordan Ritter c8fed358d1 fix(showcase): flatten hitl-steps tool format and drop stale search_flights from d5 fixtures 2026-04-27 11:19:02 -07:00
github-actions[bot] a5edd23d4f style: auto-fix formatting 2026-04-27 18:10:21 +00:00
Jordan Ritter fc0a66cd38 fix(showcase): bundle D5 fixtures for aimock and update startCommand
D5 probes hit real LLMs instead of fixtures because the 8 D5 fixture
files in showcase/ops/fixtures/d5/ were never loaded into aimock.
Railway fetches fixtures from GitHub raw URLs at boot — bundle all
D5 fixtures into d5-all.json and document the updated startCommand
with the new URL appearing before feature-parity.json for match
precedence.
2026-04-27 11:07:13 -07:00
Jordan Ritter da808ad7fb fix(showcase): add D5 multi-turn and tool-call fixtures to feature-parity
Add 20 new aimock fixtures covering all D5 e2e-deep probe prompts:
- 9 toolCallId second-leg fixtures for tool-call round-trips
- 11 new userMessage fixtures for multi-turn and feature-specific flows
- Explicit tool call IDs on existing weather/pie-chart fixtures
- Correct Alice ordering to prevent substring collision

Without these fixtures, D5 probes hit real OpenAI (401 with test key)
or loop infinitely on tool calls with no second-leg response.
2026-04-26 23:35:53 -07:00
Jordan Ritter a2a82aaf45 Merge remote-tracking branch 'origin/main' into feat/d5-d6-probes
# Conflicts:
#	showcase/ops/src/probes/loader/probe-loader.test.ts
2026-04-26 14:34:18 -07:00
Jordan Ritter 6a03c3e11d fix: add aimock fixtures for byoc-json-render chart prompts
The aimock-fixture-coverage probe scans byoc-json-render.spec.ts and was
flagging two prompts without matching feature-parity.json fixtures:

  - "Break down revenue by category as a pie chart"
  - "Show me monthly expenses as a bar chart"

Existing pie/bar fixtures used substrings ("revenue distribution by
category", "expenses by category") that don't appear in these prompts,
so the substring matcher returned no hits.

Added two new fixtures with substrings that DO appear in the prompts
("revenue by category as a pie chart", "monthly expenses as a bar
chart") so the coverage probe stays green and aimock returns
deterministic chart data when the BYOC json-render specs run against
the sidecar.
2026-04-25 19:56:41 -07:00
Jordan Ritter 5f1b4f6b9c fix(showcase-ops): add missing aimock fixtures for byoc-json-render
Adds two feature-parity.json fixtures for the langgraph-python
byoc-json-render E2E spec prompts that the
aimock-fixture-coverage probe flagged as uncovered:

  - "Break down revenue by category as a pie chart"
  - "Show me monthly expenses as a bar chart"

Responses mirror the worked examples baked into the
byoc_json_render_agent system prompt verbatim, so the fixture
output matches what the live LLM would emit (a single
@json-render/react flat-spec JSON object) and the spec
assertions on [data-testid="pie-chart"] /
[data-testid="bar-chart"] resolve deterministically when the
suite runs against the aimock sidecar.
2026-04-25 19:25:07 -07:00
Alem Tuzlak 7c66420a9f Merge remote-tracking branch 'origin/main' into feat/wave2b-multimodal-demo
# Conflicts:
#	showcase/aimock/feature-parity.json
#	showcase/packages/langgraph-python/docs-links.json
#	showcase/packages/langgraph-python/manifest.yaml
#	showcase/shell-docs/src/data/demo-content.json
#	showcase/shell-docs/src/data/registry.json
#	showcase/shell-dojo/src/data/demo-content.json
#	showcase/shell-dojo/src/data/registry.json
#	showcase/shell/src/data/constraints.json
#	showcase/shell/src/data/demo-content.json
#	showcase/shell/src/data/docs-status.json
#	showcase/shell/src/data/registry.json
2026-04-24 10:51:23 +02:00
Alem Tuzlak 1d32b0f057 feat(showcase/langgraph-python): multimodal attachments demo (Wave 2b)
Wire CopilotChat AttachmentsConfig (image + PDF, inline base64) into a
dedicated /demos/multimodal cell backed by a vision-capable LangGraph
agent (gpt-4o). Scoped to its own runtime route so the vision cost
stays on this demo.

- Dedicated /api/copilotkit-multimodal route registering the new
  multimodal graph under the multimodal-demo slug
- src/agents/multimodal_agent.py with an AgentMiddleware that flattens
  PDF content parts to text server-side via pypdf
- Try with sample image / PDF buttons that drive CopilotChats own
  hidden file input via DataTransfer + change event, so the sample and
  paperclip paths exercise the same useAttachments pipeline
- public/demo-files/ placeholder — sample.png / sample.pdf binaries
  are produced by the user at the end of the Wave 2b rollout
- Manifest + registry + constraints + docs-links wired; normalized
  multi-modal -> multimodal in constraints to match the feature
  registry id
- QA checklist + Playwright E2E spec (not run yet — pending deploy)
- aimock fixtures for sample prompts so CI runs are deterministic
- Extend the check-binaries hook allowlist to cover the shell-docs and
  shell-dojo demo-content.json mirrors (same size as shell/src/data,
  already allowed) so regenerated derived data can be committed
2026-04-23 22:15:14 +02:00
Alem Tuzlak 8fb68293f6 feat(showcase): add aimock fixtures and coverage enforcement
9 new userMessage->toolCall/content fixtures in feature-parity.json
for langgraph-python demos. New aimock-fixture-coverage.test.ts
enforces per-spec fixture presence — every E2E spec prompt must
have a matching aimock fixture or be explicitly skipped.

Tighten pie-chart fixture match to avoid beautiful-chat collision:
match specific gen-ui-tool-based prompt text instead of generic
substring that also hits beautiful-chat suggestion pills.
2026-04-23 12:46:15 -07:00
Jordan Ritter 577aec070d chore(showcase-aimock): remove wrapper Dockerfile (Phase 3)
Railway now pulls ghcr.io/copilotkit/aimock:latest directly
via a startCommand that fetches fixtures from GitHub raw URLs.
The wrapper Dockerfile that copied fixtures into the image is
no longer needed.

Also updates README.md to reflect the new architecture:
- References Railway startCommand instead of Dockerfile
- Removes the Dockerfile entry from the fixtures listing
- Updates related workflow descriptions
2026-04-23 08:58:09 -07:00
Jordan Ritter 123ae66b87 docs(showcase-aimock): add Railway service reconstruction reference 2026-04-22 20:02:15 -07:00
Jordan Ritter 78a4a6d0a6 chore(repo): root workspace config + top-level showcase docs
Bump pnpm-lock / package.json / pnpm-workspace / lefthook.yml for the
showcase-ops branch, add FRONTEND-STRATEGY / TESTING / QA-COVERAGE /
INTEGRATION-CHECKLIST top-level showcase docs + aimock README, refresh
showcase/.gitignore + showcase/shared/constraints.yaml.
2026-04-22 10:50:09 -07:00
Alem Tuzlak b691b63660 fix(showcase): narrow aimock fixture patterns + add drift guardrail
Substring-match fixtures (pie chart, bar chart, schedule, trip, etc.)
cross-fired across demos with different tool surfaces and returned tool
names the target agent never registered, causing demos to render nothing
in prod when aimock handles traffic.

Fixture changes (showcase/aimock/feature-parity.json):
- Replace generic pie-chart / bar-chart matches with per-suggestion
  specific phrases so gen-ui-tool-based gets render_pie_chart /
  render_bar_chart directly and beautiful-chat gets pieChart / barChart
  with real data (skipping the query_data two-step that caused the
  infinite loop on re-matching prompts).
- Narrow schedule+meeting to the Beautiful Chat 30-minute prompt
  returning scheduleTime.
- Narrow flight+fly to flights-from-SFO-to-JFK.
- Narrow background to sunset-themed-gradient.
- Remove trip, sales, pipeline, todo: substring-false-firing across
  unrelated demos; interrupt/A2UI demos fall through to real LLM.

Guardrail (showcase/scripts/validate-fixture-tool-surface.ts):
- Pure validate() cross-references every fixture's tool-call name
  against the tool surface of each demo whose suggestion prompt contains
  the fixture's match substring. Loud failure when the fixture returns a
  tool the demo's agent does not register.
- CLI walks packages/ collecting suggestions from page.tsx + hooks/,
  frontend tools from useComponent / useHumanInTheLoop / useFrontendTool
  / useRenderTool / useDefaultRenderTool, and backend tools via route.ts
  agentId->graphId map + langgraph.json graph->file + @tool decorators.
- 7 vitest cases written TDD-first covering the drift detection,
  content-only fixtures, case-insensitivity, and multi-tool responses.
- Current state: 33 fixtures x 191 demos, no drift. Counterfactual
  (reverting the pie-chart fix) correctly flags gen-ui-tool-based and
  declarative-gen-ui.

Also fixes a separate runtime bug in the langgraph-python package
Dockerfile: WORKDIR /app left /app owned by root; the app user could
not create the .langgraph_api cache dir LangGraph's in-memory runtime
needs, so the agent crashed on boot. Added a non-recursive chown
app:app /app (preserves the original perf intent of the explicit
--chown on COPY, which avoided a recursive chown).
2026-04-22 10:16:11 -05:00
Jordan Ritter 3f62cd45e6 fix: validate aimock fixtures at load time to prevent runtime 500s
Follow-up to #3971. aimock supports fixture schema validation at startup via
--validate-on-load, but it's opt-in. The showcase Dockerfile did not pass
the flag, so fixtures with unrecognized response keys (e.g. "text" instead
of "content") loaded silently and only failed at request time with HTTP 500.
That's what crashed crewai-crews on startup.

Changes:
- showcase/aimock/Dockerfile: pass --validate-on-load so broken fixtures
  fail the container boot, not individual requests.
- showcase/scripts/__tests__/aimock-fixtures.test.ts: new vitest spec that
  loads feature-parity.json and smoke.json via @copilotkit/aimock's
  loadFixtureFile + validateFixtures and asserts zero errors. Runs as part
  of the existing showcase-validate CI workflow.
- showcase/scripts/package.json: add @copilotkit/aimock dependency for the
  validator import.

Verified red-green: with the pre-#3971 broken "text" fixtures, validateFixtures
flags 5 errors; post-#3971 it returns zero. Docker red-green: container with
an intentionally broken fixture fails to start with "Validation failed: 1
error(s)" and non-zero exit.
2026-04-16 12:32:24 -07:00
Jordan Ritter 08c60ce5fd fix(aimock): use valid response schema for broken fixtures, unblock crewai-crews
The showcase aimock feature-parity fixture had five entries that used the
unrecognized field "text" instead of "content" in the response body.

aimock's /v1/chat/completions handler validates the response shape through
discriminators (isTextResponse, isToolCallResponse, etc.) and when none
match, returns 500 "Fixture response did not match any known type".

When CrewAI's ChatWithCrewFlow.__init__ calls generate_input_description_with_ai,
it sends a prompt containing "Reporting Analyst" / "reports" text (from the
LatestAiDevelopment crew definition). aimock substring-matches this against
the "report" fixture, which had the broken "text" shape, and returned 500.

CrewAI treats that as an InternalServerError, the crewai-crews container
crashes during module import (before uvicorn binds), Railway healthchecks
fail, and the deployment rolls back to the previous ACTIVE build.

Fix:
- Convert the 5 fixtures with "text" (plan, steps, mars, dashboard, report)
  to the valid "content" field so they return 200 OK.
- Add a targeted fixture that matches CrewAI's exact startup prompt prefix
  ("Based on the following context, write a concise") and returns a clean
  text description. This catches generate_input_description_with_ai and
  generate_crew_description_with_ai before they can fall through to the
  more generic fixtures, providing a more predictable response for CrewAI
  startup.

Verified locally against @copilotkit/aimock@1.10.0:
- Before: the CrewAI-style prompt returns 500 with the exact error message
  seen in Railway logs.
- After: returns 200 with a valid chat.completion envelope.

Note: CrewAI's blocking LLM call at module import remains fragile (any
aimock hiccup crashes the container before it can bind a port). That is
an upstream issue and will be tracked separately.
2026-04-16 12:06:41 -07:00
Jordan Ritter 5ac07d4d9b feat: CI, shell, and aimock integration for showcase starters
CI:
- Add 17 starter services to showcase deploy workflow with Railway IDs
- Add drift detection workflow (triggers on starters, packages, scripts, shared)
- Remove shared_frontend copy step for starter builds

Shell:
- Update clone command to npx degit with clipboard fallback
- Update starter content bundler for full component tree + .java support
- Add clone_command to manifest schema and registry types

Aimock:
- Expand feature-parity.json from 18 to 37 fixture rules
- Add docker-compose.packages.yml for CI aimock sidecar (strict mode)
- Add run-e2e-with-aimock.sh convenience script
2026-04-14 12:52:52 -07:00
Jordan Ritter c070403014 feat: Docker and CI support for shared showcase modules
- CI copies shared_python, shared_frontend/src, shared_typescript/tools into
  each package's build context before Docker build
- 12 Python Dockerfiles: COPY shared_python, ENV PYTHONPATH
- All Dockerfiles: npm install --legacy-peer-deps + tsconfig paths for
  @copilotkit/showcase-shared (no npm dependency needed)
- spring-ai: unpinned eclipse-temurin:17 base images
- Aimock: 18 deterministic fixtures for all demo scenarios
2026-04-13 17:39:05 -07:00
Jordan Ritter 0e53a48c19 feat: add aimock proxy for showcase smoke checks 2026-04-09 13:53:58 -07:00