Files
Alem Tuzlak dca1b9894d fix(showcase): emit reasoning events in langgraph-python and langgraph-fastapi
The agentic-chat-reasoning and reasoning-default-render cells in
langgraph-python and langgraph-fastapi were configured with
gpt-4o-mini + use_responses_api=False, which never produces AG-UI
REASONING_MESSAGE_* events: gpt-4o-mini is not a reasoning model and
the Chat Completions API does not surface reasoning summary items at
all. The frontend's reasoningMessage slot was rendering nothing,
even though the cells were billed as "reasoning" demos.

- Switch both reasoning agents to gpt-5-mini (override via
  OPENAI_REASONING_MODEL) routed through the Responses API with
  reasoning={"effort":"medium","summary":"detailed"} so the model's
  chain of thought streams as content blocks that @ag-ui/langgraph
  translates into REASONING_MESSAGE_* events.
- Update the aimock d5-all.json and harness reasoning-display.json
  fixtures to include a "reasoning" field so aimock emits
  response.reasoning_summary_text.delta SSE events deterministically
  in CI without hitting a real LLM.
- Add a "Show reasoning" useConfigureSuggestions pill on both
  reasoning demo pages so the user can trigger the fixture-matched
  prompt with one click.
- Tighten the d5-reasoning-display probe: it now also asserts a
  reasoning-role message rendered via [data-testid="reasoning-block"]
  or [data-message-role="reasoning"], so a plain text response
  containing the word "reasoning" no longer falsely passes.
- Un-skip the three streaming reasoning-block tests in
  langgraph-python's agentic-chat-reasoning.spec.ts and add a
  suggestion-pill test; expand the reasoning-default-render spec to
  cover the default reasoning slot.
- Update the langgraph-python QA doc to describe the new model +
  Responses API setup and the suggestion-pill flow.
2026-05-01 13:30:28 +02:00

4.2 KiB

QA: Agentic Chat (Reasoning) — LangGraph (Python)

Prerequisites

  • Demo is deployed and accessible at /demos/agentic-chat-reasoning on the dashboard host
  • Agent backend is healthy (/api/copilotkit GET returns langgraph_status: "reachable"); OPENAI_API_KEY is set; LANGGRAPH_DEPLOYMENT_URL points at a deployment exposing the reasoning_agent graph (registered under agent name agentic-chat-reasoning)
  • The demo overrides the reasoningMessage slot with a custom ReasoningBlock component (amber-tinted banner); it relies on deepagents.create_deep_agent with a reasoning-capable OpenAI model (gpt-5-mini by default, override via OPENAI_REASONING_MODEL) routed through the Responses API with reasoning={"effort": "medium", "summary": "detailed"} so the model's chain of thought streams as AG-UI REASONING_MESSAGE_* events. A "Show reasoning" suggestion pill above the input fires show your reasoning step by step, which the aimock fixture in showcase/aimock/d5-all.json matches with reasoning summary deltas for deterministic local/CI runs.

Test Steps

1. Basic Functionality

  • Navigate to /demos/agentic-chat-reasoning; verify the CopilotChat panel renders centered (max width 4xl, rounded-2xl corners) within 3s
  • Verify the input field is visible with placeholder "Type a message"
  • Send "Hello"; verify an assistant text response appears within 10s
  • Verify at least one reasoning block (data-testid="reasoning-block") renders above (or alongside) the final assistant text

2. Feature-Specific Checks

Reasoning Block — Streaming State

  • Send a prompt that requires multi-step thinking (e.g. "If a train leaves Boston at 3pm going 60mph and another leaves NY at 4pm going 80mph, when do they meet? Explain your approach.")
  • While the response is streaming, verify the data-testid="reasoning-block" element is visible with:
    • An uppercase "REASONING" pill badge (white background, purple border #BEC2FF)
    • A "Thinking…" label to the right of the pill (text color #57575B)
  • Verify reasoning text accumulates inside the block (italic, whitespace-pre-wrap) as tokens stream in
  • Verify the container border color is #DBDBE5 and background tint is #BEC2FF1A (amber/indigo banner)

Reasoning Block — Completed State

  • After streaming finishes, verify the label next to the REASONING pill flips from "Thinking…" to "Agent reasoning"
  • Verify the reasoning content remains visible (not collapsed) and contains the agent's step-by-step rationale
  • Verify a separate assistant text bubble renders the final concise answer below/after the reasoning block
  • Verify the final answer is distinct from the reasoning text (not duplicated verbatim)

Multi-Turn Reasoning State Retention

  • After the first exchange completes, send a follow-up prompt (e.g. "Now walk me through your work for 100mph and 120mph instead.")
  • Verify the previous turn's reasoning-block remains rendered in the transcript (not cleared)
  • Verify a NEW reasoning-block appears for the second turn
  • Count the [data-testid="reasoning-block"] elements; verify there are at least 2 after the second response
  • Send a third prompt; verify all three prior turns' reasoning blocks remain intact and render in chronological order

3. Error Handling

  • Attempt to send an empty message; verify it is a no-op (no user bubble, no reasoning-block, no assistant response)
  • Send a ~500-character prompt; verify the reasoning block and answer both render without horizontal scroll or layout break
  • With the backend stopped, send a message; verify the UI surfaces a visible error rather than hanging silently
  • Open DevTools → Console; verify no uncaught errors during any flow above

Expected Results

  • Chat loads within 3 seconds; first reasoning token appears within 5s of sending
  • The reasoning-block renders BEFORE the final assistant text (or simultaneously) — never after
  • "Thinking…" label is visible during streaming and flips to "Agent reasoning" once isRunning settles
  • Prior turns' reasoning blocks persist across new messages (no clearing)
  • No UI layout breaks, no flash of unstyled content, no uncaught console errors