mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
dca1b9894d
The agentic-chat-reasoning and reasoning-default-render cells in
langgraph-python and langgraph-fastapi were configured with
gpt-4o-mini + use_responses_api=False, which never produces AG-UI
REASONING_MESSAGE_* events: gpt-4o-mini is not a reasoning model and
the Chat Completions API does not surface reasoning summary items at
all. The frontend's reasoningMessage slot was rendering nothing,
even though the cells were billed as "reasoning" demos.
- Switch both reasoning agents to gpt-5-mini (override via
OPENAI_REASONING_MODEL) routed through the Responses API with
reasoning={"effort":"medium","summary":"detailed"} so the model's
chain of thought streams as content blocks that @ag-ui/langgraph
translates into REASONING_MESSAGE_* events.
- Update the aimock d5-all.json and harness reasoning-display.json
fixtures to include a "reasoning" field so aimock emits
response.reasoning_summary_text.delta SSE events deterministically
in CI without hitting a real LLM.
- Add a "Show reasoning" useConfigureSuggestions pill on both
reasoning demo pages so the user can trigger the fixture-matched
prompt with one click.
- Tighten the d5-reasoning-display probe: it now also asserts a
reasoning-role message rendered via [data-testid="reasoning-block"]
or [data-message-role="reasoning"], so a plain text response
containing the word "reasoning" no longer falsely passes.
- Un-skip the three streaming reasoning-block tests in
langgraph-python's agentic-chat-reasoning.spec.ts and add a
suggestion-pill test; expand the reasoning-default-render spec to
cover the default reasoning slot.
- Update the langgraph-python QA doc to describe the new model +
Responses API setup and the suggestion-pill flow.
4.2 KiB
4.2 KiB
QA: Agentic Chat (Reasoning) — LangGraph (Python)
Prerequisites
- Demo is deployed and accessible at
/demos/agentic-chat-reasoningon the dashboard host - Agent backend is healthy (
/api/copilotkitGET returnslanggraph_status: "reachable");OPENAI_API_KEYis set;LANGGRAPH_DEPLOYMENT_URLpoints at a deployment exposing thereasoning_agentgraph (registered under agent nameagentic-chat-reasoning) - The demo overrides the
reasoningMessageslot with a customReasoningBlockcomponent (amber-tinted banner); it relies ondeepagents.create_deep_agentwith a reasoning-capable OpenAI model (gpt-5-miniby default, override viaOPENAI_REASONING_MODEL) routed through the Responses API withreasoning={"effort": "medium", "summary": "detailed"}so the model's chain of thought streams as AG-UIREASONING_MESSAGE_*events. A "Show reasoning" suggestion pill above the input firesshow your reasoning step by step, which the aimock fixture inshowcase/aimock/d5-all.jsonmatches with reasoning summary deltas for deterministic local/CI runs.
Test Steps
1. Basic Functionality
- Navigate to
/demos/agentic-chat-reasoning; verify theCopilotChatpanel renders centered (max width 4xl, rounded-2xl corners) within 3s - Verify the input field is visible with placeholder "Type a message"
- Send "Hello"; verify an assistant text response appears within 10s
- Verify at least one reasoning block (
data-testid="reasoning-block") renders above (or alongside) the final assistant text
2. Feature-Specific Checks
Reasoning Block — Streaming State
- Send a prompt that requires multi-step thinking (e.g. "If a train leaves Boston at 3pm going 60mph and another leaves NY at 4pm going 80mph, when do they meet? Explain your approach.")
- While the response is streaming, verify the
data-testid="reasoning-block"element is visible with:- An uppercase "REASONING" pill badge (white background, purple border
#BEC2FF) - A "Thinking…" label to the right of the pill (text color
#57575B)
- An uppercase "REASONING" pill badge (white background, purple border
- Verify reasoning text accumulates inside the block (italic, whitespace-pre-wrap) as tokens stream in
- Verify the container border color is
#DBDBE5and background tint is#BEC2FF1A(amber/indigo banner)
Reasoning Block — Completed State
- After streaming finishes, verify the label next to the REASONING pill flips from "Thinking…" to "Agent reasoning"
- Verify the reasoning content remains visible (not collapsed) and contains the agent's step-by-step rationale
- Verify a separate assistant text bubble renders the final concise answer below/after the reasoning block
- Verify the final answer is distinct from the reasoning text (not duplicated verbatim)
Multi-Turn Reasoning State Retention
- After the first exchange completes, send a follow-up prompt (e.g. "Now walk me through your work for 100mph and 120mph instead.")
- Verify the previous turn's
reasoning-blockremains rendered in the transcript (not cleared) - Verify a NEW
reasoning-blockappears for the second turn - Count the
[data-testid="reasoning-block"]elements; verify there are at least 2 after the second response - Send a third prompt; verify all three prior turns' reasoning blocks remain intact and render in chronological order
3. Error Handling
- Attempt to send an empty message; verify it is a no-op (no user bubble, no reasoning-block, no assistant response)
- Send a ~500-character prompt; verify the reasoning block and answer both render without horizontal scroll or layout break
- With the backend stopped, send a message; verify the UI surfaces a visible error rather than hanging silently
- Open DevTools → Console; verify no uncaught errors during any flow above
Expected Results
- Chat loads within 3 seconds; first reasoning token appears within 5s of sending
- The
reasoning-blockrenders BEFORE the final assistant text (or simultaneously) — never after - "Thinking…" label is visible during streaming and flips to "Agent reasoning" once
isRunningsettles - Prior turns' reasoning blocks persist across new messages (no clearing)
- No UI layout breaks, no flash of unstyled content, no uncaught console errors