mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
c2bc645228
Investigation of the remaining bucket-C D4 cells uncovered that the 'recorded' fixtures shipped in PR #4725 are at best redundant with hand-curated canonical fixtures already in d5-all.json, and at worst override them with worse data — the recordings collapsed beautiful-chat's canonical $349/$289 search_flights payload to the LLM-of-the-day's $319/$289 (probe asserts $349 verbatim) and reasoning-custom's canonical fixture (which carries a 'reasoning' field) with a content-only recording (refused chain-of-thought disclosure). Net D5 effect of removing the recordings: - d5:langgraph-python/gen-ui-custom (catalog gen-ui-tool-based) → green (canonical d5-all.json #39/#40) - d5:langgraph-python/reasoning-display (catalog reasoning-custom) → green (canonical d5-all.json #91 with 'reasoning' field that cleanly satisfies the probe's reasoning-role testid contract) Both flips confirmed live against PocketBase via --d5 --live. The aimock record/replay infrastructure introduced alongside these fixtures (showcase/docker-compose.{record,replay}.yml, showcase/scripts/record-d5-fixtures.mjs) stays — it is reusable for future demos where canonical fixtures don't yet cover the prompts — just without the misleading initial recordings. Bucket-C cells that remain at D4 after this change require fixes outside the fixture/recording scope; see the PR description for per-cell findings (release-blocked package testids, frontend-tool dispatch defect, suggestion-bar misrender, resume-from-interrupt UI, built-in reasoning-message testid).