mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
728ed61ce8
The langgraph-python voice cell sat at D4 even when its d5-voice probe row was green. Root cause: the dashboard's CATALOG_TO_D5_KEY mirror in showcase/shell-dashboard/src/lib/live-status.ts was missing voice -> ["voice"], so computeMaxPossible capped voice at D4 regardless of probe state. The harness REGISTRY_TO_D5 already had the entry; only the dashboard mirror was out of sync. Separately, the "Play sample" button used to fetch sample.wav and POST it to /transcribe. With aimock that meant both the sample button AND the mic returned the same canned response, which made it impossible to demo the mic path locally without conflating the two affordances. Reworked the button into a synchronous static-text injector (onTranscribed(sampleText)) so: - Sample button = deterministic test/demo affordance, no runtime calls. - Mic = real Whisper transcription via /transcribe. Synced across all 18 voice-enabled integrations. Phrase stays "What is the weather in Tokyo?" so aimock's "weather in Tokyo" substring fixture still matches. Also adds the missing d5-voice.test.ts companion (every other d5-* probe script has one) and trims the langgraph-python qa/voice.md + e2e steps that depended on the now-removed async behavior.
99 lines
3.7 KiB
TypeScript
99 lines
3.7 KiB
TypeScript
import { test, expect } from "@playwright/test";
|
|
|
|
// E2E for the voice demo — sample-audio path only.
|
|
//
|
|
// The "Play sample" button is a deterministic test/demo affordance: it
|
|
// synchronously injects the canned phrase ("What is the weather in Tokyo?")
|
|
// into the chat composer without touching the runtime's `/transcribe`
|
|
// endpoint. That keeps this suite stable across environments where Whisper
|
|
// or aimock might be unavailable.
|
|
//
|
|
// The microphone path is intentionally out of scope: MediaRecorder is hard
|
|
// to exercise headlessly without mocking, and the mic is the only path that
|
|
// actually exercises real transcription. It's covered by the manual QA
|
|
// checklist at qa/voice.md.
|
|
//
|
|
// Stability expectation: 3 consecutive runs against Railway must pass.
|
|
|
|
test.describe("Voice Input", () => {
|
|
test.beforeEach(async ({ page }) => {
|
|
await page.goto("/demos/voice");
|
|
});
|
|
|
|
test("page loads with sample button, chat composer, and mic affordance", async ({
|
|
page,
|
|
}) => {
|
|
await expect(
|
|
page.getByRole("heading", { name: "Voice input" }),
|
|
).toBeVisible();
|
|
await expect(
|
|
page.locator('[data-testid="voice-sample-audio"]'),
|
|
).toBeVisible();
|
|
await expect(
|
|
page.locator('[data-testid="voice-sample-audio-button"]'),
|
|
).toBeEnabled();
|
|
await expect(
|
|
page.getByText('Sample: "What is the weather in Tokyo?"'),
|
|
).toBeVisible();
|
|
await expect(
|
|
page.locator('[data-testid="copilot-chat-input"]'),
|
|
).toBeVisible();
|
|
// The mic button is the authoritative signal that the runtime advertised
|
|
// `audioFileTranscriptionEnabled: true` — i.e. transcriptionService is
|
|
// wired on /api/copilotkit-voice. Exposed by react-core's v2 CopilotChatInput.
|
|
await expect(
|
|
page.locator('[data-testid="copilot-start-transcribe-button"]'),
|
|
).toBeVisible();
|
|
});
|
|
|
|
test("sample audio button injects the canned phrase into the input", async ({
|
|
page,
|
|
}) => {
|
|
const sampleButton = page.locator(
|
|
'[data-testid="voice-sample-audio-button"]',
|
|
);
|
|
const textarea = page.locator('[data-testid="copilot-chat-textarea"]');
|
|
|
|
await expect(sampleButton).toBeEnabled();
|
|
await expect(textarea).toHaveValue("");
|
|
await sampleButton.click();
|
|
|
|
// The button is synchronous — clicking immediately populates the
|
|
// textarea with the canned sample text. No transient "Transcribing…"
|
|
// state, no /transcribe round trip.
|
|
await expect(textarea).toHaveValue(/weather|tokyo/i, { timeout: 1000 });
|
|
await expect(sampleButton).toBeEnabled();
|
|
});
|
|
|
|
test("sending the transcribed text produces a weather tool render", async ({
|
|
page,
|
|
}) => {
|
|
const sampleButton = page.locator(
|
|
'[data-testid="voice-sample-audio-button"]',
|
|
);
|
|
const textarea = page.locator('[data-testid="copilot-chat-textarea"]');
|
|
const sendButton = page.locator('[data-testid="copilot-send-button"]');
|
|
|
|
await sampleButton.click();
|
|
await expect(textarea).toHaveValue(/weather|tokyo/i, { timeout: 1000 });
|
|
await sendButton.click();
|
|
|
|
// The voice-demo route reuses the neutral sample_agent graph, which
|
|
// doesn't itself render a weather card — but if the runtime has a
|
|
// tool-rendering configuration that handles weather, one of these will
|
|
// be visible. The assertion is permissive: we care that *some*
|
|
// agent-authored response surface appeared, not exactly which renderer
|
|
// was used.
|
|
const assistantOrTool = page
|
|
.locator(
|
|
[
|
|
'[data-testid="weather-card"]',
|
|
'[data-testid="custom-catchall-card"][data-tool-name="get_weather"]',
|
|
'[data-role="assistant"]',
|
|
].join(", "),
|
|
)
|
|
.first();
|
|
await expect(assistantOrTool).toBeVisible({ timeout: 45000 });
|
|
});
|
|
});
|