Files
Alem Tuzlak 985bebf39c fix(showcase/voice): use real Whisper for mic transcription
The voice route's OpenAI client previously fell through to OPENAI_BASE_URL,
which docker-compose.local.yml sets to http://aimock:4010/v1. Aimock has a
catchall transcription fixture that returns "What is the weather in Tokyo?"
for every audio file, so the mic button always produced that phrase no
matter what the user actually said.

Pin baseURL to real OpenAI (overridable via OPENAI_TRANSCRIPTION_BASE_URL).
The sample-audio button stays as synchronous text injection — that's the
documented design, and what the e2e + d5 probe rely on.

Also:
- Tidy the sample button label ("Try a sample question" -> "Try a sample
  audio") so the affordance matches what it does.
- Realign tests/e2e/voice.spec.ts with the shipped component (the
  voice-sample-audio container testid and Sample: "..." caption it asserted
  on never existed on HEAD) and add cold-start timeout headroom for the
  mic-button render and the agent-flow test.
- Add "env": ".env" to langgraph.json so langgraph_cli dev picks up
  OPENAI_API_KEY locally. Docker/Railway paths inject env vars directly so
  this is a no-op there.
2026-05-11 14:45:10 +02:00

100 lines
4.0 KiB
TypeScript

import { test, expect } from "@playwright/test";
// E2E for the voice demo — sample-audio path only.
//
// The "Play sample" button is a deterministic test/demo affordance: it
// synchronously injects the canned phrase ("What is the weather in Tokyo?")
// into the chat composer without touching the runtime's `/transcribe`
// endpoint. That keeps this suite stable across environments where Whisper
// or aimock might be unavailable.
//
// The microphone path is intentionally out of scope: MediaRecorder is hard
// to exercise headlessly without mocking, and the mic is the only path that
// actually exercises real transcription. It's covered by the manual QA
// checklist at qa/voice.md.
//
// Stability expectation: 3 consecutive runs against Railway must pass.
test.describe("Voice Input", () => {
test.beforeEach(async ({ page }) => {
await page.goto("/demos/voice");
});
test("page loads with sample button, chat composer, and mic affordance", async ({
page,
}) => {
await expect(
page.getByRole("heading", { name: "Voice input" }),
).toBeVisible();
await expect(
page.locator('[data-testid="voice-sample-audio-button"]'),
).toBeEnabled();
await expect(page.getByText("Try a sample audio")).toBeVisible();
await expect(
page.locator('[data-testid="copilot-chat-input"]'),
).toBeVisible();
// The mic button is the authoritative signal that the runtime advertised
// `audioFileTranscriptionEnabled: true` — i.e. transcriptionService is
// wired on /api/copilotkit-voice. Exposed by react-core's v2 CopilotChatInput.
// It renders after the /info round trip resolves on the client, which on
// a cold dev server can exceed Playwright's 5s default — give it room.
await expect(
page.locator('[data-testid="copilot-start-transcribe-button"]'),
).toBeVisible({ timeout: 15_000 });
});
test("sample audio button injects the canned phrase into the input", async ({
page,
}) => {
const sampleButton = page.locator(
'[data-testid="voice-sample-audio-button"]',
);
const textarea = page.locator('[data-testid="copilot-chat-textarea"]');
await expect(sampleButton).toBeEnabled();
await expect(textarea).toHaveValue("");
await sampleButton.click();
// The button is synchronous — clicking immediately populates the
// textarea with the canned sample text. No transient "Transcribing…"
// state, no /transcribe round trip.
await expect(textarea).toHaveValue(/weather|tokyo/i, { timeout: 1000 });
await expect(sampleButton).toBeEnabled();
});
test("sending the transcribed text produces a weather tool render", async ({
page,
}) => {
// The end-to-end flow (click → run agent → first assistant chunk) can run
// up to ~50s on a cold langgraph dev server, so override the default 30s
// suite timeout to give the locator's own 45s timeout headroom.
test.setTimeout(90_000);
const sampleButton = page.locator(
'[data-testid="voice-sample-audio-button"]',
);
const textarea = page.locator('[data-testid="copilot-chat-textarea"]');
const sendButton = page.locator('[data-testid="copilot-send-button"]');
await sampleButton.click();
await expect(textarea).toHaveValue(/weather|tokyo/i, { timeout: 1000 });
await sendButton.click();
// The voice-demo route reuses the neutral sample_agent graph, which
// doesn't itself render a weather card — but if the runtime has a
// tool-rendering configuration that handles weather, one of these will
// be visible. The assertion is permissive: we care that *some*
// agent-authored response surface appeared, not exactly which renderer
// was used.
const assistantOrTool = page
.locator(
[
'[data-testid="weather-card"]',
'[data-testid="custom-catchall-card"][data-tool-name="get_weather"]',
'[data-testid="copilot-assistant-message"]',
].join(", "),
)
.first();
await expect(assistantOrTool).toBeVisible({ timeout: 45000 });
});
});