Commit Graph

21 Commits

Author SHA1 Message Date
Maxim c1819f4932 test(banking): correct OGUI double-pill rationale comment
The .first() guard on the 'rendered on the canvas' handoff-pill assertion
was justified by an inaccurate comment (accumulation across exchanges). The
real cause is intra-turn: generateSandboxedUi has followUp:true, so aimock
re-serves the same fixture on the unchanged-userMessage follow-up turn in
replay -> a second identical pill. A terminating sequenceIndex follow-up
fixture was attempted but destabilized the suite (title-generation requests
substring-match the pill text and consume the sequence counter before the
real leg-1 turn), so the .first() guard remains. Test/fixtures only; no
source changes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-09 20:34:29 +02:00
github-actions[bot] 87b6fa330d style: auto-fix formatting 2026-07-09 17:34:36 +00:00
Maxim cd7e49e467 test(banking): OGUI routing asserts the canvas surface 2026-07-09 19:27:57 +02:00
Maxim c54148412a test(banking): deterministic OGUI routing guard over the adjacency set
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-09 16:43:47 +02:00
Maxim 2dfedb0769 test(banking): fix stale smoke-test pill assertions 2026-07-09 02:54:43 +02:00
Maxim 3b1e04275e chore(banking): rename memory kinds to Intelligence main enum
Phase 2 of the banking->Intelligence-main migration. Main's memory lib
(libs/memory/src/types.ts) closes MemoryKind to topical|episodic|operational;
the demo was authored against the legacy semantic|procedural names, which the
backend now rejects/misfiles. Rename across the whole surface:
- agent prompt (route.ts CLASSIFY + SAVE-THE-PROCEDURE): semantic->topical,
  procedural->operational
- recorder instruction (copilot-context.tsx), learning-tab dual-read dropped,
  memory-tab KIND_COLORS, memory unit-test fixture
- smokes (facts + drift) and the e2e spec seed + fixtures comment

Only true kind: values renamed; "semantic recall"/"top-k semantic search"
mechanism descriptions left intact (recall is vector search regardless of enum).
aimock fixture re-record was a no-op: the one fixture pins the recall-and-apply
arc (no kind: values); the seed is REST-side in the spec.

Verified: pnpm test:unit 47/47, tsc --noEmit clean, eslint clean on touched files.
Local-only until the full migration is verified against the main stack.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 16:20:00 +02:00
Maxim 84095d63ec test(banking): add /mcp readiness gate to memory smokes; document backend boot-window flake 2026-07-03 17:17:50 +02:00
Maxim 7257068466 feat(showcase): add A2UI Transactions catalog node with status filter
Replace the catalog's fixed PendingTable node with a general Transactions node
that takes a status filter (all | pending | approved | denied, default all);
the agent picks the slice when composing a report and the client binds live
data via useReportData(). render_report's `pendingTable: boolean` param becomes
`transactions: <status>` (presence includes the table). The shared
TransactionsList component and the chat's showPendingApprovals flow are
unchanged — only the A2UI catalog node is generalized.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:41:47 +02:00
Maxim cb50f777ef style(showcase): satisfy lint (no-explicit-any, no-children-prop, react-compiler) + prettier
Post-gate cleanup on the A2UI re-architecture: type the op-builder test helpers
instead of `as any`, disable react/no-children-prop on the RendererProps
render-callback in the StatCard test, drop the manual useMemo in useReportSurface
(the React Compiler can't preserve it; downstream consumers guard on values), and
apply prettier.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:41:46 +02:00
Maxim 128e0894aa test(showcase): rewrite A2UI canvas e2e for the render_report flow
The spec described the abandoned render_a2ui/mirror/injectA2UITool:true path.
Update its docblock + fixture to the render_report backend-tool flow
(injectA2UITool:false, ops detected from the tool result, canvas reads the
agent message stream). Still test.fixme — the aimock fixture isn't wired into
aimock-server.mjs and no headless green run is confirmed; the live path is
verified manually.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:41:46 +02:00
Maxim e17a389135 test(showcase): A2UI canvas e2e + surface testid
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:41:44 +02:00
github-actions[bot] c76a7940b1 style: auto-fix formatting 2026-07-02 19:37:34 +02:00
Maxim 1bf8c44c5d fix(showcase): load aimock e2e fixtures via addFixtures (LLMock ignored options.fixtures)
new LLMock({ fixtures }) stores options but never copies fixtures into the
server, so the mock served 0 fixtures and every agent LLM turn 404'd. Register
the loaded fixtures with server.addFixtures() and report the served count.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:37:33 +02:00
Maxim f1662bc6c8 fix(showcase): migrate durable memory kind operational -> procedural
A fresh build of the Intelligence composite from current HEAD restricts memory
kind to {semantic, episodic, procedural} and rejects the legacy "operational"
kind the demo was written against (verified: operational -> HTTP 400, procedural
-> 201 against the freshly-built local composite). Migrate the over-limit
procedure write + agent prompts + inspector + unit test/e2e seed/fixture to the
project-scoped "procedural" kind so the demo works on a fresh local build.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:37:33 +02:00
Maxim f807238250 test(showcase): open the docked chat + target input by placeholder in memory e2e
First real run of the e2e (browser was never installed) surfaced that both
tests assumed the chat was already open. The docked CopilotSidebar starts
closed, so getByRole('textbox') had nothing to fill. Open it via the 'Open chat'
launcher first, and target the input by its 'Type a message...' placeholder
(robust against the Memory tab's recall input). Selectors verified live against
the running app via Playwright.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:37:32 +02:00
github-actions[bot] 5f77e8696a style: auto-fix formatting 2026-07-02 19:37:31 +02:00
Maxim 5072792244 test(showcase): e2e — Glass Engine inspector (gate set, streams events + shows learned procedure)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:37:30 +02:00
Maxim d033b02837 test(showcase): use confirmed @copilotkit/aimock export (LLMock) in E2E launcher
Drops the non-existent AIMock/MockServer fallback names (TS-flagged once aimock was
installed); LLMock/loadFixtureFile/validateFixtures are the real exports.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:33 +02:00
Maxim cdcd5460d6 test(showcase): deterministic aimock+Playwright cross-thread memory proof (FOR-149)
AUTHORED + statically validated (playwright --list compiles spec+config, fixtures/
package JSON valid, launcher syntax OK). NOT yet green-verified — needs aimock
installed (pnpm i), the docker memory stack up, and the dev server in Intelligence
mode (multi-process; not runnable in the current sandbox). Each file carries a
'VERIFY ON FIRST GREEN RUN' checklist for the shakedown.

- e2e/memory-learning.spec.ts: seeds the procedure via REST (recall-half isolation
  per the plan), drives a fresh thread, asserts recall->unlock with no recording
  offer + the over-limit gate lifted. Save half stays HITL+LLM (drift smoke / manual).
- e2e/fixtures/memory-learning.fixtures.json: pins recall_memory -> openPolicyException
  -> finalizePolicyException -> approveTransaction.
- e2e/aimock-server.mjs: aimock launcher (programmatic, CLI fallback documented).
- playwright.config.ts: webServer array (aimock + dev in Intelligence mode, OPENAI_BASE_URL->aimock).
- package.json: +@copilotkit/aimock devDep; test:self-learning -> the spec.
- remove scripts/self-learning-smoke.mjs (dead #192 distill path).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:32 +02:00
Maxim 5263ac2e59 test(saas-demo): fail smoke test on hydration + uncaught errors
Remove the /Hydration/i pattern from the ignored-console-errors filter so
Next.js hydration mismatches surface as real failures, and add a
page.on("pageerror", ...) listener that records uncaught exceptions and
asserts none occurred. pageerrors are not filtered — any uncaught
exception fails the smoke test.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-03 11:54:02 +02:00
Maxim 8efd6ecb33 test(saas-demo): Playwright smoke test for golden path
Adds a CI-safe Playwright smoke test that verifies the banking showcase
boots and the CopilotKit v2 popup opens with its configured suggestion
pills, without invoking the agent (so it runs in CI without secrets).

Covers:
  - Document title shows the Northwind Finance brand
  - Credit-cards dashboard renders ("Credit Cards" heading)
  - CopilotPopup launcher opens the dialog (Northwind Copilot)
  - Three suggestion pills render: View transactions, Add a card,
    Assign a policy

The dev server gets a dummy OPENAI_API_KEY so the runtime route boots,
but the test never sends a chat message or clicks a suggestion.

.gitignore: ignore test-results/, playwright-report/, and
tsconfig.tsbuildinfo so they don't become clutter.
2026-06-03 11:27:25 +02:00