new LLMock({ fixtures }) stores options but never copies fixtures into the
server, so the mock served 0 fixtures and every agent LLM turn 404'd. Register
the loaded fixtures with server.addFixtures() and report the served count.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A fresh build of the Intelligence composite from current HEAD restricts memory
kind to {semantic, episodic, procedural} and rejects the legacy "operational"
kind the demo was written against (verified: operational -> HTTP 400, procedural
-> 201 against the freshly-built local composite). Migrate the over-limit
procedure write + agent prompts + inspector + unit test/e2e seed/fixture to the
project-scoped "procedural" kind so the demo works on a fresh local build.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
First real run of the e2e (browser was never installed) surfaced that both
tests assumed the chat was already open. The docked CopilotSidebar starts
closed, so getByRole('textbox') had nothing to fill. Open it via the 'Open chat'
launcher first, and target the input by its 'Type a message...' placeholder
(robust against the Memory tab's recall input). Selectors verified live against
the running app via Playwright.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Pre-existing: `next build` failed type-check on chat-inbox.tsx's
`style={{ "--inbox-width" }}` (not allowed by React.CSSProperties without a
cast). Unrelated to Glass Engine; surfaced while running the build gate.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The durable-memory demo depends on the Intelligence platform's recall_memory /
save_memory tools, which are served from `${apiUrl}/mcp` and attached to the
local BuiltInAgent run via MCP middleware ONLY when CopilotKitIntelligence is
constructed with `enableEnterpriseLearning: true` (gated in
attachIntelligenceEnterpriseLearning). The banking route never set the flag, so
in any real (non-aimock) run the memory tools never loaded: save_memory was a
no-op during teaching and recall_memory was absent, making the agent re-offer
workflow recording on every over-limit charge instead of recalling the learned
procedure.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The teach-flow's multi-step tool sequencing is the weak spot for the mini model;
the non-mini gpt-5.4 follows the recall->offer->demonstrate->save routing far
more consistently. openai/gpt-5.4 is the alias already used across the repo.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
My collapse fix branched on isRecording, but the render closure captures a stale
isRecording (false), so clicking 'Start recording' wrongly collapsed the card to
'Okay — not recording.'. Branch on the resolved `result` instead (fresh on
complete, like the save card): onDeny resolves 'declined' -> 'not recording';
onApprove resolves the directive -> 'Recording started'. Drops the now-unused
isRecording destructure.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The 'Record a workflow?' card appeared inconsistently because two intro prompt
lines still steered the agent the wrong way on an over-limit approve — 'review
... over-limit charges -> showPendingApprovals' and 'showApprovalFlow -> explain
how an over-limit charge gets cleared' — competing with the recall->offer block.
Scope both to browse/explain only and exclude them from the approve path, and
drop the agent temperature 0.3 -> 0 so it picks the same route every time.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Same lingering-button bug as the offer card: saveLearnedWorkflow and
awaitDashboardDemonstration only special-cased status 'inProgress', so after
clicking 'Save workflow' / 'I'm done' the card kept its button (confusing —
looked like nothing happened, though the save/respond went through). Add a
status 'complete' branch that collapses each to a static line, using the HITL
render's 'result' to show saved vs discarded / finished vs cancelled.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
On 'approve the over-limit charge' with no saved procedure, gpt-5.4-mini
sometimes called showApprovalFlow ('diagram of how to clear an over-limit
charge') or showPendingApprovals ('...including over-limit charges') instead of
recall_memory -> offerWorkflowRecording, so it never asked to record. Scope
those two tools to their real uses (explainer only when asked how it works;
queue only for reviewing pending) and explicitly exclude them from the
over-limit approve path. The approval-flow diagram stays as an informational
tool — just not a response to an approve request.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
After 'Start recording', offerWorkflowRecording resolved a bare 'started', so
gpt-5.4-mini often just SAID awaitDashboardDemonstration's 'go ahead and I'll
watch' line (from its description) instead of CALLING it — leaving the offer
card frozen with the Start-recording button and a confusing text reply, even
though recording (the vignette) had begun.
- Resolve a directive result telling the agent to immediately call
awaitDashboardDemonstration and not reply in prose (mirrors the working
awaitDashboardDemonstration -> saveLearnedWorkflow beat).
- Collapse the offer card to a static line once resolved (status 'complete')
so the button can't linger or be re-clicked; the live 'Recording your
workflow' card takes over.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Generalizes the row-action dropdown fix: every shadcn overlay (Select,
DropdownMenu, Popover, Tooltip) portals to <body> at z-50 and renders behind
the CopilotKit chat sidebar (z-1200) when opened in-chat — the policy-code
Select in the exception form hit this too. One globals.css rule lifts every
Radix popper wrapper to z-1300 (the file's existing 'above the chat panel'
value), fixing the whole class instead of per-component patches.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The 'More actions' menu portals to <body> at z-50, but the CopilotKit chat
sidebar is z-1200 — so the menu opened BEHIND the panel and was invisible
(clicking the three-dots appeared to do nothing). Lift this dropdown's content
to z-[1300] so it renders in front of the panel. Pairs with the earlier
pointer-events-auto fix; together the row's overflow menu works in-chat.
Verified: menu computed z-index is now 1300 vs panel 1200.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
showPendingApprovals is a display-only useComponent, which CopilotKit paints
pointer-events:none on the assistant message. The table is interactive
(Approve / Deny / File policy exception 'More actions' menu), so the inherited
none made every row action unclickable in the chat. Opt the card subtree back
into pointer events.
Verified live: the 'More actions' button's computed pointer-events flips
none -> auto with the fix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Re-running the over-limit teach demo mutates the in-memory store (approve +
filed exception flip the charge overLimit→cleared). This endpoint re-seeds the
store in place so the $5,000 Google Ads / Marketing charge (t-1) returns to
pending/over-limit without a server restart. Gated off when NODE_ENV=production.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Drops the non-existent AIMock/MockServer fallback names (TS-flagged once aimock was
installed); LLMock/loadFixtureFile/validateFixtures are the real exports.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
AUTHORED + statically validated (playwright --list compiles spec+config, fixtures/
package JSON valid, launcher syntax OK). NOT yet green-verified — needs aimock
installed (pnpm i), the docker memory stack up, and the dev server in Intelligence
mode (multi-process; not runnable in the current sandbox). Each file carries a
'VERIFY ON FIRST GREEN RUN' checklist for the shakedown.
- e2e/memory-learning.spec.ts: seeds the procedure via REST (recall-half isolation
per the plan), drives a fresh thread, asserts recall->unlock with no recording
offer + the over-limit gate lifted. Save half stays HITL+LLM (drift smoke / manual).
- e2e/fixtures/memory-learning.fixtures.json: pins recall_memory -> openPolicyException
-> finalizePolicyException -> approveTransaction.
- e2e/aimock-server.mjs: aimock launcher (programmatic, CLI fallback documented).
- playwright.config.ts: webServer array (aimock + dev in Intelligence mode, OPENAI_BASE_URL->aimock).
- package.json: +@copilotkit/aimock devDep; test:self-learning -> the spec.
- remove scripts/self-learning-smoke.mjs (dead #192 distill path).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- README: replace dead sl-worker/knowledge/annotate Phase-C section with the
memory runbook (docker compose one-command stack, host-TEI override for Apple
Silicon, 715x ports, .env, cross-thread+cross-persona FOR-149 walkthrough,
aimock E2E + drift-smoke testing notes)
- scripts/memory-drift-smoke.mjs: non-gating real-LLM tripwire that the live model
still emits recall_memory on a fresh-thread over-limit request (save half is
HITL-gated; covered by the manual walkthrough + aimock E2E)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Save action now resolves a string the agent reads as status:saved + an explicit
save_memory(scope:project, kind:operational) instruction with the demonstrated
code, preserving the already-approved/don't-re-run guard. respond() takes a
string in this file, so the result is delivered as text the recall-first prompt
(Task 3) acts on.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Surfaced while verifying Task 0 against a live stack:
- minio-init: retry mc alias set until Docker DNS resolves (idempotent bucket create)
- embedder pluggable: MEMORY_EMBEDDINGS_URL overridable + bundled tei dependency required:false
(point at host/native TEI on RAM-constrained Apple Silicon where the emulated tei OOMs)
- deps default host ports remapped to 715x so a bare `docker compose up` coexists with a dev's Intelligence stack
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Resolve the single conflict in packages/web-inspector/src/index.ts by keeping
both additions: this branch's CpkMemoryList memory-tab element and main's
ɵCpkThreadDetails back-compat alias (independent top-level declarations).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The chart components' descriptions steered the model to answer chart-shaped
questions BY rendering ("do NOT answer in plain text"), which it obeyed too
literally: "which policy is closest to its limit?" produced the right chart
and no answer. Every chart description now carries the same rule — the chart
replaces restating the raw numbers, not the answer itself; follow the render
with one or two grounded sentences. Pending-approvals card gets the matching
"point at what needs attention" line.
Verified live: the budget-usage pill now yields the chart plus "The Marketing
policy is closest to its limit…".
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Dismissal was persisted in sessionStorage, so after one engagement the
demo's opening beat — the copilot noticing breached charges before the
user types or clicks anything — never appeared again for the whole tab
session. Demo-wrong: every fresh load must open with it. Dismissal is now
per-page-load component state; within a load it still never re-nags (the
wrapper stays mounted across client-side navigation).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>