Commit Graph

230 Commits

Author SHA1 Message Date
Maxim c03c064371 feat(showcase): dev-only POST /api/v1/dev/reset to re-seed the demo store
Re-running the over-limit teach demo mutates the in-memory store (approve +
filed exception flip the charge overLimit→cleared). This endpoint re-seeds the
store in place so the $5,000 Google Ads / Marketing charge (t-1) returns to
pending/over-limit without a server restart. Gated off when NODE_ENV=production.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:34 +02:00
Maxim d033b02837 test(showcase): use confirmed @copilotkit/aimock export (LLMock) in E2E launcher
Drops the non-existent AIMock/MockServer fallback names (TS-flagged once aimock was
installed); LLMock/loadFixtureFile/validateFixtures are the real exports.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:33 +02:00
Maxim cdcd5460d6 test(showcase): deterministic aimock+Playwright cross-thread memory proof (FOR-149)
AUTHORED + statically validated (playwright --list compiles spec+config, fixtures/
package JSON valid, launcher syntax OK). NOT yet green-verified — needs aimock
installed (pnpm i), the docker memory stack up, and the dev server in Intelligence
mode (multi-process; not runnable in the current sandbox). Each file carries a
'VERIFY ON FIRST GREEN RUN' checklist for the shakedown.

- e2e/memory-learning.spec.ts: seeds the procedure via REST (recall-half isolation
  per the plan), drives a fresh thread, asserts recall->unlock with no recording
  offer + the over-limit gate lifted. Save half stays HITL+LLM (drift smoke / manual).
- e2e/fixtures/memory-learning.fixtures.json: pins recall_memory -> openPolicyException
  -> finalizePolicyException -> approveTransaction.
- e2e/aimock-server.mjs: aimock launcher (programmatic, CLI fallback documented).
- playwright.config.ts: webServer array (aimock + dev in Intelligence mode, OPENAI_BASE_URL->aimock).
- package.json: +@copilotkit/aimock devDep; test:self-learning -> the spec.
- remove scripts/self-learning-smoke.mjs (dead #192 distill path).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:32 +02:00
Maxim 177a68f7d4 docs(showcase): memory-based durable-learning runbook + real-LLM recall drift smoke
- README: replace dead sl-worker/knowledge/annotate Phase-C section with the
  memory runbook (docker compose one-command stack, host-TEI override for Apple
  Silicon, 715x ports, .env, cross-thread+cross-persona FOR-149 walkthrough,
  aimock E2E + drift-smoke testing notes)
- scripts/memory-drift-smoke.mjs: non-gating real-LLM tripwire that the live model
  still emits recall_memory on a fresh-thread over-limit request (save half is
  HITL-gated; covered by the manual walkthrough + aimock E2E)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:30 +02:00
Maxim 841f6db2fb feat(showcase): saveLearnedWorkflow resolves a save-triggering result (drives save_memory)
Save action now resolves a string the agent reads as status:saved + an explicit
save_memory(scope:project, kind:operational) instruction with the demonstrated
code, preserving the already-approved/don't-re-run guard. respond() takes a
string in this file, so the result is delivered as text the recall-first prompt
(Task 3) acts on.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:29 +02:00
Maxim bf8a01b250 feat(showcase): recall-first over-limit handling + save learned procedure to project memory
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:27 +02:00
Maxim 66cccfb8a6 refactor(showcase): remove POC /annotate recording seam (superseded by save_memory)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:25 +02:00
Maxim 7dcb5dd4d5 feat(showcase): wire banking runtime for memory (license token + locks); full .env.example
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:24 +02:00
Maxim 42a4066d45 fix(showcase): harden banking memory stack (minio-init DNS retry, pluggable embedder, non-colliding host ports)
Surfaced while verifying Task 0 against a live stack:
- minio-init: retry mc alias set until Docker DNS resolves (idempotent bucket create)
- embedder pluggable: MEMORY_EMBEDDINGS_URL overridable + bundled tei dependency required:false
  (point at host/native TEI on RAM-constrained Apple Silicon where the emulated tei OOMs)
- deps default host ports remapped to 715x so a bare `docker compose up` coexists with a dev's Intelligence stack

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:24 +02:00
Maxim f447e1a624 feat(showcase): vendor memory-enabled Intelligence stack (cloned memory-chat recipe) for banking demo
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 19:33:23 +02:00
David McKay f13776560b fix(showcase): chart gen-UI answers the question, not just renders the chart
The chart components' descriptions steered the model to answer chart-shaped
questions BY rendering ("do NOT answer in plain text"), which it obeyed too
literally: "which policy is closest to its limit?" produced the right chart
and no answer. Every chart description now carries the same rule — the chart
replaces restating the raw numbers, not the answer itself; follow the render
with one or two grounded sentences. Pending-approvals card gets the matching
"point at what needs attention" line.

Verified live: the budget-usage pill now yields the chart plus "The Marketing
policy is closest to its limit…".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-01 19:19:08 -05:00
David McKay d80aeb2e5e fix(showcase): proactive notice fires on every fresh page load
Dismissal was persisted in sessionStorage, so after one engagement the
demo's opening beat — the copilot noticing breached charges before the
user types or clicks anything — never appeared again for the whole tab
session. Demo-wrong: every fresh load must open with it. Dismissal is now
per-page-load component state; within a load it still never re-nags (the
wrapper stays mounted across client-side navigation).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 01:34:51 +02:00
David McKay 0ef4f75554 feat(showcase): clickable in-chat approvals, always-on use-case pills, interactive trend chart
Click-around feedback fixes on the banking demo:

- In-chat pending approvals are now actually usable: the dashboard's 4-column
  approval table is ~550px wide, so inside the chat card its Actions column
  rendered past the card edge — visible but unclickable. New
  PendingApprovalsChat stacks each charge as a chat-width card with labeled
  Approve/Deny/File-exception actions, with exact behavioral parity
  (over-limit gating, PolicyExceptionInline form, identical teach-mode
  recordUserAction payloads and logStep narration).

- Suggestion bubbles never disappear: the full use-case catalog (12 pills —
  the 4 self-learning arc beats plus charts, breakdowns, cash flow, approvals
  explainer, report prep, PIN change, team invite) is registered with
  available:"always", so the demo stays fully click-drivable after every
  exchange. Panel widened 440→560px so pills flow two-per-row instead of
  stacking.

- StatisticsChart (dashboard rail + spending-trend gen-UI) is a real chart
  now: y-axis dollar gridlines, x-axis labels aligned to the plot area, and
  pointer hover showing the exact month + amount with a guide line and
  highlighted point. Still hand-rolled SVG, no charting dependency.

Validated with Playwright: approvals filed and approved entirely inside the
chat (Cleared state transition + queue drains), pills persist after
exchanges, tooltip renders on hover.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 01:34:51 +02:00
David McKay 4f26181659 feat(showcase): banking wow moments — proactive copilot, chart pills, report artifacts
Make the banking demo pass the click-around test: every moment reachable by
clicking, zero typing, and the dead Analytics/Reports tabs replaced with real
views (FOR-190; claims FOR-178).

Additive modules under src/components/wow/:
- ProactiveNotice: on load, the copilot surfaces overnight policy breaches
  unprompted and offers to walk through them (tap-to-accept -> existing
  pending-approvals gen-UI). Session-scoped dismissal.
- ChartCard + AnalyticsView: the four brand charts (previously chat-only)
  rendered as the dashboard's Analytics tab, each carrying contextual
  conversation-starter pills grounded via the existing useAgentContext data.
- createReport tool + ReportsView: "prep the Q2 spend report" files a durable
  artifact (summary + highlights + live charts) in the Reports tab through a
  new /api/v1/reports REST route and store collection.
- useAskCopilot: shared open-panel + addMessage + runAgent helper (same path
  a suggestion-pill click takes).

Correctness fixes found by a full Playwright click-through sweep:
- navigateToPageAndPerform: full-page reload tore down the chat panel mid-run
  (conversation and in-flight operation lost); now client-side router.push.
  Card operations also targeted /cards, which mirrors the dashboard and has no
  card tools — they now land on / where the tools are registered.
- chat-inbox: CSS custom property cast so `next build` compiles again
  (production build was failing on main, blocking deploy).

Validated end-to-end with Playwright against a live agent, including the full
teach -> demonstrate -> save -> recall arc in OSS mode.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-02 01:34:51 +02:00
Mark 0addf1a3bd Merge branch 'main' into feat/oracle-langgraph-checkpointer 2026-06-30 01:07:51 -07:00
GeneralJerel 698f739e8a fix(examples): repair dangling HITL tool-calls + HTML-unescape persisted memory in oracle showcase
Ports two fixes from the canonical oracle-cookbook demo into the oracle-agent-memory
showcase agent (server.py was byte-identical to the demo's pre-fix version):

1. Dangling tool-calls: booking conversationally calls the book_flight HITL tool,
   which interrupts and emits an assistant tool_call awaiting the UI's Confirm/Cancel.
   If the traveler sends another chat message instead, the unanswered tool_call made
   the next turn 400 ("tool_call_ids did not have response messages"). _repair_dangling_tool_calls
   synthesizes a "not completed" tool result for any dangling call before the history
   reaches the graph. (Inverse of the duplicate-tool-block issue the history-replace
   already handled — documented in docs/known-issues/agentspec-multiturn-toolcall-correlation.md.)

2. HTML-escaped persistence: the agentspec exporter HTML-escapes streamed deltas
   (& < > -> &amp; &lt; &gt;); the SSE generator persisted them raw, so assistant
   replies were stored in Oracle Agent Memory as e.g. "fares &lt; $700".
   _clean_assistant_text html.unescapes the assembled text before persisting.

Both are reproduced + verified in oracle-cookbook (unit tests + in-browser); the ported
functions are byte-identical to the verified demo code.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 11:00:37 -07:00
GeneralJerel fac83b0ca6 feat(examples): add optional Oracle LangGraph checkpointer to oracle-agent-memory showcase
Flag-gated (LANGGRAPH_CHECKPOINTER=oracle, default memory) AsyncOracleSaver from
langgraph-oracledb for durable per-thread LangGraph graph state in Oracle,
complementing oracleagentmemory. Default-safe (in-memory fallback). Mirrors
jerelvelarde/oracle-cookbook#4. Draft, stacked on #5563.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 06:51:10 -07:00
GeneralJerel 2d5a6fe6aa Merge upstream/main into showcase/oracle-agent-memory (refresh for review) 2026-06-29 06:40:39 -07:00
GeneralJerel 54920c65be fix(showcase/oracle-agent-memory): background-persist lag + inline client-side booking
Sync the three travel-concierge UX fixes from the source repo
(jerelvelarde/oracle-cookbook#3) into the showcase:

- agent/concierge/server.py: persist memory in a serialized background task so
  the SSE stream closes at RUN_FINISHED instead of blocking ~2-13s on memory
  extraction + reconciliation (the post-generation loading lag). Adds a FastAPI
  lifespan drain + background-task failure surfacing.
- frontend/.../FlightOptions.tsx: "Select this flight" now drives confirm -> book
  entirely client-side, so the booking confirm card renders inline in view
  instead of off-screen via a fresh agent turn. The conversational book_flight
  HITL path is unchanged.
- frontend/e2e: add a covering test for the inline booking flow; rename
  sendAndPersist -> sendAndAwaitRun (stream-close is no longer a "memory written"
  signal now that persistence is backgrounded).

oxfmt --check + tsc --noEmit clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 15:53:32 -07:00
Tyler Slaton 5f9c2404d1 chore: format banking showcase 2026-06-25 15:35:45 -07:00
GeneralJerel 9ae901ed0a Merge remote-tracking branch 'upstream/main' into jerel/saas-banking-demo
# Conflicts:
#	pnpm-lock.yaml
2026-06-24 18:23:24 -07:00
Jerel Velarde 670bb17591 Merge branch 'main' into showcase/oracle-agent-memory 2026-06-24 23:03:09 +08:00
Alem Tuzlak 5ecdee36b8 feat(bot): pluggable StateStore persistence + cross-platform transcripts
Adds a durable persistence layer for @copilotkit/bot, replacing the
in-memory-only ActionStore with a pluggable StateStore.

- StateStore interface (kv/list/lock/dedup/queue) with a shared
  conformance suite; MemoryStore default plus @copilotkit/bot-store-redis
  and @copilotkit/bot-store-postgres backends.
- createBot({ store }): typed per-thread state via Standard Schema,
  action snapshots persisted through the store, per-conversation turn
  lock (onLockConflict drop|force), and inbound-event dedup keyed on a
  stable eventId. ActionStore is kept as a deprecated alias.
- Cross-platform transcripts (bot.transcripts + identity resolver) with
  age-bounded retention (prune on append + filter on read), and
  runAgent({ transcript: true }) to auto-inject history and capture the
  reply.
- createBot({ components }) re-registers components so durable actions
  re-fire after a restart; restart-durability demo in examples/slack.
- Dedup is marked seen only after the turn lock is acquired, so a turn
  dropped on lock-conflict does not burn its eventId (no lost retries).
- Release lockstep: bot-store-redis/postgres version with bot + bot-ui.
2026-06-23 18:33:38 +02:00
github-actions[bot] 790dfd0a75 style: auto-fix formatting 2026-06-19 17:33:38 -07:00
Tyler Slaton db2fd6539b fix: address merge conflicts and run formatter 2026-06-19 15:51:03 -07:00
GeneralJerel 9023efb266 chore(oracle-agent-memory): drop committed package-lock.json
The standalone showcase installs with `npm install` (Nixpacks does this
automatically when there's no lockfile; the README already uses npm install).
Removing the 16.6k-line lockfile also clears the fork-pr-monitor security
alert, which flagged it as a "large added file." Verified: a fresh
`npm install` + `next build` succeeds with no committed lockfile.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 13:25:48 -07:00
GeneralJerel be7d674964 chore(oracle-agent-memory): genericize db image namespace
build-and-push.sh defaulted IMAGE to a personal GHCR namespace; use an
OWNER placeholder (still IMAGE-overridable) so the showcase ships no
personal/private registry path. Resolves the pre-merge follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 18:52:04 -07:00
GeneralJerel 1b9ab035ce fix(oracle-agent-memory): ASSM tablespace (ORA-43853) + thread auto-titling
Synced from the cookbook dev repo (verified live on Railway + Playwright):

- db/init/01-create-user.sql: create the `cookbook` user on a dedicated
  ASSM tablespace (cookbook_ts) as its DEFAULT, instead of inheriting
  SYSTEM/MSSM (FREEPDB1's default on the Free :latest-lite image). JSON and
  VECTOR columns are SecureFile-backed and cannot live in an MSSM
  tablespace, so the old user failed schema creation with ORA-43853 and
  memory never persisted. Self-heals a user stranded on SYSTEM
  (drop+recreate); leaves a working ASSM user (local :latest USERS) alone.

- frontend: a new thread is named from its OWN first message, not the
  previous thread's. ThreadTitler now titles off the agent's message/run
  events gated on agent.threadId === activeThreadId, replacing a useEffect
  keyed on activeThreadId that read the shared, agentId-scoped transcript
  while it still held the prior thread's messages during a switch. Adds a
  Playwright regression test (e2e/thread-title.spec.ts).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 17:50:12 -07:00
Jordan Ritter 9aac779c51 fix(generative-ui-playground): use valid model id and externalize @copilotkit/runtime
The opengenui route used the nonexistent model id "openai/gpt-5.2", causing
every chat turn to error. Switch to "openai/gpt-4o", the verified-real id used
by the sibling showcase/shell copilotkit route.

The example's API routes import @copilotkit/runtime/v2 (a server-only package),
but next.config.ts lacked serverExternalPackages, so Next.js attempted to bundle
it. Add serverExternalPackages: ["@copilotkit/runtime"] (base package name covers
the /v2 subpath), mirroring showcase/shell/next.config.ts.
2026-06-18 16:48:32 -07:00
Jordan Ritter 1398fe225c chore: migrate @copilotkitnext usages to @copilotkit/*/v2 entrypoints 2026-06-18 16:37:58 -07:00
GeneralJerel 2d7926f867 feat(examples): add oracle-agent-memory showcase
Runnable companion to the "Build an Agentic Travel App with Oracle Agent
Memory, Agent Spec, and CopilotKit" cookbook recipe: a portable Oracle Agent
Spec agent on LangGraph over AG-UI, long-term memory on Oracle AI Database,
and a CopilotKit V2 frontend with generative UI + human-in-the-loop.

- agent/    Python (uv) Agent Spec agent + FastAPI AG-UI server
- frontend/ Next.js CopilotKit V2 chat (flight cards, recall chip, HITL booking)
- db/       Oracle AI Database (Free image) + cookbook-user init
- docker-compose.yml for local Oracle; per-service railway.json for deploy

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 15:35:58 -07:00
GeneralJerel 89dbb562f7 feat(showcase): summon pending approvals + analytics charts in banking chat
Add display-only generative-UI the copilot renders on request:
- showPendingApprovals — the dashboard's interactive approval table, in chat
- showSpendingTrend, showBudgetUsage, showSpendBreakdown, showIncomeVsExpenses —
  brand-styled charts (dependency-free, hand-rolled SVG/CSS)
- showApprovalFlow — a diagram of how an over-limit charge gets cleared

Registered globally in copilot-context (so they work on any route) and routed
via the agent prompt.
2026-06-17 08:37:58 -07:00
GeneralJerel be0e0b60f9 fix(showcase): banking cards scroll instead of squishing; no clipped shadows
- Cards page: the credit-card grid compressed each card (wrapping the holder /
  valid-thru text) when the chat panel narrowed the content area. It's now a
  horizontal-scroll row — each card holds a 300px minimum and grows to fill on
  wide screens, scrolling instead of squishing.
- overflow-x:auto forces overflow-y:auto, which clipped the cards' soft drop
  shadows; matching negative-margin/padding pairs give the shadows room inside
  the scroll viewport without moving the cards.
- Chat: space stacked HITL/action cards so consecutive cards no longer share an
  edge (CopilotKit stacks them with no row-gap).
2026-06-17 08:37:45 -07:00
Thierry Damiba 8143df4da8 docs(cookbook): add Arcade authenticated-tools recipe and showcase
Adds a cookbook recipe + runnable showcase that gives the Built-in Agent
OAuth-backed Arcade tools (Gmail, Google News) and renders Arcade's one-time
authorization step as a generative-UI "Connect" card in the chat.

- docs: showcase/shell-docs/src/content/docs/cookbook/arcade.mdx (+ meta.json, index card)
- app: examples/showcases/arcade-tools (Next.js App Router, single-route runtime)

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-16 23:37:48 -07:00
GeneralJerel d4fa49eed4 fix(showcase): refresh every view after a chat-driven mutation
useCreditCards() kept per-instance React state and was called independently by
the dashboard page and by copilot-context (where the chat's approve / finalize /
open-exception tools live). A mutation made through one instance refetched only
itself — so when the agent approved an over-limit charge in chat, the dashboard
pending table kept showing it as pending until a manual reload.

Add a module-level revalidation bus: each useCreditCards() instance registers a
refetch callback, and every mutation calls notifyDataChanged() to fan a refetch
out to all live instances. The dashboard now reflects agent-driven approvals
immediately (verified: recall-approve in chat drops the charge from the pending
table with no reload).
2026-06-15 12:19:52 -07:00
GeneralJerel 0b65391501 docs(showcase): clarify same-thread (OSS) vs cross-thread (Intelligence) recall
Adds a "What each mode actually recalls" subsection to the demo README: in OSS
mode the taught workflow is recalled only within the same conversation (the
saved procedure is echoed back into that thread), so a brand-new chat won't know
it — that's expected. Cross-conversation persistence is what the external
Intelligence backend provides. Names the symptom explicitly so reviewers aren't
surprised when a new conversation "doesn't know" the workflow in OSS mode.
2026-06-15 12:04:44 -07:00
GeneralJerel 7e1641ff1b feat(showcase): pending-approval table redesign + approval-gate UX fixes
Rework the dashboard's Pending approval view and fix two approval-gate UX bugs
on the banking demo (PR #5266):

- Table layout: replace the center-stacked per-row card with a scannable table
  (Merchant / Amount / Policy / Actions). Actions are check / x icon buttons plus
  a "more actions" overflow menu holding File policy exception. Status is its own
  column (Over limit / Cleared / Within limit), and Approve is gated until the
  charge is actually clearable.

- Fix the table shrinking when the more-actions menu opens: the Radix menu is
  modal by default and engaged react-remove-scroll, whose scrollbar compensation
  reflowed the table. Set modal={false} (a row menu needn't be modal) and add
  whitespace-nowrap to the status badges.

- Fix "cannot approve after filing an exception": the inline card offered all
  codes, including non-justifying ones that set activeExceptionId (flipping the
  row to Cleared) but never lift the server gate, so the approve 422'd silently.
  The card exists to clear an over-limit charge, so it now offers only justifying
  codes. The gate itself is unchanged.

Verified live in OSS dev: file (justifying) -> Cleared -> approve succeeds; the
menu opens without reflow; lint + build green.
2026-06-15 11:55:21 -07:00
GeneralJerel 4dfa6cbe90 refactor(showcase): render the workflow recorder inside the chat card
The recorder was a floating HUD overlay, which read as a different UI element
from the in-chat cards (e.g. "Save this workflow?"). Fold the live step feed
into the demonstration card itself: the `awaitDashboardDemonstration` card is
now titled "Recording your workflow" (with a REC pulse), keeps the
non-directional copy, and embeds a reactive `<RecordingSteps/>` child that
narrates each captured action inside the card. Removes the floating panel and
its CSS. Same card chrome as the other cards, so it reads consistently in the
conversation.

`RecordingSteps` subscribes to the recording context itself (not via the host
card's render closure), so it updates live as each step is logged without a
stale-closure dep.
2026-06-15 10:29:15 -07:00
GeneralJerel 3107d3c7c5 feat(showcase): live recorder HUD + non-directional teach copy; fix save-card stall
Address review feedback on the banking self-learning "teach a workflow" loop
(PR #5266):

- Recorder HUD: a floating "Recording your workflow" panel narrates each officer
  action live (Opened Dashboard -> Opened Transactions -> Opened Pending approval
  -> Opened the exception form -> Filed the policy exception -> Approved the
  charge), driven by logStep on the nav / tab / file-exception / approve call
  sites. New recording-feed.tsx; steps + logStep added to recording-context.

- Non-directional demonstration: the await card is retitled "Show me how" and no
  longer lists the steps ("go ahead and do it yourself now and I'll watch and
  learn"); the agent's spoken handoff is likewise generic.

- Fix the "nothing happens after I'm done" stall: the model sometimes asked
  "should I save this?" in prose instead of calling saveLearnedWorkflow, leaving
  no Save card to click. The await tool-result is now directive (call
  saveLearnedWorkflow; the card is how you ask), reinforced in the prompt.

- Harden the ending: after saving, the agent treats the demonstrated charge as
  already cleared and waits, instead of re-running the freshly-saved procedure
  on it.

Verified end-to-end in OSS dev (taught Google Ads, recalled AWS); lint + build
green.
2026-06-15 10:11:03 -07:00
GeneralJerel dac136d4f6 feat(showcase): demonstrate the self-learning workflow on the real dashboard
Rework the banking demo's teach-a-workflow loop so the officer demonstrates the
over-limit unlock on the actual dashboard instead of an inline chat card, and so
the first over-limit approve request no longer shows an approval card that fails.

When asked to approve an over-limit charge it has no saved procedure for, the
agent now declines and offers to record (no approval card). The officer opens the
new Dashboard -> Transactions -> Pending approval view, files a policy exception
and approves the charge there; a waiting card holds the chat until they click
"I'm done". The agent then saves the procedure and applies it itself to a
different over-limit charge.

Move the teach/recall HITL tools (offerWorkflowRecording,
awaitDashboardDemonstration, saveLearnedWorkflow, openPolicyException,
finalizePolicyException, approveTransaction) and the agent data/permission
readables into the global CopilotContext. A route-scoped registration unmounts
when the officer navigates to the dashboard, which drops the in-progress card and
prevents the followUp from firing; global registration survives navigation and
renders on every route.

The demonstrated exception code is captured via the recording context and handed
to the Save step, so the saved procedure names the exact code used. The agent
prompt is updated for the decline+offer and dashboard handoff, and still never
spells out the unlock or names a justifying code.

Verified end-to-end in OSS dev (lint + build green, gate smoke 3/3): Beat 1 shows
no card, the dashboard demonstration clears the Google Ads charge, and recall
clears the AWS charge via the learned procedure.
2026-06-15 08:43:38 -07:00
GeneralJerel 654c96f7d0 fix(showcase): keep best-effort recording from raising a dev error overlay
The self-learning recorder POSTs to the annotate endpoint, which only exists with an
Intelligence backend; in OSS mode it returns 422. Call sites logged that rejection with
console.error, which Next.js 16 surfaces as a full-screen dev error overlay mid-demo even
though the failure is expected and harmless.

Swallow the failure in the recorder seam (useRecordUserActionInCurrentThread): catch it and
log quietly via console.debug instead of letting it reject. Recording stays best-effort — a
no-op without an Intelligence backend, and unchanged (records normally) with one.
2026-06-15 06:13:54 -07:00
GeneralJerel 8e7c6e2fb6 feat(showcase): narrate the self-learning teach-a-workflow loop
Drive the FOR-137 self-learning story as an agent-orchestrated, narrated loop. When an
over-limit approval is rejected, the agent offers to record a workflow; the officer
demonstrates by filing a policy exception; the agent summarizes and saves the procedure;
then it applies that procedure itself to a different over-limit charge. Same-session recall
works by echoing the learned procedure back into the thread.

page.tsx: three new human-in-the-loop tools (offerWorkflowRecording,
recordExceptionDemonstration, saveLearnedWorkflow) plus a transactions agent-readable so the
agent resolves a charge the user names to the right id instead of guessing.
openPolicyException now returns the new exception id, and the agent-driven exception tools
are followUp:true so the recall chain (open then finalize then approve) does not stall.

route.ts: TEACH & RECALL prompt rules and an ACTION DISCIPLINE carve-out. The prompt does
not restate the unlock procedure, preserving the learning invariant.

policy-exception-inline.tsx: surface the demonstrated exception code via onFiled(code).

scripts/over-limit-gate-smoke.mjs: regression guard proving only a finalized
justifying-code exception lifts the policy-limit gate.

Verified end-to-end in OSS dev mode (lint and build green): the demonstration clears the
Google Ads charge and recall clears the AWS charge with a single successful approve.
2026-06-15 06:05:27 -07:00
GeneralJerel ceb4b48979 feat(showcase): sequence welcome chips for the self-learning demo arc
Rewire the before-first-message suggestion pills to drive the FOR-137
self-learning story: (1) the teachable over-limit ask, (2) surface the pending
charges so the officer can demonstrate the unlock, (3) recall on a different
over-limit charge on a fresh thread. Titles stay symptom-only so they do not
hint at the exception path the agent is meant to learn on its own.
2026-06-11 19:12:54 -07:00
GeneralJerel f04d9e992a fix(showcase): pass deps to HITL tools so their render uses loaded data
The banking demo's human-in-the-loop tools registered their render in a
mount-keyed effect (useFrontendTool), so without a deps array the render
closure froze on the EMPTY initial cards/policies/transactions. Those arrays
load async after mount, so the registered render kept filtering empty data:
showAndApproveTransactions painted a card with no rows (the agent-driven
approve flow appeared to do nothing), and assignPolicyToCard / setCardPin /
addNoteToTransaction showed raw ids instead of the resolved card/transaction.

Pass the data each render reads as the useHumanInTheLoop deps so it
re-registers when that data loads, mirroring the existing selectCard and
showTransactions (useComponent) deps. addNewCard / openPolicyException /
finalizePolicyException render their args only, so they are left as-is.

Verified against unmodified workspace react-core: the agent-driven approval
card now renders the Google Ads charge with its over-limit badge and the
file-exception action instead of a blank card.
2026-06-11 19:12:53 -07:00
GeneralJerel c310f4142c fix(showcase): present all pending approvals in one HITL call
The model answered 'show me the unapproved transactions' with one
showAndApproveTransactions call per pending transaction (parallel tool
calls). Parallel calls to the same useHumanInTheLoop tool wedge the
render at inProgress, nobody can respond, and the thread is then
poisoned — every later run fails with 'Tool results are missing for
tool calls …'. Make the tool take a comma-separated id list and
instruct the model to call it exactly once; the renderer's existing
string .includes() filter was already written for a combined call.
(An array schema renders an empty tool slot — react-core issue — so
the param stays a string.)

Verified live in OSS mode: single call streams complete args +
RUN_FINISHED, the approval card renders with per-row Approve/Deny,
deny fires the recording vignette (data-recording=true) and POSTs
/annotate; over-limit approve is rejected by the server gate as
designed. In Intelligence mode the BFF /annotate path records 200.
2026-06-11 06:41:29 -07:00
GeneralJerel bd0cf2a21d fix(showcase): neutral framing for the empty permissions deny-list context
The unavailable-actions agent context had an unconditional "the user does
not have permission to perform these actions" description. For Admins the
list is empty, and the model read the menacing description plus "[]" as a
blanket prohibition — refusing showAndApproveTransactions and every other
gated tool even though they were forwarded with the run. Reframe the
description so an empty list explicitly means no restrictions and refusals
are only allowed for listed actions.

Verified live in Intelligence mode: before, the agent answered "you don't
have permission" as Admin; after, it calls showAndApproveTransactions
(wire capture shows the corrected context and the tool call).
2026-06-11 06:13:06 -07:00
GeneralJerel 6a7f1187f7 test(showcase): add self-learning smoke script (record, distill, recall)
scripts/self-learning-smoke.mjs proves the banking demo's recording seam
end-to-end against a running Intelligence backend: posts four teaching
actions through the demo BFF /api/copilotkit/annotate (the platform
requires UUID clientEventIds), optionally runs one sl-worker sweep when
INTELLIGENCE_REPO is set, and asserts the distilled vendor policy reads
back through the platform /mcp knowledge tool. Wired as the
test:self-learning package script and documented in the README.

Verified live: PASS 6/6 against Intelligence @ mme/learn-from-user-activity
(records as rows 13-16, sweep editCount=0 steady-state, recall returns the
pre-cleared vendor policy).
2026-06-11 04:54:23 -07:00
GeneralJerel 3b06802423 fix(showcase): migrate banking demo lint to eslint cli (next 16 removed next lint)
next lint was removed in Next 16, so the demo's lint script failed before
linting anything. Switch to eslint . with a flat eslint.config.mjs built on
eslint-config-next's native flat exports (same shape as the other Next 16
example apps), and fix the findings the new react-hooks rules surfaced:

- actions.ts / team/actions.ts: wrap mount fetches in an async IIFE so
  set-state-in-effect can see the setState calls are asynchronous
- auth-context: derive currentUser from selection ?? team[0] instead of
  syncing state in an effect
- use-theme: lazy-init theme from localStorage (SSR-guarded) and apply the
  DOM class in an effect keyed on theme; hoist applyTheme to module scope
- threads-drawer: copy timeout maps to locals inside the effect so cleanup
  does not read refs that may have changed
2026-06-11 02:29:32 -07:00
GeneralJerel b907d6397d docs(showcase): update recording section — gap is closed
The 'Known gap' section predated the recording fix: the hook exists as
useLearnFromUserActionInCurrentThread and record-user-action.ts is now a
real adapter. Document the /annotate flow and the backend route
requirement (/connector/annotate) instead.
2026-06-10 19:05:23 -07:00
GeneralJerel 3b44d04746 docs(showcase): add banking demo screenshots to README
Dashboard, copilot chat panel, and the learning-mode recording vignette
(the violet glow shown while an officer demonstration is being recorded).
PNGs are LFS-tracked per the repo .gitattributes.
2026-06-10 19:04:33 -07:00