mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
codex/cloudplot-showcase-migration
15424 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ffc74d7412 |
feat(reskinnable-demo): land airline's beats 4, 5 and 6 on the client
Beat 4 — recall with a VISIBLE why. `showTrips` grows a REQUIRED `note` string that the agent fills with the preference it recalled and applied, rendered as a "Remembered · …" band above the trip wall. Without that band the room watches a competent summary and has no way to know anything was recalled, so the beat is invisible and does not count. The slot is required rather than optional because an optional one is the one a model silently omits; a blank note still renders no band, which is the honest state on the OSS path where there is no `recall_memory`. The prompt gains a "RECALL FIRST, THEN SAY WHAT YOU APPLIED" clause — recalling after the answer is on screen is not recalling. Beat 5 — the seeded procedure now has prompt discipline around it: resolve the booking from context and use its BOOKING ID (AV7QK2 covers two of Camila's legs, so the confirmation code is genuinely ambiguous), finding is not handling, and an EXPLICIT statement that this is a DIFFERENT procedure from beat 6's with three named things it must not do (no `offerWorkflowRecording`, no `awaitDemonstration`, no offer to record). Conflating the two is the easiest mistake available in this demo. Beat 6 — the client half of the teach loop: - `offerWorkflowRecording` → `awaitDemonstration` → `saveLearnedProcedure`, all `followUp: true`, plus `DemonstrationCard`, which OWNS the outer recording bracket from "show me" to "I'm done". The two clicks of a demonstration nest inside it; a ref count reaching zero between them clears the feed and STRANDS the demonstrated category. - `teach-mode-directives.ts` — each builder beside its reader, so a card states only what its producer reported. The save card CLASSIFIES its settle rather than testing for presence: both buttons settle with a string, so branching on presence prints "Saved" over a decline, live and on every replay. - `components/fare-exception-form.tsx` — the passenger-facing filing form, the ONE sanctioned place the waiver vocabulary appears, mounted on Your account. The menu lists justifying categories and decoys together, unmarked, in catalogue order, and the form shows the booking's own `fareNotes` prose, because the gate is GROUNDED: the learned procedure has to be "read what this booking documents, file the matching category, approve, retry", not a memorized string. The filing step carries the category as DATA (`logStep(label, code)`), which is what `getDemonstratedCode()` reads. - `components/authorizable.ts` gains `blockedByFare`, deriving the gated cases from the same clause order the server runs. It is optimistic about a linked approved exception for the same reason `offerableOptions` is — the wire cannot see `waiverGround` — so a decoy filing drops the case off the list while the server still refuses it; the form keeps the selection so the presenter can still press Retry on the case they just filed against. All SIX leak channels stay closed, and `tools.test.ts` now checks four files rather than two (the two agent-facing ones plus the directives and the seeded memories, both of whose text reaches the model). It also adds the positive assertion the negatives cannot make — that the catalogue IS imported by the form and the form IS mounted — because absence everywhere with no form is an unlearnable gate that would pass every negative case. Reskin-skill review: checked, no skill impact. No contract field, link builder, lint rule, registration site or beat mechanism changed; this is one skin catching up to mechanisms `SKILL.md` and `demo-beats.md` already describe. The teach chain and the filing form are modelled on logistics'/commerce's, which the skill already names as the references. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a12b95e004 |
feat(reskinnable-demo): give airline durable memory and a per-passenger scope
Aeronova had the full REST substrate for beats 4, 5 and 6 and none of the
Intelligence half, so its `dev/reset` restored the trip record and left a
previous run's learned procedure standing — the most demo-destroying state this
app has, because everything still works and it just proves nothing.
Adds the three modules the other memory-complete skins ship:
- `intelligence/user-id.ts` — server-safe `IdentifyRunUser`, registered as
airline's `identifyUser` in `agent-registry.ts`. The traveller→bucket map is
BUILT FROM `data/trip-seed.ts` and is a `Map`, not a plain object: the key is
client-forwarded, and a prototype-chain hit ("constructor", "__proto__") would
pass the truthiness guard and then scope memory under `undefined`.
- `intelligence/seed-memories.ts` — beat 4's standing preference (aisle, forward
of the wing, never Basic Economy, times in America/Santiago, disrupted first)
and beat 5's cancellation procedure. Beat 6's fare-exception procedure is
DELIBERATELY absent, with a comment saying so; the beat-5 text also states
out loud that it is not that procedure.
- `intelligence/forget-memories.ts` — mirrors commerce's, including the
project-scope SKIP: project rows are global to the shared Intelligence
instance, so sweeping them would delete banking's seeded memories. That skip
is why beat 5's procedure and the teach chain both scope `user` — a
project-scoped row would survive every presenter reset and open beat 6
already taught.
Both `memorySeedTargetUserIds()` and `memoryScopeUserIds()` are derived from
`resolveUserId`, so a pinned `INTELLIGENCE_USER_ID` collapses them onto the
bucket the runtime will actually read. The seed targets include the DEFAULT
bucket as well as the account holder's, because runs frequently resolve to the
default and a single-bucket seed recalls nothing while looking fine.
Client half: `runtime-properties.ts` forwards `{ userId, userRole }` as a frozen
module constant. No `RuntimeProviders` — Aeronova has one account holder and no
switcher, so the hook reads no context and nothing has to sit above
`CopilotKitProvider`. `runtime-properties.test.ts` is the drift guard for the two
duplicated literals (the seed is not imported client-side on purpose).
`skin.tsx` also picks up the teach chain's `toolLabels` here since it is the same
declaration site.
Reskin-skill review: checked. The `Skin` contract, link builders, registration
and the client/server boundary are untouched; airline now sets two optional
fields the skill already documents. `demo-beats.md` still describes banking's
`project` scope for the stored procedure, which is out of date for the reasons
above — that is a pre-existing skill gap this change does not widen, and it is
reported to the orchestrator rather than edited here (`.claude/skills/**` is
outside this slot's boundary). CLAUDE.md's "airline is the only one that omits
`identifyUser`" line is now false and is likewise reported.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
b2ab37469a | chore(deps): update astral-sh/setup-uv action to v10 | ||
|
|
e8d5fa71d6 |
fix(runtime): bump uuid off deprecated v10.0.0 (#6118)
## Summary
`packages/runtime` declares its own `"uuid": "^10.0.0"` dependency, but
nothing in the package's source actually imports it directly — id
generation in `@copilotkit/runtime` goes through `randomUUID()`
re-exported from `@copilotkit/shared`, which already depends on
`uuid@^11.1.0`. The unused v10 pin just adds an npm deprecation warning
for every consumer installing `@copilotkit/runtime`:
```
npm warn deprecated uuid@10.0.0: uuid@10 and below is no longer supported. For ESM codebases, update to uuid@latest. For CommonJS codebases, use uuid@11 (but be aware this version will likely be deprecated in 2028).
```
This bumps it to `^11.1.0` to match `@copilotkit/shared` and clears the
warning.
## Test plan
- [x] `pnpm --filter "@copilotkit/runtime^..." run build` — all
workspace dependencies build cleanly
- [x] `pnpm --filter @copilotkit/runtime run check-types` — no type
errors
- [x] `pnpm --filter @copilotkit/runtime run test` — 126 test files /
1746 tests passing
- [x] Confirmed no file in `packages/runtime/src` imports `uuid`
directly (grepped for both `from "uuid"` / `from 'uuid'` and
`require("uuid")` — zero matches)
- [x] Confirmed `pnpm-lock.yaml` now resolves `uuid@11.1.0` for this
dependency, which is not on npm's deprecated-versions list
🤖 Generated with [Claude Code](https://claude.com/claude-code)
|
||
|
|
329c8a7a22 | feat(reskinnable-demo): wire keel's skin, tools and agent, and settle runs server-side | ||
|
|
0d21adea4d |
feat(reskinnable-demo): give keel's agent the canvas tool and the screen clause
`render_impact_brief` is the SERVER tool beat 3d had no way to reach: without it
nothing ever opens the canvas for a filed brief. It takes a `briefId` and nothing
else — every string, date and citation on that surface is read out of the stored
record, so the brief on the canvas and the brief on the Register page cannot say
different things about the same document. An unknown id returns `{ error }`
naming the id and the tool that mints one, because an agent handed no surface and
no reason retries the same wrong id. The surfaceId is suffixed per call so
dismissing one brief never suppresses a later one, but rooted at the BRIEF's own
id, never the ops report's.
BEAT 3b's third leg: the SCREEN AWARENESS clause now describes THIS skin's
screen. It names the Policy Register's four levers, the admitted-versus-displayed
row counts, the rows in the order shown, and the document sections on a detail
route — and it states the two things in that context that must not be smoothed
over: a null `attestation_coverage_percent` means NOT MEASURABLE rather than 0%,
and `book` figures describe the whole register and never the filtered view. Two
new rules cover the release gate (relay a refusal, never route around it, and it
is not the same gate as a run approval) and the ingest-to-artifact arc. The nav
label is now "Register", so the prompt calls it that everywhere.
`agent.test.ts` is the drift guard `agent.ts` never had. It asserts the resolved
server-tool list EXACTLY (a dropped tool and an added one both matter), executes
`render_impact_brief` against a really-filed brief — including the uncarried-ref
row that proves the document was read — and reads the prompt for the claims a
reviewer would otherwise take on trust, including that it names NONE of the six
publication-variance codes. A tool defined and never registered compiles, lints
and renders; it fails once, on stage, as "the canvas never opened".
Reskin-skill check: no impact. The skill already tells a skin author to
co-locate a server-safe `agent.ts` and warns that it has no drift guard; this
adds one for keel without changing the contract it describes.
|
||
|
|
aa07ab5f6d |
feat(reskinnable-demo): wire keel onto the ledger and land beats 1, 2 and 3a
Mounts `KeelLedgerProvider` in `KeelRuntimeProviders`, moves every consumer onto the one `GET /ledger` snapshot through a new `useKeelDesk()`, and RETIRES `data/use-data.ts` — deleting its 900ms `setInterval(() => setRuns(tick(...)))` in the same change that flips the consumers. Not before (it was the only clock until then) and not after (that is the two-clocks bug: the client painting progress the server never heard of, and the next `refresh()` silently rewinding it). `skin.useData` is gone with it, so `useSkinData<T>()` now correctly returns undefined for keel, as it does for the four other REST-backed skins. Runs and the policy register are finally one substrate. `useKeelDesk` carries the pure derivations across unchanged — `approvals`, `approvalsForMe`, `kpis`, and the `summaryKey` churn guard, which matters MORE now that the poll hands back a fresh snapshot object every 900ms. Its mutations are POST-then-re-read, so they return promises and carry a third outcome an in-memory store could not have: `stale`, meaning the write LANDED and the re-read did not. Every caller surfaces `reason` even on success, because "this view is behind" printed as a green tick is indistinguishable from a slow network. BEAT 1 — `showRegisterHealth` renders policy-library health as a card in the transcript: the four tiles plus a per-space bar with the review debt tinted inside it. It takes NO figures; the card re-derives every one through the same `deriveRegisterKpiTiles` / `summarizeRegister` the Register page uses, so the chat and the page cannot disagree. Coverage stays a tri-state — "Not measured", never 0%, for a document nobody has been assigned. BEAT 2 — every render in `tools.tsx` now chooses its terminal branch from the recorded `result` through one `settledText` helper, and `ToolCallStatus` is no longer imported at all, so a status-keyed branch is not expressible. `countersignRelease` CLASSIFIES its result rather than merely detecting one: a refusal and a cancellation are settled results too, and rendering the success receipt for either would replay a release that never happened. A new `AwaitingCard` covers the no-result/no-respond frame, which is streaming when live and an unanswered interrupt on replay. BEAT 3a — `countersignRelease` opens a `SigningPinCard` that POSTs the six-digit e-signature PIN straight to `/countersignatures`. The agent names only the DOCUMENT (there is deliberately no `revision` parameter), the card reads the record's own pending revision, and `respond()` gets one sentence. It is NOT an authority override: the route re-runs the same `checkReleaseAuthority()` gate, so a valid PIN on an unendorsed revision is still refused, and the card RELAYS that refusal — one request, one endpoint, no fallback. Weakening it would give beat 6 a second door that nothing would fail on. Write/HITL tools read the desk through a ref with `[]` deps: a `[data]` dep would unregister a tool mid-call every time the poll produced a new snapshot. BEAT 3d's ingest half is wired too — `chatHeaderActions` (the paperclip) and `onSuggestionSelect` (claims only `BULLETIN_MESSAGE`, because the default suggestion path drops attachments), plus a `fileImpactBrief` tool that files the durable record. Keel's parameterized routes survive: `pages/parameterized-routes.test.tsx` now seeds its run into the LEDGER and still asserts both `knowledge/<docId>` and `runs/<runId>` resolve and render. Reskin-skill check: keel is no longer one of the two in-memory skins, so `.claude/skills/reskin/SKILL.md`'s `useData` guidance and the substrate lists in CLAUDE.md/README are now stale. Those files are outside this slot's boundary and are called out for the doc slot rather than edited here. |
||
|
|
8e3522b564 |
feat(reskinnable-demo): settle keel's runs on the server, on read
Time lives on the SERVER for keel's run engine. `settleRuns()` advances runs through the pure `engine.tick(runs, Date.now())` and COMMITS the result, and BOTH read routes call it: `GET /ledger` and `GET /runs/<runId>`. Settling only one is worse than settling neither — the run-detail page and the Runs table would then contradict each other about a single run in front of the room. `engine.tick` is pure and duration-driven, so a run's state at an instant is a total function of its stored steps and the clock. Settling on read therefore yields exactly what a client ticker would have converged to, which is what makes this a decision rather than a workaround: the client needs no clock of its own, only a re-read. The commit is an in-place splice into the array `store.runs()` hands back, because `data/store.ts` exports no setter and is another slot's file. That keeps the settlement durable rather than recomputed per request, so a subsequent `store.approveStep` composes on settled steps. `tick`'s fast path returns the same array reference when nothing moved, so an idle read does not touch the store at all. `settle-runs.test.ts` is the only thing that can see a regression here: a route that stopped settling still returns 200 with a well-formed run, and the sole symptom is a started run that never moves. Reskin-skill check: no impact — this is keel's own substrate, and nothing in `.claude/skills/reskin/` describes where a skin's clock lives. |
||
|
|
bbef10a3bc | feat(reskinnable-demo): wire airline's skin, tools and agent (beats 1, 2, 3a) | ||
|
|
12f4a7342c |
fix(reskinnable-demo): stop resolvePage returning Object.prototype members
`/banking/constructor` answered 500 where it owed 404, and so did
/logistics/constructor, /people/constructor and the same URL for toString,
valueOf, hasOwnProperty and __proto__.
Three skins resolved pages out of an object literal:
const PAGES: Record<string, ComponentType> = { "": Index, cards: Cards };
return PAGES[key] ?? null;
An object literal inherits Object.prototype, so PAGES["constructor"] is a
truthy Function and the `?? null` never fires. src/app/[skin]/[[...rest]]/
page.tsx then does `if (!Page) notFound(); return <Page />` -- `!Page` is
false, notFound() is skipped, and React is handed something that is not a
component. Commerce and keel were already Map-backed and unaffected.
Each of the three now uses a Map, which has no prototype keys.
The real fix is the new guard, src/shell/resolve-page-prototype.test.ts: it
walks EVERY registered skin against every own key of Object.prototype, taken
from the prototype itself rather than hand-listed, on both a top-level and a
nested segment. It pins behaviour rather than implementation, so a skin that
prefers `Object.hasOwn` passes too, and it keeps holding for skins that do
not exist yet.
Two things learned writing it, both kept in the file:
- The property is NOT "returns null for a prototype key". Keel resolves
knowledge/<docId> for ANY docId on purpose and renders an in-page
not-found body, so a non-null answer there is correct. The property is
"never returns a value INHERITED from Object.prototype".
- It carries a non-vacuity assertion (at least six skins registered) and a
per-skin companion (the index still resolves), because a resolvePage that
returned null for everything would otherwise pass every other assertion
while serving a dead skin.
Mutation-verified: reverting people to the object literal turns it red.
Found by hand while wiring a fourth skin, not by any gate -- it type-checks
(the Record's index signature says ComponentType), it lints, and no test that
walks the REAL segments ever passes a prototype key.
airline has the same defect and is fixed by the air-wire slot merged next;
this guard is what will hold it green afterwards.
|
||
|
|
d6ee68706f |
feat(runtime): add agentId to AgentRunnerConnectRequest (#6120)
Closes #5911 Adds an optional `agentId` field to `AgentRunnerConnectRequest` so custom agent runners can hydrate messages when the cache is empty. Also passes the already-available `agentId` in `handleSseConnect` through to `runner.connect()`. ## Changes - `packages/runtime/src/v2/runtime/runner/agent-runner.ts`: Added `agentId?: string` to `AgentRunnerConnectRequest` interface - `packages/runtime/src/v2/runtime/handlers/sse/connect.ts`: Pass `agentId` to `runner.connect()` ## Verification - Runtime package builds successfully (`pnpm --filter @copilotkit/runtime build`) - The change is purely additive — the field is optional and does not affect existing runners |
||
|
|
68e3712e0c |
test(reskinnable-demo): drop the import() type annotation from airline's skin test
`oxlint`'s `consistent-type-imports` warns on `typeof import("…")` inside the
`vi.mock` factory. Spread the original module through a plain record instead;
`importOriginal` still supplies the REAL `HOTEL_CONFIRMATION_MESSAGE`, so a
pill whose text drifts from the constant still fails the file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7720d5c3c7 |
test(reskinnable-demo): guard airline's wiring, replay safety and withheld gate
Four guards, each for a defect that leaves the app compiling, linting and rendering — which is why none of them was caught by the gates. `skin.test.tsx` — the mounted skin object. `Providers` missing throws only when a page that needs the ledger renders; `CanvasSurface` missing opens the canvas region and draws nothing; `nav`/`resolvePage` disagreeing 404s a sidebar link; a re-added `useData` puts a second seed of AV1423 back in the app. It also pins the Map-backed `resolvePage` against prototype-chain keys, the way keel's and commerce's do, and asserts the beat-3d pill interception fires the send rather than merely returning `true`. `tools.test.ts` — the drift guard `agent.ts` and `tools.tsx` do not otherwise have. It cross-checks that every tool the beats need is registered, that `render_trip_brief` is a SERVER tool listed on the agent (a client tool result never produces an `a2ui-surface` activity), and that the prompt names no tool nothing registers — a failure whose only other symptom is "I don't have a tool for that", live. It also holds beat 2's invariant (no terminal render keyed off `status`, no `result.match`, no `typeof result === "string"`) and beat 6's withholding across all four greppable channels, including the PROSE one no lint rule can see: the four justifying categories and the three decoys are checked by name in both files. `components/card-confirmation-card.test.tsx` — beat 3a, rendered. The digits go to `/authorizations` and nowhere else, the sentence handed to `respond()` does not contain them, an unreadable value is REFUSED with a reason on screen rather than silently stripped, a double click cannot charge twice, and a `FARE_NOT_CHANGEABLE` refusal is shown verbatim with nothing reported as authorized. That last one is the only symptom the "second door around beat 6" failure has on the client. `components/concierge-view.test.ts` — one substrate, one AV1423. No `use-data.ts` to read, no `useAirlineData`/`useSkinData` call site left, and no read of the four seed constants the REST seed duplicates field for field. Plus the two derivations that replaced stored values: the disruption follows the flight (the old seeded alert said "55 minutes" whatever the ledger held) and the seat map offers only seats the flight lists free. `readables.test.tsx` gains beat 3b's THIRD LEG, which its own header said a later slot had to add: the SCREEN AWARENESS clause in `agent.ts`, plus the `loading` flag on every page readable. Skill impact: checked, none. No skill file names any of these paths, and SKILL.md § Verification's commands (`pnpm lint`, `pnpm test:unit`, `pnpm build`) all still exist and all cover these. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9d601ed03f |
feat(reskinnable-demo): wire airline onto the ledger and land beats 1, 2, 3a
Aeronova had a complete REST substrate and no UI attached to it. This mounts the skin object, retires the second seed of Camila's AV1423, and lands the three beats the substrate could otherwise only half-prove. MOUNT AND FLIP. `skin.Providers` is now `AirlineProviders` (the one `GET /ledger` read plus the shell's teach recorder), `nav`/`resolvePage` come from `nav.ts`/`pages/index.ts` so `/airline/account` and `/airline/rebook` stop 404ing, and `CanvasSurface` is `AirlineCanvasSurface`. `resolvePage` is now Map-backed: the object literal it replaced resolved `/airline/constructor` to `Object.prototype.constructor`, a truthy Function the shell renders as a page — a 500 where a 404 belongs. ONE SUBSTRATE. `data/use-data.ts` is deleted and `useData` is dropped from the skin. `components/concierge-view.ts` projects the REST ledger onto the shapes the check-in components were written against, so the flight, passenger, seat map, disruption and rebooking options all come from `GET /ledger` and the REST seed is the only authority for AV1423. Loyalty mileage, the redemption catalogue and the bags stay seeded because the ledger models no counterpart — they are beat 5's distractors — but the member identity is overwritten from the ledger traveller so the tier appears in two places from one source. The disruption banner is DERIVED from the flight's own status and delay rather than stored, so it can no longer outlive the condition it describes. BEAT 1. `showTrips` leads with the whole account as a trip wall; nine gen-UI components in all, including the `showSeatMap` distractor `data/beat-map.md` § "Beat 5" requires and which was never registered. BEAT 2. Every terminal render reads the recorded `result` through one shared `ToolNote`; no `status === ToolCallStatus.Complete`, no `result.match`, no `typeof result === "string"`. airline can now be added to `statusKeyedTerminalRender`'s glob in eslint.config.mjs. BEAT 3a. `authorizeWithCardConfirmation` renders `CardConfirmationCard`, which POSTs the last four digits straight to `/authorizations` and hands `respond()` one sentence that does not contain them. It is offered only through `offerableOptions` — permitted change, money actually due — and it prints the server's refusal verbatim, so it can never become a second door around beat 6's gate. BEAT 3b's third leg. `agent.ts` gains a SCREEN AWARENESS clause telling the agent its context IS its view of the screen, plus the truncation and loading rules. Every readable now carries a `loading` flag, so a screen that is still spinning is not reported as empty. BEAT 3d. `render_trip_brief` is a SERVER tool emitting `buildTripBriefOps(briefId)` under `A2UI_OPERATIONS_KEY` — without it nothing ever opens the canvas — and `fileTripBrief` returns the brief id for it. Also lands beat 3c's `showRebookingSearch` (four levers plus top-N, confirmed then navigated through `useSkinHref`), beat 5's three ordered writes, and beat 6's `fileFareException` with a free `z.string()` code and the withholding stated in its `.describe()`. Skill impact: checked. The `Skin` contract, the link builders, the registration sites and the demo beats are all unchanged — this slot only implements them. `.claude/skills/reskin/` names no file this change deletes or renames (`data/use-data.ts` is airline's own, not a template path). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4383a86198 |
fix(runtime): populate request headers in runtime error context (#6287)
## Summary `CopilotRuntime` exposes an `onError` callback, but the v1-to-v2 delegation drops it before server request processing. Runtime failures therefore can't provide the incoming request headers that applications use to correlate errors with a user or tenant. ## Root cause The legacy constructor retains the callback type, while delegated runtime options omit it. Common, SSE, and Intelligence handlers consume failures inside their own boundaries, before a generic endpoint hook can reconstruct the legacy event. ## Changes - Add one internal runtime reporter that snapshots incoming Fetch request headers. - Attach the configured legacy callback to the delegated runtime instance. - Route common, SSE, and Intelligence agent-run failures, including standard `RUN_ERROR` events, through that reporter exactly once. - Redact sensitive request headers (authorization, proxy-authorization, cookie, set-cookie, x-api-key, api-key, and the CopilotKit public-key header) before they reach the `onError` event. - Preserve responses, stream close behavior, telemetry, cleanup, logging, endpoint hooks, and callback isolation. - Add production-path regressions. ## Compatibility The existing `CopilotErrorEvent` type and optional `context.request.headers` field remain unchanged. Header names and values come from the failing request's Fetch `Headers` object, with sensitive credentials stripped before the event is emitted. The callback receives a fresh record, so mutation cannot alter the request or a later event. ## Out of scope React provider propagation, Chat, CopilotMessages, the deprecated runtime-client hook, redaction policy, HTTP response exposure, v2 endpoint-hook semantics, and non-agent runtime routes remain outside this slice. ## Related PRs and Issues Addresses #2716. Scope follows https://github.com/CopilotKit/CopilotKit/issues/2716#issuecomment-5086936254. Runtime error-routing precedent: https://github.com/CopilotKit/CopilotKit/pull/2143. ## Test plan - [x] Runtime error regression and reporter tests, 12 + 4 tests passed. Covers public routing, sanitized header propagation with credential redaction, callback rejection containment, snapshots, mutation isolation, and malformed requests. - [x] SSE and Intelligence telemetry tests, 7 + 14 tests passed. Covers agent-run failures, `RUN_ERROR`, setup boundaries, exact-once reporting, telemetry, cleanup, and response preservation. - [x] Runtime preservation suites, 57 + 129 + 56 tests passed. Existing request, endpoint-hook, and runtime-library behavior remains intact. - [x] Typecheck, formatting, and lint passed on changed runtime files; lint reported four pre-existing warnings. - [ ] CI green (`static / quality`, `test / unit` on Node 20/22/24). |
||
|
|
30986613d8 | feat(runtime): add MiniMax built-in models | ||
|
|
a3ade9a52a |
fix(reskinnable-demo): make the beat-3d staging test deterministic
Two wave-2 agents reported `stage-attachment.test.ts > "stages only once the chip is queued AND finished encoding"` intermittently failing the whole suite with `expected false to be true`, always while several agents were building concurrently. The report is real, and the cause is the test harness rather than the code under test. The three SUCCESS-path tests shared the tight `FAST` budgets built for the EXPIRY branches: a 40ms `readyMs` ceiling around a real `encodeMs: 10` timer. `waitUntil` measures its budget with `Date.now()`, so a descheduled vitest worker spends the ceiling without the event loop ever running the timer it is waiting for. Measured in the real vitest/jsdom environment, reproducing the exact timer structure, 60 samples per condition: quiet p50=6ms p95=12ms max=12ms (budget 40) 60 busy loops on 10 cores p50=11ms p95=18ms max=39ms (budget 40) A sample landed 1ms inside a 40ms budget. That is the reported failure. Fixed WITHOUT weakening the property, and the property is now asserted more directly than before: - `PATIENT` budgets for the paths that succeed. A success-path wait is condition-based and returns the instant its predicate holds, so a generous ceiling costs zero wall clock — it only stops a loaded worker from expiring a budget never meant to be reached. `FAST` stays exactly as it was for the expiry tests, which still need it small. - The ordering test no longer TIMES the encode, it DRIVES it: a new `encodeMs: "manual"` fixture mode latches the `uploading` -> `ready` transition behind a `finishEncoding()` the test calls. So the two halves are asserted separately and in order — queued-but-encoding must NOT stage, and only finishing the encode may. The old shape read only the end state, inferring "waited for ready" from "ready by the time it finished". Verified by mutation: deleting the production ready wait (`const ready = true`) left the previous assertion GREEN, and leaves the new one RED. The production file is byte-identical to HEAD. Evidence: full suite 6/6 green (175 files / 1910 tests); the fixed file 6/6 green under the 60-busy-loop load that pushed the old budget to 39/40ms. `promotions.test.tsx` drives fetch through held promises and microtask flushes with no real timers, and `recording.tsx` already clears its hold timer on unmount, so neither had a wall-clock budget to lose; both stayed green across all runs, including three full-suite runs under 4x oversubscription. Reskin-skill check (per CLAUDE.md standing rule): checked, no skill impact. The skill references this file only for the fifteen-cause exhaustiveness gate, the `[attach:<cause>]` log regex, and `Beat3dTimings` being injectable to force an expiry — all three unchanged and still accurate. Nothing in `Skin`, the production module, or any surface a skin author touches changed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dbb411ee1f | feat(reskinnable-demo): give keel an impact-brief canvas (beat 3d) | ||
|
|
36d619c4de | feat(reskinnable-demo): give keel its policy register surface (beats 3b, 3c) | ||
|
|
934ed789e6 | feat(reskinnable-demo): give airline a trip-brief canvas (beat 3d) | ||
|
|
d9f74b6f6c | feat(reskinnable-demo): give airline its screen surface (beats 3b, 3c) | ||
|
|
d7f43ae1ae |
feat(reskinnable-demo): give keel a route readable and on-screen readables
BEAT 3b, both halves. `layout.tsx` registers the ROUTE readable — the path, the highlighted nav entry and, on the two parameterized routes, the id of the record open — so the agent knows WHERE the operator is. Each page then describes its own contents, so "what's on my screen?" answers differently on the Register than it does on `knowledge/<docId>`. The document page grows the register overlay beside the corpus prose, which is what gives the second ask something of its own to be about: this document's review debt, its attestation coverage, its pending revision and the bodies that have not endorsed it. That readable is registered UNCONDITIONALLY, before the not-found early return, so an unknown docId is described rather than answered with "I cannot see the screen". Unmeasurable attestation coverage travels as null, never 0. A model cannot discount what you omitted and will restate a zero as an all-clear, out loud. `pages/on-screen-readables.test.tsx` is the guard that matters: it stubs `useAgentContext`, renders each page, and asserts the readable's rows against the rows the DOM actually painted, element for element and in order — the drift no source grep can see, and the failure that survives a live demo unnoticed. It also pins beat 3c's four levers against the rendered board and the four tinted controls, on a 24-row fixture deliberately larger than the nine-document seed so any later cap is exercised. `pages/parameterized-routes.test.tsx` pins the property keel is the only skin to have: `knowledge/<docId>` and `runs/<runId>` still resolve AND still render their record, and an unknown id is still an in-page not-found body rather than a 404. Skill impact: checked. `.claude/skills/reskin/demo-beats.md` § 3b names banking, people, commerce and logistics as the skins with a route readable plus per-page readables, and tells the reader to DERIVE that list with `grep -rln useAgentContext src/skins/*/layout.tsx` rather than trust the sentence. Keel now answers that grep, so the derivation is correct without an edit — and the beat matrix in CLAUDE.md is the orchestrator's to update once every keel slot has landed, not this one's to change mid-flight. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e21b6a0689 |
feat(reskinnable-demo): make keel's knowledge page the policy register
BEAT 3c. Four levers — space, attention class, sort and top-N — all arriving from the query string, all filtering through `data/register-levers.ts`, and all four controls VISIBLY tinted when set. It is the controls that light up rather than the rows, because a filtered list alone asks the room to take the maneuver on faith. The page renders the lever module's output and reimplements none of it, so a value the schema can advertise and the view will not honour is not expressible. An unrecognised value normalizes to null: the view renders as it does with the lever absent and the control stays untinted. One pipeline publishes TWO lengths — `matching` under the levers before truncation, `visible` after it — and the caption, the rows and the readable all read that one result, so "Top N of M" cannot report that the filters did nothing. Served at the `knowledge` segment rather than a new one: the register IS the parent of `knowledge/<docId>`, the route a citation lands on. Only the nav LABEL changes, so `resolvePage`, `navigateTo`'s page enum and every citation href are untouched. Row links go through `useKeelHref`. `now` comes from the snapshot's own `asOf` rather than the wall clock — the rows, the tiles and the readable are then measured at the instant the server measured the register, a test pins the clock by pinning the fixture, and no impure clock read happens during render. Two sections are marked-but-absent at the foot of the page — beat 3d's filed Impact Briefs and beat 6's operator variance form — so the next author adds them rather than discovering the page has no room. Skill impact: checked, none. This changes one skin's page, not the `Skin` contract, the registration sites, or any gate a skin must pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5cd3a1bbe3 |
feat(reskinnable-demo): add keel's ledger context over the REST substrate
One snapshot read of `GET /api/keel/v1/ledger`, shared by every consumer under `KeelLedgerProvider`, so the register board, the KPI tiles and the beat-3b readables can never describe the register a fetch apart. Shipped UNMOUNTED on purpose: `skin.tsx` belongs to a later slot, `useKeelData` is still wired, and both parameterized routes still render from it. `useKeelLedger()` therefore falls back to a standalone read outside the provider rather than throwing — keel's own `useRole` takes the same position — so a page can adopt it before the provider lands. Decides the migration's open question, WHERE TIME LIVES. Keel's run engine ticks on a 900ms client interval today while the server holds runs as state only, and keeping both after the migration would put two clocks on one set of runs: the client's local advance would paint progress the server never heard of, and the next refresh after any write would silently rewind it. Time lives on the SERVER, and the client's only interval RE-READS. That is defensible rather than tidy because `engine.tick` is pure and duration-driven, so settling on read yields exactly the value the client interval would have converged to. The poll here calls `refresh`, never `tick`, and only while a run is running. The two follow-ups that must land with the consumer flip are written out in the module header: settle runs in both read routes, and delete `useKeelData`'s ticker in the same change. Skill impact: checked. `.claude/skills/reskin/` describes the `Skin` contract and the registration sites, none of which this touches — it adds a skin-internal data hook alongside the substrate the beat map already documents. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6f7051deba |
feat(reskinnable-demo): give airline on-screen readables and a lever board
Beats 3b and 3c, which `data/beat-map.md` records airline has never hit.
BEAT 3b — a route readable in `layout.tsx` (via `useSkinSegments`, never a
fixed-offset `pathname.split`, which reports the wrong page under LOCK_SKIN),
plus a per-page on-screen readable on all five pages. The three in-memory pages
gain theirs immediately and are LIVE — Trip, Aeronova Club and Disruptions now
answer "what's on my screen?" differently, which is the beat.
`seat-map.tsx` grows an exported `orderedSeats` + `isSelectableSeat` so the Trip
readable lists the seats the map actually painted, in paint order, rather than
re-deriving them — the commerce 5-rows-against-6 bug, which fails silently.
BEAT 3c — `pages/rebook.tsx`, a passenger's rebooking search with window, stops,
cabin and sort levers plus a top-N, all read off the query string through the
shared `readLevers`/`applyLevers` the API route also runs, and all five controls
tinted when set. ONE pipeline publishes `matching` and `visible`, so the caption
("Top 5 of N matching flights") can never say the filters did nothing while the
rows say they did. The trip picker is the search's SUBJECT, not a lever, and
never tints.
`pages/account.tsx` keeps the account visibly CAMILA'S — her name, tier and card
in the header, her own trips first under "Your trips", and the two companions
nested under their own named cards as saved travellers described by their
relationship to her. `data/beat-map.md` names this page as where the rejected
operations-desk reframe would creep back in, so a test asserts the shape: no
table, no traveller column, a per-traveller list each.
Both new pages are UNREACHABLE until a later slot wires `skin.tsx` — they read
`useAirlineLedger()`, and `nav.ts` + `pages/index.ts` exist to make that a
two-import swap. `useAirlineData` and the three existing pages are untouched.
Tests: `readables.test.tsx` guards OMISSION (source grep, anchored inside the
`useAgentContext` call so it cannot pass on a heading or an import) and
`pages/on-screen-readables.test.tsx` guards DRIFT (renders each page and
compares the readable's rows against the painted DOM, element for element and in
order). Note the third leg of beat 3b — the agent's SCREEN AWARENESS prompt
clause — is NOT yet written: `agent.ts` belongs to another slot, and
`readables.test.tsx` says so where the assertion would go.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7fc55dbf01 |
feat(reskinnable-demo): add airline's REST ledger context, unmounted
`AirlineLedgerProvider` + `useAirlineLedger()` over `GET /api/airline/v1/ledger` — one fetch, shared, plus the cross-instance revalidation bus logistics uses so a write that goes straight from a chat card to a route (beat 3a's card confirmation) still refreshes the screen. A context rather than logistics' per-instance hook: Aeronova publishes ONE cross-cutting snapshot, and beat 3b asks the agent to describe exactly what the passenger can see, so two panels disagreeing about the ledger is the specific failure this must not have. SHIPPED UNMOUNTED on purpose. `skin.tsx` belongs to a later slot, so nothing renders the provider yet and `useAirlineData` remains the live substrate for the Trip / Loyalty / Disruptions pages — `data/beat-map.md` § "It is ADDITIVE". The hook THROWS outside its provider and names the three edits that mount it, rather than returning an empty ledger: a blank account on stage is indistinguishable from a seed that failed to load. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4ecfeeaaff |
feat(reskinnable-demo): give airline a durable trip-brief canvas
Beat 3d, the "out" half — airline's first `CanvasSurface`. It renders the Trip Brief `POST /api/airline/v1/briefs` filed, read back off the app over `GET /briefs`, so what the room sees is the artifact that survives deleting the conversation rather than a replay of what the model said. The two-column split IS the argument, not decoration: the left column is what only a reader of the attachment could know (hotel, confirmation number, the 22:30 desk cutoff, the cancellation deadline), the right is what only Aeronova holds (booking, traveler, arrival station, scheduled arrival), and the headline banner on top is the collision of the two — "AV1423 gets into Lima at 23:00; Casa Miraflores stops taking arrivals at 22:30". A flat list of the same fields would render identically and prove nothing. `arrivesAfterLastCheckIn` is honoured as the tri-state it is: an unmatched document paints the muted UNCHECKED banner and prints its dropped ledger fields as "not on file", never the reassuring green one over a comparison nobody made. Why there is no `A2UIProvider`/`A2UIRenderer` here, unlike the other four skins: their briefs are compositions the agent selects, so their ops carry a component tree and their catalogs carry the renderers. Aeronova's brief is one durable server-settled record of fixed shape — the only selection left is WHICH brief, so that is all `canvas/trip-brief-ops.ts` carries. `readBriefId` is deliberately tolerant, so a later slot that emits a richer tree still resolves the right brief; a named-but-missing brief refuses to fall through to the newest, because showing a different brief under this run's headline is worse than showing none. Shipped UNMOUNTED: `skin.tsx` still omits `CanvasSurface`, and no tool emits the ops yet. Both file headers name exactly what the later slot must wire. Tests drive the real substrate end to end — the document facts go through the route, the server settles the ledger half, and the canvas renders what `GET /briefs` returns — rather than a hand-written brief object, which would pass while the two halves disagreed about field names. Reskin-skill check: no impact on the `Skin` contract, lint rules, registries, routing or beat mechanisms. Worth the orchestrator's attention though: SKILL.md and CLAUDE.md both describe `CanvasSurface` as "this skin's own a2ui report surface", and this one activates off the same `a2ui-surface` activity while rendering the durable record directly. Flagged, not edited — `.claude/skills/` is outside this slot. |
||
|
|
ad2780f3db |
feat(reskinnable-demo): stage the airline hotel confirmation as an attachment
Beat 3d, the "in" half. Aeronova now has the beat-3d attachment wrapper the other three demo-complete skins have, built on the shell primitive rather than a fourth private copy of the chain: - `HOTEL_CONFIRMATION_MESSAGE` — one string shared by the pill and the interception, so a drift cannot send the prompt WITHOUT the file (the failure that leaves the model inventing the document's contents). - `sendHotelConfirmationMessage` — the pill path. - `attachHotelConfirmationByHand` — the presenter's paperclip fallback. The booking is NAMED (`?booking=bkg-av1423`) rather than left to the route's default, so a reseed that moves Camila's Lima trip 404s the fetch and ABORTS the pill loudly instead of quietly attaching some other traveler's room. Shipped UNMOUNTED: `skin.tsx`, `suggestions.ts` and `tools.tsx` belong to a later slot, and the file header names exactly what that slot has to wire. Everything load-bearing — composer lookup, PDF byte check, the four bounded condition waits, the abort rule, the dual console/alert reporting — is `@/shell/attach`, verified once in `src/shell/attach/stage-attachment.test.ts`. The new tests cover only what is genuinely airline's and silent when wrong: the three parameter values, that the named booking still resolves a confirmation through `hotelConfirmationFor`, and that the composer chip's filename matches the one `GET /hotel-confirmation` derives from the confirmation number. Reskin-skill check: no impact. The `Skin` contract, the lint rules, the registries, routing and the beat mechanisms are untouched; `templates.md`'s "DO NOT IMPLEMENT THE CHAIN" guard is what this file follows. |
||
|
|
54e05fb3f7 |
fix(reskinnable-demo): type the empty-items case so tsc can check it
`BriefImpacts({ props: { items: [] } } as Parameters<...>[0])` is a TS2352
error: an empty array literal infers `never[]`, which does not sufficiently
overlap `RendererProps<{ items: string[] }>`, so the assertion is a mistake
rather than a widening. `[] as string[]` restores the overlap.
Worth knowing WHY this survived a green slot. Nothing in this app's gates
type-checks test files:
- `pnpm build` (next build) only type-checks what the app's module graph
reaches, and no test file is imported by the app;
- vitest does not type-check at all, so the test passed while being
ill-typed;
- there is no `typecheck` script, and CLAUDE.md names `pnpm build` as THE
type-check gate.
`tsconfig.json` DOES include `**/*.tsx`, so `pnpm exec tsc --noEmit` catches
it -- that command is the only thing in the tree that does.
Found by an explicit tsc run over the slot before merge, not by the slot's
own three green gates.
|
||
|
|
d4ed9072d7 |
feat(reskinnable-demo): render keel's filed Impact Brief on the shared canvas
Beat 3d's outbound half. `canvas/impact-brief-ops.ts` expands a brief the store already holds into a2ui operations under its own surface id (`keel-impact-brief`), and `canvas/impact-brief-components.tsx` contributes the three renderers. `canvas-surface.tsx` spreads both into keel's ONE report catalog, so the ops report is untouched and the two surfaces are told apart by surfaceId rather than by catalog. Where this deliberately departs from ops-report.ts: the figures ARE in the ops. Run KPIs tick, so the report binds live `useSkinData`; a filed Impact Brief is immutable the instant `POST /briefs` returns, so its values are read out of the stored record. The tool takes a `briefId` and nothing else, which is what keeps every string on the canvas server-sourced rather than the model's second telling of what it just filed — and it makes the surface replay-safe with no fetch. `carried` is re-derived against the LIVE register instead of stored: "the library does not hold POL-118" is a claim about the register NOW, and a reseed that adds the ref must be able to change the answer. That is the row the beat rests on, so it is drawn as a finding, and "never released" is kept distinct from "not in the library" — two facts `POST /briefs` goes out of its way not to merge. `bulletin-citations.ts` now EXPORTS its canonical-ref reduction (previously a private `canonical`) so the canvas asks the same question of the same refs; a third private copy is how the two surfaces come to disagree about POL-118. Shipped UNMOUNTED for the tool half: `renderImpactBriefParams` and `buildImpactBriefOps` are exported for a later slot's `agent.ts`. The canvas half needs no wiring — keel already sets `CanvasSurface`. Skill check: no impact on `.claude/skills/reskin/`. The `Skin` contract, the link builders, registration and the lint gates are untouched; the skill's one reference to `canvas-surface.tsx` (SKILL.md:536, naming it as the file behind `CanvasSurface`) is still accurate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7e38f97c76 |
feat(reskinnable-demo): stage keel's regulatory bulletin into the composer
Beat 3d's inbound half for keel. `attach-bulletin.ts` is the three per-skin values and nothing else — the URL, the filename and the message — over the shell-owned staging chain in `@/shell/attach`. The knowledge space is named in the URL rather than left to the route's `DEFAULT_SPACE`, so a corpus rename 404s the fetch and ABORTS the pill instead of quietly serving a different space's bulletin under a filename that says privacy. Shipped UNMOUNTED: both mount points (`chatHeaderActions` for the paperclip, `onSuggestionSelect` for the pill) live in `skin.tsx`, and the pill itself in `suggestions.ts` — a later slot's files. `BULLETIN_MESSAGE` is exported so the pill and the interception key off one value and cannot drift. Tests drive the real chain against a fake composer rather than mocking `@/shell/attach`: a mock would prove keel passes three values to a function and nothing about whether those values work. The fifteen failure causes stay covered once, in the shell. Skill check: no impact on `.claude/skills/reskin/`. This follows templates.md's "DO NOT IMPLEMENT THE CHAIN" wrapper shape verbatim and adds no new pattern; nothing the skill references was renamed or deleted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1dabe60372 |
feat(reskinnable-demo): give keel a policy-register REST substrate
Merges blitz slot keel-substrate. Harbor Point Health keeps its identity as a knowledge and operations desk; the substrate adds the half every beat needed -- the policy REGISTER, the lifecycle state of the nine corpus documents. The corpus/register split is load-bearing: prose only changes when an author edits it and stays a static server-safe module, while review dates, attestation coverage and the pending revision are operational state the demo mutates. GET /documents/[docId] joins the two server-side, which is how the parameterized knowledge/<docId> route survives the migration. - Beat 6 gate: 403 UNENDORSED_REVISION on releasing a revision its governance committee has not signed. Symptom only -- it names the document, the revision and the body that has not signed, never a way through. Unlock is a ratified publication variance under a justifying code. Justifying: PATIENT_SAFETY_ALERT, ACCREDITATION_FINDING, REGULATORY_MANDATE, INCIDENT_CONTAINMENT. Decoys: COMMITTEE_CALENDAR, EDITORIAL_CLEANUP. COMMITTEE_CALENDAR is load-bearing -- it is the real reason a person reaches for an interim release, so a bluffing agent picks it and stays blocked. - Beat 3a is not an override: /countersignatures re-runs the same checkReleaseAuthority(), pinned by a four-test block whose companion case proves a justifying variance DOES lift the same block, so a route that refused everything could not pass. - Scope discipline: no beat here decides anything clinical. The gate is about who may RELEASE a revision, not what the policy says. dev/reset returns reset: ["store"] and reports memoryBeats "unarmed" rather than a bare ok, because keel has no intelligence/ module yet -- the incompleteness is legible in the artifact instead of hidden in a comment. Reskin skill impact: deferred to the docs slot on purpose -- airline is landing the same shape in parallel and both would conflict on the same paragraphs. |
||
|
|
992d578476 |
feat(reskinnable-demo): give airline a passenger-facing REST substrate
Merges blitz slot air-substrate. Aeronova stays a passenger concierge and reaches beats 3a, 3c and 6 through ENTITLEMENT rather than hierarchy: a fare rule refuses as hard as an approval authority and needs no operator. - Beat 6 gate: 422 FARE_NOT_CHANGEABLE, naming the fare condition only. Unlock is a fare exception filed under a waiver category and approved. Justifying: SCHEDULE_CHANGE_TRIGGERED, MEDICAL_DOCUMENTED, BEREAVEMENT_DOCUMENTED, MILITARY_ORDERS. Decoys: CHANGED_PLANS, FOUND_LOWER_FARE, ELITE_COURTESY -- what everyone actually tries. - Grounding: an exception lifts only when its category matches what the booking documents, so the unaided replay is a procedure rather than a memorized literal. This adds a SIXTH leak channel, Booking.waiverGround, which store.snapshot() strips; the passenger reads the same fact as prose. - Beat 3a is not an override: /authorizations re-runs the same fare check, so the card's last 4 cannot release what the gate blocks. Three gated bookings, one (AV5KD1) unlockable by nothing, which is what makes the decoys real rather than theoretical. Reskin skill impact: deferred to the docs slot on purpose -- keel is landing the same shape in parallel and both would conflict on the same paragraphs. |
||
|
|
3da30eec82 |
feat(reskinnable-demo): add the airline skin's v1 API routes
Twelve routes over the passenger-facing substrate: one `/ledger` snapshot, the booking reads, beat 5's three ordered writes, beat 3a's card-confirmed change, beat 6's gate plus its two-step unlock, beat 3d's generated document and the brief it files, and a gated `dev/reset`. Three properties are load-bearing and each is pinned by a test that goes red alone when the property is removed: - `POST /authorizations` runs the SAME `checkFareChange()` as the ordinary change route, so a valid card confirmation on a non-changeable fare is still 422. Its test walks EVERY option on all three gated bookings, discovered from the live ledger rather than hardcoded, and a companion assertion says what the unlock path IS so the block cannot pass by refusing everything. - The change route stops at 402 when money is due, which is what makes the card the only path to a paid change rather than a step the demo could skip. - Neither the filing route nor the approve route ever reveals whether an exception lifts. Their responses are asserted field-for-field identical between a justifying category and a decoy — a `lifts` flag would hand over the withheld catalogue one probe at a time. `/hotel-confirmation` generates the document per request and fails loud with a logged cause, because a silent 404 there aborts beat 3d's pill with nothing to debug. `POST /briefs` settles the ledger's facts in every direction — overwritten on a unique match, dropped when there is none, and reported in `settled`/`unmatched` so the tool can tell the agent rather than silently overrule it — while leaving the document's own facts model-authored, which is the beat's whole proof. `dev/reset` reports `memoryBeats: "unarmed"` on purpose: airline has no seed-memories module yet, and a reset route without one looks identical to a reset route with one. Reskin-skill check: no impact. All twelve routes are new; no shared route helper, contract, lint rule, gate or registration changed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
88be5ea8ce |
feat(reskinnable-demo): give the airline skin a passenger-facing REST substrate
Aeronova keeps its in-memory concierge store and gains a server. ADDITIVE throughout: `use-data.ts` is untouched, every existing page still reads it, and the new seed is written to AGREE with it about Camila's AV1423 (same flight, route, cities, aircraft, gate, times, delay, and the same three alternatives) because both substrates are on screen in the same demo. The gate is on the FARE, not on a person's rank — see `data/beat-map.md`. `checkFareChange` refuses a Basic Economy or promo ticket naming the fare condition and nothing else, short-circuits to a free involuntary change on a cancelled flight or a schedule move past four hours (which is what keeps beat 5 clear of beat 6), and lifts only for an approved exception whose category matches the circumstance the booking's record documents. That grounding is what makes the two taught cases genuinely unlike each other: without it, a demonstration on the schedule-change booking replays on the medical one as a memorized literal. A third booking documents nothing at all, so the decoys are real rather than theoretical — every category in the catalogue files honestly against it and releases nothing. The catalogue is WITHHELD from the agent, and `store.snapshot()` strips `waiverGround` for the same reason: it is a code-shaped token mapping one-to-one onto a justifying category, and no other skin has a grounded gate, so it is a sixth leak channel the failure-modes list does not name. The passenger reads the same fact as prose in `fareNotes`. Beat 3a's card confirmation is a second factor and never an entitlement override; beat 5's vocabulary lives in a module sharing no token with beat 6's, asserted both ways. 30 rebooking options on the cancelled return leave ten rows under beat 3c's own lever set, asserted directly so a reseed that thins the board fails a test instead of a demo. Reskin-skill check: no impact. Every file here is new; no contract field, link builder, lint rule, test, gate, registration, routing boundary, beat mechanism or skin identity changed. The obligations this creates for LATER slots — the `withheldGateVocabulary` glob, the seed-memories module, the exception form — are recorded in `data/beat-map.md` rather than in the skill, because they are facts about this skin and not about authoring skins in general. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
949de51de8 |
docs(reskinnable-demo): map the airline skin's passenger-facing demo beats
Aeronova's beat map, written before any substrate code, per demo-beats.md. The premise it is built on: authority does not have to be ORGANIZATIONAL, it can be ENTITLEMENT. A gate that says "your fare does not permit this" refuses exactly as hard as one that says "you lack approval authority", and it needs no hierarchy at all — so Aeronova stays a passenger-facing concierge and still reaches beats 3a, 3c and 6. Everything the beats wanted from an operator has a passenger analogue that is more familiar, not less: fare rules for an approval authority, a fare exception for an escalation, the account's own bookings for a board, the rebooking search everyone has used for a work queue, the card confirmation on a paid change for a sign-off PIN. Also records, for the slots that follow: the withheld-catalogue lint glob does not list airline; the reset route deliberately reports its memory beats as unarmed; and the two substrates (this one and `useAirlineData`) are both live and deliberately agree about Camila's AV1423. Reskin-skill check: no impact. This commit adds a file and changes no contract, lint rule, test, gate, route shape or skin identity. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7e59e92568 |
feat(reskinnable-demo): add keel's REST API over the policy register
Puts keel on the same substrate as banking, logistics, people and commerce.
The route tree, one line each:
GET /api/keel/v1/ledger the one snapshot read
GET /api/keel/v1/documents/:docId { doc, record }, joined server-side
POST /api/keel/v1/documents/:docId/release BEAT 6 — the gated write
POST /api/keel/v1/documents/:docId/flag BEAT 5 step 1
POST /api/keel/v1/documents/:docId/notices BEAT 5 step 2
POST /api/keel/v1/documents/:docId/notes BEAT 5 step 3
POST /api/keel/v1/variances BEAT 6 — file a draft
POST /api/keel/v1/variances/:id/ratify BEAT 6 — the other half
POST /api/keel/v1/countersignatures BEAT 3a — the e-signature release
GET /api/keel/v1/bulletin?space= BEAT 3d — the ingested document
GET /POST /api/keel/v1/briefs BEAT 3d — the durable artifact
GET /POST /api/keel/v1/runs, /runs/:runId, and the step approve/reject/cancel writes
POST /api/keel/v1/dev/reset presenter reset, gated
`knowledge/:docId` is served by `GET /documents/:docId`, which returns the
corpus prose and the register row together so the two can never be fetched a
moment apart; `runs/:runId` by `GET /runs/:runId`, which returns the run and
its playbook together.
Three separations the tests exist to hold:
- `countersignatures` re-runs the SAME `checkReleaseAuthority()` the release
route runs, so a valid e-signature PIN on an UNENDORSED revision is still
refused with `UNENDORSED_REVISION`. Without that, beat 3a is a second door
around beat 6, the agent takes it, the teach arc never fires and nothing
fails.
- `POST /variances` refuses an uncatalogued code WITHOUT enumerating the
catalogue, while `flag` and `notices` deliberately DO name their valid sets
— beat 5's vocabulary is given to the agent, beat 6's is withheld.
- `dev/reset` reports `reset: ["store"]` and never "memory". Keel has no
intelligence/seed-memories.ts yet, so it cannot re-arm the memory beats,
and a route that claimed otherwise would stop a presenter looking for the
reason beat 6 opened already taught.
Skill impact: checked, none. Adds a skin-local API tree; no `Skin` contract
field, registration site, link builder or lint/test gate the reskin skill
documents is affected. Keel is still correctly absent from eslint.config.mjs's
`withheldGateVocabulary` glob — it has no agent-facing gate file yet, and the
slot that writes `tools.tsx` owns adding it.
|
||
|
|
5bbcf93323 |
feat(reskinnable-demo): add keel's policy-register data substrate
The half of Harbor Point Health that every demo beat needs and that the in-memory skin never had: the LIFECYCLE state of the nine corpus documents. The corpus keeps supplying a document's words; the register supplies its review dates, attestation coverage, effective revision and the revision waiting to be released. Strictly ADDITIVE. `useKeelData`, its 900 ms run ticker and the pages that read it through `useSkinData` are untouched — `types.ts` gains the register types below a marked divider and `personas.ts` gains one strict `findPersona` (the falling-back `getPersona` is right for a render and wrong for a mutation, which would otherwise attribute an unknown id to Ana Reyes). The load-bearing modules: - `release-authority.ts` is beat 6's gate and the ONLY gate. Its refusal names the symptom — document, revision, and which body has not endorsed — and nothing about a way through. - `variance-codes.ts` is the unlock catalogue, four justifying codes and two decoys, WITHHELD from the agent; the header states the withholding the way logistics' escalation-codes.ts does. - `signing-pin.ts` is beat 3a's second factor, and says out loud that its validity is FORMAT-ONLY and is not an authentication control. - `attention.ts` models attestation coverage as a tri-state, so a document nobody is assigned to reads as unmeasurable rather than as 0%. - `register-levers.ts` owns beat 3c's four levers in one normalized record, with an explicit "not pulled" sentinel inside each enum. - `bulletin-citations.ts` / `bulletin-pdf.ts` build beat 3d's ingested document, carrying exactly one policy ref the register does NOT hold — re-checked against the live register, and dropped rather than misattributed once a reseed adds it. Skill impact: checked, none. This adds a skin's internals under src/skins/keel/data/; it touches no `Skin` contract field, no link builder, no registration site, and no lint rule or test gate the reskin skill names. |
||
|
|
a32af92121 |
docs(reskinnable-demo): map keel's demo beats before the substrate
Authored per .claude/skills/reskin/demo-beats.md, before any substrate code: the tools, pages, prompt and pills of every later keel slot come out of this file. Records the two design decisions that are easiest to get wrong later — beat 3a's e-signature PIN is a SECOND FACTOR that re-runs beat 6's release gate rather than a second door around it, and beat 6's publication-variance catalogue is withheld from the agent while its human filing form stays open. Skill impact: checked. This adds a beat map in the shape .claude/skills/reskin/demo-beats.md already prescribes; it changes no contract, command or gate the skill documents. |
||
|
|
8d29771164 | Merge remote-tracking branch 'origin/main' into feat/reskinnable-demo-beat-parity | ||
|
|
8c670653ce |
fix: align Angular 20 support and resolve packed smoke paths (#6452)
## What broke The regression was introduced by commit |
||
|
|
04c4198a14 |
docs: fix Copilot Runtime reference links (#5296)
## What does this PR do? Fixes stale Copilot Runtime documentation links that still point to `/concepts/copilot-runtime` and now route users to the existing `/backend/copilot-runtime` page. This updates both the source JSDoc and the generated reference MDX so the current docs content and future regenerated reference docs stay aligned. ## Related PRs and Issues - Closes #2082 ## Testing - `rg -n "concepts/copilot-runtime" packages showcase/shell-docs/src/content` returns no matches - `rg -n "backend/copilot-runtime" packages/runtime/src/lib/runtime/copilot-runtime.ts packages/react-core/src/components/copilot-provider/copilotkit-props.tsx showcase/shell-docs/src/content/reference/v1/classes/CopilotRuntime.mdx showcase/shell-docs/src/content/reference/v1/components/CopilotKit.mdx` - `git diff --check` ## Checklist - [x] I have read the [Contribution Guide](https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md) - [x] If the PR changes or adds functionality, I have updated the relevant documentation - [x] "Allow edits by maintainers" is checked (lets us help iterate on your PR directly — faster turnaround for everyone) |
||
|
|
62aa7e944e |
docs(reskinnable-demo): record logistics as the fourth demo-complete skin
The beat matrix in CLAUDE.md had logistics ticking beats 4, 5 and 6 while still showing ❌ on beats 2 and 3a-3d, which had shipped earlier on this same branch. Every cell in the logistics column is now derived from the tree rather than from the commit log: Gen-UI count grep -A3 'useComponent(' src/skins/logistics/tools.tsx | grep -c 'name:' -> 6, not 5 Replay-safe the write tools carry structured results Planner PIN beat 3a Readables grep -rln useAgentContext src/skins/*/layout.tsx Four levers status, class, sort, top-N all arrive from the query string Rate sheet beat 3d Prose that contradicted the matrix is corrected in the same pass, in the five places a skin author actually reads: README's "three of the six skins", the skill's routing table ("the only three at 9/9"), its "do not use airline, logistics or keel as demo references" warning, SKILL.md's short version, and templates.md's pill-count paragraph. Pill counts are derived too and the real spread is 4 to 10 (banking 8, people 9, commerce 9, airline 5, logistics 10, keel 4), so the paragraph now says plainly that neither end of that range predicts coverage. Where a number was load-bearing it is replaced with the command that derives it, continuing the convention this branch established: the matrix note now carries the gen-UI derivation alongside the three that were already there, and says why -- this table is prose, it rots silently, and logistics spent two releases showing red on beats it had already shipped. Reskin skill impact: YES, and it is fixed here. Four claims in .claude/skills/reskin/ named logistics as a non-demo reference or counted three demo-complete skins. Left standing they would have routed the next skin author away from the only worked example of retrofitting beats onto an existing skin. src/shell/skin-roster-docs.test.ts (27 tests) passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b5800c35b6 |
docs(reskinnable-demo): fold logistics into the teach-mode roster
The standing rule at the top of CLAUDE.md, answered out loud: YES, this change made things in `.claude/skills/reskin/` and its neighbours wrong. Logistics is now a FOURTH teach-mode skin, and four separate places were written against three. The DERIVATIONS the previous commit introduced all still read correctly — `grep -l offerWorkflowRecording src/skins/*/tools.tsx`, `grep -rln useAgentContext src/skins/*/layout.tsx` and `ls src/skins/*/intelligence/seed-memories.ts` each pick logistics up automatically, which is exactly what they were converted for. What rotted was the PROSE around them: - `docs/teach-mode/README.md` enumerated a role-by-role verdict for every skin the roster grep returns. The grep now returns four and the paragraph named three, so a reader would have found logistics uncertified with no way to tell whether that was an omission or a verdict. Logistics satisfies all five roles; it is also a second worked example of the pinned replay directives, so "so copy commerce" now names both. - That file's memory-SCOPE divergence note gave only half the argument for `user` over `project`. The half it was missing is the decisive one: a skin whose `forget-memories.ts` skips project-scoped rows — logistics and commerce both do — has a presenter reset that physically cannot un-teach a project-scoped memory, so the second run of the demo opens already taught and the beat proves nothing while looking perfect. - `failure-modes.md` § 10 described withholding the vocabulary from the agent and never said to BUILD the surface the operator demonstrates on. Withholding perfectly with no filing form is an unlearnable gate — which is the state logistics was actually in. Points at the worked example and the two load-bearing properties of it. - `demo-beats.md` said the pre-bar skins "ship four or five" pills. Logistics ships ten. Fixed, and the sentence now says so as the reason to run the count rather than trust the prose. `CLAUDE.md`'s beat matrix gains logistics' teach-a-procedure row, and its logistics bullet is corrected: it claimed the skin omits `chatHeaderActions` and `onSuggestionSelect`, which beat 3d set two commits ago. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
eb88fa6139 |
feat(reskinnable-demo): wire Meridian's teach chain and beat-6 pill
The client half of beat 6. The REST gate, the decoy catalogue and the pure-REST proof already shipped; what was missing was the loop that lets the agent learn the unlock by watching. HITL chain, in order: offerWorkflowRecording -> awaitDemonstration -> saveLearnedProcedure. The REPLAY chain is deliberately not new — a later request on a different gated shipment goes through fileEscalation then commitMitigation, the very write that was refused. Nothing is special-cased for the replay, which is the point: the agent applies ordinary tools in an order it was never told. `agent.ts` gains an ACTION DISCIPLINE clause that recalls FIRST and branches on what comes back, so a taught agent does not keep offering to be taught, and that names the plausible substitutes it must not reach for — including the PIN card, which confirms who is acting and never how much they may spend. The five leak channels, each checked by hand because lint sees only two: readable (none registered), schema (fileEscalation's `code` stays a free z.string()), tool description (states the withholding), prompt (names escalation codes as the one thing NOT in context), refusal body (asserted against the whole catalogue in data/blocked-by-authority.test.ts). The pill's comment states only figures `blocked-by-authority.test.ts` re-derives from seed.json, and says plainly which parts of the beat no test in this repo asserts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f1f0329a0c |
feat(reskinnable-demo): pin Meridian's teach-chain replay directives
`docs/teach-mode/README.md` names two replay invariants and says to copy commerce, because commerce is the only skin whose behaviour on them is pinned rather than hand-rolled. This is Meridian's sibling of that module. Both failures are invisible at runtime — the app compiles, the cards render, the demo runs, and the card simply states something the thread does not support: - SURVIVES REPLAY. The recorder REPORTS its step count inside the directive and the card prints the reported number. A card recounting `\d+\.\s` matches in its own prose miscounts every step label containing a numeral, and freight labels carry them constantly. - A SETTLE IS NOT AN ANSWER. Both buttons on the save card settle with a string, so `typeof result === "string"` says the card was answered and nothing about the answer. `classifySaveProcedureResult` never guesses "saved": an unrecognized settle renders as "already answered", never as a receipt for a durable write that may not have happened. Every test round-trips a builder through its reader, so the wording and the matcher cannot drift apart. Scope is `user`, not `project`, and the module says why: `intelligence/forget-memories.ts` deliberately SKIPS project-scoped rows so a Meridian reset cannot delete a sibling skin's seeded memories — which means a project-scoped beat-6 procedure would survive every presenter reset and the second run of the demo would open already taught. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6cf392b43d |
feat(reskinnable-demo): add Meridian's planner escalation filing form
Beat 6's part 5. `data/escalation-codes.ts` reserved `ESCALATION_CODE_LABELS` for "a planner's filing form" and then noted that no such form existed — so the gate was withheld from the agent and unlearnable by it, because there was nowhere for a human to demonstrate the unlock. This builds the form: the ONE surface in the skin where the code vocabulary legitimately appears, because a HUMAN reads it and the agent learns the code by WATCHING the planner choose. - `data/authority.ts` grows `blockedByAuthority`, deriving the gated cases from the same pure option calculation the mitigate route recomputes with, so the form can never advertise a cost the gate would not check. - `components/escalation-form.tsx` lists justifying codes and decoys TOGETHER, unmarked and in catalogue order. A form that flagged the working ones would make the demonstration a guided tour rather than an exercise of knowledge only the operator has. - Both write paths bracket themselves with the shell recorder, and the filing step carries the code as DATA (`logStep(label, code)`) — that is what `getDemonstratedCode()` reads. It records the code the planner ACTUALLY filed, decoy included: a recorder that silently corrected them would report a procedure nobody demonstrated. `blocked-by-authority.test.ts` asserts against `seed.json` itself that the network still offers TWO gated shipments — one to teach on, a different one to replay on — because the case taught on stage is released by the demonstration. A fixture cannot notice that going to one. `on-screen-readables.test.tsx` gains a planner-auth stub: `usePlannerAuth` throws outside its provider by design and these pages render bare. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
48f523b72b |
feat(reskinnable-demo): mount the shell recorder in Meridian's providers
Beat 6 needs the teach-mode recorder, and it has to wrap BOTH the app card and the chat card: the demonstration happens on the Control Tower while the card that reads `steps` / `getDemonstratedCode()` lives in the transcript. A provider around only one of them makes every `logStep` from the other a silent no-op — `useRecording` returns inert fallbacks outside a provider, so nothing throws and the feed is simply empty, discovered on stage. Imported from `@/shell/teach` rather than re-implemented. Three skins each shipped a private copy and they diverged; this is the one module. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6d59e02ab6 |
fix(reskinnable-demo): correct the beat-5 pill's distractor count
The comment said "the seven other registered write tools" and then listed the three real writes among them, which is both wrong and self-contradictory — and it was a hardcoded count in exactly the place this repo keeps learning not to put one. Replaced with the distractors by name plus the grep that enumerates the real set, so the next tool added to `tools.tsx` cannot silently falsify it. Checked, no reskin-skill impact: this is a comment inside one skin's pill list. |
||
|
|
c05463a63a |
docs(reskinnable-demo): retire the "three skins have memory" count claims
THE STANDING RULE, ANSWERED. Adding beats 4 and 5 to logistics falsified five
sentences in `.claude/skills/reskin/` and two in CLAUDE.md — every one of them a
hardcoded count of which skins ship `intelligence/seed-memories.ts`, and every
one true when written:
- demo-beats.md § Presentation requirements ("the only three with
seed-memories.ts + forget-memories.ts; logistics restores its data store only,
so it cannot reset beats 4–6")
- demo-beats.md § Which skin to copy for what ("the only three with a seed file")
- demo-beats.md's "do not use these as demo references" paragraph (logistics has
"no memory prompts, no memory tools and no seed file")
- templates.md § seed-memories ("two of the only three in the repo")
- templates.md § copy forget-memories from COMMERCE ("All three skins have one")
- SKILL.md § intelligence ("`banking`, `commerce` and `people`")
- CLAUDE.md's agentRegistry bullet and its beat matrix
Each is replaced with its DERIVATION rather than a corrected number —
`ls src/skins/*/intelligence/seed-memories.ts`, `ls -d src/app/api/*/v1/dev/reset`
— which is the fix failure-modes.md § "counts rot" already prescribes and which
the next skin cannot re-falsify. The two skin-count guards in
`skin-roster-docs.test.ts` caught both halves of this: the new prose, and the now
dead COUNT_EXEMPTIONS entry for a phrase that no longer exists. That exemption
list is now empty, with a comment saying to prefer a derivation over adding a new
one — an exemption is a promise to hand-recount a subset forever, and that file
exists because the promise is not kept.
The one substantive addition rather than a correction: the memory half of a reset
is now called out as a SEPARATE, narrower question from having a reset route at
all, because a skin with a route and no seed file has a Reset button that looks
identical and cannot re-arm beats 4–6.
Nothing else in the skill went stale. The `Skin` contract is untouched, no link
builder or hook changed, no lint rule or gate moved (`withheldGateVocabulary`
already covers logistics' `tools.tsx` and `agent.ts`, and no file's selector set
changed), registration and the client/server boundary are unchanged, and beat 6's
mechanism is exactly as it was.
|