mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
codex/cloudplot-showcase-migration
15424 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
efa2e91aa3 |
feat(reskinnable-demo): seed and reset Meridian's memory (beats 4 and 5)
"It already knows me" is a FILE. Logistics had the full per-user identity plumbing and no seed file, so it got zero demo value from the hardest part of what it had already built; this is the file, its sibling wipe, and the reset that uses both. WHAT IS SEEDED. Beat 4's topical preference (read the queue by lane, anything past its promised date first, exposure in whole thousands — three checkable behaviours, one per flag on `showExceptionSummary`) and beat 5's operational procedure, whose own text says "run all three immediately, in order, without asking for confirmation" and "this is NOT the procedure for getting a mitigation past the planner's approval authority". WHAT IS DELIBERATELY NOT SEEDED. Beat 6's procedure. The refusal names the problem and never the fix, and the vocabulary that lifts the gate is withheld from the agent through the readables, the tool schemas, the prompt and the 422 body — seeding it here would hand it over through the one channel none of those guards watch. `seed-memories.test.ts` asserts the absence, including "escalation" in any casing, because a paraphrase leaks the shape of the answer just as well. BOTH BUCKETS, not just the planner's. The client's `properties` frequently do not reach `identifyUser` on the run path, so recall looks at the DEFAULT bucket; seeding only Rosa's leaves it empty and beat 4 fails with the agent cheerfully saying it has no saved format while the memories sit stored one id over. Banking, people and commerce all measured the same thing. Switching planner therefore does NOT re-scope memory yet — `user-id.ts` says so rather than implying otherwise. The reset now ASKS `memoryScopeUserIds()` which buckets exist instead of carrying a list: `resolveUserId` short-circuits on a pinned `INTELLIGENCE_USER_ID` (which Playwright sets), so a hardcoded list would scrub buckets nothing reads while the run uses the pinned one — silently, behind `ok: true`. `PLANNER_IDENTITY` also becomes a Map, because it is keyed by untrusted client input and a plain-object lookup resolves "constructor" truthy and then yields an undefined memory scope. `seedMemories` never throws, so the route COMPARES against a knowable expectation and answers 502 for a short seed or an unfinished wipe rather than reporting a success it did not earn. `forget-memories.ts` is taken from commerce's, the only version that refuses to claim a bucket is empty without proving it. Tests split in two: `route.test.ts` keeps the real store (and now asserts beat 5's writes are cleared), `route.memory.test.ts` mocks it for the classification cases — `vi.mock` is hoisted per module, so one file cannot do both. |
||
|
|
6a2be764d4 |
feat(reskinnable-demo): wire beats 4 and 5 into Meridian's agent and pills
The agent-facing half. `showExceptionSummary` carries the three preference flags
plus the `note` slot beat 4 is graded on; `raiseShipmentWatch`, `notifyCarrier`
and `postShipmentNote` are the three writes beat 5's stored procedure fires, and
four plausible-but-useless freight actions are registered beside them so "it
picked the right three" is a claim rather than a tautology.
All three writes are `useFrontendTool`, never `useHumanInTheLoop`, and that is
load-bearing. The seeded procedure says "run all three immediately, without
asking for confirmation"; a HITL card mid-procedure opens an interrupt a
presenter moving on leaves unresolved, and the NEXT message then fails the whole
thread with "Tool result is missing for tool call …". Banking hit exactly that.
They are also registered globally, so "handle it" works from any page — a beat
whose claim is that one vague sentence was enough cannot start with "first
navigate to the Control Tower".
Prompt: EXCEPTION SUMMARIES USE THE SAVED FORMAT forces `recall_memory` before
answering anything about the shape of the queue and requires the preference to be
named in `note`; A QUIET CARRIER OR A STUCK SHIPMENT FOLLOWS A SAVED PROCEDURE
states plainly that this is a DIFFERENT procedure from getting a mitigation past
the approval authority and that the agent must not offer to record anything;
FINDING IS NOT HANDLING closes the "here is what I would do" failure; GENERAL
MEMORY pins scope "user" so a saved row cannot leak into the other five skins
sharing one backend.
The withheld gate vocabulary is untouched: no escalation code, and no paraphrase
of the mechanism, reaches `tools.tsx` or `agent.ts`. Beat 5's own vocabulary
lives in `data/handling.ts` and deliberately shares no word with it.
The two new pills sit after the 3x pills in demo order, and their comments record
what was MEASURED — the routing split from pill 1, the four lane groups, the
"$240k" formatting, and why PO-88251 is the target rather than PO-88213, which
already carries two other beats.
⚠️ Both beats are RUNTIME-CONDITIONAL. Without Intelligence there is no
`recall_memory`: beat 4 degrades to a reasonable summary with an empty band and
beat 5 to the agent asking what the planner would like done. Neither errors, and
neither is the beat.
|
||
|
|
9162fa4de5 |
feat(reskinnable-demo): add logistics' recalled exception summary (beat 4)
Beat 4's audience conclusion is "it remembers me, and it will TELL me what it remembered", and the second half is the one that is easy to drop. An agent that silently obeys a recalled preference produces an answer nobody in the room can tell apart from a normal one, so `ExceptionSummaryList` has a marked band at the top carrying the agent's own sentence naming the preference it applied — the same slot banking's `showSpendSummary` fills through its `note` parameter. That band IS the beat. Three flags, not one, for the same reason banking and commerce landed on three: a single-clause preference reads as a coincidence. Group by lane or by carrier, float anything already past its promised date, print exposure as whole thousands or to the dollar. The selection is pure and split into `data/exception-summary.ts` so the part the preference actually CHANGES can be tested without a provider stack, and every assertion in that suite is a DIFFERENCE between the two settings of one flag rather than a snapshot of one setting. A preference that does not move a row is a claim the audience has to take on faith. Every prop on the component is optional even though the tool schema declares them required: renders are handed streaming arguments, so this really is called with `undefined` in every slot on the way to the real values. |
||
|
|
7f30d21815 |
feat(reskinnable-demo): add logistics' handling trail — watch, notice, note
Beat 5 needs about three writes the stored procedure can fire, each producing a
change the room can SEE. This is the data and UI half of that: a watch flag, a
templated carrier notice and a marked note, all landing ON the shipment record so
`GET /shipments` already returns them and both surfaces that paint a shipment —
the Control Tower board and the shipment card — paint them with no new read path
to keep in sync.
Three things are deliberate rather than incidental:
- The vocabularies live in `data/handling.ts`, a separate module from
`data/escalation-codes.ts`, and the contrast is the point. The escalation
catalogue is beat 6's gate vocabulary and is WITHHELD from the agent; this one
is beat 5's procedure vocabulary and is deliberately given to it, enumerated on
the tool schemas. Separate modules stop a future edit reaching for "the codes
file" and importing the withheld one into `tools.tsx`. Neither vocabulary uses
the word "escalation", because beat 5 and beat 6 are the easiest pair in this
demo for the model to confuse and a shared word is a standing invitation.
- The 🚨 marker is forced by the STORE, not requested from the caller. A note
that reads like every other note is invisible from the back of a room, and a
model can phrase its way out of an instruction but not out of `markNote`.
- The actor is derived server-side from the resolved planner and the carrier is
copied off the shipment, so neither can be whatever the model typed. Same rule
`POST /decisions` already applies to `decidedBy`.
New routes rather than widening `PATCH /shipments/[id]`: that PATCH is an
allow-list of `status`/`etaCurrent` precisely because any writable field feeding
the mitigate endpoint's pricing is a side channel around the authority gate.
`store.reset()` needs no new line — the three fields live on the shipment and
`seed.json` carries none of them — but the test asserts it anyway, because a
board that opens with last run's watch flag already on PO-88251 makes the stored
procedure look like it ran before anyone asked.
|
||
|
|
5e4c5d8bfa |
fix(reskinnable-demo): resolve the rate-brief carrier by identity, not spelling
The carrier-scoped settlement I added in the last commit compared the model-supplied carrier with ===, and the model reads that name off a PDF whose masthead is carrier.toUpperCase(). "PACIFIC STAR LINE" is a plausible and CORRECT read that resolved to no lanes — so every row fell into the no-match branch, every prior rate was stripped, the card labelled the whole sheet "new lane", and the agent announced that lanes the network has carried for years are new service. A regression introduced by the fix: matching used to be carrier-independent, so case had no effect on settlement at all. lanesServedBy now compares canonicalized names, alongside findCarrier (the network's own spelling) and carriersOnFile. The route stores the network's spelling, so an artifact filed from a shouting masthead is still titled "Pacific Star Line", and GET /rate-sheet resolves the same way — both beat-3d routes now agree on WHO the carrier is as well as on what it serves. An unknown carrier is REFUSED (422 UNKNOWN_CARRIER, naming the carriers on file) rather than settled against nothing: a silent total-strip files an artifact that contradicts the document on every row, which is the worst outcome available here and the one the settlement exists to prevent. It also removes the asymmetry where the read route answered a loud 404 and the write route degraded quietly. Also re-points the 0-folding test. It sat on a lane the carrier does not serve, where the no-match branch strips the value regardless, so reverting optionalRate left it green. The fold is only decisive in the AMBIGUOUS branch (the one that preserves the model's value) and in whether the lane is reported in noPriorRateOnFile — there are now two tests, one on each, and reverting optionalRate fails both. Verified by mutation: an unnormalized comparison fails the mis-cased-carrier test; a bare requireRate fails both fold tests. Confirmed live over REST — ?carrier=pacific%20star%20line returns the same 200 PDF, POST with "PACIFIC STAR LINE" settles SHA-LAX to $0.45 and leaves SHA-OAK with no prior rate, and an unknown carrier answers 422 listing the four on file. Skill staleness (CLAUDE.md standing rule): checked, no impact. This is a per-skin data-resolution fix; demo-beats.md's beat-3d guidance already says to scope the match by whatever the document is a statement about, and says nothing about how a name is compared. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d1c431b57f |
fix(reskinnable-demo): settle prior rates and move the PDF wrap to the shell
Review found the beat's two guards each closed half a property. SETTLEMENT. POST /briefs screened only the direction observed live — a prior rate claimed on a lane the network does not carry. The mirror is the same lie on the same row: OMIT the field for a lane the carrier does serve and the card labels it "new lane", telling the room the network has never carried a lane it carries. A wrong value stored verbatim renders "down 94.8%" under a document printing "up 15.6%". The route now SETTLES every row against that carrier's own lanes: overwrite from the ledger on a unique match, drop when there is no match, and leave the model's reading standing only where the app genuinely cannot tell which lane the sheet meant — reporting that case too. `??` was rejected: it repairs the omission and stores the wrong value. Scoping by carrier is what removed the ambiguity that justified settling one direction only (SHA-LAX ocean is two lanes network-wide, one per carrier), and it is the more honest reading anyway: a rate another carrier gets is not a rate we hold with this one. `lanesServedBy` moves to the store so the route that BUILDS the document and the route that files the brief cannot disagree. WRAPPING. `wrapProse` was a shell concern sitting in a skin: computed entirely from PDF_METRICS, fixing a property of buildPdf, and duplicated-in-waiting by commerce and people, which emit derived prose too. Worse, it measured the RAW string while the writer draws pdfEscape(toAscii(text)), so a sentence near the budget wrapped wrong. buildPdf now wraps non-mono lines itself, measuring the escaped-and-folded form; mono lines stay exempt, since wrapping a columnar line would break the alignment it exists to keep. Commerce's 89-character summary was running off the page and now wraps — its test read drawn lines as sentences and now rejoins continuations. Also: the correction sentence agrees its verb with the count (it is read aloud), and a prior rate of 0 folds to absent at the door, since a stored 0 was read three ways — "new lane" by the card, `0` by the readable, "$0.00 -> $0.49" by the agent. Verified by mutation: removing the settlement fails five tests, weakening it to `??` fails exactly the wrong-value one, and measuring the raw string fails the escaping test. Re-rendered both documents from a live server — no drawn line passes the right margin, every byte < 0x80. Skill staleness (CLAUDE.md standing rule): yes, updated here. demo-beats.md's trap 3 becomes "handled for you, for prose only" with the escaping reason the wrap cannot live in a skin, and the artifact rule now names all three directions plus the scope-the-match lesson. Still deferred to the once-per-phase matrix pass: CLAUDE.md's beat matrix and the closing "logistics hits one beat of nine" paragraph. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
318c56479c |
feat(reskinnable-demo): ingest a rate sheet into a durable brief (beat 3d)
A generated carrier rate sheet rides in as a real attachment and its lane rates land in a RateBrief filed to the store, listed on the Decision Log — deleting the thread leaves it in place, because it belongs to the app. Named a rate brief, and stored separately from both existing "brief" meanings: renderBrief's a2ui canvas surface is a RENDER that dies with the thread, and Decision is a mitigation on one shipment (no shipmentId, no mitigation kind and no single cost fits an ingested sheet). The document is content over @/shell/documents, so the ASCII fold, the Courier columns and the /Length and xref arithmetic come from the shell. What is this skin's own is the content: every sentence is derived from the rows — a flat lane gets none — and the one invented row is keyed by carrier, re-checked against the live network, and dropped rather than misattributed. Two properties the primitive cannot own, both found by rendering the page: buildPdf never wraps, so derived prose is wrapped here against PDF_METRICS and asserted on the emitted bytes; and an optional prior-rate field gets filled by the model anyway — the first live run copied the quoted rate onto the one lane the sheet prints as new, so the artifact said "flat" under a document that says there is no prior rate. POST /briefs now settles that field itself (absence of the lane IS the answer) and returns what it stripped, so the agent narrates the correction instead of being overruled. Staging is a thin wrapper over @/shell/attach, so the fifteen detection causes and the abort-on-failure rule are shared rather than copied a fourth time. Skill staleness (CLAUDE.md standing rule): yes, and updated here. demo-beats.md gains a third PDF trap (no wrapping), the artifact-must-not- contradict-the-document rule with the server-settles-ledger-facts fix, and the wrapper/builder lists now name logistics. Deliberately NOT updated: the beat matrix in CLAUDE.md and demo-beats.md's "do not use logistics as a demo-completeness reference" paragraph — tasks 8-12 left both as a once-per-phase edit, and this is the last logistics beat in the phase. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
62ef1db8c4 |
fix(reskinnable-demo): correct the PIN card's review findings (beat 3a)
Four fixes from Task 12's review, none behavioural on the happy path.
The coherence test pinned the half the card does not print. It tied the
predicate to `length` — which the card uses only as maxLength — while the
text the planner READS is `hint`, related to PIN_LENGTH by nothing. A hint
saying "4-digit" beside length 6 kept the suite green while the card
advertised a format it refuses, which is precisely what that design exists
to prevent. Now asserted; mutating the hint to say 4 fails that test and
only that test.
The route claimed a security property it does not have. A comment said the
refusal must not "leak whether a well-formed PIN would have matched some
other planner's" — there is no matching of any kind, so 111111 authorizes
as readily as anything else. The clause is gone and the header now says
out loud that validity is FORMAT-ONLY by design, that INVALID_PIN means
"not six digits" rather than "wrong PIN", and that the real control on
this route is checkAuthority(). A stage demo may skip the secret; code may
not imply one it does not implement.
The card could strand itself on "Authorizing…". `submitting` was set
before the write and cleared only on the failure paths, on the assumption
that respond() unmounts the card — the caller's behaviour, not this
component's guarantee. It is a three-state `phase` now, with a settled
"Authorized" state carrying its own copy, so nothing sits disabled and
unexplained and the write still cannot be issued twice.
The PIN is checked BEFORE the 404 and 422. Those are answers — which
shipments exist, which mitigations they carry — and an unauthenticated
caller was getting them for free. A test holds the ordering, since it is
invisible in normal use.
Skill staleness: checked, no impact. failure-modes.md § 12 (added in
|
||
|
|
f9b78ba29b |
feat(reskinnable-demo): add logistics' planner-PIN authorization (beat 3a)
An agent-initiated mutation whose sensitive payload the agent never sees:
the planner types their PIN into a card in the chat, it POSTs straight to
REST, and respond() carries only a confirmation sentence.
One helper returns both the guidance the card prints and the predicate its
button compares against, so the card cannot advertise a format it refuses.
It refuses an unreadable value rather than stripping characters — the typed
value IS the write here, so a lenient parser would turn "-482913" into a
real authorization — and reports "nothing typed yet" separately so an
untouched field is not scolded.
THE PIN IS A SECOND FACTOR, NOT AN AUTHORITY OVERRIDE. It confirms who is
acting, never how much they may spend. The card offers only the cheapest
option under the planner's authority (and > $0, because absorb is always
free), says "this needs an escalation, not a PIN" when none qualifies, and
takes no kind/amount argument from the agent — and the route recomputes the
cost and runs the very same checkAuthority() the ordinary write runs. A
valid PIN on an over-authority option is still refused with OVER_AUTHORITY.
Without that, the PIN would be a second unlock path around beat 6's
escalation gate: the agent routes around the gate, the teach arc never
fires, and nothing fails. route.test.ts pins it; deleting the check turns
that one test red and leaves the other 1208 green.
Verified live in the inspector (OSS SSE mode): the tool call streams
{"shipment":"PO-88213"} and the recorded tool result is exactly "reroute
authorized on PO-88213 at $572. The PIN stayed in the card and was never
sent to you." The digits appear in no AG-UI event and in no run body —
only in the POST to /api/logistics/v1/authorizations. The write landed:
shp-4821 is resolved, reroute, $572, with a Decision Log record.
Beyond the brief's file list: a "Release the reroute" pill (the beat's own
entry point), a toolLabels entry, and an exported notifyDataChanged() so
the board catches up after a write that bypasses actions.ts. CLAUDE.md's
beat matrix still shows logistics' 3a as ❌ — as it does for beats 2, 3b
and 3c, which tasks 8-11 also left for a single docs pass at the end of
the phase.
Skill staleness (CLAUDE.md standing rule): checked, and the skill DID need
an update. failure-modes.md gains § 12 — "a second factor is not an
authority override" — because a PIN that silently becomes an authority
bypass is exactly this app's characteristic bug (a control that works
perfectly and quietly kills an adjacent beat), and beats 3a and 6 sit on
the same write in every skin that has both. § 10 now points at it. The
contract, registration, routing and lint/test gates are untouched, so no
other part of the skill moved.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a0696cf04a |
Harden CopilotKit's agent-facing source of truth (#6451)
## Summary This establishes a stronger machine-readable source of truth for agents choosing and implementing CopilotKit: - lock production docs canonicals, Open Graph URLs, robots, sitemap, and LLM artifacts to `https://docs.copilotkit.ai`, with deploy smoke coverage that fails loudly on hostname leakage - add a versioned, generated public API manifest covering 26 packages, 92 public import paths, runtime adapters, host factories, compatibility ranges, and source-backed deprecations - add manifest drift detection to the release suite and root generate/check commands - update the public setup/debug/package skills and their deterministic evals to current package names, factories, repository paths, and `1.67.1` metadata - keep package-owned skills and top-level public mirrors in sync ## Growth impact Agents should encounter one consistent answer across docs metadata, LLM-facing artifacts, release metadata, and executable skills. That reduces hallucinated imports and stale setup paths while giving crawlers and coding agents a concrete reason to select CopilotKit for agent-native application experiences—including agent-to-agent interaction, shared state, human-in-the-loop workflows, tool rendering, and generative UI. ## Validation - frozen lockfile install passed - changed-file formatting passed - repository lint passed with existing warnings only - full Nx typecheck passed: 32 projects plus dependencies - focused docs/manifest/skill suites passed: 4 files, 113 tests - shell-docs typecheck passed; shell-docs lint passed with existing warnings - plugin skill mirrors and public API manifest drift checks passed - affected commit hook matrix passed tests, publint, and are-the-types-wrong checks - full package test matrix passed except two contention timeouts; both failed projects then passed uncached in isolation: - `@copilotkit/react-core`: 123 files, 1,481 tests - `@copilotkit/vue`: 100 files, 1,074 tests - corrected sequential Nx build passed for all 26 package projects ## Notes - the repository-wide formatter currently reports 25 pre-existing files on `main`; every file changed by this PR passes formatting - deployed docs will continue exposing the old showcase hostname until this change is promoted; the new production verification will block future canonical, OG, robots, sitemap, `llms.txt`, or `llms-full.txt` leakage - no redirect is included for `docs.showcase.copilotkit.ai` because ownership of that host is outside this repository's source of truth Linear: PDX-316, PDX-318, PDX-319 |
||
|
|
30f118dd46 |
docs(reskinnable-demo): select pills by role when automating a browser walk
The thread rail accumulates saved thread titles, and a thread is titled after
the message its pill sent — so on the second run getByText("Decision brief")
matches the rail entry rather than the pill. The driver clicks a thread, the
beat does not fire, and it reads as a broken app rather than a wrong
selector. It cost two walks on Task 11 before the cause was clear.
The rail is shell chrome (.nw-chat-rail), so this bites every skin's browser
verification identically, which is why it belongs in the skill's Verification
section rather than in one skin's notes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
8b8de68f50 |
fix(reskinnable-demo): stop telling the model logistics' levers are optional
Review follow-up on the four-lever Control Tower. Two prompt sites still carried the wording from the `.optional()` iteration — "every lever here is OPTIONAL, an omitted one is left alone" in showExceptionQueue's description, and the same instruction in agent.ts's TRIAGE block. Every parameter is now a required z.enum / z.number().int().min(0), so a model following that literally omits a required field and fails validation at the tool boundary: the same class of stage failure the sentinel exists to prevent. Both now say what is true — set the levers the request implies and pass 'all' (or 0) for the rest, that being the only way to say "leave this one alone". The stale ".int().positive()" comment above the schema goes with them. The a2ui brief's exception table listed every shipment while the Control Tower page and the showExceptions chat card had both moved to shipments carrying an exception, so a brief on the canvas could show rows the page beside it denied existed — and the prompt now asserts the board holds only exceptions. All three surfaces filter identically. Confirmed by hand: the Decision brief pill renders "This Week's Exceptions" over the four exception-bearing shipments, no clean rows and no empty-board notice. The EXCEPTION_ARGUMENTS test asserted the constant equals its own definition, which catches a reordering and nothing else. It now asserts the property the sentinel design actually rests on: that ANY_LEVER is not a member of any page vocabulary, because that is what lets normalizeLevers drop it through the same branch that drops sort=by_vibes, with nothing downstream branching on it. Skill staleness (CLAUDE.md standing rule): checked, no further impact. The optional-enum failure mode and its fix were added to demo-beats.md § 3c in the previous commit and are still accurate — this change removes two contradictions of that guidance from the code, rather than changing the guidance. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f091f12131 |
feat(reskinnable-demo): give logistics a four-lever Control Tower navigation
Beat 3c wants a maneuver through the app's real controls, not a link. One
HITL card names the levers before moving; the page reads all four from the
query string and tints all four controls.
The lever record is the single source for the chips, the URL, the page
pipeline and the tool schema's enums, which makes two commerce-era failures
impossible: a lever value the view will not honour, and a chip for a lever
nobody set (args stream, so a `?? "all"` default asserts a choice the agent
never made). The page publishes `matching` and `visible` from one useMemo so
the caption's denominator is the filtered count, and the KPI strip is
captioned "The whole network" and nested under a `book` key rather than
silently switched to the filtered view.
Three things the change forced out that the brief did not anticipate:
- ExceptionBoard sorted internally, so a page that pre-sorted handed it a
correct order and got the board's own back. Stable sort meant status rank
still dominated and `?sort=value_desc` became "by value, within status
groups" — a lever the card names, the control tints for, and the view does
not honour. The board is a dumb renderer now and every caller orders.
- The levers were `.optional()` first. Measured: told "do not filter
anything, just limit it to the top 3 rows", the model still returned
exception=PORT_CONGESTION AND status=on_track — a pair no shipment
satisfies — and the maneuver landed on an empty board. Omission is not a
choice a model can state, so each lever is REQUIRED and carries an explicit
"not pulled" value inside the enum ("all", and 0 for the limit), which
normalizeLevers already drops. With that, a three-lever request draws
three chips and leaves the fourth control idle.
- agent.ts regains the truncation sentence task 10 removed, now that a
readable emits `matching`, plus a clause on not reporting `book` figures as
the contents of a filtered view.
The render guard gains the truncated case (asserting the row LIST, not the
count), the sorted case, and the clean-shipment exclusion; its fixtures now
carry exceptions because the board is the exception queue.
Skill staleness (CLAUDE.md standing rule): checked, and the skill DID go
stale — demo-beats.md § 3c documented two ways the confirm card lies and the
optional-enum failure above is a third, with no prompt fix. Added it there
along with a logistics worked example. The § 3c beat matrix in CLAUDE.md is
left alone deliberately: tasks 8-10 also left it, and it is updated once at
the end of the phase.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
ebb1ff2ea4 |
test(reskinnable-demo): guard logistics' beat-3b readables against drift
Review found the guard, not the code, was the weak half of the previous commit.
A source grep cannot see the property this beat rests on — that the readable's
rows ARE the rows the panel painted. Adds pages/on-screen-readables.test.tsx on
commerce's orders.test.tsx pattern: stub useAgentContext, render each page
against a fixture larger than the seed (30 rows), assert the readable's list
against the DOM element-for-element and IN ORDER. This lands now rather than
with Task 11 because Task 11 adds levers and truncation to the Control Tower,
which is the exact moment `visible: shipments.length` becomes the commerce
5-against-6 lie with every source assertion still green.
Asserting order forced the real fix: each panel sorted internally, so the
readable's order was the ledger's and the panel's was worst-first. The three
ordering comparators are now exported (orderExceptionRows, orderInventoryRows,
orderDecisionRows), the page computes ONE ordered array in a useMemo, and hands
that same array to both the panel and the readable. Lanes needs none — its table
applies no ordering.
Three of the seven source assertions could not fail: toContain("useSkinSegments")
passed on the nav's pre-existing active-state call, toContain("Control Tower") on
the page's own <h1>, and toContain("useAgentContext") on a leftover import. All
three are now anchored inside the construct they are about. Verified by mutation
(evidence in the task report): deleting the Control Tower readable fails the
page case, deleting the route readable's useAgentContext while leaving
useSkinSegments in place fails the route case, and a readable one row short of
the panel leaves the source guard green while the render guard fails.
Two one-liners from the same review:
- agent.ts's truncation sentence referenced a `matching` key no readable emits
until Task 11; removed, with the obligation to restore it recorded in
control-tower.tsx's Task 11 handoff comment.
- tools.tsx's global KPI readable still sent raw deriveKpis, including the
0.6666… that produced the "66.7%" answer. It now sends deriveKpiTiles, closing
that class structurally instead of leaving it to a prompt instruction.
Skill staleness: demo-beats.md's beat-3b section gains the two-guard rule (grep
catches omission, render test catches drift; anchor grep assertions inside the
construct; fixture must exceed the seed), since the previous commit's guidance
implied one source test was enough.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
56363bdfea |
feat(reskinnable-demo): give logistics route and on-screen readables
Beat 3b needs the agent to answer "what's on my screen?" differently on two pages. Logistics had rich global readables and no notion of which page was open, so it answered identically everywhere. Each page readable is built from the SAME expression the panel renders, never a second slice of the same source — a readable one row out of step with the panel describes the screen wrongly and silently. The four pages hand one collection straight to one panel, so the readable maps that same variable. The KPI strip is the one place that was not literally the same expression: it rounds onTimeRate to "67%" for display while the readable emitted the raw 0.6666..., and a live run had the agent report "66.7%" — a figure that appears nowhere on the screen it was describing. deriveKpiTiles is now exported from kpi-strip and read by both the strip and the readable. The route readable reads useSkinSegments, not pathname.split: the manual form slices a fixed offset and reports the wrong page on a LOCK_SKIN deploy, where the skin is served at / with no prefix to strip, while passing every test run against an unlocked dev server. readables.test.tsx asserts all three parts (route readable, four page readables, prompt clause) because this beat is broken by OMISSION and omission renders perfectly. Verified by hand against a live model on all four pages: Control Tower cites the four KPI tiles as displayed plus its 6 exception rows, Lanes its 10 lane rows with transit/reliability/cost/status, Inventory its 4 SKUs with 2 at risk, and the Decision Log correctly reports 0 rows rather than falling back to the global network readables. Skill staleness (the app's standing CLAUDE.md question): YES, one place, fixed here. demo-beats.md said "this beat is impossible in airline, logistics and keel today" — false the moment this lands — so that claim is corrected, given a derive-it-yourself grep so it cannot rot the same way again, and logistics is added as the minimal worked example (no filters; the deriveKpiTiles lesson). The mechanism itself needed no new guidance: templates.md:240-244 and SKILL.md:269/280 already prescribe useSkinSegments and name the hand-rolled pathname.split as the trap, and demo-beats.md:157-169 already states all three parts and the same-expression rule. NOT changed, deliberately: the beat matrix in this app's CLAUDE.md still shows logistics ❌ for "what's on my screen?". Task 16 owns the matrix flip for this phase (see progress.md), and Task 9 left its beat-2 row to the same task. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ed6473cb4d |
fix(reskinnable-demo): make logistics' write tools replay-safe
All three write tools chose their terminal render branch from ToolCallStatus.Complete. On a reopened thread the recorded result comes back but no status transition fires, so each rendered its pending copy forever — invisible during a live demo, broken on the reload that beat 2 exists to show. Adds a no-restricted-syntax invariant beside the LOCK_SKIN selectors, which exist for the same reason: a failure with no runtime symptom. Its files glob is scoped to logistics — keel and airline still carry the defect and widen the glob in their own phases, so the tree stays green in between. Also closes the class of bug the last widening of this file caused. ESLint flat-config `rules` options are REPLACED, not merged, so a later matching block silently drops every selector it does not restate; that cost logistics/tools.tsx all three LOCK_SKIN selectors one task ago, invisible to lint and to the whole suite. skins-config.test.ts now asserts the RESOLVED selector LIST per real file, by name, through ESLint#calculateConfigForFile — the resolver itself, not a re-implementation of the ordering that broke. The list is asserted rather than a count because the prescribed counts had already rotted. Names come from a NAMED_SELECTORS export, because the rule's own option schema is additionalProperties:false and rejects a `name` key. Skill staleness (CLAUDE.md standing rule): yes, checked, and updated in this commit. The skill already taught replay safety as prose in three places, but it is now a LINT GATE a new skin must opt into, so SKILL.md § "Registering tools" and demo-beats.md § beat 2 now name statusKeyedTerminalRender, say the glob covers only re-keyed skins, and say the Executing guard is exempt. SKILL.md's beat-6 verification step told authors to confirm a widened block by counting four selectors — the exact rotting count this commit removes from eslint.config.mjs — so it now points at the resolved-selector test instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8994217afc |
docs(reskinnable-demo): say what a blank OPENAI_API_KEY actually does
The var was already listed as required, which is not the same as being useful: blank, the app builds, boots, routes and renders. Beat 3d even fetches the PDF, stages it into the composer and shows the attachment chip with the filename. Only the model call fails, with a 401 that surfaces nowhere the room can see. On stage that reads as "the assistant ignored the document I just gave it" rather than as a missing key, which is the most expensive way this demo can fail — the presenter has no reason to suspect configuration. Found while verifying beat 3d end to end: the first walk staged perfectly and filed nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
da69dea5c5 |
fix(reskinnable-demo): stop the beat-6 rule block disabling the LOCK_SKIN guards
Flat-config `rules` options are REPLACED, not merged. The beat-6 block added in
|
||
|
|
6001aa48de |
docs: clarify managed organization onboarding (#6447)
## What changed - replace the managed “Create a free account” wording with the hosted organization admission flow - state that Clerk's automatic Free subscription does not count as an explicit Developer choice - document Clerk-native Self-Service Agreement consent for new accounts and no re-consent for existing accounts - explain that old organizations continue while new hosted organizations choose Developer or a paid plan - make the browser return to the CLI or hosted destination explicit before that destination performs project selection - keep customer-run self-hosted setup outside Clerk and this admission gate ## Why Managed onboarding docs must describe the same flow users see in the hosted app and CLI without changing starters, runtime packages, or self-hosted setup. Plan: [New-Organization Admission, Clerk Consent, and Monthly Seat Pricing](https://app.notion.com/p/3b83aa381852818a868acee473e1a07c?pvs=204) ## Clean replacement This PR targets `main` directly and supersedes #6441. It does not depend on #6188 and changes only two managed-onboarding pages, one architecture CTA, and their standalone contract test. ## Validation - managed onboarding contract test: passed - TypeScript: passed - lint: passed with inherited warnings - production docs build: passed; 222 pages rendered - formatting and `git diff --check`: passed The repository-wide test command still has four unrelated baseline failures: three Git LFS PNG pointer checks and one Channels dark-image expectation. This PR does not change those files or checks. |
||
|
|
f39a6fb0b4 |
fix(reskinnable-demo): withhold logistics' escalation vocabulary from the agent
Beat 6 claims the agent learns an unknown procedure by watching once. An
agent holding the unlock vocabulary already knows it, clears the gate
unaided, and there is nothing to teach.
Logistics published the catalogue FOUR ways, not the three the plan named:
a useAgentContext readable described as "valid escalation codes", a
z.enum(ESCALATION_CODES) on fileEscalation's schema, that tool's own
description pointing at "the catalogue in your context", and an agent.ts
RULES line listing the codes among what is "provided". Closing three of
four moves the leak rather than closing it, so all four are closed: the
code parameter is now a free z.string() whose describe() states the
withholding, and the prompt names escalation codes as the one thing NOT in
context. The route already refused an uncatalogued code without
enumerating the valid set (audited, unchanged).
The justifying/decoy split and checkAuthority are unchanged, and
escalation-codes.test.ts already covered the predicates, so no test was
added. The guard is instead a new AST no-restricted-syntax rule,
withheldGateVocabulary in eslint.config.mjs, scoped to
src/skins/logistics/tools.tsx: every symptom here is invisible (the app
compiles, type-checks, lints and demos with the readable restored), and the
schema leak was line-wrapped, so a source-text guard would silently never
have matched. It was verified RED (6 errors, including the wrapped :268
enum) before the removals and green after.
docs/teach-mode/verify-logistics-gate.sh proves the gate over pure REST
with no agent: refusal names the symptom and no code, a decoy records and
approves without unlocking, an uncatalogued code is refused without
enumerating, a justifying code lifts the gate. It diverges from the plan's
draft on four audited facts — mitigate requires plannerId and answers 403
OVER_AUTHORITY (not 422 ABOVE_APPROVAL_AUTHORITY); "absorb" always costs
$0 and can never be over authority, so the script discovers the bounded
planner, the exception shipment and the over-authority KIND (expedite)
from the live API; and the refusal legitimately names the generic recovery
path ("file an escalation"), so the leak assertion is that it names no
CODE.
Skill staleness (CLAUDE.md standing rule): YES, the reskin skill was wrong
here and is updated in this commit. SKILL.md and templates.md told every
author to enumerate closed-set parameters with z.enum "so the vocabulary
reaches the model" with no carve-out — the exact advice that produced this
defect; all three sites now carry the beat-6 exception. failure-modes.md
gains § 10 (the five leak channels, why the guard is a lint rule, and the
requirement to append your skin to its files glob), renumbering the old
§ 10 to § 11. demo-beats.md's beat-6 part 3 and SKILL.md's beat-6 walk step
now name the channels and the REST proof.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
1c50e3f977 |
docs(reskinnable-demo): write the attach snippets in the form the hook keeps
The two `@/shell/attach` call-site snippets used an inline `type AttachmentDocument` in the value import. The commit hook's `oxlint --fix` (consistent-type-imports, `stage_fixed: true`) rewrites that form — it did so to all three skin wrappers inside the previous commit — so a skin author copying the snippet verbatim would watch their file change under them. Both snippets now show the separated `import type` the hook leaves alone, and say why. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
fce12788c5 |
refactor(reskinnable-demo): stage all three attachments through the shell chain
Each skin's file reduces to a document URL, a filename and a message; the
detection, the bounded waits and the abort-on-failure rule come from
@/shell/attach.
Banking's and people's chains covered only the first rows of the detection
table, so both silently gain the causes they lacked — including the two that
let a prompt go out WITHOUT the file: no wait for the attachment chip, and no
wait for the chip to print the filename while the file is still base64-encoding.
FIXES BANKING'S AND PEOPLE'S UNGATED SEND. Both had their send half in
skin.tsx, not in their attach file, in this shape:
const staged = await stageOfferLetterAttachment();
if (staged) await wait(500); // `staged` gated ONLY the wait
setTextareaValue(textarea, PACKET_MESSAGE);
sendButton?.click(); // ...and sent REGARDLESS
So a failed stage still sent the prompt, the model invented the document's
contents, and the tool filed a plausible-looking artifact off them — the beat
proving the opposite of its claim with nobody in the room able to tell. Four
defects in those ~20 lines, all now handled by the shell: the ungated send, a
fixed 500 ms sleep racing an async encode, a silent `if (!textarea) return`,
and a `click()` that never checked the STOP state (mid-run that click CANCELS
the run instead of sending). None of that behaviour is preserved.
Commerce's `launchBeat3d` wrapper is deleted as redundant, not unfashionable:
both shell entry points are wholly inside their own `try` and report cause
"unexpected" before resolving `false`, so neither can reject.
Commerce's attachment test drops the cause enumeration (26 tests -> 6; the
fifteen causes are driven once in src/shell/attach/stage-attachment.test.ts)
and keeps only what is still commerce's: that the wrapper passes the right URL,
filename and message — asserted against hardcoded literals, since a test that
imports the constant it checks cannot notice the constant changing — and that a
failure reports and does not send. It imports composer selectors from
@/shell/attach/stage-attachment, because the barrel deliberately omits them.
Verified in a REAL browser (fresh dev server, forced-new thread), which the
chain had never had: on /banking, /people and /commerce the pill stages a chip
that PRINTS the filename, the chip is consumed by the send, and the artifact is
filed citing document content — banking's report carries the invoice's five
Meridian line items, people's packet the offer letter's 15 Aug start date and
week-one schedule, commerce's plan the sheet's Net 45 / 30% deposit terms and
the first-time Alder Crewneck quote. With the document route forced to 500 all
three abort: composer never driven, `[attach:http-error]` logged, alert raised,
artifact count unmoved.
Skill staleness (CLAUDE.md standing rule): YES, this made
.claude/skills/reskin/ wrong, and all of it is fixed here. demo-beats.md § 3d
now says the chain is shell-owned and shows the ~10-line call site instead of
instructing the author to build it; its dangling `stageInvoiceAttachment`
reference and its "copy commerce, not banking" caveat are gone, and the
"copy from" table row points at @/shell/attach. failure-modes.md § 4 is
repointed at src/shell/attach/stage-attachment.ts and records that the taxonomy
is now complete (all fifteen causes emitted AND driven, which was not true of
the file it described). templates.md's build-it-by-hand comment block becomes
the call site, and its `launchBeat3d(...)` lines become bare `void`s. Three
facts new to the skill: the `[attach:<cause>]` log tag is load-bearing (17 cases
across 10 blocks parse it by regex), `Beat3dTimings` is injectable so a test
never sleeps a production budget, and NOTHING_SENT_LEDE is overridden in
exactly two places. Stale comments naming `stagePriceSheetAttachment` /
`sendRestockRequestWithPriceSheet` in the price-sheet route and its test are
updated too — one of them claimed the prompt is sent regardless, which is the
behaviour this commit removes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
766ac3afd3 | test(showcase): refresh unavailable backend fixture | ||
|
|
4151807154 | test(showcase): update partner baseline count | ||
|
|
cbf79ef52a | fix: align Angular 20 support and resolve packed smoke paths | ||
|
|
79f56d9f8e | fix(showcase): finalize CrewAI D6 on official bridge | ||
|
|
18cd9534b9 | fix(docs): link managed onboarding CTAs | ||
|
|
310c2f1410 |
test(reskinnable-demo): make the attachment cause taxonomy exhaustive
`observed.size === 15` only counted causes some test DROVE, so a sixteenth union member that no code ever constructs never entered the map, the size stayed 15, and the assertion passed. That is precisely the rot this module was extracted to fix — commerce declared fifteen causes and constructed eight, and nothing could see the difference. `Record<AttachmentFailureCause, true>` is exhaustive by construction, so the expected count now comes FROM THE UNION: a new member is a tsc error until it is listed, and once listed it raises the count until a test drives it. Verified both directions by mutation. Also: the paperclip kept its own lede on the failures it anticipated and then fell back to "nothing was sent" in its catch, contradicting itself depending on which failure it hit — and that catch was the only one of the module's two `unexpected` emissions with no test. Both fixed. `staging-threw` joins the send-path table, which covered seven of the eight ways staging can fail. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bd0eb9cf60 |
feat(reskinnable-demo): add the shell attachment staging chain
Every step of staging a file into CopilotKit's composer is a request made of framework code we do not own, so an unobserved step is an assumption. This centralizes all fifteen detection causes, the bounded condition waits, and the rule that any failure ABORTS the send. That last rule is what makes beat 3d honest rather than merely defensive: a chain that lets the prompt go out without the file makes the model invent the document's contents, the tool files the artifact anyway, and the result reads plausibly — so the beat proves the opposite of its claim and nobody in the room can tell. Moved from commerce's `attach-price-sheet.ts` with only three values lifted into parameters — the document's url, its filename, and the pill's message. `reportAttachmentFailure` now takes the cause and tags the log line with it, which is what lets the send path's failures be pinned by tests at all: those entry points return a bare boolean, so the log is the only place the cause is observable. Commerce keeps its own copy until task 7 migrates the skins. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
703d649a03 |
refactor(reskinnable-demo): build both PDFs on the shell primitive
Commerce's builder loses its byte layout to @/shell/documents and keeps its
content. People's gains the ASCII fold it never had — a live defect, not a
hypothetical: the seed carries Ines Vidal, Sasha Bergstrom and Montreal
(accented in the data) and the offer-letter route reaches all three, which
produced mojibake AND a /Length that disagreed with the byte count.
Commerce's duplicated byte-level assertions move to the shell's test, where
they cover every skin once.
toAscii now transliterates instead of blanking. The fold Task 4 moved verbatim
covered punctuation only, so every accented LETTER became "?" — the price sheet
printed "?MILE & FILS" and "Cr?me Br?l?e Tee", and commerce's suite asserted
those strings, pinning the defect as though it were the spec. It now
NFD-normalizes and drops combining marks, keeping a visible "?" only where
there is no ASCII base (CJK, currency symbols): a dropped character is a silent
corruption, a "?" is a legible one. The class is \p{M}, NOT the \p{Diacritic}
the plan proposed — Diacritic also covers ASCII "^" and "`" (deleted out of
ordinary prose) and U+00B7 MIDDLE DOT, which the fold map deliberately turns
into "-". This deliberately breaks the byte-identity Task 4 established; that
claim existed to prove the move was faithful, and it did.
Two audited differences change the offer letter's rendered page, both accepted
because one shared primitive is the point of the extraction:
- Margin. People used MARGIN = 64, the shell is fixed at commerce's 58, so
the letter's text now starts 6pt further left and 6pt higher.
- Body type. People defaulted to 11pt with 4pt extra leading; the shell
defaults to 10.5pt with 3.5pt. Unlisted in the plan, found while migrating.
Parameterizing either would hand every caller a knob nobody asked for.
The letter's page dictionary now declares Courier /F3 and /F4 without
referencing them. Left alone: declared-but-unreferenced is inert, whereas
referenced-but-undeclared renders blank.
Three of commerce's eleven layout cases were deleted, not five: the xref,
/Length and font-resolution cases are pure mechanism and the shell's test now
covers them for every skin. "Draws every columnar line in a monospaced face"
and "keeps every mono line inside the drawable width" STAYED — the shell proves
that a `mono` line comes out in Courier, but only commerce's own test can prove
that commerce marks its rows `mono` and picks column widths that fit. Both
would stay green in the shell while the sheet rendered ragged.
Skill impact (the standing CLAUDE.md question): yes, and fixed here. Six
locations in .claude/skills/reskin/ told a new author to copy commerce's
builder and reimplement both traps — demo-beats.md's copy-from table, its
"Writing a PDF by hand?" block (all four paragraphs, including the one naming
people as the unfixed counter-example) and its alignment paragraph, plus
failure-modes.md's decorative-assertions entry. They now point at
@/shell/documents, describe the traps as centralized, and say which half of
the alignment trap a skin still owns.
Verified: npx tsc --noEmit clean; pnpm lint clean; pnpm vitest run 102 files /
1153 tests (was 101/1151). offer-letter-pdf.test.ts red before the migration
(3 failures, including /Length 1740 vs 1736) and green after. Both documents
rendered and inspected: all four PDFs pure ASCII with correct /Length and xref,
the letter reads "Ines Vidal"/"Sasha Bergstrom"/"Montreal", and the price sheet
for a non-accented vendor is byte-identical to its pre-migration output.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
ab00e284d2 |
test(reskinnable-demo): make the shell PDF font and xref guards able to fail
Both cases passed unconditionally. The font test scanned the WHOLE document for /Fn, so it harvested the page's own /Font dictionary and its size check was always 4, then asserted the presence of the substring it had just matched with a \d+ that accepted any object number. It now reads the referenced faces from the content stream alone and follows stream -> /Font dictionary -> emitted font object, checking the BaseFont at the end of the chain. The xref case checked only that startxref landed on the literal "xref", never that an entry pointed at the object it claims. It now walks every entry and checks /Size, per commerce's expectCorrectXref. Both verified by mutation: FIRST_FONT_OBJECT 5 -> 6 (every glyph blank, whole suite otherwise green) and one recorded offset +1 now fail, and only these two cases fail. Also: assert drawableWidth is 496 rather than positive, cover the middot fold, assert %PDF-1.4 with its trailing newline unmasked by trimEnd, and drop a font-set assertion implied by the per-line check above it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
478e200a88 |
feat(reskinnable-demo): add the shell PDF document primitive
Extracts commerce's price-sheet builder into a content-agnostic primitive: the toAscii fold, the Helvetica + Courier font dictionary, the /Length and xref arithmetic, and the mono page-fit bound. Callers supply a flat list of lines; a section is a heading line carrying a gap. Two traps this centralizes, both of which produce a VALID PDF that is wrong on screen, and neither of which anything type-checks: a UTF-8 content stream under base-14 WinAnsi fonts (mojibake plus a desynced /Length), and padEnd alignment drawn in a proportional face. A verbatim move: commerce's builder is untouched and keeps its own copy until the migration lands. Byte-identity was verified by running commerce's own content half over both implementations, ASCII and non-ASCII, and comparing the emitted bytes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
190be901d5 |
refactor(reskinnable-demo): migrate people and commerce onto the shell teach module
Both gain the visible-duration floor, the feed de-dupe and the automatic fresh-window reset they were missing; both lose their manual reset() and their per-skin @keyframes, which the shell's token-driven .recording-vignette replaces. Six copies of the ref-counted state machine are now one. Skill impact: checked. `.claude/skills/reskin/demo-beats.md` item 4 already points at `@/shell/teach` (Task 2), and neither SKILL.md nor templates.md teaches a per-skin recording context or per-skin vignette keyframes, so nothing there went stale. One live doc DID: `docs/teach-mode/README.md`'s role table still named the banking copy Task 2 deleted — repointed at `src/shell/teach/`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7a45003d37 |
refactor(reskinnable-demo): migrate banking onto the shell teach module
Deletes banking's recording context, feed and vignette in favour of @/shell/teach. noteDemonstratedCode collapses into logStep's second argument, so the feed line and the code that lifted the gate are recorded by one call at the moment the operator files it. Behaviour is unchanged; banking already had the union's stricter half. Two additions beyond a straight import swap, both to keep coverage flat: - banking's recording-context.test.tsx held three cases, and only two were genuinely re-covered by the shell module's own test. The third was a PendingApprovalsChat integration test proving that component's handleApprove brackets in the right order (logStep before beginRecording silently drops the line — it has been wrong once). Moved to components/wow/pending-approvals-chat.test.tsx rather than deleted. - pinned the invariant the API collapse now depends on: the outer bracket held open from "start recording" until "I'm done" is what keeps the derived code alive across the nested bracket the filing opens. The old ref-based noteDemonstratedCode did not need this; getDemonstratedCode does, and nothing tested it. Reskin-skill review (per CLAUDE.md): NOT clean — demo-beats.md item 4 pointed a new skin's author at banking's three deleted files as the pattern to copy. Rewritten to send them to @/shell/teach instead, with the outer-bracket rule and the silent failure modes spelled out. Also refreshed the six HITL-chain line anchors, which were already ~8 lines stale before this change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5c86863aec |
fix(reskinnable-demo): make the recording feed readonly so it type-checks
Object.freeze([]) is readonly never[], and the `as RecordedStep[]` assertion on the frozen no-provider fallback was not a legal widening — TS2352. It went unnoticed because this app has no typecheck script, so `pnpm lint` passing said nothing about it; `npx tsc --noEmit` (~4s) is the cheap authoritative gate. Fixed by changing the contract rather than satisfying the cast: RecordingValue.steps is now `readonly RecordedStep[]`, so the frozen fallback needs no assertion at all. That is also the honest type — consumers only read the feed (.length, .map), and every out-of-provider consumer shares ONE frozen array, so a mutable type invited splicing a singleton out from under all the others. No consumer needed a mutable array: nothing outside the module imports @/shell/teach yet, and none of the three per-skin copies Tasks 2/3 migrate onto this contract mutates its feed in place. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a218e326f3 |
fix(reskinnable-demo): pin the recording module's untested properties
Review of the shell-owned recording module found three of its four load-bearing properties had no failing assertion behind them, and one undisclosed change. - getDemonstratedCode is back to the intended [...x].reverse(). It had shipped as x.toReversed() because lefthook's lint-fix runs `oxlint --fix` with stage_fixed and unicorn/no-array-reverse autofixed it into the commit after the last local test run. toReversed() is ES2023 against an ES2017 target with lib:esnext, so it type-checks clean and silently raises the browser floor to Chrome 110 / Safari 16.4. Suppressed with oxlint-disable-next-line, which is the only spelling that holds: the eslint- form fails `pnpm lint` because no unicorn plugin is loaded. - The ref-count test advances past the MIN_VISIBLE_MS hold on both sides of the inner bracket. Previously the 1200ms floor alone kept the flag true, so the assertion passed with ref-counting deleted. Note ref-counting is guarded in two places (the early return and the timer callback's count re-check); only removing both fails the test. - Added a two-coded-step case so the reversal is pinned. The other code tests all pass under a forward find(), which is precisely the demo path: file a decoy, get refused, file the real code. - Added the first coverage of RecordingVignette's data-recording attribute, the entire contract with the .recording-vignette CSS. Also hoists the no-provider fallback to a frozen module constant so consumers outside a provider get stable function identities rather than a fresh literal per render, and documents that a beginRecording() during the hold inherits the existing feed on purpose. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2ddea0b403 |
feat(reskinnable-demo): add the shell-owned teach-mode recording module
Three skins had shipped three copies of the recording context and they had diverged: banking alone had the minimum-visible-duration floor, the feed de-dupe and the fresh-window reset; commerce and people alone derived the demonstrated code from what the operator actually filed. This is the union. Every failure mode here is silent — useRecording returns inert no-ops outside a provider and logStep returns early when idle — so a broken copy does not throw, the feed is simply empty and the glow never appears. The vignette CSS moves from banking/theme.css to globals.css. It reads the shared --brand-violet / --brand-indigo tokens, which all six skins define, so the glow now reskins without a per-skin copy. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9303fabcdd |
fix(channels-slack): let an image use a file already in the workspace
Slack's image block takes either an external `image_url` or a `slack_file` pointing at a file that already exists in the workspace. The required-field check demanded `image_url` unconditionally, so the `slack_file` form could not be built at all — an image sourced from Slack itself was unreachable through the catalog. An image now needs alt text plus *either* source. Passing neither is still an error: the check moved, it did not disappear, and the test covers both halves. |
||
|
|
078116ee8d |
fix(channels-slack): stop tagging untyped composition objects with a type
Slack's option object is `{text, value}` — it has no `type` field, and neither do
confirm, option_group, conversation_filter, dispatch_action_config, slack_file,
trigger or workflow. The codec tagged every catalog entry with its manifest type
regardless, so each of those carried an unknown field and Slack refused the
entire message containing it.
That took out every select, multi-select, checkbox, radio group, overflow menu
and confirmation dialog authored through `Slack.Object.*` — the whole interactive
surface. Measured against a real workspace: 1 of 26 block elements was delivered
before this change, 23 after.
It went unnoticed because a refused payload produces no error. There is no log
line and no exception; the message simply never arrives, which looks exactly like
a bot that had nothing to say.
The existing catalog test asserted the very assumption that was wrong — that
every entry serializes its discriminator — so it was green while the product was
broken. It now asserts the corrected rule and guards against a silent relapse.
|
||
|
|
81e7fefb18 |
fix(channels-slack): correct the container slot and drop the unpostable file block
Two catalog errors, both surfaced by delivering every documented block into a real Slack workspace rather than reading the reference again. container serialized its children into `blocks`; Slack reads `child_blocks`, so every container an app sent was refused — silently, because a refused payload produces no error anywhere, just a message that never arrives. file leaves the manifest. Slack states you cannot add it to app surfaces directly; it only appears when *reading* messages that contain remote files, and the same sentence appears verbatim in `@slack/types`' own doc comment. Keeping it in an authorable catalog offered a component that could never succeed. Sending a file remains thread.postFile(). alert stays out for the same class of reason (modals only), and both exclusions now carry their citation so "missing" and "deliberately not authorable" stay distinguishable. |
||
|
|
54c32a1945 |
fix(reskinnable-demo): guard the page-level HITL approval renders (#6453)
## What
`0a9b99aae3` added the durable `resolved` prop to `ApprovalButtons` and
wired the three call sites in `tools.tsx` that need it. The five
`useHumanInTheLoop` renders that live on the pages were missed:
| File | Actions |
|---|---|
| `pages/team.tsx` | `removeMember`, `changeMemberRole`,
`changeMemberTeam` |
| `pages/cards.tsx` | `addNewCard`, `assignPolicyToCard` |
All five guarded only on `status === "inProgress"` and then fell through
to `ApprovalButtons`, whose local `responded` state dies with the
component. Each now destructures `result` and passes the same signal the
three wired renders use:
```tsx
resolved={status === "complete" || !!result}
```
## Why it matters
Reload a thread containing an already-answered `removeMember` and the
local state is gone, so live Approve/Deny buttons come back on a settled
call. A second click fires a duplicate write. This is reachable without
the multi-step chain that motivated the original fix.
The other three renders in `tools.tsx` still need nothing — they return
a terminal card when complete, so they never reach the buttons.
## Correctness
The render props are a discriminated union: `result` is `undefined`
while `Executing` and a `string` once `Complete`, so the buttons stay
live in exactly the one state where the user should act.
`ToolCallStatus` is a string enum, so the `status === "complete"`
comparison deserved a check rather than an assumption. Verified against
the repo's tsc 5.9.2 in isolation: it compiles clean, and a negative
control (`status === "definitelyNotAMember"`) errors with TS2367, so the
check isn't vacuous.
`|| !!result` is redundant given the union, but kept for consistency
with the three existing call sites.
## Verification
- `nx run reskinnable-demo:lint` — clean
- `nx run reskinnable-demo:test:unit` — 1130 tests across 99 files pass
- lefthook pre-commit + commitlint — passed, no `--no-verify`
Two things I did not do, stated plainly:
- **Not exercised in a browser.** The original three were verified live;
these five are structurally identical renders extended by inspection.
- **`nx run reskinnable-demo:build` fails locally**, on `Module not
found: Can't resolve '@copilotkit/shared'` from
`packages/web-inspector/dist`. It fails identically on clean `main`, so
it is a stale local workspace link and not a regression from this change
— but it does mean CI is the first real type-check of this diff.
## Out of scope
Only the banking skin's `ApprovalButtons` sites were audited. The other
skins (commerce, keel, logistics, people, airline) were not checked for
their own approval UIs carrying the same local-state-only pattern.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
|
||
|
|
24c6c5dafe |
fix(reskinnable-demo): guard the page-level HITL approval renders
|
||
|
|
ea587ec828 |
feat(reskinnable-demo): add the commerce skin (Bellwether) (#6442)
Adds a sixth skin to the reskinnable demo: **`commerce` / "Bellwether"**, a storefront operations console for a DTC retail brand. REST-backed over `/api/commerce/v1/*`, and demo-complete against all nine demo beats plus the presenter-reset requirement — which puts it alongside `banking` and `people` rather than the earlier wiring-only skins. Its teachable gate is approving a markdown that would trade **below the category margin floor** (422 `BELOW_MARGIN_FLOOR`), unlocked by a margin waiver filed under a justifying code. Two below-floor markdowns are seeded, so the case taught on stage and the unaided replay are different products. ## The signature visual: the margin ladder One rail per category, each anchored to that category's own margin floor, so *"how far from the line I may not cross"* is comparable across categories at a glance. The decisions inside it are all recorded in the commit bodies, because all of them were wrong first: - Every floor line sits at **one** rail height, so the rails read as a set — the per-category floor is encoded by dot position, not by moving the line. - A dot that overflows its rail is **never hidden**. Hiding it lent the dot to the *next* rail, silently misattributing a below-floor product to a compliant category. - An **empty** ladder renders as "nothing plotted", not as an all-clear. - A category with **no margin floor on file** renders as visibly *unchecked* — not green, not red, and not a bare figure in neutral ink, which reads as "checked, fine" on the exact question the ladder exists to answer. ## Also the reference for a four-lever view (beat 3c) Status, exception class, sort and top-N all arrive from the query string and **all four controls tint**. The queue count is computed against the filters actually applied rather than the unfiltered set, and an unusable top-N lever is ignored instead of collapsing the view to one row. ## Review This branch went through an unusually long review: **95 findings** fixed in the first cycle, then a **24-agent unbiased confirmation round** (all seven slot types, verbatim prompt) which returned ~233 more. Those resolved into **11 recurring defect classes**, closed by class sweeps rather than per-instance fixes. **Every sweep found more instances than the finding that triggered it** — which is the argument for the approach, as data rather than opinion: | Class | Named | Found | | ----- | ----- | ----- | | Secret redaction covering the URL but not the key | 2 secrets | **5** (the API key appeared verbatim 5× in one response body) | | Beat 3d attachment reporting success it never verified | 6 | 6 | | Silent input coercion on the refund path | 2 | **9 of 22 probed inputs became issuable figures** | | Gen-UI renderers dereferencing streaming args | 4 | **9**, from enumerating all 19 renders | | Guards whose premise was wrong | 4 | **8** | | Unchecked-floor claims | 5 | **9**, from 33 consumers enumerated | | Write controls that could fire twice or latch | 4 | 6, across three pages | | Doc fixes that had falsified their neighbours | 4 | 4 + 21 verified sound | Four post-convergence audits (2 → 2 → 1 → 0) then caught a class no diff-scoped reviewer could: **claims this PR made about code outside the diff.** That is how the two most valuable findings in the whole review surfaced — see below. Tests: **372 → 1130**. Full evidence, per-finding, in the review ledger. ## Two live defects in *other* skins, found by this review Neither is in scope here; both are on the follow-up list and worth prioritising: 1. **`people`'s offer letter generates a structurally corrupt PDF today.** Same `/Length`-from-string-length byte math commerce's price sheet had, but with **no ASCII fold at all** — live-reachable for the seeded employees `Inés Vidal`, `Sasha Bergström` and the city `Montréal`. Commerce's fix here is the template. (Commerce's own guard was found to work only by accident: with `toAscii` neutered, **8 pre-existing tests stayed green on a structurally corrupt document**. The invariant is now pinned.) 2. **`banking` and `people` leak more secrets than commerce did** — an unredacted `memoryError` plus `apiUrl` echoed into a presenter-facing reset body, where commerce leaked one API key. `src/lib/redact-secrets.ts` added here is skin-agnostic and adoptable as-is. Also: `banking/attach-invoice.ts` and `people/attach-offer-letter.ts` carry beat 3d's identical unverified-attachment defect — **wiring `onUploadFailed` in `chat-panel.tsx` would close it for all three skins at once.** ## Shipping with a known, documented parked list This is deliberate and recorded, not an oversight. The parked items are in-subject and real but not demo-blocking — a conditional cross-skin memory deletion (needs a pinned `INTELLIGENCE_USER_ID`), a latent ladder-rail inversion (unreachable at current seed floors), two prototype-chain lookups needing a hostile URL, 500-vs-400 on malformed request bodies, an unguarded presenter Reset, and ~20 test-hygiene items that belong in a test-hygiene PR of their own. ## Skill-staleness rule `CLAUDE.md` requires every change to existing code to answer whether it leaves `.claude/skills/reskin/` wrong, incomplete or misleading. **Answered in full in the final commit body: yes, in many places, and that commit is the fix.** The skill gained a new `failure-modes.md` encoding the cross-cutting lessons — chiefly that **a demo skin's characteristic bug is not a crash but a confident falsehood**, and that a skin publishes facts to three audiences (the screen, an agent readable, generated UI) with three different obligations. Existing files were corrected throughout: a demo-complete skin scores out of nine beats not six, the required pill count is now derived from the beat map, the Reset link no longer teaches an href the app's own lint bans, and the `theme.css` scaffold is written so prettier cannot mangle its selector. ## History 109 commits reorganised into 12 by area of concern. The tree was proven identical to the pre-rewrite state three ways — empty diff, identical root tree object, and nothing staged after the move — and all 109 original SHAs are cited across the 12 bodies, so the per-fix reasoning stays reachable. 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
0c3a4962ac | Merge branch 'main' into feat/reskinnable-demo-commerce-skin | ||
|
|
d718d9d875 |
docs(reskinnable-demo): update the reskin skill and add failure-modes.md
THE REPO'S STANDING QUESTION, answered here as CLAUDE.md requires of every change to
existing code: does this change make anything in `.claude/skills/reskin/` wrong,
incomplete or misleading for the next person authoring a skin?
YES -- in a lot of places, and this commit is the fix. The skill is the only
instruction a new skin's author reads and it goes stale SILENTLY: nothing
type-checks it, no test imports it, and a skin built from a stale template still
compiles, lints and renders.
NEW: `failure-modes.md`. Building commerce surfaced the same class of defect over
and over, in code that compiled and looked right, so the lessons are now written
down as a checklist rather than left implicit in one skin's diff. It is about the
ways a skin LIES: publishing a verdict it never checked, claiming a write that never
happened, reporting success it has not earned, narrating a partial failure as a
complete one, and counting rows it silently truncated.
CORRECTED throughout `SKILL.md`, `demo-beats.md` and `templates.md`:
- A demo-complete skin scores out of NINE beats, not six, and the required pill
count is DERIVED from the beat map rather than stated as a magic number.
- The URL-contract section stopped calling this a four-skin demo, and now tells a
new skin to register in `skinIds` AND `skinIdentities`, not just the registry.
- All three skins that seed memories are credited; `keel` is credited in the
`useData` contract row; the presenter-reset beat names the skins that ship it.
- The template's Reset link taught a hardcoded `/${skin.id}/…` href -- exactly the
pattern the app's own lint bans and that breaks silently under a LOCK_SKIN deploy.
- The multi-page route template taught the record lookup the skill forbids elsewhere.
- The layout template coupled sidebar width to an inset the shell no longer applies.
- The `theme.css` scaffold is written so prettier cannot mangle the `.theme-<id>`
selector when an author runs the formatter over it.
Subsumes: ae2a8dde6e 933c7f1f38 4644354f65 628f27a1ca d329643c84 08e6692cbf
777014d288 a6667cfbdf 8dc393addf f205ee0c04 89a58d7bd1 8147676dd9 d72db4a20b
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d9dcf69893 |
docs(reskinnable-demo): update the app docs for a six-skin roster
`CLAUDE.md`, `README.md`, `.env.example` and `docs/teach-mode/README.md`. The roster went from four skins to six, and almost every count derived from it was wrong. Corrected here: - The demo-beat matrix, which now lists `commerce` and reports gen-UI counts that match the registrations actually made. - The `useData` contract row, which names BOTH in-memory skins rather than one. - The list of skins that identify their user, and the list that ship `intelligence/seed-memories.ts` -- `commerce` was missing from both. - `.env.example`, which capped `LOCK_SKIN` at four legal values. - The claim that `people` re-scopes memory per operator, stated more strongly in CLAUDE.md than the code supports. Two decisions about HOW the docs were fixed, because they cost the most time: - CROSS-SKIN CLAIMS ARE SCOPED TO THE SKINS ACTUALLY CHECKED. Several sentences asserted a property of "every skin" on the evidence of one or two. They now name the skins verified. - EACH FIX WAS CHECKED AGAINST THE SENTENCE BESIDE IT. Correcting one count repeatedly left its neighbour false, because the counts are stated redundantly in adjacent prose. That pattern is what motivated the roster test in the shell. Subsumes: 7f5813d7ea 1600591918 8071d05812 f78f6f7cd4 ca2145228d ad76437e1c 4308c02135 3fab172fb1 0ca1564406 f0afa4acce Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0cc2a00218 |
feat(reskinnable-demo): register commerce and repair the LOCK_SKIN lint guard
Registration is deliberately duplicated across three places -- the client
`registry.ts`, the server `agent-registry.ts` (as `{ createAgent, identifyUser }`),
and `skins-config.ts`'s `skinIds` -- because the server registry must never pull
client-only modules and the config must be importable from an RSC and the proxy.
This commit adds commerce to all three, plus `eslint.config.mjs` and
`src/lib/locked-skin.ts`.
Decisions:
- `LINTED_SKIN_IDS` IN `eslint.config.mjs` HAD ROTTED. It is a hand-copy of
`skinIds` -- an ESLint flat config is loaded by Node and cannot import a `.ts`
module -- and it still named four skins two releases after `people` and `commerce`
shipped, so the LOCK_SKIN skin-prefix guard was blind to both. `skins-config.test.ts`
now lints a synthetic prefixed link for EVERY registered skin through the real
selectors, so the copy cannot silently rot again.
- `skin-roster-docs.test.ts` is new and FAILS THE BUILD when a doc miscounts the
skin roster. The "four skins" claim was wrong in several places at once; a test is
the only thing that keeps prose counts honest, since nothing else type-checks them.
- The registries' own comments were lying: they credited `logistics` with durable
memory it does not have, promised `commerce` a memory switch it does not ship, and
said Rowan re-scopes memory in a way it does not.
Reviewer note: `skin-roster-docs.test.ts` was largely rewritten by `3f994daf0f`,
which is cited in the theme commit.
Subsumes: 00514f678d 229fcdee66 951a8090c8 bc3a15f677 7ea21b37c0 724e518ca8
0f18a24d74
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
989016e01c |
feat(reskinnable-demo): assemble the commerce Skin contract object
`skin.tsx` implements the frozen `Skin` contract, with `identity.ts` (brand, logo, favicon), `suggestions.ts` (one pill per beat, in demo order, with the skin's beat map written out at the top of the file) and `providers.tsx` (the teach-mode recording stack). Decisions: - IT OMITS `useData`. The ledger is read through the skin's own `useCommerceLedger()` context, mounted in `RuntimeProviders` rather than `Providers`, so the single ledger fetch also feeds `useRuntimeProperties`. That is the same shape banking, logistics and people use. - Beyond the required fields it sets `Providers`, `CanvasSurface`, `sandboxFunctions`, `toolLabels`, `chatHeaderActions`, `onSuggestionSelect`, `RuntimeProviders` and `useRuntimeProperties` -- the full optional surface, which is what a demo-complete skin needs. - `resolvePage` 404s a segment named after something on `Object.prototype` instead of resolving an inherited property to a page component. Subsumes: eb6c516d3f Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1f7a0eae65 |
feat(reskinnable-demo): add Bellwether's theme block and its contrast guard
`theme.css` is a single `.theme-commerce` block that RE-VALUES the shell's shared design tokens -- it invents no token names, which is what keeps a reskin a pure value swap. `theme.test.ts` is the guard that the values are actually usable. Decisions: - Bellwether's dark-mode buttons were unreadable: the button foreground was resolved against the wrong background, so the pair that shipped had contrast far under the bar the token values were chosen to hit. - THE GUARD NOW MEASURES WHAT IT NAMES. Several assertions were checking a different token pair from the one in their own description -- passing tests that proved nothing about the thing they claimed. Two further "coverage" claims promised checks that were never run at all. Reviewer note: `3f994daf0f` and `e2f3d61489` also touched `src/shell/skin-roster-docs.test.ts` and `docs/teach-mode/README.md`; those files land in the shell and docs commits respectively, but the SHAs are cited only here. Subsumes: 511e173d81 3f994daf0f e2f3d61489 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4598666200 |
feat(reskinnable-demo): scope commerce's memory and add its presenter reset
The long-term-memory and stored-procedure beats are not emergent behaviour; they need a per-user memory scope and a reset that puts the demo back to a known state. This area is that machinery: `intelligence/user-id.ts` (the server-safe `identifyUser`), `intelligence/seed-memories.ts`, `intelligence/forget-memories.ts`, the gated `dev/reset` route, the skin `layout.tsx` that hosts the Reset control, and a shared `src/lib/redact-secrets.ts`. Decisions: - RESET BUCKETS ARE DERIVED FROM `resolveUserId`, not hand-listed. A hand-listed set drifts from the identity the runs actually use, and the failure mode is a reset that reports success while wiping a bucket nobody writes to. - THE IDENTITY MAP REFUSES INHERITED KEYS. A role named after something on `Object.prototype` resolved to a function and produced a nonsense scope. - THE RESET NEVER REPORTS SUCCESS IT HAS NOT EARNED. It fails when the memory WIPE did not finish, does not claim memories it never seeded, and when it throws it reports what it had measured up to that point rather than a bare error. - A PARTIAL RESET NO LONGER LEAVES TWO NARRATORS DISAGREEING. The route's summary and the on-screen confirmation are now driven from one result, so a half-completed reset cannot be described as complete by one of them. - THE RESPONSE BODY CARRIES NEITHER THE INTELLIGENCE BACKEND URL NOR THE API KEY. `src/lib/redact-secrets.ts` is the shared scrubber, and it scrubs the KEY, not just the URL -- scrubbing only the URL left the credential in the body it was embedded in. Subsumes: 085ca5484e c3f6012ddb 4dd9bcd1c9 4b5b70e3e7 6f2ca38f88 fec06be733 c4d1ef71af 691c6e84e7 d1d6570deb Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |