Commit Graph

3770 Commits

Author SHA1 Message Date
Atai Barkai 0d55bfc2e2 docs(shell-docs): document frontend applicability policy 2026-06-18 09:54:05 -07:00
Atai Barkai b6cb1d0866 chore(docs): drop unrelated formatting churn 2026-06-18 08:29:18 -07:00
Atai Barkai cab27e2b55 Merge remote-tracking branch 'origin/main' into codex/docs-local-iteration
# Conflicts:
#	showcase/shell-docs/src/components/framework-selector.tsx
2026-06-18 08:27:46 -07:00
Atai Barkai e3d8ea9587 fix(shell-docs): deploy correct Vercel preview 2026-06-18 08:22:00 -07:00
Atai Barkai d5004cabaa fix(shell-docs): resolve frontend folder-root docs 2026-06-18 07:33:09 -07:00
Atai Barkai ddb941c075 docs(shell-docs): route universal frontend docs in context 2026-06-18 07:21:56 -07:00
Tyler Slaton 406df02991 feat(bot-discord): Discord PlatformAdapter for @copilotkit/bot (#5524)
## Summary

Adds **`@copilotkit/bot-discord`** — a Discord `PlatformAdapter` for
`@copilotkit/bot`, built on **discord.js v14**, mirroring the existing
`@copilotkit/bot-slack`. It lets users build Discord bots on the shared
`@copilotkit/bot` + `@copilotkit/bot-ui` primitives, running on agents
via the AG-UI protocol.

Ships three things: the **package**, a runnable **example**, and
**docs**.

## Package (`packages/bot-discord`)

- **Gateway ingress** via discord.js — intents `Guilds`,
`GuildMessages`, `MessageContent` (privileged), `DirectMessages`,
`GuildMembers` (privileged). Both privileged intents must be enabled in
the Discord Developer Portal; `GuildMembers` powers user lookup.
- **Components V2 egress** — IR → `ContainerBuilder` with
`MessageFlags.IsComponentsV2`; full
`Message`/`Header`/`Section`/`Markdown`/`Fields`/`Context`/`Actions`/`Button`/`Select`/`Image`/`Divider`/`Table`
mapping, budget-clamped via `DISCORD_LIMITS`.
- **Streaming replies** — 1100 ms edit throttle; rollover at a 1900-char
soft limit / 2000 hard limit.
- **Ack-first interactions** (3000 ms deadline) — slash commands ack
with an ephemeral reply; buttons/selects ack via `deferUpdate`.
- **Slash-command registration** — per-guild (instant) when `guildId` is
set, else global (propagates within ~1h).
- **HITL**, **typing + reactions**, **built-in tools**
(`lookup_discord_user`) + **context**, in-memory conversation store.
- Capabilities: `{ supportsModals: false, supportsTyping: true,
supportsReactions: true, supportsStreaming: true, maxBlocksPerMessage:
40 }`.

## Example (`examples/discord`)

Runnable Discord bot — `BuiltInAgent` + Linear/Notion MCP (model
`openai/gpt-5.5`), app tools/components/commands/context/HITL, plus a
**`DISCORD_E2E`-gated manual e2e harness**.

```
DISCORD_E2E=1 pnpm --filter discord-example exec tsx e2e/run.ts
```

## Docs

- `showcase/shell-docs` reference under `reference/bot/discord/` —
`index`, `defaultDiscordTools`, `defaultDiscordContext`,
`renderComponents`, `DISCORD_LIMITS`; bot reference index + nav.
- Package `README.md` + `ARCHITECTURE.md`.

## Review summary

Built and reviewed via a multi-module adversarial CR process:

- **Module 1 (package):** 4 CR rounds, ~22 bugs fixed, **156 tests**.
- **Module 2 (example):** 1 round, ~10 fixes (stale Slack-port residue +
missing error guards), **38 tests**.
- **Module 3 (docs):** 1 round — fixed a blocker (doc invented three
non-existent adapter options) + majors (missing `GuildMembers`, wrong
streaming throttle/rollover numbers).
- **Module 4 (cross-cutting, 7 lenses + promotion audit):** caught &
fixed a **blocker — slash-command interactions were never
acknowledged**, so every slash command showed the user Discord's "The
application did not respond" error. Now acked within the 3 s window.
Also fixed: frozen `_thinking…_` placeholder on mid-stream throw, an
unhandled-rejection window in the chunked stream, an accidental NUL byte
that made a test file binary, and `ARCHITECTURE.md` intent-list
accuracy.

## Deferred follow-ups (not blocking)

- **True modals / multi-step forms** — `supportsModals: false` in v1.
- **Inbound-attachment auto-wiring** — `buildFileContentParts` is
exported + tested but not yet wired into the listener, so a user
attaching a file gets nothing delivered to the agent. Two reviewers
rated this blocker-severity; the promotion audit confirmed there's no
false promise (README claims no upload support) and nothing depends on
it. **Strong fast-follow candidate.**
- **`@copilotkit/bot-slack` port** — bot-discord fixed a latent
`flushNow` ordering bug and a select `custom_id` JSON encode/decode
asymmetry that still exist in bot-slack.
- **Latent (not reachable by shipped example):** select-value JSON
round-trip type coercion; `>5` action-row silent drop (no overflow
marker).
- **Adapter-glue test coverage** —
`post`/`update`/`delete`/`lookupUser`/`getMessages`/`postFile`/ack-ordering/`interruptEventNames`/`custom_id`
round-trip are covered indirectly; direct tests would harden them.
- **Minor:** `getMessages` placeholder filtering (shared with
bot-slack); `buildFileContentParts` config-key/default divergence vs
bot-slack; type-cleanliness (`ChannelLike` `as never` bridge,
barrel-surface trim).
- **e2e harness** is a `DISCORD_E2E`-gated **manual** tool — its
synthetic interaction lacks an Ed25519 signature, so it works only
against a test shim, not real Discord.

## Caveat (pre-existing, unrelated to this PR)

`@copilotkit/core` (tsc error) and `@copilotkit/runtime` (check-types
JS-heap OOM) are already broken on `main`. They are not touched by this
branch; typecheck/build were scoped to `@copilotkit/bot-discord` +
`discord-example`.

## Test plan

- [x] `nx run @copilotkit/bot-discord:test` — 156 passing
- [x] `nx run discord-example:test` — 38 passing
- [x] `tsc --noEmit` (both `tsconfig.json` + `tsconfig.check.json`)
clean
- [x] `nx run @copilotkit/bot-discord:build` succeeds
- [x] oxfmt `--check` clean, oxlint 0 errors, `pnpm-lock.yaml` in sync
- [ ] Manual smoke against a real Discord app (mention, slash command,
button, streaming reply)
2026-06-18 01:52:30 -07:00
Atai Barkai 3b2ac441c4 docs(shell-docs): share concepts across frontend docs 2026-06-17 19:02:33 -07:00
Atai Barkai 67241bafa0 docs(shell-docs): add agent prompt copy action 2026-06-17 18:54:50 -07:00
Atai Barkai 6333636b9d docs(shell-docs): refine react docs proxy sidebar 2026-06-17 18:42:09 -07:00
Alem Tuzlak 29d5a9d61d chore(discord): remove unrelated PR changes 2026-06-17 18:32:26 -07:00
Atai Barkai 312f3cd38f docs(shell-docs): move react guidance into sidebar popup 2026-06-17 18:23:31 -07:00
Jordan Ritter 14138b7d18 fix(showcase/aimock): add generate_a2ui d6 fixtures for 8 slugs (Sales Dashboard probe)
Mirrors LGP's `generate_a2ui` outer-emit fixture into 8 non-LGP slugs to close
the production 503 gap on the "Show me my sales dashboard for this quarter."
userMessage. Per /tmp/staging-journal-diff.md, aimock-staging returns 503 on
shape-B traffic (model=gpt-4.1, stream=true, tools=[generate_a2ui],
UA=AsyncOpenAI/Python) for all 8 slugs because no fixture matched.

Slugs patched (one entry each in d6/<slug>/gen-ui-declarative.json):
  - llamaindex
  - built-in-agent
  - ag2
  - langroid
  - claude-sdk-typescript
  - claude-sdk-python
  - ms-agent-dotnet
  - ms-agent-python

Each entry matches `userMessage` + `context: "<slug>"` and emits a
`generate_a2ui` toolcall with no args, mirroring LGP's sales-dashboard outer
entry. Per-slug unique toolCallId.

Red-green proof (local aimock, ghcr.io/copilotkit/aimock:latest):
  RED (baseline main, 8 slugs):  HTTP=404 no_fixture_match
  GREEN (this branch, 8 slugs):  HTTP=200 tool=generate_a2ui id=call_d6_decl_dash_outer_<slug>_001
  LGP regression (baseline+fix): HTTP=200 (unchanged)

aimock fixture validation: 737/737 tests pass.

PR scope is intentionally narrow per CLAUDE.md "Scope PRs to flagged
findings": this closes ONE userMessage gap. Shape-A 503s (no tools key) and
other userMessage gaps remain as separate follow-ups.
2026-06-17 16:02:37 -07:00
Jordan Ritter 6652552d32 fix(showcase/aimock/d6/ag2): add Excalidraw create_view turn-1 fixture
Staging strict-mode replay confirmed ag2 was the only integration missing a
fixture for `Open Excalidraw and sketch a system diagram...` + `create_view`
(the D5 mcp-apps probe's turn-1). All 11 other contexts (langgraph-python,
google-adk, ms-agent-{dotnet,python}, strands, llamaindex, built-in-agent,
claude-sdk-{typescript,python}, langroid, mastra) returned 200 against
aimock-staging with the same payload; ag2 returned 503
`{"code":"no_fixture_match"}`.

Add the missing turn-1 entry to ag2/mcp-apps.json, mirroring the LGP
gold-standard at langgraph-python/tool-rendering-reasoning-chain.json (same
fixture id `call_d5_mcp_apps_create_view_001`, same arguments payload, same
chunkSize). Match keys: userMessage + toolName + context.

Local GREEN proof: iso9 aimock (port 6010) with new fixture mounted returns
200 with the canonical tool_call payload for the exact staging-replay request
body. LGP regression check (same probe with X-AIMock-Context: langgraph-python)
stays green.

Scope intentionally narrow per "scope PRs to the originally-flagged findings":
this PR closes only the one genuine fixture gap identified by the staging
verification audit. The 9 other 503s in that audit (sales-dashboard across 9
slugs + google-adk Excalidraw) are NOT fixture gaps — they are backend
tool-array forwarding issues and are tracked separately.
2026-06-17 15:28:44 -07:00
Atai Barkai 092994f680 docs(shell-docs): clarify frontend sdk docs guidance 2026-06-17 12:18:55 -07:00
Alem Tuzlak f4e00eab8b chore(discord): merge main into discord branch 2026-06-17 11:58:33 -07:00
Jordan Ritter c540734143 chore(showcase): bump @copilotkit/* + @ag-ui/* across all integrations (#5523)
## Summary

Aligns showcase integration dependencies to current released minor
versions for the **1.60.2** cycle. Closes the version-coherence gap that
was blocking the `react-core@1.60.2` resume-path / gen-ui-interrupt
fixes from taking effect on staging.

## Package families bumped

| Family | From | To | Scope |
|---|---|---|---|
| `@copilotkit/{a2ui-renderer, react-core, react-ui, runtime, shared,
sdk-js, voice}` | `1.59.4` (18 integrations) / `1.57.2`
(ms-agent-harness-dotnet) | **`1.60.2`** | 19 integrations |
| `@ag-ui/{client, core, encoder}` | `0.0.55` | **`0.0.57`** | 15
integrations |
| `@ag-ui/mastra` | `0.2.1-beta.2` | **`0.2.4`** | mastra only |

`@ag-ui/mastra@1.0.x` (major) **deliberately not** bumped — major jump
held back per the broad-scope dep-bump policy.

## Integrations covered (19/19)

`ag2`, `agno`, `built-in-agent`, `claude-sdk-python`,
`claude-sdk-typescript`, `crewai-crews`, `google-adk`,
`langgraph-fastapi`, `langgraph-python`, `langgraph-typescript`,
`langroid`, `llamaindex`, `mastra`, `ms-agent-dotnet`,
**`ms-agent-harness-dotnet`** (newly added — was missed by the prior
18-integration staging and jumps two minor lines), `ms-agent-python`,
`pydantic-ai`, `spring-ai`, `strands`.

`built-in-agent`, `langgraph-fastapi`, `langgraph-python`,
`langgraph-typescript` carry no `@ag-ui/*` deps directly (the langgraph
trio gets the protocol via `@copilotkit/runtime`'s nested resolution,
which has been verified at `0.0.57` post-install).

## Coherence note

`@copilotkit/react-core@1.60.2` does not declare a hard peer-dep on
`@ag-ui/core` at the package-manifest level; it bundles its own copy via
nested `node_modules`. Lockfile inspection confirms nested
`@copilotkit/{react-core,runtime,shared}/node_modules/@ag-ui/client`
resolved at `0.0.57` across every integration that ships them,
satisfying the `feedback_agui_client_bump_scope` rule (`@ag-ui/core >=
0.0.48`).

## Reconciliation method

Per-integration `npm install --package-lock-only --legacy-peer-deps` (no
`node_modules` mutation). The `--legacy-peer-deps` flag is required to
step past a **pre-existing** `cmdk@0.2.1` ↔ `react@^19` peer conflict
that long predates this bump; lockfile contents are otherwise unchanged
in shape (only dep-version touchups). 30 files in the visible diff
because 8 of the 38 changed files were identical between staged-index
and working-tree.

## Red / Green proof

**RED** (pre-merge staging):
- `gen-ui-interrupt` cells on the langgraph trio (LGP / LGTS /
LG-FastAPI) and across the broader integration matrix are RED on the
showcase dashboard pending the `@copilotkit/react-core@1.60.2`
resume-path fix landing on every integration container.

**GREEN** (expected post-merge):
- `showcase_deploy` will rebuild every integration container on push to
`main`. The dashboard `gen-ui-interrupt` + `resume-path` cells should
flip GREEN once those containers redeploy.
- Per-cell empirical value-test (`bin/showcase test
<slug>:gen-ui-interrupt --d6 --direct`) on ≥3 cells across LGP / LGTS /
LG-FastAPI is queued for the post-merge verification window.

## Out of scope

- No source-code changes (TS/Python/.NET/Java/Go).
- No fixture changes.
- No frontend page changes.
- No `packages/`, `examples/`, `showcase/shell-*` touched.
- No Railway worker restart (separate effort).
- No langgraph-typescript backend agent changes (separate effort).

## Test plan

- [ ] CI green on this PR (lint / format / build / publint / attw on the
affected workspaces)
- [ ] Admin-merge once cr-loop converges to zero findings
- [ ] Post-merge: confirm `showcase_deploy` rebuilds integration
containers
- [ ] Post-merge value-test: `gen-ui-interrupt` + `resume-path` cells
flip GREEN on staging dashboard for LGP, LGTS, LG-FastAPI (≥3 cells)
2026-06-17 11:55:22 -07:00
Jordan Ritter 0f1fbeac23 Showcase promote-notify workflow + canonical fixtures (PR1) (#5522)
## Summary
- New `.github/workflows/showcase_promote_notify.yml` — Slack notify
workflow (workflow_dispatch only) for promote results. Posts initiation
+ threaded reply to `#team-showcase`; cross-posts to `#oss-alerts` on
partial/total failure.
- New `showcase_promote_notify.dry-run.sh` — local render-logic mirror
(no Slack API calls); validates payload + emits expected Slack messages
for manual review.
- New `showcase/test-fixtures/promote-notify/` — three canonical
fixtures (success / partial / total-failure), a strict schema validator,
and a README.
- New `docs/runbooks/showcase-promote-notify-pr1-checklist.md` —
pre-merge runbook with Slack-membership checks, dispatch-fixture
commands, and schema-mismatch test.

## Why this is PR1
PR1 lands the notify workflow with no callers. The CLI (`bin/railway
promote --notify`) lands in PR2. Splitting is required because `gh
workflow run` resolves the workflow file from the repo's default branch
— PR2's CLI cannot dispatch a workflow that doesn't yet exist on `main`.

## Test plan
- [ ] Pre-merge checklist passes (see
`docs/runbooks/showcase-promote-notify-pr1-checklist.md`)
- [ ] Required `SLACK_BOT_TOKEN` secret is set with scopes `chat:write`,
`chat:write.public`, `users:read.email`
- [ ] Bot is a member of `#team-showcase` AND `#oss-alerts` (or scope
sufficient for public posts)
- [ ] All 3 canonical fixtures dispatched via `gh workflow run --ref
<pr-branch>` produce the expected Slack messages
- [ ] Schema-mismatch test (step 5 of runbook) produces `::warning::`
annotation with NO Slack API call

## Follow-up
A bucket(d) list of defensive-hardening items was deferred to a
follow-up PR (notify workflow defensive hardening). The CR loop (4
rounds, 7 unbiased agents each) identified ~50 items across categories:
silent-Slack-failure exit codes, validator strictness gaps (ISO-8601
fractional seconds, enum constraints), dry-run/workflow parity
(lookup-failure simulation, permalink-empty fallback), payload
defensive-validation (cross-field, run_id source-of-truth). Per cr-loop
convergence-audit: these are PR1-adjacent but their own subject; they'll
land in a focused follow-up PR.

Additionally, local `actionlint` (not currently in CI for this workflow)
flags one SC2034 (`failed_count` retained for workflow/dry-run parity
but not echoed) and two SC2016 (intentional single-quoted Slack-mrkdwn
backticks in `trigger_label`). Both fold into the bucket(d) follow-up.

## CR rounds
4 rounds × 7 agents = 28 reviews. Final convergence: ALL bucket(a)
findings fixed; bucket(b)/(c)/(d) preserved for follow-up. Integration
HEAD: `b65b8452d9b847f73c885854c146547ce82c3706`.
2026-06-17 11:54:10 -07:00
Jordan Ritter c308ddff16 chore(showcase): ratchet canonical pin set to 1.60.2, revert @ag-ui/mastra bump
- showcase-canonical-pins.json: bump canonicalCopilotKitVersion 1.59.4 -> 1.60.2;
  remove ms-agent-harness-dotnet override (caught up to canonical in prior commit).
- fail-baseline.json: re-ratchet validatePinsFailCount 39 -> 38 and hash to match
  the one-item drop (ms-agent-harness-dotnet override no longer counted).
- @ag-ui/mastra: revert 0.2.4 -> 0.2.1-beta.2. 0.2.4 imports
  '@mastra/core/runtime-context' which the pinned @mastra/core@1.41.0 does not
  export, breaking 'next build' (failing mastra build-check in CI). Holding
  @ag-ui/mastra at the prior pin until a coordinated @mastra/core upgrade lands.
2026-06-17 11:46:19 -07:00
Alem Tuzlak 817f8c718d docs(bot-discord): add Discord adapter reference docs 2026-06-17 11:36:42 -07:00
Jordan Ritter b957f955e0 chore(showcase): align @copilotkit/* + @ag-ui/* deps across integrations
Aligns dependency versions across all 19 showcase integrations to current
released minor versions for the 1.60.2 release cycle.

Package families:
- @copilotkit/{a2ui-renderer, react-core, react-ui, runtime, shared, sdk-js, voice}
  1.59.4 -> 1.60.2 (18 integrations already staged; ms-agent-harness-dotnet
  catches up from 1.57.2)
- @ag-ui/{client, core, encoder} 0.0.55 -> 0.0.57
- @ag-ui/mastra 0.2.1-beta.2 -> 0.2.4 (stable on 0.x; 1.0.x major held back)

Includes the previously-missed ms-agent-harness-dotnet integration in the
@copilotkit/* bump, plus the @copilotkit/web-inspector override pin.

Lockfile-only reconciliation via npm install --package-lock-only
--legacy-peer-deps (cmdk@0.2.1 pre-existing react^18 peer-dep is unaffected).
2026-06-17 11:33:43 -07:00
Atai Barkai e5ecf00329 docs(shell-docs): clarify frontend docs guidance 2026-06-17 11:31:36 -07:00
github-actions[bot] 076a572c56 style: auto-fix formatting 2026-06-17 18:30:10 +00:00
Jordan Ritter a70236391c test(showcase): add promote-notify canonical fixtures + schema validator
Adds three canonical fixtures under showcase/test-fixtures/promote-notify/:
success.json (29 services green), partial.json (26 green / 3 red with
mixed exit codes + categories), and total-failure.json (29 red fleet
abort with truncation-suffix sentinel).

Includes validate.sh — strict schema enforcement (rc propagation, enum +
regex assertions for run_id, pre_staging, abort_reason, category, exit
codes). And a README documenting the fixture contract + how to dispatch
them via gh workflow run.
2026-06-17 11:28:52 -07:00
Atai Barkai 104dbf918e docs(shell-docs): simplify react docs sidebar section 2026-06-17 11:27:56 -07:00
Atai Barkai 7970440d1c docs(shell-docs): integrate frontend guide references 2026-06-17 11:14:04 -07:00
Atai Barkai c172982c93 docs(shell-docs): add building emoji to frontend guidance link 2026-06-17 10:59:55 -07:00
Atai Barkai 65e5a0312f docs(shell-docs): add read more cue to frontend note 2026-06-17 10:57:35 -07:00
Atai Barkai 47163b6a51 docs(shell-docs): rename react parallels section 2026-06-17 10:53:35 -07:00
Atai Barkai 2037ead870 docs(shell-docs): refine frontend guidance note 2026-06-17 10:51:27 -07:00
Atai Barkai ebb2c3cefd docs(shell-docs): add frontend docs in-progress guidance 2026-06-17 10:33:50 -07:00
Sam Julien 2d3bafd61c docs: clarify shell-docs hybrid authoring 2026-06-17 10:18:13 -07:00
Tyler Slaton 506997f4f9 showcase(docs): update premium features to be enterprise (#5511)
Updating the premium features section to instead be "enterprise" with
some reworked documentation pages.
2026-06-17 10:06:04 -07:00
Tyler Slaton 8ec7d0d4e5 docs(shell-docs): restore enterprise intelligence product name 2026-06-17 09:16:29 -07:00
Jordan Ritter c8a9053bee test(showcase/harness): red-green gate for Railway-GQL resilience + 3-tick silence threshold
Adds an integration-style test file that pins the BOTH layers of the
2026-06-17 Cloudflare-WAF-burst incident fix together:
  - the enumerator retries 3× on a 429+Cloudflare-1015 burst before
    bubbling (the original bug let one 429 abort the whole enumerate);
  - the silence monitor requires THREE consecutive silent evaluation
    cycles before posting an alert (the original bug fired on tick #1).

These assertions were the LITERAL red proof on `main`:
  - on main both `it` blocks PASSED while asserting the buggy behavior
    (calls === 1, posts.length === 1 after a single silent tick),
  - on this branch the same gates re-pin the fixed behavior (calls === 4
    after retries; posts === [] until the third silent tick).

Run-time output captured for the PR body confirms the inversion.
2026-06-17 08:21:01 -07:00
Jordan Ritter 94f6d87f54 fix(showcase/harness): require 3 consecutive silent ticks before family-silence alert
Layer a per-family consecutive-silent-tick counter ON TOP of the existing
3×period elapsed-time gate (`SILENCE_PERIOD_MULTIPLIER`). The silence
alert now requires BOTH:
  - `now - lastSuccessAt > 3 × period` (existing elapsed-time gate), AND
  - `SILENCE_CONSECUTIVE_TICK_THRESHOLD = 3` consecutive evaluation cycles
    observed silent (new — the counter resets on any successful evaluation).

Without the new gate a single bad cron tick on a family whose
`lastSuccessAt` was already stale (e.g. after a long quiet window or a
deploy gap) tripped the alert immediately — the failure mode the
2026-06-17 Cloudflare-WAF-burst incident exposed where one ~25 min flap
on backboard.railway.com/graphql/v2 paged every family at once.

The counter is NOT incremented during boot grace, so a cold-start cycle
can't alone push it to threshold. The meta-alert (`family-silence-eval`)
path keeps its own clock and is unaffected. Existing tests advance
through three consecutive silent ticks before asserting the post.
2026-06-17 08:20:51 -07:00
Jordan Ritter 7f118c5955 fix(showcase/harness): retry+cached-catalog producer enumerate (Railway-GQL resilience)
Three-retry exponential backoff (1s/4s/16s) on `source.enumerate` against
Railway-GQL when the underlying error is transient (HTTP 429, 5xx, or a
Cloudflare 1015/1020/1022 WAF marker, or a transport-level reject). On
persistent failure, fall back to the last successful catalog from a
per-enumerator in-memory cache — LOUDLY logged via
`fleet.producer.enumerate-failed-using-cache` so the cache-use shows up
in observability.

A fresh-boot process with no cached entry preserves the current
hard-fail behavior (the producer's `enumerate-failed` short-circuit) —
without a catalog there is nothing to enqueue. Real config errors
(`DiscoverySourceAuthError`, non-429 4xx, schema rot) are NOT retried so
operator-actionable failures surface immediately.

Context: 2026-06-17 Cloudflare WAF burst-blocked
backboard.railway.com/graphql/v2 for ~25 min, hard-failing the producer
enumerate on every cron tick and zeroing out D4/D5/D6 writes — the
entire staging dashboard went red within one tick. Retries + cache ride
out the burst on the same tick and preserve job production across
longer outages.

Pre-existing repo-wide lefthook test failures (@copilotkit/core,
@copilotkit/runtime, @copilotkit/shared, etc.) are unrelated to the
harness; verified by stashing my changes and running the same hook on
main with identical failures. Harness suite (131 files / 2812 tests),
typecheck, and build pass on this branch.
2026-06-17 08:20:40 -07:00
Atai Barkai 281c0c6156 docs(shell-docs): add frontend and backend docs picker 2026-06-17 08:19:00 -07:00
Atai Barkai 083723771a docs(shell-docs): refine frontend sidebar link 2026-06-17 00:35:42 -07:00
Atai Barkai 1d0ba8f105 docs(shell-docs): add React parallels sidebar link 2026-06-17 00:30:08 -07:00
Atai Barkai 3392b850a2 docs(shell-docs): simplify frontend quickstart nav 2026-06-17 00:26:59 -07:00
Jordan Ritter 5a62acbf72 docs(showcase): cell red→green SOP + agent-tiered fanout from README.md (#5512)
## Summary

Two-commit docs PR sequenced AFTER #5495 — it references CLI semantics
introduced there (control-plane `:demo` scoping, `--isolate` rebuild
scope).

1. **SOP + CLI reference + prune stale.** `showcase/TESTING.md` gains:
   - The cell red→green SOP (10-step procedural workflow for agents)
- `bin/showcase test` CLI invocation table (control-plane vs `--direct`
semantics, post-A18 / post-A21+A21b)
- Operational gotchas added to `showcase/GOTCHAS.md` (aimock fixture
caching, `--isolate` slot collisions)
- Stale invocation guidance pruned across
RUNBOOK/README/DEBUGGING/TESTING (9 items)

2. **Consolidation + agent-tiered fanout from README.md.**
- DELETE `showcase/QA-COVERAGE.md` → folded into `TESTING.md` as
Per-Demo Coverage Matrix
- DELETE `showcase/RUNBOOK.md` → unique ops content merged into
`DEBUGGING.md`; duplicated `--isolate` mechanics/CLI rules already
covered in `TESTING.md`
- README.md re-tiered as agent entry point: top-of-file fanout table
("when X, see Y.md") routing to procedural docs
- Each remaining doc gains a one-line tagline answering "what does this
answer"
   - Cross-refs use relative `./<file>.md` paths

3. **style: auto-fix formatting** — oxfmt applied locally during
pre-push to prevent CI auto-format-bot from firing.

## Test plan
- [x] All cross-refs resolved (no dangling links after deletions)
- [x] Pre-push quality on docs branch (oxfmt clean, commit hygiene
clean)
- [ ] CI gates pass (CI is the only gate for doc-only PRs per
`feedback_cr_rigor_scales`)

Note: depends on #5495 for accurate CLI semantics references.
2026-06-16 23:37:12 -07:00
Jordan Ritter cf615df21a fix(showcase): disjoint catchall userMessages + content-asserting probes (supersedes #5465) (#5495)
## Summary

Restores **all 4 D5 custom-catchall cells** (LGP, crewai-crews,
built-in-agent, claude-sdk-typescript) to green via the
production-equivalent control-plane pipeline, plus 4 harness honesty
fixes that turn `--isolate` and `:demo` invocations into
apples-to-apples staging mirrors.

**Cells GREEN locally (verified post-CR via `bin/showcase test
<slug>:tool-rendering-custom-catchall --d5 --isolate`):**
- `langgraph-python` ✓
- `crewai-crews` ✓
- `built-in-agent` ✓
- `claude-sdk-typescript` ✓

## Harness honesty fixes (the load-bearing ones)

1. **A11 — probe-scan inline-needle.** `page.evaluate(fn, arg)` was
passing `undefined` to the browser closure →
`customContentPhrasePresent` was permanently false fleet-wide, masking
every other failure mode. Fix: inline the canonical phrase literal in
the closure. A25a propagated the same fix to
`d5-tool-rendering-default-catchall.ts` (still had the broken pattern).
2. **A18 — control-plane honors `:demo`.** `bin/showcase test
<slug>:<demo> --d5/--d6 --isolate` previously ignored the demo qualifier
(d5 hardcoded to `agentic-chat`; d6 aggregate-only). Now per-demo
scoping flows through `buildLocalServicesJson` + `expectedKeys`.
Eliminates a class of silent false-positive PASS.
3. **A21 + A21b — `--isolate` rebuild scope.** A21 scoped `--build` to
target slug (BuildKit contention unblock); A21b corrected an A21
regression where positional-after-`up` restricted which services
started. Result: two-call compose split — `compose --profile infra up
-d` (no build, uses cached images), then `compose --profile infra
--profile <slug> up -d --build <slug>` (rebuild target only). Cold-build
~30s–2 min instead of 10+ min full-stack rebuild.

## Cell fixes

- **A14 (crewai-crews)**: backend defect —
`tool-rendering-custom-catchall` agentId routed to shared
`LatestAiDevelopment` flow at `/` with no
`get_weather`/`get_stock_price` handlers; tool-loop never closed. Added
`get_stock_price_impl`, re-routed agentId to `/tool-rendering`.
- **A19b (built-in-agent)** + **A20 (claude-sdk-typescript)**: backend
ID-rewrite (TanStack `fc_*`, Anthropic `toolu_*`) broke
`toolCallId`-gated narration fixtures. Swapped to `turnIndex`
discriminator (backend-id-invariant). `response.content` (canonical
phrase) preserved verbatim.
- **A10 (langgraph-python)**: parent commit \`9491b8934\` disjoined the
d5 probe userMessages but didn't add matching LGP-gold fixture entries.
Added 4 entries (turn-1 emit + turn-2 narration for Tokyo + AAPL).
- **A1-A9 (R1+R2 fixture hygiene)**: cleanups to dead `turnIndex:0`
fallbacks, `hasToolResult:false` gates, AAPL value drift, csdkts
copy-paste bug — pre-A11 era.

## Test coverage (A25 round, addresses post-A21b CR findings)

- A11 inline-needle invariant (regression test — fails if the fix is
reverted)
- A7 `requireContentPhrase=true` branch end-to-end
- A18 `buildLocalServicesJson` + `expectedKeys` + `dedupeScopes` +
`runViaControlPlane` error surfacing (17 new tests in
\`control-plane-run.test.ts\`)
- A21+A21b two-call compose argv contract (lifecycle.test.ts, 7 tests)

Full harness vitest suite: **2790/2790 pass**. Typecheck clean. oxfmt
clean.

## Test plan
- [x] Local control-plane (`--d5 --isolate`) on all 4 cells — GREEN
- [x] LGP regression check across every fix-round — GREEN throughout
- [x] vitest harness suite (2790 tests) — GREEN
- [x] tsc --noEmit — clean
- [x] oxfmt --check — clean
- [x] Pre-push quality + 7-agent cr-loop + 3-slot post-A25 confirmation
round — converged ZERO blockers
- [ ] CI gates (gh pr checks 5495) on push — to be observed

## Notes for reviewers
Docs PR (\`docs/showcase-sop-tiering\`, off main) sequences AFTER this
one — it documents the new CLI semantics (\`:demo\` scoping,
\`--isolate\` rebuild scope) introduced here.
2026-06-16 23:30:43 -07:00
github-actions[bot] ac85ac4f96 style: auto-fix formatting 2026-06-17 06:20:33 +00:00
Jordan Ritter 38aa931c71 fix(showcase/harness): add A18 test coverage + tighten control-plane error surfacing
- Add control-plane-run.test.ts (17 tests) covering buildLocalServicesJson,
  expectedKeys, dedupeScopes, and runViaControlPlane error surfacing
- Export SlugScope, buildLocalServicesJson, expectedKeys, dedupeScopes
  for unit-test coverage (factored inline dedup loop into dedupeScopes
  helper at the same time)
- runViaControlPlane: surface scopeLabel (demo-aware) in the 0-enqueue
  error instead of the bare-slug join, with an empty-targets guard so
  the error never renders with a double-space gap
- runViaControlPlane: treat tick.enqueueFailures > 0 as fatal — partial
  enqueue used to silently proceed and either mask missing cells or
  hang the poll loop to timeout
- Eliminate a stray literal NUL byte in the source by switching the
  dedup key separator to a \x00 escape
- lifecycle.up(): name the compose call (infra-up vs target rebuild)
  in the health-fail error so an operator can tell which call left a
  service unhealthy

(cherry picked from commit 9f35c64adfdf7f5ff2bf0a5ae57ed03818cea607)
2026-06-16 23:14:51 -07:00
Jordan Ritter 61edef02a6 fix(showcase/harness): apply A11-style inline-needle to default-catchall + add regression tests for inline-needle + requireContentPhrase
CR Finding 1 (BLOCKER): d5-tool-rendering-default-catchall.ts used the
broken page.evaluate(fn, arg) second-arg form to pass the leak-phrase
needle into the browser-side closure. A11 proved empirically that the
arg arrives as undefined inside the closure, making 'if (needle)' guard
the entire leak-detection cascade as dead code — customLeakPhrasePresent
stayed false forever, rendering validateDefaultCatchall's leak branch
dead code as well. Mirrors the A11 fix on the sibling custom-catchall
probe by inlining the needle as a JS string literal inside the closure;
no page.evaluate(fn, arg) dependency at all. Both probes now share the
same inline-needle pattern and keep the canonical literal in lock-step
with their exported phrase constant.

CR Finding 2 (MAJOR): the A11 inline-needle fix on the sibling
custom-catchall probe had no regression test — fake Page.evaluate in
makePageReturning never executes the probe closure, so reverting the
fix would not be caught. Added regression tests that capture the
probe's function source via toString() and assert (a) the canonical
phrase appears as a literal inside the page.evaluate(...) closure and
(b) the closure takes no parameter / the evaluate call has no
second arg. Added the same coverage to default-catchall to protect
the new A25a fix.

CR Finding 3 (MAJOR): A7's requireContentPhrase=true branch in
validateCustomCatchall / assertCustomCatchall had zero coverage —
tests omitted the third arg and exercised only the default false
branch. Added coverage for the true branch (pass on phrase present,
fail on phrase absent, fail on phrase undefined, default-branch
preserved) plus assertCustomCatchall plumbing through the options
form. Also added coverage for default-catchall's customLeakPhrasePresent
branch in validateDefaultCatchall for symmetry.

Local proof:
- RED (fix reverted via git stash): 2 inline-needle regression tests
  fail on d5-tool-rendering-default-catchall.test.ts
- GREEN (fix restored): 34/34 tests pass across both files

Out of scope (NOT touched this commit): showcase/harness/src/cli/
control-plane-run.ts and lifecycle.ts (A25b's scope).

(cherry picked from commit c409e3a8d99e16ad0bb05ee3c2e2051e5792049f)
2026-06-16 23:14:50 -07:00
Jordan Ritter 15f828915a style(showcase/harness): apply oxfmt to lifecycle.ts compose argv call 2026-06-16 22:39:29 -07:00
Jordan Ritter 19b65882d1 style: auto-fix formatting 2026-06-16 22:34:44 -07:00
Jordan Ritter 6cc1803f37 docs(showcase): consolidate + re-tier for agent navigation (README fanout entry)
Re-tier the showcase docs tree to be an agent entry point: README.md
opens with a 'when X, see Y' fanout table that routes to the right
procedural doc; each procedural doc gets a one-line tagline answering
'what does this answer'.

Consolidation:
- DELETE showcase/RUNBOOK.md — operational content merged into DEBUGGING.md
  (Integration Patterns, Docker Compose Environment, Production Debugging,
  Anti-Patterns, Aimock Fixture Deployment, Dev Iteration Speed). The
  --isolate mechanics + CLI rules were already duplicated in DEBUGGING.md.
- DELETE showcase/QA-COVERAGE.md — per-demo coverage matrix + starter hero
  matrix + probe depth + infra locations + gaps folded into TESTING.md as
  the 'Per-Demo Coverage Matrix' section.

Taglines added (no behavioral change to content): TESTING.md, DEBUGGING.md,
GOTCHAS.md, INTEGRATION-CHECKLIST.md, STYLING-GUIDE.md, FRONTEND-STRATEGY.md,
RAILWAY.md, bin/README.md, aimock/README.md, aimock/RAILWAY.md,
harness/README.md, harness/docs/rotation-drill.md.

Cross-link fixups: FRONTEND-STRATEGY.md (was QA-COVERAGE.md →
TESTING.md#per-demo-coverage-matrix), TESTING.md (removed dangling RUNBOOK
companion reference), README.md (rewritten as fanout entry + retained
from-scratch setup + dashboard SOPs below the fanout).

PARITY_NOTES.md × 12 left alone (per-slug context, not redundant).

(cherry picked from commit 75c9d9755c9118c8abc1fa52deda2012b768cab1)
(cherry picked from commit b64189bae0fe2c9e3a5e3ca440013deb4121f23b)
2026-06-16 22:30:08 -07:00
Jordan Ritter 423167d12e docs(showcase): SOP for cell red→green + control-plane vs --direct CLI reference; prune stale invocation guidance
New content:
- TESTING.md: add 10-step cell red→green SOP + bin/showcase test invocation
  table (control-plane vs --direct, per-demo scoping matrix); retain
  existing CI gating matrix below.
- GOTCHAS.md: add operational gotchas — aimock caches fixtures at container
  startup (warm-slot reuse needs docker restart) + --isolate slot collisions
  with foreign Docker projects.
- README.md: cross-link to TESTING.md SOP from CLI section; flesh out
  --isolate / --direct in test options table; update use cases.
- RUNBOOK.md: update Verifying a Slug's D6 State to use auto-named --isolate;
  note A21+A21b per-slug rebuild scoping; rewrite Fixture Matching to teach
  picking the backend-id-invariant discriminator (turnIndex post-A12/A13/A20);
  modernize Debugging Sequence to --isolate flow.
- DEBUGGING.md: lead with TESTING.md SOP cross-link; update Phase 1 to
  --isolate canonical; soften turnIndex-only log-line description; note
  aimock startup caching in Phase 5; switch Strategy 5 gold-standard check
  to --isolate.

Pruned/updated stale claims (post-A11/A12/A13/A18/A20/A21/A21b):
- RUNBOOK.md "Do not use turnIndex in new fixtures" — turnIndex is now
  the canonical backend-id-invariant alternative when toolCallId is fragile
  (Anthropic / TanStack Responses API ID rewrites). Replaced with discriminator
  selection guidance.
- RUNBOOK.md anti-pattern "NEVER use turnIndex" — replaced with NEVER
  anchor on toolCallId strict equality against ID-rewriting backends, and
  NEVER use --direct for value-tests.
- RUNBOOK.md bin/showcase test <slug> --d5 (no --isolate) as canonical SOP
  — replaced with --isolate canonical, no manual name required.
- README.md --d5 option description claiming "subagents/tool-rendering/agentic-chat"
  fixed slate — replaced with "defaults to agentic-chat representative; :demo
  qualifier honored post-A18".
- DEBUGGING.md Phase 1 "showcase up aimock <slug> && showcase test <slug> --d5"
  as primary — kept as legacy alternative; --isolate is now lead.
- DEBUGGING.md Phase 5 "fixtures baked into Docker image" — clarified that
  aimock additionally caches fixtures in memory at startup (volume-mounted
  isolated stack still requires docker restart for warm-slot edits).
- DEBUGGING.md Strategy 5 "showcase test langgraph-python --d5" — replaced
  with :demo + --isolate so the gold-standard check exercises the same cell.

(cherry picked from commit 0e548455043396972f7fb5b96f8c0ea8abdf1d98)
(cherry picked from commit 592c02d392350d02cc5e17544e663a6605b8da65)
2026-06-16 22:30:08 -07:00