Commit Graph

4085 Commits

Author SHA1 Message Date
Ran Shem Tov f404791ecd chore(showcase): bump @ag-ui/aws-strands 0.2.0 -> 0.2.2 (strands-typescript)
0.2.2 ships the adapter fixes for the two upstream defects this hunt surfaced:
empty tool-result content (render-tool demos -> OpenAI 400) and RUN_FINISHED
emitted before parallel tool calls drain (INCOMPLETE_STREAM). With the bump,
beautiful-chat, the chart demos, and the catch-all 'Chain tools' pill complete
cleanly. Verified on real gpt-4o.
2026-06-23 19:46:43 -07:00
Ran Shem Tov a5dea4164c feat(showcase): render a2ui-fixed-schema via dedicated agent + A2UIMiddleware
Both integrations: a dedicated `a2ui_fixed_schema` backend agent exposes a
`display_flight` tool that returns the A2UI `a2ui_operations` envelope
(createSurface -> updateComponents -> updateDataModel) built from a fixed,
pre-authored flight layout targeting the page's `copilotkit://flight-fixed-catalog`
catalog. Mirrors the upstream ag-ui dojo a2ui_fixed_schema demo (ag-ui#2021)
adapted to the showcase's existing catalog/renderers.

- TS returns the envelope as an object (lands in a json content block); Python
  returns it as a JSON string (text block) — per each SDK's tool-return shape.
- Enable the runtime A2UIMiddleware on the a2ui-fixed-schema route
  (injectA2UITool: false, defaultCatalogId pinned) so it detects the envelope
  and paints; the agent emits the envelope itself, so no generate_a2ui injection.
- Mount the agent on the /a2ui-fixed-schema sub-path; point the route there.

Verified on real gpt-4o: the flight card paints on both integrations.
2026-06-23 19:46:43 -07:00
Ran Shem Tov e223fe0fb9 feat(showcase): register strands-typescript as a dashboard baseline partner
Add strands-typescript to BASELINE_PARTNERS so it gets its own coverage
column alongside its Python sibling (mirroring how langgraph-typescript sits
beside langgraph-python). Bump the partner-count assertion 26 -> 27.
2026-06-23 19:46:43 -07:00
Ran Shem Tov 0978b4f5ad fix(showcase): make strands + strands-typescript context, open-gen-ui, hitl, and dice demos work
Apply the same fixes to both the strands (Python) and strands-typescript
integrations for 1:1 parity:

- Lift RunAgentInput.context into the prompt (buildStatePrompt /
  build_state_prompt) so useAgentContext values (readonly-state-agent-context)
  and the openGenerativeUI design-skill / sandbox-function context actually
  reach the model. The adapter does not surface context on its own; this
  mirrors langgraph's lift-context-into-prompt pattern.
- open-gen-ui: prepend an imperative to the visualization design skill (and add
  a design skill to the advanced cell) so the model calls generateSandboxedUi
  instead of answering in plain text, and clarify that sandbox functions are
  iframe->host bridges, not LLM tools.
- hitl-in-chat: sharpen the book_call description so it wins scheduling intents
  over the shared backend schedule_meeting tool (which renders no picker).
- Add a roll_dice tool (shared python tools + TS tools) so the tool-rendering
  catch-all 'Roll a d20' and 'Chain tools' pills work.
2026-06-23 19:46:43 -07:00
Jordan Ritter 543a700c54 fix(showcase): record pydantic-ai generate_a2ui fixture
Add the missing "Show me my sales dashboard for this quarter." pill to
the pydantic-ai gen-ui-declarative D6 fixture set: an outer turn
(generate_a2ui, no args) plus the matching _design_a2ui_surface inner
turn carrying the dashboard component payload (KPI metrics row + revenue
pie + monthly-revenue bar), mirrored from the langgraph-python canonical
and the ms-agent-dotnet equivalent.

Closes the staging pydantic-ai D-chat 503 (no_fixture_match): the backend
hits aimock with tools=[generate_a2ui] for this userMessage under
x-aimock-context: pydantic-ai, but only ms-agent-dotnet had the fixture.

Deterministic canonical mirror (no real-LLM recording); contains only the
request-match shape and the A2UI response — no credentials.
2026-06-23 17:57:52 -07:00
Jordan Ritter ec646bbf4f Merge remote-tracking branch 'origin/main' into chore/remove-harness-legacy-ssot
# Conflicts:
#	showcase/scripts/railway-envs.generated.json
#	showcase/scripts/railway-envs.ts
2026-06-23 17:56:18 -07:00
Mark Fogle 06e6de87c7 feat(showcase/google-adk): add ADK-only a2ui-recovery demo (OSS-158 recovery made visible)
A new demo that surfaces the ag_ui_adk >= 0.7.0 validate->retry recovery loop
+ recovery-exhausted hard-fail envelope, rendered as the building/retrying/
failed lifecycle by @ag-ui/a2ui-middleware >= 0.0.10. Two pills, driven by
aimock fixtures that key the inner render_a2ui call by recovery attempt
(sequenceIndex):
- HEAL: attempt 0 structurally-invalid (unresolved child) -> heals to valid
- EXHAUST: invalid on every attempt -> a2ui_recovery_exhausted fallback

Backend-owned (get_a2ui_tool, recovery.maxAttempts=3, injectA2UITool=false),
reusing the declarative-gen-ui catalog + sales context. ADK-only by design:
the recovery loop lives in the ADK middleware; langgraph-python's runtime A2UI
path (@ag-ui/a2ui-middleware single-pass) has no equivalent, so there is no LP
parity reference (exempt from LP e2e-parity). Ported from ag-ui
examples/.../a2ui_recovery.py.

Wires: recovery_agent.py + registry; copilotkit-a2ui-recovery route;
demos/a2ui-recovery page/chat/suggestions; manifest demo+feature; qa spec;
aimock fixture; feature-registry + constraints (constrained-explicit) entries.

Verified in-sandbox: import-smoke (41 agents, recovery cfg), generate-registry +
validate-constraints + validate-parity (20/20 pass) + aimock-fixtures (818) +
python suite (62) all green. Browser e2e (lifecycle render + sequenceIndex
routing) + e2e spec are the documented outside-sandbox steps.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 00:49:18 +00:00
Mark Fogle 0e7b803c3e chore(showcase/google-adk): remove dead hand-rolled A2UI planner + SalesPipelineAgent
The hand-rolled google.genai generate_a2ui planner (and the orphaned
SalesPipelineAgent that consumed it) in main.py are superseded by the
ag_ui_adk 0.7.0 middleware (get_a2ui_tool), now wired backend-owned in
declarative_gen_ui_agent.py / beautiful_chat_agent.py. main.py is reduced to
the shared tool wrappers + before_model/before_agent callbacks still covered
by tests; dead A2UI imports pruned.

- delete tests/python/test_generate_a2ui.py (tested the removed planner)
- manifest declarative-gen-ui: drop stale src/agents/main.py highlight + fix description

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 00:32:32 +00:00
Jordan Ritter 5859ae4057 docs(showcase): sync promote-notify README counts after harness-legacy removal
PR #5658 dropped harness-legacy from the fixtures but the README still
showed pre-removal counts, contradicting the edited fixtures.
2026-06-23 17:19:29 -07:00
Jordan Ritter 0809c25dac chore(showcase): remove redundant harness-legacy service from SSOT and fixtures
The pool-fleet migration is complete: the control-plane harness plus the
prod workers (deployed 2026-06-19, HARNESS_ROLE=worker, pool count 2) now
cover every probe dimension the interim `harness-legacy` fleet-migration
bridge was holding live, so `harness-legacy` is dead config.

This is the code-side cleanup only:
- Remove the `harness-legacy` entry from the railway-envs SSOT and the
  now-dead `key === "harness-legacy"` special-case in computePromoteClosure.
- Regenerate railway-envs.generated.json (41 -> 40 services).
- Drop harness-legacy from the golden snapshot, the gateIgnore expectation
  sets, the promote-notify fixtures, and the redeploy-env doc comments;
  update the service-count assertions (41 -> 40).
- Fix the stale "STAGING-ONLY" harness-workers comment: prod workers are
  live on Railway, though this SSOT entry still models the staging
  instance only (no prod env backfilled here yet).

The live Railway `harness-legacy` service is torn down separately as a
follow-up infra step.
2026-06-23 17:05:26 -07:00
Jordan Ritter 59e18693eb feat(showcase): honest promote-notify message + durable healthcheckPath SSOT tracking
Promote-notify Slack message: name the promoted AND failed services (one
Failed: header + bullets), legible "(N): <names>" count, real wall-clock
elapsed (integer-coerced), and drop the constant verify-prod legend line.

Durable healthcheckPath: track it per-service/env in the SSOT (railway-envs),
re-assert it on the promote pin path (omit-when-absent, never null), and route
deploy-to-railway provisioning through isTrackedService/resolveProvisionHealthcheck
so a tracked-null service omits the healthcheck while an untracked one keeps the
/api/health default — fixing the silent prod-healthcheck drift that refused aimock.

Tests: ruby pin-reassert spec + deploy-to-railway healthcheck spec + emit/golden/accessor.
2026-06-23 16:56:16 -07:00
Mark Fogle 55f4f43afd feat(showcase/google-adk): surface 0.7.0 A2UI middleware (backend-owned) in declarative-gen-ui + beautiful-chat
Replace the hand-rolled google.genai A2UI planners (in main.py and
beautiful_chat_agent.py) with the published ag-ui-adk >= 0.7.0 middleware
sub-agent via get_a2ui_tool(), surfacing OSS-158 (forced render_a2ui
sub-agent + toolkit validate->retry recovery loop + recovery-exhausted
hard-fail envelope + render_as_llm_instructions / parse_and_fix healing).

Wiring is BACKEND-OWNED (injectA2UITool: false), matching the AWS Strands /
ag2 external-framework convention rather than langgraph-python's
runtime-driven injectA2UITool: true. Backend-owned is required: the planner
now lives in the ADK middleware, so letting the runtime also inject would
double-bind the tool slot. The explicit  is load-bearing post
CopilotKit#5611 (a provider catalog otherwise defaults injectA2UITool to true).

- declarative_gen_ui_agent / beautiful_chat_agent: tools include
  get_a2ui_tool({model, default_catalog_id}); beautiful keeps its other tools.
- shared_chat: add get_a2ui_model() to resolve a concrete Gemini BaseLlm for
  the sub-agent (mirrors get_model's aimock-proxy wiring).
- routes: injectA2UITool stays false (declarative + beautiful-chat).
- registry/agent_server: no spec-level a2ui config needed (tool is agent-owned).

Verified vs published 0.7.0 + real Gemini: backend-wired generate_a2ui emits
a2ui_operations; OSS-158 gate (subagent + recovery invalid->valid + hard-fail
a2ui_recovery_exhausted) all retained through this exact wiring.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 23:19:45 +00:00
Mark Fogle ebd87a23f3 chore(showcase): bump google-adk to ag-ui-adk 0.7.0 (OSS-158 A2UI middleware)
Pins the published middleware that ships the A2UI auto-inject + toolkit
recovery + render_as_llm_instructions/parse_and_fix path. Foundational
step; the A2UI agent re-wire onto that path follows under verification
once 0.7.0 is on PyPI.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 23:19:45 +00:00
Sam Julien e8c2cd5005 docs(shell-docs): show threads in authored framework nav 2026-06-23 15:42:01 -07:00
Sam Julien 7e54dbe746 docs(shell-docs): place threads with chat UI guides 2026-06-23 15:42:01 -07:00
Tyler Slaton 75611b272c chore: merge main into PR 5480 2026-06-23 15:32:09 -07:00
Tyler Slaton 6d9397c42e Preserve quickstart CTA styling in shell docs reference pages (#5654)
## Summary
- Add a dedicated `shell-docs-primary-cta` style for the hero quickstart
link so it keeps the primary CTA color inside reference content.
- Cover the new class usage in the hero and framework overview tests.

## Testing
- Updated unit tests to verify the quickstart CTA class and the matching
global CSS override.
- Updated unit tests to verify the framework overview markup includes
the primary CTA class.
2026-06-23 15:27:06 -07:00
Tyler Slaton 307a52c8eb Preserve quickstart CTA styling in reference content 2026-06-23 15:26:36 -07:00
Jordan Ritter 9f2b859450 fix(showcase): stop ag2 generate_a2ui empty-arg validation loop
The ag2 declarative-gen-ui route pointed its HttpAgent at the root
catch-all mount (agents/agent.py) instead of the dedicated
/declarative-gen-ui mount (a2ui_dynamic.py), and generate_a2ui declared
a required context arg that the model emits as {}. pydantic rejected
every call with "context Field required" and AG2 retried without bound —
a 630-iteration hot loop per pill that flooded logs and starved the
frontend.

Fix: route to the dedicated mount with injectA2UITool:false (the
dedicated agent owns generate_a2ui and emits a2ui_operations itself);
make generate_a2ui a no-arg tool matching the D6 fixtures and the
langgraph-python gold standard, with a constant inner system prompt
(per-pill distinctness comes from the captured user message). Regenerated
the gen-ui-declarative fixture and ported the LP definitions/renderers
catalog (all 7 driver testids) for parity. Eliminates the validation
loop: runsFinished=1, zero validation errors.
2026-06-23 15:12:09 -07:00
Jordan Ritter 7cb43581aa fix(showcase): forward request-time tools in llamaindex AG-UI adapter
The LlamaIndex AG-UI adapter never forwarded RunAgentInput.tools, so
page-injected frontend tools (useFrontendTool / useComponent /
useHumanInTheLoop) were invisible to the LLM — runs the model couldn't
satisfy ended in RUN_ERROR (no RUN_FINISHED), which the harness reported
as sse-missing.

New RequestAwareAGUIChatWorkflow re-implements the chat step as the
upstream body plus three additions: forward RunAgentInput.tools as no-op
FunctionTool stubs carrying the verbatim injected JSON schema; re-role
tool-result messages to role="tool" on the LLM-bound message copy only
(so the hasToolResult second-leg fixture matches) without mutating
stored/snapshot history; and for frontend tool calls dispatch the
ToolCallEvent but skip the duplicate TOOL_CALL_CHUNK (the bare snapshot
already delivers the call). beautiful-chat 0/5 -> 5/5; mcp-apps fixed.
2026-06-23 15:12:09 -07:00
Sam Julien 81c716a4aa docs(shell-docs): add framework-scoped threads callouts 2026-06-23 15:07:13 -07:00
Jordan Ritter 160ba5a4aa fix(showcase): skip non-probe-eligible services in staging precondition instead of crashing 2026-06-23 14:51:16 -07:00
Jordan Ritter 7a00e2373d feat(showcase): make starter fleet always-on + verify-deploy starter driver (#5641)
## Summary

The 12 `starter-*` services were deliberately `sleepApplication=true`
(sleepable), `probe.staging=false` (held out of the verify-deploy
staging matrix). That class-difference is the root of three recurring
symptoms: the `service=all` promote false-fails them (SLEEPING / no
running-digest), the deployed-starter smoke test 404s a cold container,
and they're excluded from staging validation. Decision: **bring them
into the normal managed-fleet flow — always-on + staging-probed.**

- **`verify-deploy.drivers.starter.ts` (new)** — real `probeStarter`
baseline driver (deployment SUCCESS + HTTP 200 on `/`, mirroring
`probeShell`); replaces the fail-loud `case "starter"` stub.
- **`railway-envs.ts`** — `staging.probe: false→true` for all 12
starters (prod already true); stale "staging probe OFF" comments
rewritten. `railway-envs.generated.json` regenerated (in-sync via
`--check`).
- **`provision-starter-fleet.ts`** — `sleepApplication: true→false` +
comments/logs; test updated.

## ⚠️ Deploy ordering (must hold)
This PR flips `probe.staging=true`, which routes starters into the
staging matrix. It must **not** merge until the live services are
flipped always-on, or a `service=all` promote would probe still-sleeping
starters and fail. Sequence: **(1) live-flip 24 instances
`sleepApplication=false` + redeploy (12 staging + 12 prod), (2) verify
awake + serving `/`, (3) merge this PR.**

## Verification
- Red-green: starter driver (stub→probe `/`), SSOT golden (probe.staging
true), provisioner (sleep false) — all RED→GREEN.
- `showcase/scripts`: 2092 tests pass; `tsc --noEmit` clean on changed
files.

## Test plan
- [x] vitest suites green (2092)
- [ ] CR
- [ ] live 24-instance always-on flip + redeploy (staging + prod)
- [ ] CI green
- [ ] post-flip: `service=all` dry-run shows starters pass staging probe

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-23 14:43:55 -07:00
Jordan Ritter d0497649c7 test(showcase/harness): turn-done signal completion/backstop matrix
Cover the data-copilot-running turn-done signal in waitForTurnComplete:
true->false transition completion, stayed-stopped quiescence, the
attr-gated early backstop, pre-send run-start baseline, and the
integration wait-for-turn-complete behavior.
2026-06-23 14:32:05 -07:00
Jordan Ritter 494d836005 feat(showcase/harness): reliable data-copilot-running turn-done signal in waitForTurnComplete
Make waitForTurnComplete use the page-side data-copilot-running attribute
as the primary turn-done signal: detect the running true->false transition,
require stayed-stopped quiescence, and gate the early backstop on
attrPresent + runningNow to avoid headless false-RED. Capture a pre-send
run-start baseline so fast turns keep the primary signal alive, read
surfaceReady once per poll, and add computeMaxTurnDurationMs.
2026-06-23 14:31:59 -07:00
Jordan Ritter 1696445da3 feat(showcase/harness): page-side data-copilot-running MutationObserver
Add buildCopilotRunningObserverScript to sse-interceptor.ts and wire it
via addInitScript so the page exposes a data-copilot-running attribute
that the harness can observe for turn-completion signaling.
2026-06-23 14:31:50 -07:00
Nathan 🔶 Tarbert 97d1de4682 fix(showcase): use agent.setState in built-in-agent shared-state demo
The built-in-agent shared-state-read-write demo wrote agent state via
direct property assignment (`agent.state = {...}`), which sets the value
but does not fire `onStateChanged` — so subscribed components never
re-render off a UI write. Every other framework demo (including the
langgraph-python gold standard) already uses `agent.setState(...)`, which
both updates state and notifies subscribers. This was the lone divergence.

Switch the seed and the title write to `agent.setState(...)` so the demo
matches the documented pattern (docs/shared-state.mdx says writes go
through `setState`) and the rest of the showcase.

Verified: built-in-agent D5 e2e-deep suite green after rebuild.
2026-06-23 17:21:49 -04:00
Jordan Ritter bc784a1412 feat(showcase): make starter fleet always-on (sleepApplication=false, staging probe on)
Flip the 12 starter-* services to behave like every other managed
showcase service: always-on and staging-probed.

- railway-envs SSOT: staging probe false -> true for all 12 starters
  (prod was already on); refresh the stale "staging probe OFF" /
  fail-loud-placeholder doc-comments to reflect baseline-driver verification.
- provision-starter-fleet: sleepApplication true -> false so the
  provisioner creates always-on (not sleepable) Railway instances; update
  the "the whole point — sleepable" header/inline comments + log lines.
- Regenerate railway-envs.generated.json (emit-railway-envs-json.ts) and
  the golden fixture; update the SSOT probe + provisioner sleep assertions.
2026-06-23 12:00:53 -07:00
Jordan Ritter cb8f9ebea2 feat(showcase): add verify-deploy baseline driver for the starter fleet
Replace the fail-loud `case "starter"` stub in verify-deploy's dispatch
switch with a real baseline liveness driver. The starter-template fleet
EXPOSEs only its Next.js frontend (serving `/` and `/api/copilotkit`,
NO `/api/health`), so the driver healthchecks `/` via probeBaseline,
exactly like the Next.js shells: deployment-SUCCESS + HTTP 200 on `/`.
2026-06-23 12:00:37 -07:00
Alem Tuzlak 107cdad5af feat(slack-example): reaction/ephemeral/modal feature demos
Add per-feature demos to examples/slack that narrate per-platform degradation
explicitly rather than failing silently:
- emoji triage — 🐛/🔥/✅ reactions file/escalate/ack via the agent
- /preview — ephemeral draft issue (native only-you on Slack, DM fallback on
  Discord/Telegram)
- /file-issue — modal form (Slack rich, Discord text-only, Telegram
  conversational fallback)

Also updates the Slack frontend guide (slack.mdx) with the capability matrix.
2026-06-23 20:49:38 +02:00
Mike Ryan db09796809 fix: gate thread endpoints by runtime capability 2026-06-23 11:32:30 -07:00
Alem Tuzlak 43fdba74aa feat: AG-UI standard interrupt support in useInterrupt + BuiltInAgent
Adds the AG-UI standard interrupt flow (RUN_FINISHED outcome:interrupt + resume array) alongside the legacy on_interrupt path.

- core: forward the standard resume array through runAgent.
- react-core / vue / react-native: useInterrupt handles standard interrupts with resolve()/cancel(), surfaces the primary + full interrupt set, and persists each resolved tool-backed interrupt as a tool-result message so multi-turn conversations stay well-formed (no dangling tool call -> no tool-call loop).
- runtime BuiltInAgent: native interrupts for the aisdk + tanstack factory paths via each SDK's needsApproval primitive (tool-approval-request / CUSTOM approval-requested -> outcome:interrupt); classic interrupt-tool emission + ctx.interrupt() factory primitive; idempotent resume injection mapped to each SDK's native tool-result; getCapabilities advertises humanInTheLoop.interrupts.
- docs: document standard interrupt support.

Verified across core/react-core/runtime unit suites and a real-model multi-turn run on both aisdk and tanstack.
2026-06-23 20:14:17 +02:00
Jordan Ritter 951c20e329 fix(showcase): wire promote-notify renderer — emit results JSON + dispatch three-variant Slack message
promote-fleet.sh now emits a base64 results JSON (schema_version=1) with
both succeeded[] and failed[] alongside the existing succeeded_csv /
staging_drift outputs. showcase_promote.yml's notify job replaces the old
inline two-state (success/failure) notifier — which dumped the full
requested CSV and mislabeled any partial promote as a blanket Failed — with
a single dispatch of the previously-orphaned showcase_promote_notify.yml
renderer (success / partial / total). The dispatch enriches promote-fleet's
results blob with run context (6-hex run_id, trigger=workflow, operator,
pre_staging) and authenticates via the devops-bot App token (actions:write),
mirroring canary.yml — the default GITHUB_TOKEN cannot start new workflow runs.

promote-fleet has no failure taxonomy, so each failed entry uses the default
category "promote-failed".
2026-06-23 10:49:17 -07:00
Alem Tuzlak 7393df22ff feat(bot): pluggable StateStore persistence + cross-platform transcripts (A+C) (#5613)
## Summary

Adds the **pluggable persistence layer** for the bot packages (spec
slices **A + C**): a single `StateStore` interface with
`memory`/`redis`/`postgres` backends, the existing in-memory action
store folded onto it (durable actions across restarts for free),
turn-locking + inbound-event dedup + thread subscriptions/state, and
**cross-platform per-user transcripts** (`bot.transcripts`). These are
the open-source primitives a future managed `IntelligenceStore` slots
into -- that backend is intentionally **out of scope** here.

Surface parity (reactions / ephemeral / modals -- slice **D**) and
concurrency strategies + serverless webhook ingress (slice **B**) are
**deferred** to follow-on PRs.

## What's included

**A -- `StateStore` foundation**
- `StateStore` interface: grouped `kv` / `list` / `lock` / `dedup` /
`queue` primitives (`packages/bot`).
- Backends, each its own package, all passing a shared conformance
suite: `MemoryStore` (in `@copilotkit/bot`),
**`@copilotkit/bot-store-redis`** (Lua-scripted lock/queue atomicity),
**`@copilotkit/bot-store-postgres`** (row-based TTL lock, `SKIP LOCKED`
dequeue).
- `ActionStore`/`InMemoryActionStore` reimplemented over `state.kv`
(`kvActionStore`) and marked `@deprecated`; `ConversationStore` gains a
`StateStore`-backed helper.
- **All persistence config lives under one `store` object** --
`createBot({ store: { adapter, state, identity, transcripts,
onLockConflict, lockTtl, dedupTtl } })` (default backend `MemoryStore`;
legacy root `actionStore` still honored). `store.state` is an optional
**Standard Schema** that makes `thread.state()` / `thread.setState()`
**typesafe** (inferred type + runtime validation on write).

```ts
const bot = createBot({
  adapters: [slack(/*...*/)],
  agent,
  store: {
    adapter: createRedisStore({ url: process.env.REDIS_URL }),
    state: z.object({ step: z.enum(["ask_name", "done"]) }), // -> thread.state()/setState() typed
    identity: ({ author }) => author.email ?? null,
    transcripts: { retention: "30d", maxPerUser: 200 },
    onLockConflict: "drop",
    lockTtl: 60_000,
    dedupTtl: 300_000,
  },
});
```
- **Durable actions** fall out of the fold -- the minted `ck:` action id
is the `kv` key, so a click after a restart re-renders from the
persisted snapshot when a durable backend is configured.
- Consumers wired: action snapshots, thread subscriptions + per-thread
state (`kv`), per-conversation **turn lock** with `onLockConflict:
"drop" | "force"` (`lock`), **inbound-event idempotency** by platform
event id (`dedup`). `queue` ships fully implemented + conformance-tested
but unconsumed (its consumer is the deferred concurrency work).

**C -- Cross-platform transcripts**
- `bot.transcripts` (`append` / `list` / `delete`) over `state.list`,
keyed by a developer-chosen `identity` resolver (paired-or-throw with
`transcripts` config), `retention` + `maxPerUser`, GDPR delete. Each
entry records its `platform` + `ts`, so an agent can reference "you said
on WhatsApp X" across surfaces.
- **`thread.runAgent({ transcript: true })`** auto-bridges the
transcript into a run: injects prior cross-platform history
(platform-labeled) as agent context, appends the current user turn, and
captures the streamed assistant reply back into the transcript. Manual
`append`/`list` remain available for full control.

**Surface plumbing**
- `eventId?` on the three ingress types + Slack adapter wiring
(Events-envelope `event_id` -> `client_msg_id` -> `channel:ts`;
interaction/command ids) so the dedup guard engages on Slack.

## Testing
- Shared `runStateStoreConformance` suite (kv/list/lock/dedup/queue
contract incl. stale-token fencing, TTL expiry, queue eviction,
**mixed-ttl whole-list expiry**, **kv/lock keyspace isolation**) runs
against all three backends. **Redis and Postgres validated live (16/16
each)** against local containers; MemoryStore runs in CI.
- Focused unit tests for `kvActionStore`, `Transcripts`,
`createStateBackedConversationStore`, thread subscribe/state, the
`createBot` orchestration (lock-release-on-throw, `onLockConflict`
callback, identity throws/null, dedup+lock ordering, dedup fail-open),
and Slack `eventId` derivation.

## Code review
Reviewed by a multi-agent pass (correctness / silent-failures / tests /
types / comments / backends / packaging) and **fixed to zero**. Notable
fixes: a Postgres data-retention bug (non-TTL append leaving immortal
rows -- now whole-list expiry), a Redis kv/lock keyspace collision
(added `lock:` infix), fail-open-but-now-fail-**loud** logging on
identity/dedup failures, eager `retention` validation, and required
`ThreadDeps.state`.

## Follow-ups (not in this PR)
- **CI: provision Redis + Postgres services and set `TEST_REDIS_URL` /
`TEST_POSTGRES_URL`** so the backend conformance suites run in CI
(currently `describe.skip` without them -- validated locally here).
- Discord/Telegram `eventId` wiring (Slack done; `// TODO` left in
place).
- Slice **B** (concurrency strategies, serverless webhook ingress) and
slice **D** (reactions/ephemeral/modals).
- `IntelligenceStore` managed backend (the product layer).

## Notes
- Dedup records on **receipt** (drops duplicate deliveries within the
TTL window); a handler that throws still leaves its event marked seen
(no retry-recovery) -- intentional, documented inline.
- **Squash-merge friendly** -- intra-branch commits are individual green
slices from a gated wave execution (one commit, `d73732cf14`, absorbed a
few wave-3 files via the auto-staging pre-commit hook).
- Pre-existing failures unrelated to this branch: strict `check-types`
in `src/threads.ts` / `src/utils/phoenix-observable.ts` /
`@copilotkit/core`; eslint warnings on Slack `chat.postMessage` and the
`pg` named export.

 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-23 18:58:01 +02:00
Alem Tuzlak 5f2d7102d5 docs(bot): StateStore, persistence, and transcripts reference + guides 2026-06-23 18:38:47 +02:00
Alem Tuzlak 5ecdee36b8 feat(bot): pluggable StateStore persistence + cross-platform transcripts
Adds a durable persistence layer for @copilotkit/bot, replacing the
in-memory-only ActionStore with a pluggable StateStore.

- StateStore interface (kv/list/lock/dedup/queue) with a shared
  conformance suite; MemoryStore default plus @copilotkit/bot-store-redis
  and @copilotkit/bot-store-postgres backends.
- createBot({ store }): typed per-thread state via Standard Schema,
  action snapshots persisted through the store, per-conversation turn
  lock (onLockConflict drop|force), and inbound-event dedup keyed on a
  stable eventId. ActionStore is kept as a deprecated alias.
- Cross-platform transcripts (bot.transcripts + identity resolver) with
  age-bounded retention (prune on append + filter on read), and
  runAgent({ transcript: true }) to auto-inject history and capture the
  reply.
- createBot({ components }) re-registers components so durable actions
  re-fire after a restart; restart-durability demo in examples/slack.
- Dedup is marked seen only after the turn lock is acquired, so a turn
  dropped on lock-conflict does not burn its eventId (no lost retries).
- Release lockstep: bot-store-redis/postgres version with bot + bot-ui.
2026-06-23 18:33:38 +02:00
Jordan Ritter 43c5220398 feat(showcase): whitelist promote parity policy — drop env-key-set-diff, add advisory tier
Drop the env-var key-set-diff WARN in check_p6_parity (environment-specific keys
like NODE_ENV/CVDIAG_*/SHOWCASE_BACKEND_HOST_PATTERN/BROWSER_POOL_* no longer
block a promote) and introduce a non-blocking ADVISORY disposition demoting
region/replicas/restartPolicy, missing-expected-prod-domains, and concurrency-key
divergence from WARN. check_critical_env_key_parity stays staging-gated
((CRITICAL_ENV_KEYS & staging_keys) - prod_keys): it flags a critical key staging
carries but prod is missing, and tolerates infra/operator tokens absent from both.

Adds red-green spec coverage in test_promote_p6.rb. No source symbols renamed/removed.
2026-06-23 01:56:19 -07:00
Ran Shemtov e0d6eeaebe Merge branch 'main' into claude/trusting-babbage-f4d48a 2026-06-23 09:21:35 +02:00
Jordan Ritter f1b2595dcd feat(showcase): on-demand prod-vs-staging reconcile tool (reconcile-prod CLI + manual workflow) (#5623)
## Summary

An **on-demand** tool to answer "is prod caught up with staging right
now?". The showcase deploy model is **staging = mutable `:latest`**
(continuously rebuilt) and **prod = immutable `@sha256:`** (advances
only on an explicit promote), so a prod column can sit **behind** a
green staging.

**There is no scheduled drift alert — by design.** Prod lagging staging
is **often intentional**: changes are batched and promoted deliberately,
so a recurring "N columns stale" alert would be pure noise. This tool is
therefore manual-only: a maintainer runs it when they want to check, and
it tells them the current state.

- **`bin/railway reconcile-prod`** — for every prod-eligible
(`probe.prod == true`) service, compares the **prod serving digest**
(the `@sha256:` from
`SnapshotCommand.build_snapshot(PRODUCTION_ENV_ID)`) against the
**staging running digest** (reuses
`PromoteCommand#staging_running_digest`, the same source the promote pin
uses). Classifies each:
  - `green` — prod == staging (in sync)
- `stale` — prod != staging **and** staging is resolvable (prod is
behind a green staging)
- `gray` — staging running digest not resolvable, or the service has no
prod snapshot entry yet — informational, **not** stale
- Prints a readable per-service table + summary; **exits nonzero iff any
service is stale**; `--json` for machine output. **Read-only — no
promotes/mutations.**
- **`showcase/scripts/reconcile-prod-gate.sh`** — wrapper mirroring
`lint-prod-gate.sh`: surfaces the table into `$GITHUB_STEP_SUMMARY`,
optionally captures `--json` to `RECONCILE_JSON`, and propagates the
exit-code verdict (never swallows a non-zero).
- **`.github/workflows/showcase_reconcile.yml`** — **`workflow_dispatch`
only** (no cron). Regenerates the SSOT JSON (`EMIT_SKIP_OXFMT=1`, same
as the promote workflow's resolve/promote jobs), runs the gate with the
Railway/GHCR auth env, renders the reconcile table to the **GH step
summary**, and uploads the `--json` as a `reconcile-json` artifact. **No
Slack.** The run exits nonzero on a stale column so a manual run visibly
flags drift. `timeout-minutes: 10`.
- **Tests** — Ruby minitest (`test_reconcile_prod.rb`: classification +
exit-code + `--json` shape + dispatcher registration) and a bats gate
test (`reconcile-prod-gate.bats`). Wired the gate script into the
`showcase_validate.yml` shellcheck list.

### What changed from the original scheduled-alert design

The first cut of this PR shipped a daily cron + auto-post to #oss-alerts
on any stale column. Per owner feedback, that was reshaped to on-demand
only: the `schedule:` trigger and the Slack-on-stale step were removed
(intentional/deliberate staleness is not a bug, so an unsolicited
recurring alert is noise). The CLI command, the gate wrapper, and all
tests are unchanged.

## Gates

- `ruby showcase/bin/spec/test_reconcile_prod.rb` → **9 runs, 20
assertions, 0 failures**
- `bats showcase/scripts/__tests__/reconcile-prod-gate.bats` → **6 ok**
- `shellcheck -s bash showcase/scripts/reconcile-prod-gate.sh` →
**clean**
- `actionlint .github/workflows/showcase_reconcile.yml` → **clean** (the
pre-existing `depot-ubuntu-24.04-4` custom-runner-label warning is on
`showcase_validate.yml`, predates this PR — my only change there is one
line in the shellcheck list)

## Test plan

- [ ] CI green (Ruby suite, bats suite, actionlint/shellcheck,
commitlint)
- [ ] Optional: read-only `workflow_dispatch` run of
`showcase_reconcile.yml` to confirm it runs against live prod/staging
(safe — no mutations)
2026-06-22 23:08:17 -07:00
Jordan Ritter 9cb9260dca fix(showcase): preserve reconcile gate stderr diagnostics, skip blank artifact 2026-06-22 23:00:08 -07:00
Tyler Slaton b77d2710aa feat(bot-whatsapp): WhatsApp Cloud API platform adapter + example + docs (#5449)
## What

Adds **`@copilotkit/bot-whatsapp`** — a WhatsApp Business **Cloud API**
`PlatformAdapter` for the platform-agnostic `@copilotkit/bot` engine —
plus a runnable **`examples/whatsapp`** app and docs. This brings
WhatsApp to the bots ecosystem alongside the existing Slack support,
reusing the engine, the `@copilotkit/bot-ui` IR, and the pluggable
`ActionStore` untouched.

## How it works

- **Ingress:** the adapter owns its own HTTP server — GET verification
handshake (`hub.challenge`) + POST intake validated by
`X-Hub-Signature-256` HMAC (timing-safe), acked `200` immediately then
processed async.
- **No streaming:** WhatsApp messages are immutable, so the run renderer
**buffers** text and sends once on `TEXT_MESSAGE_END`
(`supportsStreaming: false`; `update()` posts fresh, `delete()` no-ops).
- **Interactive mapping:** text/section → text; ≤3 buttons →
reply-button message; `Select` or 4–10 actions → list message; >10 →
numbered-text fallback. A control's `value` round-trips by encoding it
into the reply id (`ck:…::<json>`), since WhatsApp replies carry no
value field; oversized encodings fail loud rather than corrupt silently.
- **Memory:** WhatsApp exposes no readable history, so a pluggable
**`HistoryStore`** (default `InMemoryHistoryStore`) holds it and replays
it into `agent.messages` each turn (fresh threadId per turn, mirroring
`bot-slack`). Swap in a durable backend to persist across restarts.
- **Commands:** leading-keyword matching (`commandPrefix`, default `/`);
the command text is injected via the engine's `runAgent({ prompt })`
path (not persisted at ingress).
- **Inbound media** → AG-UI multimodal content parts; **HITL** via
interactive replies.

## Example

`examples/whatsapp` mirrors `examples/slack`: a CopilotKit
`BuiltInAgent` over MCP (Linear + Notion), with `issue_list`, an
interactive `show_incident`, and a `confirm_write` HITL gate.

## Tests & verification

- 62 unit tests across the package (render mapping, markdown→WhatsApp,
signature verification incl. wrong-but-equal-length, interaction
decode/round-trip, buffered renderer, webhook listener/server, stores,
media, adapter).
- `build` ✅, package `check-types` ✅, `publint`/`attw` (ESM-only) ✅,
example `check-types` ✅. Full `nx run-many -t test
--projects=packages/**` passes.
- Two rounds of code review (APPROVE) — fixed slash-command history
double-append and silent value-truncation; minors (HMAC over raw bytes,
conversationKey invariant, offset-correct Blob, unused-dep pruning,
added tests).

## Docs

Package `README.md` + `ARCHITECTURE.md`, example setup guide (Meta app +
webhook + tunnel), and a `shell-docs` WhatsApp guide page (registered in
`meta.json` + early-access gate).

## Notes / out of scope (v1)

- No template-send path for messaging outside WhatsApp's 24-hour
customer-service window (documented limitation).
- Pre-existing, unrelated `@copilotkit/core` `phoenix-observable.ts`
typecheck error exists on the branch base (missing `@types/phoenix`) —
not introduced here.
2026-06-22 20:34:54 -07:00
Jordan Ritter d10aa858ec feat(harness): complete cvdiag observability on the d5/d6 probe path (#5625)
## What
Completes the cvdiag flap-observability instrument so **d5/d6/e2e probe
runs are readable from `cvdiag_events`** — previously only the d4 driver
emitted probe-layer rows, leaving the d5/d6 path (where the dashboard
flaps live) a blind spot.

- Extract `CvdiagProbeSession` (+ event shapes, constants,
`turnCompleteReason`) from `d4-chat-roundtrip.ts` into shared
`cvdiag/probe-session.ts`; d4 re-imports (behavior-preserving).
- Wire `CvdiagProbeSession` into `d6-all-pills.ts::runFeature`: emit
`probe.start/navigate.complete/message.send/firstToken` + an
exactly-once guarded `probe.exit` carrying `terminal_outcome` +
`failure_classifier`, and close the probe↔backend `test_id` join
(X-Test-Id already injected).
- Thread `cvdiagPbWriter` through `orchestrator.ts` (fleet worker) +
`cli/runner.ts` (`--live`) so probe events persist to `cvdiag_events`.

## Why
The d5/d6 probe path emitted **zero `probe.exit` rows** (verified: 2
days of cvdiag retention = 100% `d4-`, zero d5), so a flapping d5 cell
(e.g. `dom-missing` reds) could not be diagnosed from staging data. This
closes that gap.

## Verification
- Local red-green on the real surface (`langgraph-typescript --d5`):
`probe.exit` rows **0 → 38** (terminal_outcome + failure_classifier
populated).
- Full harness suite: **3167/3167 passing**; tsc/oxlint/oxfmt/build
clean.

## Code review (2-round cr-loop, 7 agents/round, converged + Procedure 3
clean)
- **A1** — `messageSend` char-count ran unguarded in the probe path; a
non-string input could throw and red a green probe. Fixed: computation
moved inside the `if (cvdiag)` guard + non-string coerced. (red-green)
- **A2** — `parseFailureClassifier` used a stale hardcoded 4-member
allow-list omitting `selector-mismatch`; reconciled both membership
checks to the canonical `FAILURE_CLASSIFIER_SET` (derived from
`CVDIAG_FAILURE_CLASSIFIERS`). (red-green)

## Follow-ups (not in this PR)
- `--headed` CLI launcher lacks the `goto`-wrap that installs the SSE
interceptor, so cvdiag (newly wired into `--headed`) mislabels
`failure_classifier` there; the supported fleet/`--direct` path is
correct. (dev-path quality)
- d6 feature-concurrency: 4 concurrent features share `test_id`
`d6-<slug>-<runId>`; per-feature reads work via `demo=featureType`, but
the §5 sequence_num join + replay buffer would benefit from per-feature
keying.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-22 17:14:32 -07:00
Jordan Ritter 4d953e5ed2 fix(showcase): standalone promote services (no deps, never gated) — flag docs
A `shell-docs` promote expanded its dependency closure to include the
always-on Tier-1 verification set (harness, dashboard, …); harness then
tripped a P6 env-divergence WARN-refusal and the tier barrier gated docs
as NOT-ATTEMPTED. docs has no runtime dependency on the control plane.

Introduce a declarative `standalone` service class: a leaf that neither
depends on anything nor gates on anything. computePromoteClosure (TS) and
the resolve-promote-targets jq both skip the Tier-1 union when the request
is entirely standalone, so a standalone request promotes ONLY itself; the
fleet runner promotes `s:`-marked services ungated (never NOT-ATTEMPTED on
an unrelated failure, never gating a tier). Flag `docs` standalone.

Red-green: railway-envs.test.ts (4), resolve-promote-targets.bats (5),
promote-fleet.bats (3) fail on origin/main (docs closure pulls harness;
`s:` is an unknown tier) and pass after the fix.
2026-06-22 17:05:25 -07:00
Jordan Ritter a618bdce35 feat(harness): emit cvdiag probe-session boundaries from the d5/d6 probe path
- open a per-feature CvdiagProbeSession for each d5/d6 pill probe
- emit exactly-once probe.exit and failure_classifier per session
- join probe-session output to its run via the X-Test-Id header
- thread cvdiagPbWriter through the orchestrator and CLI runner

Call-Site Enumeration: FAILURE_CLASSIFIER_SET is exported from
cvdiag/probe-session and consumed by d6-all-pills (classifier validation
against the canonical set). The export has no other call sites; any future
classifier addition must update the canonical set in probe-session and the
validation in d6-all-pills together.
2026-06-22 16:42:43 -07:00
Jordan Ritter 6929ad1c33 refactor(harness): extract CvdiagProbeSession into shared cvdiag/probe-session
Behavior-preserving extraction of the CvdiagProbeSession lifecycle from the
d4 chat-roundtrip driver into a shared cvdiag/probe-session module, so the
d5/d6 probe path can reuse the same session boundaries. d4-chat-roundtrip
now imports the extracted session instead of defining it inline.
2026-06-22 16:42:31 -07:00
Sam Julien b9b08381ba docs: backport Threads guide to authored integrations (#5620)
## Summary

- Backports the generated/root Threads guide content into the shared
authored Threads snippet.
- Adds the CLI “Choose your starting point” path, manual path, thread
lock options, Enterprise Intelligence CTA, and corrected next-step links
to authored Threads docs.
- Standardizes authored integration Threads pages to explicitly import
the shared snippet with `components={props.components}` so authored
routes stay aligned.

## Authored routes covered

- AG2
- Agno
- AWS Strands
- Built-in Agent
- CrewAI Flows
- LangGraph
- LlamaIndex
- Mastra
- Microsoft Agent Framework
- PydanticAI

## Validation

- `npm run pretypecheck` in `showcase/shell-docs`
- `npm run lint` in `showcase/shell-docs` (passes with existing
warnings)
- `npm run test` in `showcase/shell-docs`
- `npm run typecheck` in `showcase/shell-docs`
- `npm run build` in `showcase/shell-docs` (passes with existing
Next/Turbopack warnings)
- `git diff --check`
- Manual MDX link sweep for changed docs links (`/premium/self-hosting`,
`/premium/threads-explained`, `/reference/hooks/useThreads`, and
`http://localhost:3000`)

## Formatter note

- `pnpm run check-format` currently fails on unrelated existing files
under `examples/showcases/arcade-tools/*`,
`examples/v2/react/demo/tsconfig.json`, `migrations.json`, and
`nx.json`.
- Scoped `oxfmt --check` does not treat the changed MDX files as target
files, so there is no formatter-owned MDX change to apply here.
2026-06-22 15:47:07 -07:00
Sam Julien 2fe292e4c7 docs(shell-docs): clarify tailored content selectors (#5621)
## Summary
- restyle TailoredContent selectors as native segmented path buttons
- add a selected check indicator plus hover/focus affordances
- add a rendering test for button semantics and selected state

## Verification
- pnpm exec oxfmt --check
showcase/shell-docs/src/components/react/tailored-content.tsx
showcase/shell-docs/src/components/react/__tests__/tailored-content.test.tsx
- npm run lint (showcase/shell-docs; existing warnings only)
- npm run typecheck (showcase/shell-docs)
- npm run test (showcase/shell-docs)
- npm run build (showcase/shell-docs; existing Turbopack NFT trace
warning)

## Notes
- Repo-wide pnpm check-format currently fails on unrelated pre-existing
files outside this PR; the touched files pass scoped oxfmt check.
2026-06-22 15:46:53 -07:00
Jordan Ritter c2c19c0853 perf(showcase): within-tier parallel fan-out for promote-fleet so service=all fits the timeout
promote-fleet.sh fans out promote_one within a tier up to PROMOTE_FANOUT
(default 5) via a bash-3.2-safe PID-array bounded launcher (plain `wait`, no
`wait -n`/`declare -n`); per-service results to temp files + reap_tier
repatriates (subshell-safe); tier boundaries are hard barriers (cross-tier
serial); showcase_promote.yml wires CLOSURE_PLAN + bumps timeout-minutes 20->35;
adds bats fan-out tests (deterministic rendezvous-barrier concurrency proof;
present-but-empty .rc treated as failed; tier-barrier boundary-inclusive `>=`);
SC2317/SC2329 shellcheck disable for ubuntu-24.04 0.9.0.
2026-06-22 14:50:16 -07:00
Jordan Ritter cd3844da65 refactor(showcase): make prod-vs-staging reconcile on-demand only (no cron, no Slack)
Prod sitting behind staging is often intentional (changes are batched and
promoted deliberately), so a recurring drift alert is noise. Reshape the
reconcile workflow to manual-only:

- Remove the daily `schedule:` cron trigger — leave only `workflow_dispatch`.
- Remove the auto-Slack-on-stale step (and its SLACK_WEBHOOK env / stale_line
  output derivation) — no unsolicited #oss-alerts post on mere staleness.
- A manual run surfaces the reconcile table to the GH step summary, keeps the
  cheap `--json` capture as an uploaded artifact, and still exits nonzero on a
  stale column so a manual run visibly flags drift.
- De-noise the gate script + bats comments that referenced the removed
  scheduled/Slack behavior.

The on-demand CLI (`bin/railway reconcile-prod`), the gate wrapper, and the
Ruby + bats tests are unchanged.
2026-06-22 14:20:02 -07:00
Jordan Ritter 90bcd66b09 feat(showcase): detect prod columns stale vs green staging (reconcile-prod drift gate)
Lever 1 of the promote-reliability hardening plan. The showcase deploy
model is staging=mutable :latest (continuously rebuilt), prod=immutable
@sha256: (advances only on explicit promote), so a prod column can
silently fall BEHIND a green staging — drift today is only noticed by
eyeballing a dead column. This adds proactive, automatic detection.

- bin/railway reconcile-prod: for every prod-eligible (probe.prod==true)
  service, compares the prod SERVING digest (LintProd snapshot path) vs
  the staging RUNNING digest (reuses PromoteCommand#staging_running_digest).
  Classifies green/stale/gray, prints a table + summary, exits 1 iff any
  stale. --json for machine output. Read-only: no promotes/mutations.
- scripts/reconcile-prod-gate.sh: wrapper mirroring lint-prod-gate.sh —
  surfaces the table to the GH step summary, captures JSON for the Slack
  builder, propagates the exit-code verdict.
- .github/workflows/showcase_reconcile.yml: daily cron + workflow_dispatch;
  runs the gate; on stale services posts the stale-column list to
  #oss-alerts (SLACK_WEBHOOK_OSS_ALERTS) via the fromJSON('"\n"') idiom.
- Tests: Ruby minitest (classification + exit-code, RED-anchored on a
  drift-blind classifier) and a bats gate test. Wired the gate script
  into the showcase_validate.yml shellcheck list.

Post-promote convergence verification is deferred to a fast-follow.
2026-06-22 14:14:18 -07:00