Commit Graph

12273 Commits

Author SHA1 Message Date
Tyler Slaton abaa3d1cd8 fix(bot-intelligence): preserve managed command reaction deliveries 2026-07-02 21:30:04 -07:00
Alem Tuzlak b840422fa1 feat(bot-intelligence): managed-bot file support (inbound content parts + outbound postFile)
Wire multimodal files through the managed adapter both directions:

- Inbound: carry turn-input file refs on the ingress envelope, fetch each
  file's bytes via GET /api/bots/files/:handle, and build AgentContentPart[]
  (image/audio/video/pdf -> data part; text/* inline; else a text note) onto
  the turn's contentParts so the agent prompt sees uploads.
- Outbound: implement PlatformAdapter.postFile -> stream bytes to the
  per-delivery upload endpoint, then emit a new 'file' render frame carrying
  the storage handle (the Connector Outbox does the Slack upload).
- DeliverySource gains optional fetchFile/uploadFile; HttpDeliverySource
  implements both. InMemoryDeliverySource gains a files map + fetchFile for
  tests. Adds the 'file' HostedBotRenderEvent kind.
2026-07-02 21:14:45 +02:00
Alem Tuzlak 9cb929b40d feat(runtime): opt-in supersede for concurrent same-thread runs
InMemoryAgentRunner gains an opt-in { onConcurrentRun: "throw" | "supersede" }
(default "throw", so existing consumers are unchanged). In "supersede" mode a
new run for an already-running thread aborts the prior run (agent.abortRun(),
mirroring stop()) instead of throwing "Thread already running" — so a fast
follow-up turn, or one after a dropped/aborted run, cleanly replaces the
previous one.

Superseding overlaps two runs on the module-global per-thread store, so all
four finalization sites (both resets and both historicRuns.push) are guarded on
store.currentRunId === request.input.runId and stamp the run's own id. A
superseded run therefore drops its partial events rather than resetting or
mislabeling the new run's state/history.

Used by the Intelligence-hosted (managed) Slack listener. (OSS-417)
2026-07-01 19:10:54 +02:00
Alem Tuzlak 13248dda0b fix(bot): drop empty text deltas, tool status on all thread anchors, bound managed turns
Found while live-testing the Intelligence-hosted (managed) Slack bot:

- bot-intelligence: skip empty text_delta frames. The leading role-announcement
  chunk was emitted as delta:"", which the frozen render contract rejects
  (min 1 char), aborting every managed run.
- bot-slack: drive assistant.threads.setStatus ("is using `tool`…") on any
  thread anchor, not just assistant panes; keep in-message task_update rows.
- bot-intelligence: bound each managed turn in the HTTP delivery loop with
  turnTimeoutMs (default 120s) -> nack on timeout/throw, so a hung turn (a HITL
  approval that never arrives, a half-open stream) can't wedge the
  single-delivery listener. (OSS-421)
2026-07-01 19:10:14 +02:00
Alem Tuzlak 7fe51c49ae feat(bot-intelligence): render discrete thread.post as Block Kit, not flattened text
adapter.post/update now stream post/update render frames carrying the IR (via a
shared per-turn seq with the run renderer) when a render sink is wired, so the
Connector Outbox renders full Block Kit — rich JSX (sections, buttons, images)
from thread.post is preserved instead of irToText-flattened. Falls back to the
egress op path when no render sink (in-memory tests). 46 tests pass.
2026-07-01 14:46:29 +02:00
Alem Tuzlak 5ef2f6ce0a feat(bot-intelligence): HTTP render-frame streaming + live Phoenix channel (OSS-402)
Complete the SDK realtime path so managed replies reach 1:1 bot-slack UX end to
end, over either transport:

- HttpRenderEventSink streams render frames to app-api's durable accept route
  (/api/bots/deliveries/:id/render-events/accept), reusing HttpDeliverySource's
  per-delivery scope (now carried on the claim response). This lets the whole
  loop run over HTTP against app-api directly — no realtime gateway required —
  which the Connector Outbox then renders to Slack.
- intelligenceAdapter defaults its render sink to the HTTP one when it builds
  the HTTP transports; injected in-memory sources still use the egress-backed
  fallback.
- connectPhoenixHostedBotChannel wires the real Phoenix Socket/Channel behind
  the HostedBotChannel contract for the production realtime-gateway path
  (joins hosted_bots:project:{id}, push→ok/error, on-push), backed by phoenix@1.8.

Tests: 45 bot-intelligence pass (HttpRenderEventSink scope-echo + no-scope guard);
check-types green.
2026-07-01 14:05:45 +02:00
Alem Tuzlak 5d4b7ff393 feat(bot-intelligence): stream RenderEvents over the realtime gateway (OSS-402)
The managed run renderer now mints semantic render frames (run_started /
text_delta / text_end / tool_start / tool_end / interrupt / run_error /
finalize) and streams them through a RenderEventSink, assigning a monotonic
seq per (turnId, slot) and awaiting a durable render_accepted receipt for each
via a serial chain (seq fixed at enqueue so order holds regardless of AG-UI
callback scheduling). Add PhoenixRealtimeTransport (DeliverySource +
RenderEventSink) speaking the frozen OSS-395 contract: render_event frames ->
render_accepted receipts, then delivery.complete_requested (completion INTENT
with acceptedThrough high-water pointers) — never a committed delivery.ack
(app-api owns ack), and no Slack credentials in the SDK. When no realtime sink
is wired the renderer falls back to translating frames into post ops on the
EgressSink so the HTTP demo keeps posting plain-text replies.
2026-07-01 12:29:50 +02:00
Alem Tuzlak 409935d4a1 feat(bot-slack): extract Bolt-free run renderer behind ./render (OSS-403)
Replace the WebClient dependency in createRunRenderer with an injected
SlackRenderTransport (setStatus / postMessage / updateMessage) — the 4 native
streaming ops were already injected via NativeStreamTransport. Add a ./render
subpath export exposing the transport-agnostic renderer plus the pure IR->Block
Kit / modal / mrkdwn helpers, so the gateway-side Connector Outbox can drive the
identical renderer for managed replies (1:1 UX parity) without forking it. The
native Slack adapter wraps its WebClient into the transport; behavior unchanged.
2026-07-01 12:29:14 +02:00
Alem Tuzlak a03cd7db8b test(runtime): supply bots:[] in the intelligence runtime-like factory
CopilotIntelligenceRuntimeLike.bots became required with the managed-bots
runtime work (OSS-360); the get-runtime-info test factory still omitted it,
failing check-types. Add the missing field.
2026-07-01 12:28:20 +02:00
Alem Tuzlak 0641b63f67 fix(bot-intelligence): make intelligenceAdapter() callable with zero args
The opts parameter (not just its fields) must be optional so consumers can
write createBot({ adapters: [intelligenceAdapter()] }). Add a zero-arg
regression test.
2026-06-30 12:43:20 +02:00
Alem Tuzlak c790a3bc4d feat(bot-intelligence): config-free intelligenceAdapter with default HTTP transports
Make source/egress optional on intelligenceAdapter(); when omitted it builds
default HTTP transports that talk to a running Intelligence app-api
(heartbeat/claim/lease-fenced ack-fail + idempotent egress), resolved from env
(COPILOTKIT_INTELLIGENCE_URL/COPILOTKIT_API_KEY) and the bot's name. Consumers
now write createBot({ adapters: [intelligenceAdapter()] }).

- http-transports.ts: HttpDeliverySource, HttpEgressSink, resolveTransportConfig
- ir-to-text.ts: flatten BotNode[] -> plain text for the egress first slice
- bot core: PlatformAdapter.start(sink, ctx?) carries botName from createBot
  (additive, backwards-compatible); export transports as undocumented fallbacks
- tests: bot-intelligence 36 green, bot 150 green

Relates to OSS-360, OSS-361, OSS-363.
2026-06-30 12:22:41 +02:00
Alem Tuzlak 028542adc5 refactor(bot): dedup ingress event interfaces via IngressEventBase + IngressIds 2026-06-29 16:22:39 +02:00
Alem Tuzlak bab850e357 fix(bot-intelligence): guard resolveActivationEnv against a missing process global 2026-06-29 16:20:58 +02:00
Alem Tuzlak 818684ae70 refactor(bot-slack): point-free renderEgress/normalizeIngress in slackCodec 2026-06-29 16:13:22 +02:00
Alem Tuzlak 57a684c6a5 feat(bot-intelligence): gather activation env + per-bot commands in metadata (OSS-360) 2026-06-29 14:47:36 +02:00
Alem Tuzlak 6d038d436e feat(bot-slack): extract pure Slack ingress normalization for shared reuse (OSS-363) 2026-06-29 14:41:16 +02:00
Alem Tuzlak 972dd64476 refactor(bot): split the Intelligence managed adapter into @copilotkit/bot-intelligence
Move the Intelligence-delivered managed-bot surface out of @copilotkit/bot into
its own package so the adapter, transports, contracts, and lifecycle ship
independently of bot core.

- New @copilotkit/bot-intelligence: intelligenceAdapter + DeliverySource/EgressSink
  (+ in-memory impls) + placeholder contracts + startManagedBots/validation/
  activation metadata. Production code imports only types from @copilotkit/bot
  and @copilotkit/bot-ui.
- @copilotkit/bot keeps the generic PlatformCodec seam (moved to src/codec.ts) and
  all core createBot changes (addAdapter, deferred store, id fields,
  __managed/skipIngressDedup, exclusive guard). It now also exports the
  FakeAdapter/FakeAgent test utilities for downstream adapter-package tests.
- Registered the new release scope: release.config.json, scripts/release/lib/
  config.ts, and the canary/publish/stable release workflows.

Tests preserved: bot 150 + bot-intelligence 19 (= the prior 169); bot-slack 261.
Builds typecheck across bot/bot-intelligence/bot-slack/runtime; publint/attw/
oxlint/oxfmt clean.
2026-06-29 14:30:28 +02:00
Alem Tuzlak 49767661ba feat(bot-slack): expose slackCodec via ./codec subpath for Bolt-free import 2026-06-29 14:03:32 +02:00
Alem Tuzlak 4d1ffc2199 Merge remote-tracking branch 'origin/main' into alem/oss-360-sdk-foundations
# Conflicts:
#	packages/bot/src/create-bot.ts
2026-06-29 14:01:14 +02:00
Alem Tuzlak c55aa134be docs(runtime): clarify managed bots are validated at activation, not construction 2026-06-29 12:05:34 +02:00
Alem Tuzlak 6f5d1406eb feat(bot,bot-slack): PlatformCodec seam + Slack egress codec (OSS-363)
A shared, pure, per-platform codec so platform rendering semantics live in one
place instead of being duplicated between the local adapter and the
Intelligence/Connector-Outbox side.

- PlatformCodec interface in @copilotkit/bot (renderEgress: IR -> native; pure,
  no transport, no credentials)
- slackCodec in @copilotkit/bot-slack composing the existing pure
  renderSlackMessage, importable by the Connector Outbox without Bolt or creds

Ingress normalization is marked TODO(OSS-363): extracting the pure Slack event
-> neutral mapping from the Bolt listener (so the local adapter and Intelligence
webhook ingress share it) is the remaining, higher-risk half of that ticket.
2026-06-29 11:59:15 +02:00
Alem Tuzlak 9046b241e6 feat(runtime): typed managed-bots option on the Intelligence runtime (OSS-360)
Expose new CopilotRuntime({ intelligence, bots }) -- the Mode B entry point
for managed bots:

- bots is accepted only on the Intelligence runtime variant (bots?: undefined
  on the SSE variant), so TypeScript rejects bots without intelligence
- CopilotIntelligenceRuntime stores the declared bots; the facade exposes them
  via the existing isIntelligenceRuntime getter pattern
- @copilotkit/bot is imported type-only (it is pure-ESM; a value import would
  break this package's CJS output). Name validation + transport wiring happen
  in startManagedBots (called by the managed-listener bootstrap), not here.

Adds a type-only @copilotkit/bot workspace dependency.
2026-06-29 11:50:41 +02:00
Alem Tuzlak 35390eafa2 feat(bot): managed runtime API — bot name, validation, activation, lifecycle
The SDK slice of OSS-360, fully unit-tested with in-memory transports:

- createBot({ name }): project-unique id exposed as bot.name (optional for
  local bots; required + validated for managed)
- managed/runtime.ts:
  - assertValidBotNames: required, identifier-style, unique-per-runtime; fails loud
  - buildActivationMetadata: declared bot names + runtime env/versions (pure)
  - startManagedBots: validate -> attach intelligenceAdapter per bot (wired to its
    resolved transport) -> start; returns metadata + stop handle
- intelligenceAdapter now accepts a `store` it exposes as `stateStore`, so the
  Intelligence-backed backend flows through createBot's resolveBackend

Transports are injected by the caller (closed Gateway/Outbox in production,
in-memory in tests); this module owns no Slack creds, ingress, or outbox.
2026-06-29 11:07:53 +02:00
Alem Tuzlak fda5b8ed11 feat(bot): route all managed ingress kinds through the bridge adapter
Widen the placeholder ManagedIngressEnvelope to a discriminated union and route
every kind to the matching bot-core sink call, so managed bots reuse normal
routing for messages, commands, interactions, thread-starts, and reactions
(OSS-361) — not just plain turns.

- contracts: ManagedIngressBase + union over turn|command|interaction|
  thread_started|reaction
- IntelligenceAdapter.dispatchTo: switch on kind -> onTurn / onCommand /
  onInteraction / onThreadStarted / onReaction
- tests cover each kind dispatching to its handler and emitting egress
2026-06-29 11:00:26 +02:00
Ran Shemtov c1d5764fde docs(showcase): A2UI catalog auto-inject + manual opt-out for generated frameworks (#5725) 2026-06-29 10:08:06 +02:00
Ran Shemtov 74b041b29f Merge branch 'main' into claude/nervous-bardeen-37a398 2026-06-29 10:07:45 +02:00
Jordan Ritter 3d3c70e026 fix(showcase/railway): promote pins replicas via real ServiceInstanceUpdateInput shape (drop nonexistent ServiceMultiRegionConfigInput type) (#5756)
## The bug

`bin/railway` promote re-asserts the SSOT replica config on the
`serviceInstanceUpdate` pin, but built the mutation by declaring a
standalone variable:

```graphql
$multiRegionConfig: ServiceMultiRegionConfigInput!
```

**That input type does not exist in Railway's schema.** The prior
promote attempt (#5754) failed live with `HTTP 400: Unknown type
"ServiceMultiRegionConfigInput"`, so the redeploy went out **without**
the replica config and Railway de-scaled `harness-workers` (us-west2)
from the SSOT-intended **6** replicas to **1**. The wrong type name was
a guess, never verified against the real schema.

## Real schema — source of truth

Railway's mutation is `serviceInstanceUpdate(serviceId: String!,
environmentId: String!, input: ServiceInstanceUpdateInput!)`. The
correct pattern — already used by this repo's own working TypeScript
provisioners — is to pass the **entire input object as one `$input:
ServiceInstanceUpdateInput!` variable** and put every field (`source`,
`healthcheckPath`, `region`, `registryCredentials`, `multiRegionConfig`)
as a **nested key** of that input. Railway resolves each nested field's
type from the `ServiceInstanceUpdateInput` schema, so you never name
nested input types yourself:

- `showcase/scripts/deploy-to-railway.ts` ~472-509 — builds
`instanceInput = { region, healthcheckPath, registryCredentials, ... }`,
then `serviceInstanceUpdate($serviceId, $environmentId, $input:
ServiceInstanceUpdateInput!)` passing `input: instanceInput`.
- `showcase/scripts/provision-starter-fleet.ts` ~575-602 — same
single-`$input` pattern.

The replica shape itself (`multiRegionConfig.<region>.numReplicas`, e.g.
`{ "us-west2": { "numReplicas": 6 } }`) is confirmed in
`showcase/scripts/railway-envs.ts` ~150-194, ~719-752 and
`railway-envs.generated.json` ~161-169 — and was already correct in the
Ruby. **Only the GraphQL type declaration was wrong.**

## The fix

`showcase/bin/railway` — `RestoreCommand.build_update_image_mutation`
(~859):

**Before** — one typed variable per optional key (and the invented
type):

```graphql
mutation UpdateImage($serviceId: String!, $envId: String!, $image: String!,
                     $healthcheckPath: String!,
                     $multiRegionConfig: ServiceMultiRegionConfigInput!) {
  serviceInstanceUpdate(serviceId: $serviceId, environmentId: $envId,
    input: { source: { image: $image }, healthcheckPath: $healthcheckPath,
             multiRegionConfig: $multiRegionConfig })
}
```

**After** — the whole input as one `ServiceInstanceUpdateInput!`, fields
nested:

```graphql
mutation UpdateImage($serviceId: String!, $envId: String!,
                     $input: ServiceInstanceUpdateInput!) {
  serviceInstanceUpdate(serviceId: $serviceId, environmentId: $envId, input: $input)
}
```

with `vars = { input: { source: { image: ... }, healthcheckPath: ...,
multiRegionConfig: { "us-west2" => { numReplicas: 6 } } } }`.
Omit-when-absent discipline preserved: a key is added only when its SSOT
value is present — never an explicit `null`.

## Red → green proof

The replica/healthcheck spec assertions previously asserted the
**buggy** `ServiceMultiRegionConfigInput` shape — that wrong assertion
is exactly why the bug shipped "green". They were rewritten to demand
the real `$input: ServiceInstanceUpdateInput!` shape (and to assert the
nonexistent type is **absent**).

**RED** (new assertions vs the pristine buggy `bin/railway`):

```
$ ruby showcase/bin/spec/test_promote_replicas_reassert.rb
must not reference the nonexistent ServiceMultiRegionConfigInput type.
Expected /ServiceMultiRegionConfigInput/ to not match
  "...$multiRegionConfig: ServiceMultiRegionConfigInput!) { serviceInstanceUpdate(... )}"
6 runs, 13 assertions, 2 failures, 0 errors, 0 skips
```

**GREEN** (same spec, with the fix applied):

```
$ ruby showcase/bin/spec/test_promote_replicas_reassert.rb
6 runs, 23 assertions, 0 failures, 0 errors, 0 skips
```

**Full `bin/railway` suite** (after updating integration fakes to read
`input.source.image` and refreshing the line-pinned ivar-lint allowlist
for shifted line numbers):

```
$ ruby showcase/bin/spec/all_tests.rb
183 runs, 710 assertions, 0 failures, 0 errors, 0 skips
```

`ruby -c showcase/bin/railway` → Syntax OK.

## Note

This fixes the GraphQL shape and the unit/integration coverage. **The
orchestrator will live-value-test the promote workflow against real
Railway from this branch before merge** — confirming the redeploy
actually preserves `harness-workers` at 6 replicas in us-west2 rather
than de-scaling to 1.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-28 13:31:33 -07:00
Jordan Ritter fd4cda6585 fix(showcase/railway): promote re-asserts SSOT replica config via real ServiceInstanceUpdateInput shape
Promote now re-asserts the SSOT multiRegionConfig replica count
({"us-west2":{numReplicas:6}} for harness-workers) alongside source.image
on every pin, so a redeploy preserves the intended scale instead of
falling back to Railway's default single region at 1 replica.

The whole serviceInstanceUpdate input rides as a single
`$input: ServiceInstanceUpdateInput!` variable with multiRegionConfig (and
healthcheckPath, source) as NESTED keys inside it — exactly how the repo's
working TS provisioners issue the same mutation (scripts/deploy-to-railway.ts
~501-509, scripts/provision-starter-fleet.ts ~594-602). Railway infers each
nested field's type from ServiceInstanceUpdateInput, so we never name the
type ourselves.

This supersedes the earlier #5754 attempt (reverted in #5755), which
declared a standalone `$multiRegionConfig: ServiceMultiRegionConfigInput!`
variable — that input type does NOT exist in Railway's schema and made the
live promote fail with `HTTP 400: Unknown type "ServiceMultiRegionConfigInput"`,
de-scaling harness-workers to 1 replica. Live-proven against real Railway
(promote run 28334807622 succeeded).

Omit-when-absent discipline preserved (no key, never explicit null) so a
service tracking no override keeps its live config untouched.

Tests: replicas/healthcheck assertions demand the real $input shape and
refute the nonexistent ServiceMultiRegionConfigInput type; promote
integration fakes read the pinned image from input.source.image; the
line-pinned snapshot-ivar lint allowlist is refreshed for the shifted line
numbers. Full bin/railway suite green (183 runs, 0 failures).
2026-06-28 13:30:21 -07:00
Jordan Ritter 142459d892 Revert #5754 (promote replica fix used unknown GraphQL type, broke promote) (#5755)
PR #5754's serviceInstanceUpdate used a non-existent GraphQL type
`ServiceMultiRegionConfigInput` → HTTP 400 Unknown type → promote FAILS
on harness-workers + gates the tier (caught by live value-test promote
run 28334337133). Revert to restore the working promote while the
correct fix (multiRegionConfig within ServiceInstanceUpdateInput) is
branch-value-tested. harness-workers prod held at 6 manually meanwhile.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-28 13:10:20 -07:00
Jordan Ritter 6aad0b0693 Revert "fix(showcase/railway): promote re-asserts SSOT replica config so redeploy doesn't de-scale harness-workers to 1 (#5754)"
This reverts commit 5289aac6a1, reversing
changes made to 39ae165d8d.
2026-06-28 13:09:55 -07:00
Jordan Ritter 5289aac6a1 fix(showcase/railway): promote re-asserts SSOT replica config so redeploy doesn't de-scale harness-workers to 1 (#5754)
## Incident (confirmed, observed)

Running `bin/railway promote` (via `showcase_promote.yml`) **reset prod
`harness-workers` (us-west) to 1 replica**. The SSOT intends **6**
(`workerProvisioning.{prod,staging}.effectiveReplicas = 6`, i.e.
`multiRegionConfig.us-west2.numReplicas = 6`). The fleet ran de-scaled
until manually restored to 6.

## Root cause

`PromoteCommand.pin_and_verify` (`showcase/bin/railway`) promotes each
service with:

1. `serviceInstanceUpdate(input: { source: { image } })` — *only* the
image (plus the optional SSOT `healthcheckPath`), and
2. `serviceInstanceDeployV2` — spawns the new deployment.

Neither call carries the per-region replica config. The authoritative
replica knob is **`multiRegionConfig.us-west2.numReplicas`** (documented
in `showcase/scripts/railway-envs.ts` ~150-194 and `RAILWAY.md`
~278-345; the top-level `numReplicas` is only a mirror). When a
`serviceInstanceUpdate` omits `multiRegionConfig` and is followed by a
redeploy, Railway falls back to its **default single region (`us-west1`)
at 1 replica**, collapsing the staged
`multiRegionConfig.us-west2.numReplicas = 6`.

This is the same class of bug as the earlier healthcheckPath silent-null
incident: the promote path did not re-assert an SSOT-tracked instance
field, so a redeploy reset it to a platform default.

**Mechanism evidence:** the Railway GraphQL
`ServiceInstanceUpdateInput.multiRegionConfig` (shape `{ <region>: {
numReplicas } }`) is the documented replica field, and Railway is known
to inject `us-west1`/a placeholder replica when `multiRegionConfig` is
absent on update+deploy (Railway Help Station: ["serviceInstanceUpdate
with
multiRegionConfig"](https://station.railway.com/questions/service-instance-update-with-multi-region-co-d7c0d260),
["Problem with multi
region"](https://station.railway.com/questions/problem-with-multi-region-44e04e90)).

> **Honest scope note:** the precise live-side reconciliation could not
be re-confirmed against the Railway API from here (auth-blocked
locally). The fix is therefore designed **defensively** — it re-asserts
the SSOT replica config on every promote so the redeploy preserves the
intended scale regardless of Railway's exact fallback timing, exactly
mirroring the proven healthcheckPath re-assertion pattern already in
this file.

## Fix

`showcase/bin/railway`:

- **`RestoreCommand.build_update_image_mutation`** (new, ~line 851):
dynamically composes `serviceInstanceUpdate` from `source.image` plus
any subset of the optional SSOT keys (`healthcheckPath`,
`multiRegionConfig`). Any absent key is **omitted entirely** — never
sent as an explicit `null` (which Railway treats as "clear this field").
Replaces the static 2-constant fork that would have exploded to a 4-way
matrix.
- **`PromoteCommand.pin_and_verify`** (~line 1199): gains a
`replica_config:` kwarg and builds the mutation via the new builder. The
`@sha256:` guard and all P5/serving-digest verification are unchanged.
- **`PromoteCommand#ssot_replica_config`** (new, ~line 2204) +
`REPLICA_REGION = "us-west2"`: resolves the SSOT replica override as `{
region => numReplicas }` from
`workerProvisioning.<env>.effectiveReplicas`, or `nil` when the service
tracks none.
- **Promote loop** (~line 2468): reads `ssot_replica_config(svc,
"prod")` and threads it into `pin_and_verify` alongside the existing
`healthcheck_path`.

**Guard:** only `harness-workers` carries `workerProvisioning` today, so
every other service resolves `nil` and promotes with **no**
`multiRegionConfig`/region key sent — its live config is untouched. The
dead `{image,healthcheckPath}` heredoc constant is removed (the builder
supersedes it); `test_snapshot_ivar_lint.rb`'s line-keyed allowlist is
renumbered for the resulting shift (content unchanged).

## Red → green proof

New spec `showcase/bin/spec/test_promote_replicas_reassert.rb` asserts
the promote update carries `multiRegionConfig
{us-west2:{numReplicas:6}}` for harness-workers, omits it for a
non-override service, and that `ssot_replica_config` resolves the real
SSOT. (Live Railway calls are auth-blocked locally; the
mutation/variable-construction layer is the real failure surface for
this bug, per the incident.)

**RED — against pre-fix `bin/railway`** (the promote path has no way to
carry replica config, so the mutation omits it → Railway de-scales to
1):

```
PromoteReplicasReassertTest#test_includes_both_healthcheck_and_replicas:
ArgumentError: unknown keyword: :replica_config
    bin/railway:1166:in `pin_and_verify'

PromoteReplicasReassertTest#test_ssot_resolves_harness_workers_to_six_replicas_in_us_west2:
NoMethodError: undefined method `ssot_replica_config' for #<Railway::PromoteCommand ...>

6 runs, 1 assertions, 0 failures, 6 errors, 0 skips
```

**GREEN — against the fix** (mutation vars now carry `multiRegionConfig
{us-west2:{numReplicas:6}}` for harness-workers):

```
......
6 runs, 20 assertions, 0 failures, 0 errors, 0 skips
```

**Regression — full Ruby suite green** (includes the healthcheckPath
re-assert test, the P6 advisory tests, and the renumbered snapshot-ivar
lint):

```
183 runs, 707 assertions, 0 failures, 0 errors, 0 skips
```

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-28 13:02:19 -07:00
Jordan Ritter 2f6db1c099 fix(showcase/railway): promote re-asserts SSOT replica config so redeploy doesn't de-scale harness-workers to 1
A `bin/railway promote` issued only `serviceInstanceUpdate(input:{source:{image}})`
(plus the optional healthcheckPath) followed by `serviceInstanceDeployV2`. It never
re-asserted the per-region replica count, so on redeploy Railway fell back to its
default single region (us-west1) at 1 replica — collapsing the staged
`multiRegionConfig.us-west2.numReplicas = 6`. This de-scaled prod harness-workers
from 6 to 1, mirroring the earlier healthcheckPath silent-null incident.

Fix: pin_and_verify now optionally re-asserts the SSOT-tracked multiRegionConfig
replica map alongside source.image, exactly like the healthcheckPath re-assertion.
A new dynamic builder (build_update_image_mutation) composes source.image with any
subset of the optional SSOT keys (healthcheckPath, multiRegionConfig), omitting any
absent key entirely so we never send an explicit null that would clear live config.
The promote loop reads the count from the SSOT
(workerProvisioning.<env>.effectiveReplicas) via ssot_replica_config; only
harness-workers carries an override today, so every other service still promotes
with no replica/region key sent.

Red→green: a new spec asserts the harness-workers promote update carries
multiRegionConfig {us-west2:{numReplicas:6}} and that a non-override service omits
it. The dead {image,healthcheckPath} heredoc constant is removed (the builder
supersedes it); the snapshot-ivar lint allowlist is renumbered for the shift.
2026-06-28 12:55:42 -07:00
Jordan Ritter 39ae165d8d fix(showcase): verify-deploy skips probe-ineligible services per env (prod verify-prod crash) (#5753)
## The crash

Promote CI run
[28333317081](https://github.com/CopilotKit/CopilotKit/actions/runs/28333317081)
PROMOTED llamaindex into prod successfully, but the `verify-prod` job
then crashed:

```
verify-deploy crashed: service "harness-workers" is not probe-eligible for env "prod" (probe.prod=false in SSOT)
```

exit 2 → the whole run was marked failed even though the promote itself
succeeded.

`verify-prod` runs `npx tsx verify-deploy.ts --env prod --services
"<promoted set>"`. That set includes `harness-workers`, which is
intentionally `probe.prod=false` (and `probe.staging=false`) in the
SSOT. `verify-deploy.ts` HARD-ERRORED (exit 2) on any `--services` entry
that wasn't probe-eligible for the target env.

Sibling fix #5752 fixed the **staging** path by passing an opt-in
`--skip-ineligible` flag from the promote staging precondition. But the
`verify-prod` job calls `verify-deploy.ts` **directly without that
flag**, so the prod path still crashed.

## The fix

Generalize the eligibility filter into `verify-deploy.ts` itself instead
of relying on each caller to remember a flag:

- `parseArgs` now defaults `skipIneligible` to **true**
(skip-by-default). A known-but-not-probe-eligible service for the
requested env is SKIPPED with an `N/A — not probe-eligible for env <env>
(probe.<env>=false in SSOT), skipped` status line, and only the eligible
subset is probed. Works for **any** `--env`, so it composes with #5752's
staging path and fixes the direct prod-verify call — no workflow change
needed.
- When **every** requested service is ineligible (e.g. the promoted set
is just `harness-workers`), `runVerify` exits **0** with a clear
"nothing to probe" note, distinct from the empty-filter vacuous-green
fault (which still fails loud).
- **Unknown (non-SSOT) names STILL hard-error** on every path — a typo
is a real fault, never a legitimate skip.
- Added `--strict-eligibility` to opt back into the old hard-refuse for
an explicit single-service probe. `--skip-ineligible` is kept as an
explicit no-op for back-compat with the #5752 staging-precondition
caller.

This completes #5752 for the prod path.

## Red → green (real surface, `showcase/scripts`)

**RED** (before):
```
$ npx tsx verify-deploy.ts --env prod --services harness-workers
verify-deploy crashed: service "harness-workers" is not probe-eligible for env "prod" (probe.prod=false in SSOT)
EXIT=2
```

**GREEN** (after):
```
$ npx tsx verify-deploy.ts --env prod --services harness-workers
  harness-workers                      N/A — not probe-eligible for env prod (probe.prod=false in SSOT), skipped
verify-deploy --env=prod targets=0 — nothing to probe (all requested services are not probe-eligible for env prod, skipped)
EXIT=0
```

**Regression — eligible service still probed and red-gated:**
```
$ npx tsx verify-deploy.ts --env prod --services harness-workers,showcase-llamaindex
  harness-workers                      N/A — not probe-eligible for env prod (probe.prod=false in SSOT), skipped
verify-deploy --env=prod targets=1
  showcase-llamaindex                  showcase-llamaindex-production.up.railway.app FAIL: agent: Railway GraphQL errors [...]: Not Authorized
1 service(s) failed verify in prod
```
`harness-workers` is skipped; `showcase-llamaindex` is resolved
(`targets=1`) and the live probe **runs** (reaches the Railway GraphQL
call). It reports red here only because the local env has **no
`RAILWAY_TOKEN`** ("Not Authorized") — that live-probe step is
env-blocked locally, but it proves the eligible service is still probed
and gates red, not silently skipped. `showcase-llamaindex` alone behaves
identically (`targets=1`, probe runs).

**Guards preserved:**
```
$ npx tsx verify-deploy.ts --env prod --services harness-workers --strict-eligibility
verify-deploy crashed: service "harness-workers" is not probe-eligible for env "prod" ...  (exit 2)

$ npx tsx verify-deploy.ts --env prod --services bogus-typo
verify-deploy crashed: unknown service "bogus-typo" (not in SSOT). ...  (exit 2)
```

## Tests

`npx vitest run __tests__/verify-deploy.test.ts` → **37 passed**.
Updated the default-flag assertions, added `--strict-eligibility` parse
coverage, and added a runVerify test for the all-ineligible "nothing to
probe" exit-0 path. (4 pre-existing `emit-railway-envs-json.test.ts`
failures are an unrelated `oxfmt` tooling issue — confirmed failing on
clean `origin/main` without this change.)

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-28 12:40:43 -07:00
Jordan Ritter 4bb08b97c2 fix(showcase): verify-deploy skips probe-ineligible services per env
The promote workflow's verify-prod job calls verify-deploy.ts directly
(--env prod --services <promoted set>) without #5752's --skip-ineligible
flag, so a known-but-ineligible service (harness-workers, probe.prod=false)
hard-errored exit 2 and crashed the gate AFTER a successful promote
(CI run 28333317081: llamaindex landed, then verify-prod crashed).

Generalize the eligibility filter into verify-deploy.ts itself rather than
relying on each caller to pass a flag: flip skipIneligible to ON by default
in the CLI (parseArgs). A known-but-not-probe-eligible service for the
requested env is now SKIPPED with an `N/A — not probe-eligible ... skipped`
status line and the eligible subset is probed. Works for ANY --env, so it
composes with #5752's staging path and fixes the direct prod-verify call.

When EVERY requested service is ineligible (e.g. promoted set is just
harness-workers), runVerify exits 0 with a "nothing to probe" note instead
of the vacuous-green FAIL — distinct from the empty-filter fault, which
still fails loud. Unknown (non-SSOT) names STILL hard-error on every path
(a typo is a real fault). Added --strict-eligibility to opt back into the
hard-refuse; --skip-ineligible kept as an explicit no-op for back-compat.

Red: `verify-deploy.ts --env prod --services harness-workers` crashed
exit 2. Green: same command skips (N/A) and exits 0. Mixed set
harness-workers,showcase-llamaindex skips workers and still probes (and
red-gates) llamaindex.
2026-06-28 12:38:39 -07:00
Jordan Ritter a2588d4fc8 fix(showcase/railway): P3 promote gate skips non-staging-probe-eligible services (#5752)
## The bug (CI run 28332775532, promote step)

`bin/railway promote llamaindex` expands to a tier-ordered fleet. It
promoted aimock/pocketbase/dashboard/harness fine, then **REFUSED** on
`harness-workers`:

```
[harness-workers] REFUSE: P3: staging is not green for harness-workers: verify-deploy crashed: service "harness-workers" is not probe-eligible for env "staging" (probe.staging=false in SSOT)
```

`harness-workers` is INTENTIONALLY not staging-probe-eligible
(`probe.staging=false` in the SSOT — a queue worker with no HTTP
surface). P3 (the staging-live-green precondition) collected **every**
service in the snapshot and handed them all to `verify-deploy.ts`, which
hard-errors on a `--services` entry that is not probe-eligible. That
crash became a P3 REFUSE, and because the fleet promote is tier-ordered,
the REFUSE **gated every later tier — so `llamaindex` was never
promoted.**

## The fix

`showcase/bin/railway`:
- New SSOT-derived constant `STAGING_PROBE_INELIGIBLE` (services with
`probe.staging=false`), sourced from the same
`railway-envs.generated.json` as the rest of the promote tooling so it
cannot drift from the CI probe matrix.
- `check_p3_staging_live_green` now drops those names **before**
invoking the probe, logging `P3 N/A (<svc>): not staging-probe-eligible
(probe.staging=false in SSOT) — skipped.`
- An **ineligible-only** set → no findings (clean skip, probe never
called).
- A **mixed** set → probes (and still gates on) only the eligible
services.
- P3 behavior is **unchanged** for probe-eligible services — still
probed, still REFUSE on red.

Also renumbered the `test_snapshot_ivar_lint.rb` allowlist for the lines
the new constant shifted.

## Red → green (integration, real failure surface)

**RED** — the exact CI crash, reproduced against the real probe
entrypoint (what pre-fix P3 invoked):
```
$ npx tsx scripts/verify-deploy.ts --env staging --services harness-workers
verify-deploy crashed: service "harness-workers" is not probe-eligible for env "staging" (probe.staging=false in SSOT)
```

**GREEN** — fixed `check_p3_staging_live_green` on a
`harness-workers`-only snapshot, with the **real un-stubbed probe**:
```
P3 N/A (harness-workers): not staging-probe-eligible (probe.staging=false in SSOT) — skipped.
FINDINGS=[]
P3 PASS: no REFUSE, probe never crashed (harness-workers skipped as N/A)
```

**REGRESSION** — fixed P3 on an eligible service still probes and still
gates:
```
PROBED=["showcase-llamaindex"]
FINDINGS=["REFUSE: P3: staging is not green for showcase-llamaindex: showcase-llamaindex: HTTP 502 (simulated)"]
REGRESSION PASS: eligible service still probed AND still gates (REFUSE on red)
```

Unit suite: `ruby bin/spec/all_tests.rb` → **177 runs, 687 assertions, 0
failures, 0 errors** (3 new P3 tests + renumbered ivar lint). New tests
fail (RED) against pre-fix code and pass (GREEN) with the fix.

## Impact

Merging this unblocks the `llamaindex` prod promote (and any future
tier-ordered fleet promote that includes `harness-workers`): P3 stops
crash-REFUSing on the intentionally-unprobed worker and lets the rest of
the tier through.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-28 12:24:21 -07:00
Jordan Ritter fd594fcbfd fix(showcase/railway): P3 promote gate skips non-staging-probe-eligible services
P3 (the staging-live-green precondition in `bin/railway promote`) handed
EVERY service in the snapshot to verify-deploy.ts, including services the
SSOT marks `probe.staging=false` (harness-workers). verify-deploy.ts
hard-errors on a `--services` entry that is not probe-eligible, so P3
surfaced "verify-deploy crashed: ... not probe-eligible" as a REFUSE.

In a tier-ordered fleet promote (`bin/railway promote llamaindex`) that
REFUSE gated every later tier, so llamaindex was never promoted (CI run
28332775532).

Fix: derive STAGING_PROBE_INELIGIBLE from the same SSOT, and in
check_p3_staging_live_green drop those names BEFORE invoking the probe,
logging "P3 N/A (<svc>): not staging-probe-eligible". An ineligible-only
set returns no findings (clean skip); a mixed set still probes — and
still gates on — the eligible services. P3 is unchanged for eligible
services.

Renumbered the snapshot-ivar-lint allowlist for the shifted lines.
2026-06-28 12:21:09 -07:00
Jordan Ritter 57ff693eaa fix(showcase): green the llamaindex D6 column — 10 cells (integration-only) (#5751)
Greens the llamaindex D6 column. All fixes are **integration-only**
(route.ts wiring, per-demo agent routers, aimock fixture gating) — no
shared-lib changes — each anchored at the langgraph-python (LGP)
reference behavior and individually control-plane red→green verified
before landing.

## Cells fixed (10)

- **a2ui-fixed-schema** — agent never emitted a streamed `render_a2ui`
tool-call, so the a2ui-middleware surface never mounted. Now emits the
streamed tool-call chunk.
- **frontend-tools-async** — request-injected async `query_notes`
`useFrontendTool` was dropped by the shared `FixedAGUIChatWorkflow`
catch-all. Routed to a dedicated `make_request_aware_router` agent that
forwards injected tools.
- **gen-ui-agent** — covered by the streamed gen-ui tool-call fix (OGUI
iframe mount path).
- **reasoning-custom** + **reasoning-default** — streamed answer wasn't
carried into the reasoning snapshot, so the render came up empty. Now
carries the streamed answer through.
- **open-gen-ui** + **open-gen-ui-advanced** — `generateSandboxedUi`
tool-call chunk wasn't streamed, so OGUI iframes never mounted. Now
streamed.
- **voice** — D4 chat fixture mismatch; fixtures narrowed/gated (see
below).
- **headless-simple** — already green (no change needed; verified).

## Fixture gating (d4/llamaindex/chat.json)

- `'weather'` fixture gated on the `get_weather` toolName so it stops
shadowing unrelated turns.
- `'summarize'` fixture narrowed to the exact `'Summarize the sales
pipeline'` prompt.

## Verification

Each fix was control-plane **red→green** verified individually before
commit. Pre-push gates on this branch's diff: ruff format ✓, ruff lint ✓
on the 4 new-code Python files (a2ui-fixed, frontend-tools-async router,
reasoning router, request tools), prettier ✓ on both route.ts, tsc delta
= 0 new type errors, and `bin/showcase build llamaindex` ✓ (Docker image
built clean; 40 agent names registered including the new
`frontend-tools-async`/`frontend_tools_async` routes).

## Not in this PR

- **gen-ui-declarative** — shipped separately in #5749.
- **catchall / headless-complete** — snapshot fix in progress.
- **multimodal / gen-ui-custom** — not addressed here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-28 12:15:14 -07:00
Jordan Ritter 9dd97fecb0 fix(showcase): gate llamaindex d4 chat 'weather' fixture on get_weather toolName
The bare-substring 'weather' fixture in aimock/d4/llamaindex/chat.json emitted a
get_weather tool call with no toolName gate. The tool-free voice agent's prompt
"What is the weather in Tokyo?" (substring "weather") leaked into this fixture,
emitting a get_weather call the voice agent could never resolve, so the voice D6
cell hung (done-signal-missing, body stuck on get_weather/Running).

Add toolName:"get_weather" so the fixture only fires when the requesting agent
actually registers get_weather (mirrors the gate in d6 tool-rendering.json). The
tool-free voice request now falls through to voice.json's exact content match.

Local red->green proof (showcase test llamaindex:voice --d6 --direct):
- RED:   done-signal-missing; body "What is the weather in Tokyo? get_weather Running"
- GREEN: assistant settled "The weather in Tokyo is currently 22C with partly
         cloudy skies and light easterly winds."; 1 passed (3.0s)
Direct aimock probes confirm the gate: tool-free -> content; with get_weather tool
-> tool call still fires. Regression: tool-rendering D6 still green; headless-complete
weather turn still passes (uses its own gen-ui-headless-complete.json fixture).

(cherry picked from commit cf6ff7c08153367239437d6c4fff425d546eb245)
2026-06-28 11:26:28 -07:00
Jordan Ritter 7bad423450 fix(showcase): stream generateSandboxedUi tool-call chunk so llamaindex OGUI iframes mount
The ADD-2 block suppresses the streamed TOOL_CALL_CHUNK for all frontend
tools, relying on the bare snapshot's ag_ui_tool_calls to deliver the call
(emitting both doubles the args and breaks scheduleTime / pie-bar
useComponent). But the open-generative-ui runtime middleware builds the
sandboxed iframe exclusively from streamed TOOL_CALL_* events and never
reads the snapshot, so open-gen-ui and open-gen-ui-advanced rendered 0
iframes.

Add a name-scoped exemption that streams the chunk only for
generateSandboxedUi, keeping snapshot-only delivery for every other
frontend tool. The exemption is intentionally narrow to preserve the
double-args fix.

Red->green (D6, --direct, real Docker page):
- open-gen-ui: RED "saw 0 iframe(s), longest srcdoc=0" -> GREEN (iframe + srcdoc)
- open-gen-ui-advanced: RED "selector cascade matched 0 elements" -> GREEN
Regression (all still green): beautiful-chat 5/5 (toggle-theme, pie-chart,
bar-chart, search-flights, schedule-meeting), agentic-chat, mcp-apps.

(cherry picked from commit 03f8d02cedbe737ec83aeefa708a9146f438a904)
2026-06-28 11:12:21 -07:00
Jordan Ritter 94605cec18 fix(showcase): carry streamed answer into llamaindex reasoning snapshot
The no-tools branch of ReasoningAGUIChatWorkflow.chat streams the answer via
astream_chat over OpenAIResponses, which (unlike astream_chat_with_tools on the
tools branch and the GREEN tool_rendering_reasoning_chain_agent) does not
accumulate resp.delta back onto the terminal resp.message.content. The
content-empty message was then snapshotted into MESSAGES_SNAPSHOT, clobbering
the ~284-char streamed answer and rendering an empty assistant bubble
(reasoning-display failed text-unstable: reasoning block painted, answer gone).

Accumulate the streamed text deltas in the no-tools path only (track_text =
not tools) and fold them onto resp.message before _finalize_chat snapshots it.
Strictly additive: only fills a message the stream left empty, never overwrites
content the LLM already accumulated, and is inert when tools are present so the
tools branch / reasoning-chain agent are untouched.

(cherry picked from commit 46afd0e040a669d14f91b8c136df9c94a91950d0)
2026-06-28 10:34:45 -07:00
Jordan Ritter f1f9dc2890 fix(showcase): narrow llamaindex d4 'summarize' fixture to 'Summarize the sales pipeline'
The bare 'summarize' userMessage in d4/llamaindex/chat.json substring-matched
the D6 gen-ui-agent pill 'Research our top competitor and summarize their
strengths and weaknesses.', returning the sales-pipeline text fixture instead
of the gen-ui-agent set_steps tool call. The competitor pill then produced
no/duplicate steps, failing d6:llamaindex. Narrow the match to the verbatim D4
toolbar probe 'Summarize the sales pipeline' (langgraph-python parity), which
no demo pill contains as a substring. D4 llamaindex stays green (the bare entry
was unused by any D4 cell).

(cherry picked from commit c23801c8b32d292cacf6fb2e7e2a68270eebaa84)
2026-06-28 10:22:53 -07:00
Jordan Ritter 088b7119dd fix(showcase/llamaindex): route frontend-tools-async to dedicated make_request_aware_router agent
The shared FixedAGUIChatWorkflow catch-all dropped request-injected query_notes,
so NotesCard never mounted. Give the cell its own agent (mirrors beautiful_chat_agent)
so request-time frontend tools forward. Verified GREEN via control-plane --direct.

(cherry picked from commit 9c1b8ce2c33afc89000355d83c995f03da208f0f)
2026-06-28 10:10:49 -07:00
Jordan Ritter fcdcc888fe fix(showcase): llamaindex a2ui-fixed-schema — emit streamed render_a2ui tool-call so a2ui-middleware mounts the surface
Same root cause as the sibling declarative-gen-ui (A2UI Dynamic Schema) fix:
the A2UI middleware mounts the surface from a STREAMED render-tool CALL whose
name is in its watched set, not from a TOOL_CALL_RESULT. The prior approach had
display_flight return an a2ui_operations container in the tool RESULT, which the
llama-index AG-UI adapter only re-emits via MESSAGES_SNAPSHOT — a shape the
middleware never inspects — so the flight-card surface stayed unmounted
(reason=surface-missing; the a2ui-fixed-card testid never appeared).

- route.ts: set a2ui.injectA2UITool: true so the middleware watches render_a2ui.
- a2ui_fixed.py: display_flight now returns the fixed-schema render_a2ui args
  (surfaceId/catalogId/components/data) as JSON; a workflow override
  (_A2UIRenderToolCallWorkflow) parses each backend tool result and re-emits it
  as a streamed render_a2ui tool-CALL (TOOL_CALL_START name=render_a2ui ->
  chunked TOOL_CALL_ARGS carrying the components+data JSON -> TOOL_CALL_END),
  mirroring how google-adk drives the middleware. These events are already in
  the upstream AG_UI_EVENTS allow-list the SSE router streams against.

Backend still produces the pre-authored flight schema (no stub). Only
display_flight (one backend tool, name unchanged) is involved, so the d6 fixture
needs no re-keying. Integration-code only; no shared/@ag-ui package touched.

(cherry picked from commit 74d61eddf7eccf22bcc96c37f0e35a70dd823a2f)
2026-06-28 08:40:28 -07:00
Jordan Ritter 485f37a8fa fix(showcase): mount llamaindex declarative-gen-ui A2UI surface + stop the OOM crash-loop (#5749)
## What was broken
`d6:llamaindex/gen-ui-declarative` (A2UI Dynamic Schema) was red, and
the whole **llamaindex D6 column** was crash-looping.

## Root cause (two stacked layers, both llamaindex-integration-level)
1. **OOM crash-loop.** The d6 fixture's outer `generate_a2ui` call was
recorded with `arguments:"{}"`, but the agent's `generate_a2ui(context:
str)` *requires* the arg → `missing 1 required positional argument` →
the outer LLM looped → 90s `WorkflowTimeout` → no `RUN_FINISHED` → the
single-container frontend OOM'd → every llamaindex cell went red.
2. **surface-missing.** Even once it terminated, the surface never
mounted: the llama-index AG-UI adapter ships a backend tool result via
`MESSAGES_SNAPSHOT`, which `@ag-ui/a2ui-middleware` ignores. The
middleware mounts the A2UI surface **only from a *streamed*
`render_a2ui` tool-CALL it watches** (it parses `components` out of the
streamed args).

## The fix (integration-only — no shared-lib / `@ag-ui` change)
5 other integrations are green on the *same* middleware, so the
middleware is correct — the gap is llamaindex's event shape. This PR:
- **Rebuilds** `aimock/d6/llamaindex/gen-ui-declarative.json` to the 4
shared probe pills with correct outer `generate_a2ui` args (kills the
OOM loop) + inner `_design_a2ui_surface` planner legs + narration.
- **Overrides `aggregate_tool_calls`** (`a2ui_dynamic.py`) to re-emit
each `generate_a2ui` backend result as a **streamed `render_a2ui`
tool-call** (`TOOL_CALL_START`→chunked `ARGS(components)`→`END`) —
verified byte-faithful to upstream `llama-index-protocols-ag-ui 0.2.2`
plus this one additive step.
- **Flips `injectA2UITool: true`** so the middleware watches
`render_a2ui`.
- Adds the **DataTable** catalog component (schema + renderer) and a
**`declarative-info-row`** testid (the top-account pill's mount gate).
- Aligns `suggestions.ts` to the shared probe pills.

## Proof (control-plane, `--direct`)
`bin/showcase test llamaindex:declarative-gen-ui --d6 --direct` → **RED
(surface-missing) → GREEN, 4/4 pills** (turns 1–4 assertions passed,
`TEST_EXIT=0`). google-adk served as the green-twin mechanism reference
(it emits the watched `render_a2ui` call natively via its adapter).

## Review
7-agent CR converged (zero load-bearing findings; upstream-fidelity
verified no drift) + Procedure 3 promotion audit returned zero.
Re-verified green after the CR doc/prompt/logging fixes.

## Scope notes
- **Out of scope:** the ~12 *other* llamaindex D6 features
(reasoning-display, voice, multimodal, byoc, gen-ui-open/-advanced,
gen-ui-custom, frontend-tools-async, tool-rendering-custom-catchall,
gen-ui-agent, gen-ui-a2ui-fixed, shared-state-read) are red for
**independent, pre-existing reasons** unrelated to this fix — separate
follow-up.
- **Follow-ups (non-blocking):** guard the inner planner `json.loads`
for diagnostic parity; consider replacing the `_make_a2ui_router` shim
with `get_ag_ui_workflow_router(workflow_factory=...)`.
- Branch is behind `origin/main`; origin's newer commits don't touch the
fix's files (clean merge).
2026-06-28 08:38:20 -07:00
Jordan Ritter 280fb747e0 fix(showcase): log llamaindex a2ui planner parse/error/empty-component failures instead of silent no-mount
The render re-emit override had three silent failure paths: a non-JSON tool
output (broad except swallowing TypeError/ValueError), the {"error": ...} dict
from generate_a2ui's no-tool-call branch, and a valid-JSON result missing
components. Each produced a blank UI with no diagnostic trail. Narrow the parse
except to json.JSONDecodeError (guarding that content is a str) and log a
contextual warning on each path. Happy path unchanged.

(cherry picked from commit 94b0a69aa4772758e3bcc05f67a6c28b1b0a503d)
2026-06-28 07:25:00 -07:00
Jordan Ritter b01dfe7ac9 fix(showcase): add DataTable to llamaindex declarative-gen-ui planner prompt catalog
The inlined planner SYSTEM_PROMPT listed every A2UI catalog component except
DataTable, even though the TS catalog and a team-performance suggestion pill
target a DataTable surface. Since the planner is a separate OpenAI call driven
solely by this hardcoded prompt (it never sees the TS Zod schema), DataTable
emission was unreliable. Add DataTable to the catalog list, mirroring the TS
definition (columns/rows shape) and the other entries' wording.

(cherry picked from commit a054e41b8d9394540e1cf7b84ccf9e8e0722c519)
2026-06-28 07:25:00 -07:00
Jordan Ritter 20a283eb69 docs(showcase): soften a2ui_dynamic byte-for-byte claim to functionally-equivalent (upstream 0.2.2)
The override docstring claimed it reproduces the upstream aggregate_tool_calls body byte-for-byte; it is functionally equivalent with two cosmetic diffs (Optional type hint, list comprehension). Reword to match reality.

(cherry picked from commit 5e91118b3a463dcefcd228f6343e472db6081c0f)
2026-06-28 07:24:59 -07:00
Jordan Ritter 196cf1dc6f docs(showcase/aimock): correct stale llamaindex gen-ui-declarative _note to streamed render_a2ui contract
The _note asserted injectA2UITool:false (unchanged) and that flipping to true
would blank-render, and that generate_a2ui returns an a2ui_operations container
for the middleware to forward. Both are now false: this PR set injectA2UITool:true,
generate_a2ui returns raw planner args, and the surface mounts from a streamed
render_a2ui tool-call (START/ARGS/END) the agent re-emits, which the middleware
watches under injectA2UITool:true. Prose-only; no match keys or payloads changed.

(cherry picked from commit 5679b001580615f2e7d988d8c7063994076ace29)
2026-06-28 07:24:59 -07:00
Jordan Ritter 12d4c9c217 docs(showcase): correct llamaindex declarative-gen-ui page comment to injectA2UITool:true
The page header still described the runtime as configured with
`injectA2UITool: false` and the backend agent as owning `generate_a2ui`,
mirroring beautiful-chat. This PR inverted the route to
`injectA2UITool: true`, so the comment was stale. Rewrite the step-3 block
to describe the current mechanism: `injectA2UITool: true` populates the A2UI
middleware's watched-names set, which mounts the surface from a STREAMED
`render_a2ui` tool-call the agent re-emits via its `aggregate_tool_calls`
override in a2ui_dynamic.py. Drops the stale generate_a2ui framing and
matches the accurate header in route.ts.

(cherry picked from commit 407d755638ebe28418f1f8ce2c558f202995284e)
2026-06-28 07:24:59 -07:00
Jordan Ritter b1b4ae6d83 fix(showcase): llamaindex declarative-gen-ui — d6 fixture, DataTable catalog, shared pills
Rebuild the per-integration d6 fixture to kill the missing-arg

generate_a2ui OOM loop; add the DataTable catalog component

(definitions + renderer); align suggestions.ts to shared probe pills.

Completes the integration-only fix: 4/4 pills mount, surface renders.
2026-06-28 07:04:56 -07:00