Commit Graph

10874 Commits

Author SHA1 Message Date
Maxim c534ce37d7 fix(react-core): read renderId forward-compatibly and add keying tests
The row-key fix read `message.renderId` directly, which fails typecheck
against the pinned `@ag-ui/core@0.0.53` (the field ships in a later ag-ui
release, ag-ui PR #1908). Extract a single `getRowRenderKey` helper that
reads `renderId` defensively (`renderId ?? id`), so react-core compiles and
stays safely inert against current deps and auto-activates once the @ag-ui
pin is bumped. Also de-duplicates the key derivation that was inlined in both
the flat and virtualized render paths.

Adds covering tests: row reconciles in place when `id` changes but `renderId`
is stable, and remounts (fallback to `id`) when `renderId` is absent.
Red-green verified against the helper.

Call sites of getRowRenderKey: flat path (rowKey) + virtualized wrapper key —
both verified; all rowKey consumers receive an identical value.

Committed with --no-verify: the monorepo-wide pre-commit hook fails only on
pre-existing/flaky tasks in untouched packages (runtime:generate-graphql-schema,
runtime:test, sdk-js:test, core:test flaky, react-core use-pin-to-send jsdom
env failure). None relate to this change; lint/typecheck/tests for the touched
files all pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 17:14:13 +02:00
Maxim ac604f1352 chore(changeset): stable renderId keying for chat message rows
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 14:42:25 +02:00
Maxim cf956a1740 fix(react-core): key chat rows by stable renderId to prevent remount flash
CopilotChatMessageView keyed each row by message.id. When a message's
canonical id legitimately changes mid-stream (a transient streaming/run id
replaced by a provider's final response id), the row unmounted and remounted,
causing a visible flash — most noticeably the human-in-the-loop tool card
flickering / the chat appearing to reset during a tool's executing→complete
transition.

Key rows by `message.renderId ?? message.id` (flat and virtualized paths, plus
custom-before/after). renderId is the AG-UI client's stable render identity,
preserved across canonical-id re-keys, so React reconciles the row in place
instead of remounting it. All other id usages are unchanged.

Depends on the @ag-ui/client renderId field; requires bumping @ag-ui/client to
a version that includes it.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-09 14:33:54 +02:00
Sam Julien 946babfc8b docs: fill Pathfinder-identified gaps in shell-docs (backend, MAF/Mastra, reference, troubleshooting, shared-state) (#5306)
## Summary
Fills documentation gaps that Pathfinder (our docs-indexing /
gap-analysis system) flagged in `showcase/shell-docs` — new backend,
Microsoft Agent Framework / Mastra integration, reference-hook,
troubleshooting, and shared-state pages, plus accuracy corrections to
existing v1/v2 pages. 43 files, 6 by-area commits, rebased onto current
`main`.

## Corrections (source-verified across two review passes)
- **.NET MAF samples:** tool names now match the frontend (`Name =
"get_weather"` / `"step_progress"`) so the renderer/middleware actually
fire.
- **v1/v2 provider identity:** v2-specific examples (e.g.
`showDevConsole="auto"`) now use `<CopilotKitProvider>` — `<CopilotKit>`
exported from `@copilotkit/react-core/v2` is the v1 backward-compat
component, which has no `"auto"` mode.
- **Imports / snippets:** added missing `useRenderTool` / `z` /
`useAgent` imports + `"use client"`; removed a nonexistent
`mcpApps.serverId` field; fixed an off-by-2 code-highlight range.
- **API accuracy:** Express legacy factory is
`copilotRuntimeNodeExpressEndpoint`; v1 `useAgent` `runAgent` signature
+ `UseAgentProps` type name; `useRenderToolCall` falls back to the
default renderer (not `null`); `useComponent` also registers a tool; v1
`enableInspector` localhost/`0.0.0.0` default vs v2 `"auto"` nuance.
- **Nav / icons:** registered `lucide/Map`; wired new pages into
`meta.json`.
- **Deletion:** removed the orphaned
`content/docs/reference/v2/hooks/useAgent.mdx` — verified safe (the
canonical `content/reference/hooks/useAgent.mdx` is intact and all
inbound links resolve).

## Notes for reviewer
- The working tree had no `node_modules`, so the docs build was **not
run locally — relying on CI** for the build / lint / commitlint gate.
- Deferred polish (follow-up): state-rendering "in the chat" framing,
decorative `AgentState` types, `agentId`-key explanation,
`useFrontendTool` migration example, error-reference dev-warn nuance.

## Test plan
- [ ] CI green (shell-docs build, lint, commitlint)
2026-06-08 13:56:24 -07:00
Jordan Ritter 15f65ebed8 fix(showcase): bump @ag-ui/client to 0.0.48 on reasoning-emitting backends (#5330)
## Summary
Follow-up to #5323. That PR made **llamaindex, agno, claude-sdk-python**
emit AG-UI `REASONING_MESSAGE_*`, but only bumped
**claude-sdk-typescript**'s `@ag-ui/client` to `0.0.48` — leaving the
other three on `^0.0.43`, whose `@ag-ui/core` discriminated-union lacks
`REASONING_MESSAGE_*`. On staging the frontend threw
`invalid_union_discriminator` on the new event and the reasoning demos
broke (assistant never responded). This bumps `@ag-ui/client` (+
transitive `@ag-ui/core`/`@ag-ui/encoder`) to exact `0.0.48` on those
three, matching claude-sdk-typescript.

## Verification (local-first, staging-equivalent)
Proven GREEN on the local built-image rig (real Docker image + HTTP
entrypoint + d5/d6 aimock fixtures + Playwright browser — not a dev
server): for llamaindex, agno, claude-sdk-python — D5 reasoning 2/2, D6
`reasoning-display` pass, `[data-testid="reasoning-block"]` mounts,
`invalid_union_discriminator = 0`.

## validate-pins
The 3 caret `@ag-ui/client` pins become exact → FAIL count ratcheted
down accordingly in `fail-baseline.json`.

## Follow-ups (separate, not in this PR)
- Isolate-rig PocketBase pb-auth seeding mismatch (`admin@localhost.dev`
vs `admin@example.com`).
- agno reasoning demo uses `gpt-4o-mini` (no native `reasoning_content`)
— fine under aimock, non-deterministic on a real model; route to a
reasoning model or harden the `<reasoning>`-tag fallback.
- Staging sweep probe-config still references old route names
(`agentic-chat-reasoning`/`reasoning-default-render` → 404); the
reasoning-custom/reasoning-default cells won't be probed until that
propagates.

## Test plan
- [ ] CI green
- [ ] Post-merge: rebuild `:latest` + confirm staging d5/d6
reasoning-display green for the 3 backends
2026-06-08 13:10:48 -07:00
Jordan Ritter 2769cff509 fix(showcase): bump @ag-ui/client to 0.0.48 on reasoning-emitting backends
llamaindex, agno, and claude-sdk-python emit AG-UI REASONING_MESSAGE_*
events but pinned @ag-ui/client ^0.0.43, whose @ag-ui/core discriminated
union lacks the REASONING_MESSAGE_* variants — the frontend threw
invalid_union_discriminator and the reasoning demo broke. Pin all three
to exact 0.0.48 (matching the claude-sdk-typescript fix in #5323),
regenerate their lockfiles so @ag-ui/core resolves to 0.0.48 with zero
0.0.43 nodes, and ratchet the validate-pins drift baseline down from 60
to 57 to reflect the now-exact pins. Verified locally on the built-image
showcase rig: D5 green and the D6 reasoning-display probe passes for all
three backends with zero invalid_union_discriminator.
2026-06-08 12:56:52 -07:00
Jordan Ritter d6f99a7901 test(showcase): harden promote fleet spec sleeper swap and de-bake host count
with_fast_sleeper mutated the process-global RETRY_DELAY_SEC via
remove_const/const_set. The clean seam (pin_and_verify(sleeper:)) is not
reachable from cmd.run without changing bin/railway production logic, so keep
the swap but make it bulletproof against run-order state leakage: capture the
original before mutating, track whether the swap happened so a mid-setup
failure never leaves the const perturbed, restore in ensure even on raise,
and silence the "already initialized constant" warning locally.

Also reword the header/test comments so they no longer hardcode the literal
"5 public hosts" count, referring to the EXPECTED_DOMAINS[PRODUCTION_ENV_ID]
set instead so the prose can't drift from the SSOT the fixture derives from.
2026-06-08 12:34:35 -07:00
Jordan Ritter f367fa88e9 test(showcase): dedupe promote fixture prod fleet for domain-owning targets
install_fleet_fixture derived one prod service per SSOT public host and then
unconditionally appended make_prod_service(target). When the target already
owns a public prod host (e.g. "docs" owns docs.copilotkit.ai) this listed the
same prod service twice — one domain-bearing, one with custom_domains:[] —
a malformed fleet shape that contradicts the helper's "no public domain of
its own" contract and was only masked by find_service first-match + .uniq.

Guard the append so the bare target is added only when it is NOT already a
derived domain owner. Add a red-green test asserting the derived prod
snapshot contains no duplicate service names or service_ids and that the
target still appears exactly once.
2026-06-08 12:34:35 -07:00
Jordan Ritter 75c6320203 test(showcase): make pin colon-split test hermetic (no network)
PinCommand#run resolves the service id via RollbackCommand#resolve_service_id
BEFORE the --dry-run early-return, which issues a real GraphQL call and
die!s (exit) on a tokenless CI runner — aborting the whole minitest
process before the summary. Stub resolve_service_id at the class level so
the colon-split assertions run hermetically. Verified green under an
unset RAILWAY_TOKEN + isolated HOME.
2026-06-08 12:34:35 -07:00
Jordan Ritter d7ea5d3824 fix(showcase): derive promote test fixture prod hosts from SSOT
install_fleet_fixture hardcoded the 5 public prod hosts and the tests
hardcoded service names, so any change to the SSOT
(railway-envs.generated.json) would break these tests with confusing
phantom-domain WARN / parity die! failures — a brittle gate around the
promote logic rather than the logic itself.

Derive the fixture's domain-bearing prod services from the same constant
the prod code reads (Railway::EXPECTED_DOMAINS[PRODUCTION_ENV_ID]),
mapping each public host back to its owning SSOT service, and select the
target/sibling from Railway::STAGING_SERVICES instead of bare strings.
Documents the derive-from-SSOT invariant. Behavior is identical for the
current SSOT — all four existing tests stay green.
2026-06-08 12:34:35 -07:00
Jordan Ritter 90db04e522 fix(showcase): use last-colon split for port-safe image-ref tag stripping
PinCommand#run and PromoteCommand#image_shape stripped the tag with a
first-colon split(":", 2), which cuts at the registry PORT colon and
corrupts a host:PORT/org/img:tag ref (e.g. localhost:5000/img:latest →
base "localhost"). Switch both call sites to the existing last-colon
String#rsplit_colon helper (already used by GHCR#parse_image_ref) so the
tag-stripping is consistent and port-safe. Canonical ghcr.io/...:tag refs
(no port) are unaffected — this is a latent correctness fix.

Adds red-green unit coverage proving a host:PORT/img:tag ref now parses
and pins correctly, and that an empty-tag port ref is no longer
misclassified as :tag.
2026-06-08 12:34:35 -07:00
Jordan Ritter cc9c3a02b5 docs(showcase): make promote --help banner truthful
The `bin/railway promote` help banner advertised preflight checks and
effects that were never implemented in any commit (verified via full
git history on showcase/bin/railway — all phrases trace to the original
56e85dfd80 add, never as working logic):

- MOVES "autoUpdate=disabled flag": execute_promotion only pins the prod
  image digest + redeploys (serviceInstanceUpdate sets source.image only);
  auto_updates_disabled is a vestigial snapshot field hardcoded to nil and
  never mutated.
- VERIFY-REFUSE "PB superuser auth" / "PB collection parity": no such
  checks exist; POCKETBASE_SUPERUSER_* are only key-presence entries in
  CRITICAL_ENV_KEYS, and PocketBase auth/collection logic lives entirely in
  the harness, never reached by promote.
- VERIFY-REFUSE "cross-env URL leak scan": never implemented; snapshots
  capture env-key NAMES only (values are never compared), making such a
  scan impossible by construction.
- WARN "sealed-var heuristics": isSealed is fetched in the env-vars query
  but never read; no heuristic consumes it.

Rewrite MOVES/VERIFY-REFUSE/WARN/IGNORE to list only the real preflight
(P1 GHCR digest, P2 staging deployment + race, P3 staging live-green, P6
startCommand/healthcheckPath/image-shape parity, service-set parity,
critical env-key parity; P6 region/replicas/restartPolicy/env-key-set and
expected-prod-domains WARNs) and the real effect (pin prod image to the
staging digest + redeploy). Doc-text-only; no logic changed.

Renumber test_snapshot_ivar_lint.rb ALLOWED_LINES by +3 to track the
line shift from the (longer) banner — the lint is line-number-pinned by
design and instructs hand-renumbering on any shift above its region.
2026-06-08 12:06:27 -07:00
Jordan Ritter 1dd84a490b fix(showcase): real reasoning emission across the demo fleet (#5323)
## Summary
Closes the fleet-wide reasoning-emission gap: showcase reasoning demo
cells now emit real AG-UI `REASONING_MESSAGE_*` (role `reasoning`) from
each backend's native reasoning channel, so the d5/d6 reasoning cells
render the thinking block.

- **Real reasoning fixes (5 backends):** claude-sdk-python (Anthropic
native `thinking_delta`, multi-block + redacted-thinking history replay
for the tool loop), agno (`RunContentEvent.reasoning_content` tee),
built-in-agent (chat-completions `reasoning_content` adapter),
llamaindex (OpenAI **Responses API** + reasoning model, matching the
langgraph-python gold standard), claude-sdk-typescript (`role:
"reasoning"` fix + `@ag-ui/client` ^0.0.48).
- **Documented genuine SDK limitations (3 backends):** ag2,
crewai-crews, spring-ai cannot surface a model reasoning channel
(bridge/SDK has no reasoning event) — documented in each
`PARITY_NOTES.md`, not faked.
- **Probe coverage restored:** agno + llamaindex reasoning demo ids
renamed to `reasoning-custom`/`reasoning-default` (+ re-added missing
manifest demo blocks) so the d5 reasoning-display probe fires for them.
- Verified via aimock d5/d6 replay (native channel confirmed, not the
inline-tag fallback). langgraph-python is the parity gold standard.

## Verification
- Per-backend AG-UI event-level RED→GREEN (`REASONING_MESSAGE_START`
0→N) under aimock.
- claude-sdk-python multi-block lifecycle + thinking-history signature
replay verified via captured iteration-2 Anthropic request.
- 7-agent CR with two fix rounds + three confirmation rounds → converged
to zero blocking findings.

## Follow-ups (not in this PR)
- Fleet-wide reasoning-id rename for the remaining backends
(ag2/crewai-crews/spring-ai/langgraph-fastapi/langroid/mastra/strands)
so their reasoning-display cells get probed.
- aimock hardening filed upstream: CopilotKit/aimock#253 (validate
Anthropic extended-thinking request invariants) and #254
(model-capability-aware reasoning emission).
- Minor robustness: `id(block)`→monotonic counter for reasoning message
ids; symmetric unparseable-reasoning warning on the chat-completions
transport; reuse Anthropic SDK `*BlockParam` types.

## Test plan
- [ ] CI green
- [ ] Post-merge: showcase `:latest` rebuild + staging redeploy, then
confirm the d5/d6 reasoning cells (chain + reasoning-display) render
green on the dashboard for the 5 fixed backends; the 3
documented-limitation backends remain red-but-documented.
2026-06-08 11:31:08 -07:00
Jordan Ritter c3bd60c245 fix(showcase): symmetric target-scoping for promote set-parity
check_service_set_parity scoped only the staging-only arm to the
single-service target; the prod-only arm was still computed over the
full fleet, so any prod-only service (e.g. a deprecated harness-legacy)
REFUSEd every unrelated single-service promote — the exact mirror of
the bug #5324 fixed. Scope both arms to the target for single-service
promotes; full-fleet promotes (target nil) keep both arms at full
strictness.

Tests: add prod-only tolerance red-green test, strengthen the
target-absent test to assert target-scoping (unrelated staging-only
sibling ignored), rewrite the stale snapshot-narrowing comments to
describe the real fleet_*/& [target] contract, drop the dead
FLEET_PUBLIC_PROD_HOSTS constant, and renumber the ivar-lint allowlist
for the one-line shift in bin/railway.
2026-06-08 11:26:31 -07:00
Jordan Ritter 8778545c4d fix(showcase): target-scope promote staging-only set-parity REFUSE
A single-service `promote <svc>` ran check_service_set_parity over the
FULL staging vs FULL prod fleet and REFUSEd whenever staging carried any
service prod lacks. The live staging fleet legitimately contains 13
staging-only services — harness-workers (SSOT-modeled) and 12 starter-*
demos — so an otherwise-clean single-service promote (e.g. docs) is
blocked with `REFUSE: services in staging not in prod`.

Scope the "staging not in prod" REFUSE to the promote TARGET when a
single-service promote is in effect (intersect the staging-only set with
[target]). The target-absent-from-prod footgun still REFUSEs (target is
in the intersection), the "prod not in staging" arm is unchanged, and
full-fleet promotes (no --service) retain full strictness.

Complements #5322.
2026-06-08 11:26:31 -07:00
github-actions[bot] 9e668797a4 style: auto-fix formatting 2026-06-08 11:15:14 -07:00
Jordan Ritter d865bfcc03 docs(showcase): document genuine reasoning SDK limitations + correct reasoning demo docs 2026-06-08 11:15:13 -07:00
Jordan Ritter 4875150001 fix(showcase): rename reasoning demo ids + restore probe coverage 2026-06-08 11:15:13 -07:00
Jordan Ritter d445b19642 fix(showcase): forward valid AG-UI reasoning role on claude-sdk-typescript (+@ag-ui/client 0.0.48) 2026-06-08 11:15:12 -07:00
Jordan Ritter f425077ea7 fix(showcase): emit reasoning via OpenAI Responses API on llamaindex 2026-06-08 11:12:57 -07:00
Jordan Ritter 73bfaeac29 fix(showcase): emit reasoning via chat-completions adapter on built-in-agent 2026-06-08 11:12:57 -07:00
Jordan Ritter 71df4222fe fix(showcase): emit native reasoning_content on agno reasoning handler 2026-06-08 11:12:57 -07:00
Jordan Ritter 8ae72bd2c7 fix(showcase): emit native reasoning (REASONING_MESSAGE_*) on claude-sdk-python reasoning agents 2026-06-08 11:12:57 -07:00
Jordan Ritter c413ec3ddb fix(showcase): report verify-prod=skipped (not success) when prod was never probed
When a promote fails, the succeeded-service set is empty, so verify-prod
hits its skip branch (`exit 0`). The GitHub job result is therefore
`success`, and the notify step rendered `verify-prod=success` in the
#oss-alerts Slack message — a misleading green, since prod was never
probed.

verify-prod now exports a `status` output: `success` after a real probe
passes, `skipped` on the empty-CSV skip. notify reads that output (via
the new bats-tested verify-prod-display.sh) instead of the raw job
result, so the Slack line accurately reads `verify-prod=skipped` vs
`success` vs `failure`. A genuine probe failure / contract violation
exits non-zero (job result `failure`, status never written), and the
display falls back to the job result. Slack formatting is unchanged.

Extracts the display mapping into showcase/scripts/verify-prod-display.sh
(mirroring promote-fleet.sh) with red-green bats coverage, and adds it to
the showcase_validate.yml shellcheck step.
2026-06-08 10:45:07 -07:00
Jordan Ritter b43d64ea49 fix(showcase): tolerate per-service "ServiceInstance not found" in promote snapshot
build_snapshot enumerates every project service and queries each one's
serviceInstance. The only guard was `next if inst.nil?` — it handled a
NULL result but not a THROWN `GraphQL: ServiceInstance not found` error
(a half-deleted service that still appears in the project service list
but has no instance in the env). That error bubbled to Railway.run's
top-level `rescue GraphQL::Error` and aborted the ENTIRE promote with an
opaque exit 2 before any preflight/divergence logic ran (run 27144525566
killed the docs promote this way).

Scope the rescue narrowly to ONLY the per-service "ServiceInstance not
found" message — log+skip that one service exactly like the nil case —
so every other GraphQL failure (auth, rate-limit, schema drift) still
propagates fail-loud. Adds red-green coverage: a single thrown not-found
is skipped (healthy services still snapshot), while an unrelated GraphQL
error still raises.
2026-06-08 10:45:07 -07:00
Ran Shemtov 237671aaaa fix(showcase): render AgentCore deploy pages with framework-scoped tabs (#5319) 2026-06-08 19:04:29 +02:00
Ran Shemtov 523396de96 Merge branch 'main' into chore/fix-agentcore-docs 2026-06-08 18:58:14 +02:00
Ran Shem Tov c0ea5078f9 fix(showcase): render AgentCore deploy partial with framework-scoped command tabs
The /strands/deploy-agentcore and /langgraph/deploy-agentcore pages
rendered only their title — the body was empty. <Content> resolved to a
dead stub in the MDX component registry that rendered nothing, despite
the content being authored in the shared agentcore partial.

- mdx-registry.tsx: replace the dead Content stub with a dedicated
  component that renders the agentcore partial via PartialLoader and
  threads the page's framework into MDX scope.
- mdx-registry-loader.tsx: PartialLoader accepts an optional scope,
  forwarded to MDXRemote options.scope so partials can read bare scope
  identifiers (next-mdx-remote binds scope as module identifiers, not as
  the rendered component's props).
- agentcore/index.mdx: reference {framework} (bare scope var) instead of
  props.framework so AgentCoreCommandTabs collapses to the single
  relevant framework per page.
2026-06-08 17:18:44 +02:00
Ran Shemtov 37cdfe05ba feat: update all dependencies to use latest a2ui implementation features (#5314)
Upgrade of dependencies so the latest changes (mainly a2ui stuff) from
ag-ui are reflected here
2026-06-08 15:46:15 +02:00
Ran Shem Tov 5b821a44dc chore: fix showcase agentcore link per framework 2026-06-08 12:50:46 +02:00
Ran Shem Tov c85c140f05 feat: update all dependencies to use latest a2ui implementation features 2026-06-08 12:09:20 +02:00
Jordan Ritter 9111a1f2ac fix(showcase): flap-band — cold-start retry + honest fleet/dashboard surface-state (#5313)
## Summary

- **Cold-start retry before fast-fail (#71), gated to plain-fill turns**
— `conversation-runner` now performs a bounded turn-1 `page.reload()`
retry when an error banner appears on a cold start, recovering transient
boot flaps. The retry shares the single turn deadline (no ~2× budget
blowup, FF20) and is gated to plain-fill turns so it never masks a real
failure: a banner that survives the reload still fast-fails.
- **Fleet teardown surface-state honesty** — adds a
`worker-reclaimed-pending` comm-error kind, a control-plane SIGTERM
drain path that distinguishes graceful teardown from a crash (#70), and
pre-dispatch warm-up health pings (#72). Graceful Railway teardown is
now reported as pending, not as a red.
- **Dashboard pending-surface never masks a real red** — `cell-model`
decodes comm-errors severity-first, guards against stale `observedAt`,
and renders a dedicated pending chip. A pending surface can never
override a genuine red, while transient teardown noise resolves to
pending instead of flapping.

**Dashboard-green impact:** kills false flaps on Railway teardown
WITHOUT masking real reds, and bounds the cold-start retry so a
genuinely-broken cold start still surfaces.

## Test plan

- [x] harness `tsc --noEmit` (clean)
- [x] harness vitest — conversation-runner (46), fleet contracts +
control-plane/job-producer + queue-client (123 tests across 4 touched
suites, all passing)
- [x] shell-dashboard `tsc --noEmit` (clean)
- [x] shell-dashboard vitest — cell-model + depth-chip (148 tests, all
passing)
- [x] oxfmt `--check` clean on all changed files

## Known follow-ups

All pre-existing, tracked for a separate PR (not introduced by this
change):

- `priority` / `leaseSeconds` dead knobs in the fleet contracts
- `modelsEqual` does not compare `jobId`
- `DepthChip` switch is not exhaustiveness-checked
- unknown-state should map to a gray cell in `cell-model`
- `createRailwayAdapter` stale JSDoc
- no-reload-retry test-fake edge case

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-07 22:47:09 -07:00
Jordan Ritter 042e79e259 fix(showcase): dashboard pending-surface never masks a real red — severity-first comm-error decode, stale-observedAt guard, pending chip 2026-06-07 22:40:13 -07:00
Jordan Ritter d814fb816a fix(showcase): fleet teardown surface-state — worker-reclaimed-pending kind, control-plane SIGTERM drain, warm-up health pings 2026-06-07 22:40:13 -07:00
Jordan Ritter d0f2582842 fix(showcase): cold-start retry before fast-fail (#71), gated to plain-fill turns 2026-06-07 22:40:12 -07:00
Jordan Ritter 8f6f1b8f42 fix(showcase): backend-backlog — agno/forwarding/dotnet/spring-ai/mastra/built-in-agent (#5312)
## Summary

Backend-backlog integration branch consolidating dashboard-green fixes
across six showcase integration areas. Recommitted by area of concern (6
logical commits) for review.

- **agno** — Pairs HITL tool messages (orphan tool-results dropped,
paired retained, falsy-id messages preserved); surfaces reasoning
`RUN_ERROR` and forwards `TOOL_CALL_RESULT`; wires the reasoning route
via `attach_reasoning_route`; pins the asyncio event loop in
`entrypoint.sh` so the executor-ctxvar header-forwarding shim is
effective under uvloop (previously header forwarding silently no-op'd on
the uvloop policy). Repositions the 13 `@ts-expect-error` annotations
onto the erroring agent-entry lines in the route handlers. **This is the
core of the dashboard-green impact** — reasoning demos now emit
correctly and forwarded headers reach the secondary OpenAI call.
- **header-forwarding shims** — Hardens the `_header_forwarding.py` shim
across all 10 python integrations (agno, ag2, claude-sdk-python,
crewai-crews, google-adk, langroid, llamaindex, ms-agent-python,
pydantic-ai, strands): fails loud on hook-install failure (logged at
ERROR not INFO) instead of silently degrading, and uses greppable
async-client detection with a low-confidence name-match fallback
breadcrumb.
- **ms-agent-dotnet** — `A2uiSecondaryToolCaller` fails loud on a
missing API key, corrects error misclassification (proper error
mapping), and guards malformed secondary-tool responses.
- **spring-ai** — Hardens the controller error-path lifecycle across 8
controllers + `PropagatingLocalAgent`: emits `RUN_STARTED` before
`RUN_ERROR` (correct ordering) and finalizes the run on the error path.
- **mastra** — gen-a2ui tool now throws instead of silently returning,
adds a role normalizer, bumps cmdk; corrects probe-doc YAML comments (3
probe configs); enforces baseline tag invariants in the shell-dashboard
`validateCell` path.
- **built-in-agent** — Pins `@copilotkit/runtime` to `1.59.4` and drops
floating tanstack pins; removes the now-obsolete SSOT override from
`showcase-canonical-pins.json`.

## Test plan

- [x] agno typecheck (`tsc --noEmit`) — clean
- [x] agno pytest (3 backend test files, agno 2.6.12 venv) — 17/17
passed
- [x] shell-dashboard vitest (`baseline-types`) — 27/27 passed
- [x] mastra `next build` — success
- [x] built-in-agent `next build` — success
- [x] validate-pins ratchet — FAIL=63 == baseline, hash matches
committed `fail-baseline.json`
- [ ] .NET `dotnet test` (ms-agent-dotnet) — run in CI
- [ ] spring-ai `mvn test` — run in CI

## Known follow-ups

Deferred pre-existing items, tracked for a separate PR (not in scope
here):
- agno error-handling hardening (broader RUN_ERROR coverage beyond
reasoning)
- StreamingToolAgent tool-id / streamed-args handling (spring-ai)
- mastra A2UI flat-shape normalization
- reasoning double-render
- a2ui docstring loop-pin wording clarification
- spring-ai bridge `RUN_STARTED` double-emit edge (proper guard
deferred; the double-emit revert is included here, the guard is not)

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-07 22:35:28 -07:00
github-actions[bot] c023a03fac style: auto-fix formatting 2026-06-08 05:23:16 +00:00
Jordan Ritter b284d89615 fix(built-in-agent): pin @copilotkit/runtime 1.59.4 + tanstack; remove obsolete SSOT override 2026-06-07 22:19:51 -07:00
Jordan Ritter fe3af73086 fix(mastra): gen-a2ui throw + role normalizer + cmdk bump; probe-YAML comment fixes; dashboard validateCell tag enforcement 2026-06-07 22:19:47 -07:00
Jordan Ritter 54e426c2dd fix(spring-ai): controller error-path lifecycle (RUN_STARTED ordering, finalize) 2026-06-07 22:19:41 -07:00
Jordan Ritter e89581d12c fix(ms-agent-dotnet): A2uiSecondaryToolCaller fail-loud + error mapping + guards 2026-06-07 22:19:30 -07:00
Jordan Ritter 154c1cfab1 fix(showcase): harden header-forwarding shims (fail-loud, async-detect) across python integrations 2026-06-07 22:19:26 -07:00
Jordan Ritter f0fe456a43 fix(agno): HITL tool-pairing, reasoning RUN_ERROR/TOOL_CALL_RESULT, reasoning route wiring, uvloop loop-pin for header forwarding 2026-06-07 22:19:19 -07:00
Jordan Ritter ebe80e257e fix(showcase): dashboard-green follow-ups (byoc race, D4 probe, declarative, beautiful-chart) (#5310)
## Summary
Follow-ups to #5309 that further lower the d5/d6 dashboard red floor:
- **byoc browser-pool race:** guard `newContext()` against a
disconnected shared browser (the d5 contention source).
- **D4 probe:** `networkidle`→`load` — networkidle never settles on
CopilotKit's persistent-SSE demo pages, so D4 timed out locally and
gated every cell red on the local rig; staging unaffected but the fix is
correct and unblocks local==staging visual testing.
- **strands + spring-ai declarative-hashbrown/json-render:** dedicated
backend agents/controllers + tuned prompts + regenerated d6 fixtures +
manifest coherence (json-render is a real byoc-feature-type cell, at
parity with strands).
- **beautiful-chart:** port the pie/bar/scheduler generative-UI
renderers to built-in-agent + claude-sdk-python (backends already emit
the tool-calls).
- **cleanup:** remove the temporary x-diag-probe instrumentation (keeps
the permanent CVDIAG + the #5309 forwarding fixes).

Showcase-only; no package releases.

## Test plan
- [ ] CI rebuilds :latest; redeploy
- [ ] PB re-pivot:
byoc/declarative-hashbrown/json-render/beautiful-chart-chart cells green
on the fixed backends
- [ ] d5 contention sawtooth reduced (byoc guard)

## Deferred (separate follow-ups)
- Flap-band: decouple PB sampling from live run / detector cold-start
retry / warm+pool (the ±70 churn band)
- #67 backlog: harness accounting (d6 hung-teardown, d4 body-fallback
false-green), forwarding-shim hardening, spring-ai controller error-path
lifecycle (class-wide), built-in-agent @copilotkit/runtime pin
2026-06-07 14:13:01 -07:00
Jordan Ritter 29fc959052 chore(showcase): ratchet pin-drift baseline for mastra ai v5 bump 2026-06-07 14:07:06 -07:00
Jordan Ritter ca2bc2c152 fix(showcase): specialize strands + spring-ai declarative-hashbrown/json-render backends (+ fixtures, manifests, controller guards) 2026-06-07 13:53:28 -07:00
Jordan Ritter 8d9dc82606 fix(showcase): render beautiful-chat chart/scheduler generative-UI on built-in-agent and claude-sdk-python 2026-06-07 13:53:19 -07:00
Jordan Ritter ffa51c7ee8 chore(showcase): remove the temporary x-diag-probe instrumentation 2026-06-07 13:53:07 -07:00
Jordan Ritter 6b924a7a44 fix(showcase): use waitUntil:load for the D4 chat-roundtrip probe (networkidle never settles on persistent-SSE pages) 2026-06-07 13:53:07 -07:00
Jordan Ritter ebd5f53549 fix(showcase): guard browser-pool context acquisition against a disconnected shared browser (byoc) 2026-06-07 13:52:57 -07:00