Commit Graph

4085 Commits

Author SHA1 Message Date
Jordan Ritter a0ef5e7ef2 fix(showcase): close env-registry cross-wire/orphan/case/prototype gaps, trim REDEPLOY_SUMMARY_JSON
assertEnvRegistryConsistent gains four clauses: (iv) a key present in
both ENV_IDS and ENV_ID_BY_NAME must carry the same env-id (ENV_IDS.prod
drifted to the staging id previously passed every clause while
resolveEnv("prod") silently returned staging); (v) every ENV_IDS env-id
must be carried by a canonical name (was only caught lazily in
resolveEnv); (vi) registry keys must be trim().toLowerCase()-normalized
(resolveEnv lowercases input, so a non-lowercase spelling is registered
but unreachable); (vii) no registry key may be an Object.prototype
property name.

expandImageConsumers' per-entry env skip-check becomes an own-property
test (Object.hasOwn) for uniformity with every other lookup in the file.

REDEPLOY_SUMMARY_JSON is trimmed before the set-but-empty branch so a
whitespace-only value hits the loud warn path instead of attempting a
JSON write against a garbage path.
2026-06-10 09:28:09 -07:00
Jordan Ritter e56c4db155 docs(showcase): correct SSOT field/comment claims in railway-envs
- ciBuilt field doc: pocketbase IS showcase-CI-built — only webhooks
  remains out-of-band; keep the MUST-NOT-touch claim for webhooks only.
- gateValidated field doc: true for every service EXCEPT the two
  gateIgnore entries (harness-workers, harness-legacy), not "every
  service".
- CI_BUILT_SERVICES comment: also names the excluded non-CI-built
  harness-workers and harness-legacy alongside webhooks.
- legacyJsonCompat doc: it is bin/railway's EXPECTED_DOMAINS derivation
  that filters out *.up.railway.app hosts (no "parity test rejects"
  claim contradicting the placeholder data below).
- webhooks entry comment: a manual service=all build dispatch MAY
  bounce webhooks staging (the skip_build slot still reports success,
  entering the matrix ∩ success-set redeploy scope); only the
  push-driven default scope is guaranteed to leave webhooks untouched.
2026-06-10 09:06:55 -07:00
Jordan Ritter 3cf06a7e49 fix(showcase): uniform own-property accessors, new SSOT invariants, fail-loud exit default (redeploy scripts)
- railway-envs: route envsFor/instanceIdFor/domainFor/probeEnabled (and
  repoNameFor) through shared getEntry/getEnvCfg own-property helpers so
  inherited Object.prototype keys on either axis produce the curated
  error (or probeEnabled's contract false) instead of raw TypeErrors,
  silent undefined, or a spurious probe=true.
- railway-envs: two new module-load invariants (synthetic-map
  injectable): assertEnvRegistryConsistent (per-service env keys are
  registered canonical names; ENV_ID_BY_NAME env-ids unique; every
  canonical name has an ENV_IDS spelling) and
  assertServiceAndInstanceIdsUnique (serviceId unique per entry,
  instanceId globally unique).
- redeploy-env: invert the exit-code policy to fail-loud by default —
  any env except the documented staging carve-out exits non-zero on
  per-service failure (a future preview/canary env inherits fatal
  semantics instead of silently swallowing failures).
- redeploy-env: sanitize per-service THROWN error messages through
  sanitizeErrorBody before recording; flatten bare \r in the summary
  table escape.
- redeploy-env: warn on set-but-empty REDEPLOY_SUMMARY_JSON; reject
  flag-like --services CSV parts (both forms) and a flag-like first
  argument (missing env); derive usage env lists from ENV_IDS.
- tests: prototype-key sweep across all accessors, invariant
  positive/negative coverage, third-env exit-code pin, sanitization and
  CLI-guard coverage, makeLiveRedeploy !res.ok and non-true mutation
  branches.
2026-06-10 09:06:55 -07:00
Jordan Ritter 171943469a fix(showcase): guard reaper against reserved 'showcase' record + isolate review fixes
Cross-session review fixes for the --isolate machinery (one concern:
source + test + docs).

1) Reaper reserved-name guard (critical): _reap_isolate_slot trusted
   slot records — a record naming 'showcase' (corrupt, or written by an
   older CLI version before apply_isolation reserved the name) passes
   the charset regex, so the reap ran `docker compose -p showcase down
   --remove-orphans --volumes` against the LIVE default stack,
   destroying the PocketBase named volume. The reserved name now gets
   the same treatment as the path-traversal guard: warn (naming the
   record and why it is dangerous) and leave the slot intact for manual
   inspection — no compose-down, no state removal.

   Call-site enumeration: _reap_isolate_slot's sole caller is
   _sweep_isolate_slots, at 3 sites (dead-PID reap, project-recorded/
   no-owner reap, age-fallback reap), all passing
   "$slot_entry" "$slot_proj" — all three flow through the new guard
   identically.

   Red-green: the new bats test ("a slot whose project record reads the
   RESERVED 'showcase' is left intact...") was run against the UNFIXED
   code first and FAILED — the sweep logged "Attempting to reclaim
   stale slot 0 (project showcase has no live containers and no
   recorded owner)" and reaped the slot. It passes with the guard.

2) .iso-bak restore race: two concurrent runs can both see a stale
   backup; the loser's mv is the FINAL command of its `[ -f ] && mv`
   AND-list, so its failure trips set -e and kills the CLI pre-claim
   with a raw error. Both mv's now carry `2>/dev/null || true` — the
   survivor's restore wins, the loser proceeds with restored originals.

3) Keep-test absence regexes greped only the `--project-name <name>
   down` spelling; the reaper's own downs use `-p <name> down`, so a
   keep-branch regression via the -p form passed undetected. Both keep
   absence assertions now match `(--project-name|-p) <name> down`.
   Mutation-verified: a temporary -p-form compose-down added to the
   keep branch made BOTH broadened tests FAIL; reverted, suite green.
   (All other absence assertions use the word-matched generic
   `compose ... down` regex, which already covers both spellings.)

4) RUNBOOK.md/DEBUGGING.md contradicted shipped code: the manual
   teardown was quoted without --volumes plus notes claiming
   `down --remove-orphans` leaves named volumes (the shipped survival
   notice and every teardown path include --volumes), and the name rule
   was documented as `[a-z0-9_-]+` (actual: starts with [a-z0-9], then
   [a-z0-9_-], uppercase normalized with a warn, 'showcase' reserved).
   Both updated to the shipped semantics; the now-redundant separate
   `down --volumes` snippets removed.

Verification: full `bats showcase/scripts/__tests__/` green (60 tests);
shellcheck on _common.sh shows no new warnings vs baseline
(pre-existing SC2034/SC2115 only, line-shifted).
2026-06-10 08:58:28 -07:00
Jordan Ritter bee835c4b0 chore(showcase): SSOT invariant hardening + doc accuracy (redeploy scripts) 2026-06-10 08:28:16 -07:00
Jordan Ritter c71d183247 fix(showcase): unify env-name authority + fail-loud SSOT accessors (redeploy scripts)
Confirmation-CR bucket-(a) fixes for PR #5353 — one coherent concern:
env-name resolution has exactly ONE authority (the ENV_IDS /
ENV_ID_BY_NAME registries) and SSOT accessors fail loud instead of
silently returning wrong values.

- runRedeploy: resolve envId via ENV_ID_BY_NAME with an Object.hasOwn
  guard + fail-loud throw listing the registered envs. Removes the
  hardcoded `prod`/`staging` pair check and PRODUCTION/STAGING ternary
  that contradicted the SSOT's documented open-env contract ("a new env
  needs only a registry entry"); the registry lookup subsumes it.
- resolveEnv: derive resolution entirely from the registries (ENV_IDS
  spellings -> env-id -> canonical ENV_ID_BY_NAME name) instead of its
  own hardcoded synonym chain. Behavior identical for
  prod/production/staging; still throws on unknowns, and now also
  throws on a mis-wired registry (a spelling whose env-id has no
  canonical name).
- serviceForDispatchName: fix the docstring's false "CI-built service"
  claim — it does no ciBuilt filtering and tests pin the unfiltered
  behavior (the non-CI-built webhooks resolves).
- repoNameFor: fail loud (consistent with instanceIdFor/domainFor)
  instead of silently echoing the service name — the exact
  silently-wrong-GHCR-name class this PR's hardening targets. Throws on
  unknown service, on an env not registered in ENV_ID_BY_NAME
  (unnormalized synonyms like "production"), and on a registered env
  the service does not declare. Keeps the documented default (the
  service name) for declared envs without an override.
  Call-site enumeration confirming nothing relies on the old fallback:
    - verify-railway-image-refs.ts:523 — iterates the entry's DECLARED
      environments keys, registry-filtered, SSOT-matched service names
    - __tests__/railway-envs.golden.test.ts:81 — iterates envsFor(name)
      (declared envs only) over real SSOT keys
    - railway-envs.test.ts repoNameFor cases — dual-env services,
      prod/staging only
    - __tests__/verify-railway-image-refs.test.ts:264 — FIVE_NEW keys,
      all dual-env
- resolveTargetServices: throw when an explicitly-provided services
  list resolves to zero entries (whitespace-only programmatic input)
  instead of letting runRedeploy exit 0 having redeployed nothing; the
  default undefined -> full CI-built scope is unchanged.
- makeLiveRedeploy: add signal: AbortSignal.timeout(30s) so a hung
  Railway API records a per-service FAIL instead of stalling CI, and
  pass GraphQL errors[].message through sanitizeErrorBody for
  consistency with the HTTP-error path. Exported for direct unit tests.

Red-green: 11 new tests (open-env registry resolution incl. a
runtime-registered hypothetical env, repoNameFor negatives,
empty-resolution throw, abort-signal presence, GraphQL error
sanitization) all failed against the old code; full showcase/scripts
suite green (50 files, 1807 tests) + tsc --noEmit clean.
2026-06-10 08:28:16 -07:00
Jordan Ritter c4a5f9d57c fix(showcase): @ag-ui 0.0.55 currency, reasoning ports, google-adk demos (#5348)
## Summary

Combined landing of three browser-verified showcase work waves (37
commits, per-integration grouping preserved), plus a full multi-round
code review with all mandatory fixes folded in.

**Wave 1 — @ag-ui currency (9/9 integrations).** Bump `@ag-ui/*`
frontend deps to exact `0.0.55` for: llamaindex, agno,
claude-sdk-python, pydantic-ai, strands, claude-sdk-typescript,
ms-agent-python, ms-agent-dotnet, google-adk. Exact pins land across 13
`package.json` files.

**Wave 2 — reasoning emission ports (4/4, ref #76).** Port
reasoning-message emission to ag2, crewai-crews, langroid, and spring-ai
(each also bumped to `@ag-ui` 0.0.55). Adds per-integration
`reasoning_agent` (Python) / `ReasoningController` (spring-ai Java) and
wires the copilotkit route.

**Wave 3 — google-adk demo parity (3/4).** Port `hitl`,
`threadid-frontend-tool-roundtrip`, and `gen-ui-interrupt` demos to
google-adk for parity with the gold reference, plus d6/aimock fixtures.
The `interrupt-headless` demo is intentionally kept `not_supported`
(needs aimock `customEvents` support + an ADK interrupt route — tracked
as a known upstream gap).

## Code review

A 5-round unbiased multi-agent CR loop ran against the combined diff and
converged at zero mandatory findings; the bucket-(c) promotion audit
came back clean. Key fixes folded into the branch:

- **Client error-leak hardening** in the crewai, langroid, and spring-ai
reasoning routes — internal exception details no longer leak to the
client, with `errorId` correlation between client response and server
logs (also applied to google-adk).
- **Protocol-correct reasoning error paths** — on failure the emitters
now close any open frames, emit a generic `RUN_ERROR`, and never emit
`RUN_FINISHED` after `RUN_ERROR`; verified against `@ag-ui/client`
`verifyEvents` semantics, with red-green tests.
- **`x-aimock-context` propagation** across spring-ai's async hop, so
fixture replay stays correct through the thread boundary.
- **Bounded reasoning executor** (no unbounded thread growth) and
**parse-failure observability** (failures surface in logs instead of
being swallowed).
- **Full multi-turn history threading** in the reasoning agents — prior
turns are now forwarded to the LLM, with single-turn byte-equality
preserved so existing aimock fixtures stay valid; covered by tests.
- **Phantom deps declared** — `@copilotkit/shared`/`@ag-ui` deps that
were imported but undeclared in claude-sdk-python and
ms-agent-dotnet/python are now in their `package.json`.
- **`@copilotkit/core` `"latest"` overrides pinned** to `1.59.4`.
- **Reasoning parity token-identity test** across the 3 ported Python
integrations (ag2/crewai-crews/langroid) so their reasoning streams
can't silently drift.
- **Ratchets tightened/updated**: shadow-collision ceiling ratcheted
down 151 → 123 (actual count), duplicate-fixture ceiling 288 → 290 (new
google-adk fixtures reuse standard prebuilt-probe pills,
runtime-disambiguated by demo route), and the `validate-pins` baseline
drops 57 → 48 (currency bumps retire 9 exact-pin FAIL lines).

The final commit is the `style: auto-fix formatting` bot commit
(formatting on the new Python test files).

## Test plan

- [x] Per-integration browser verification (each touched integration's
demos exercised end-to-end) with aimock fixtures added/updated.
- [x] crewai pytest — 113/113 pass.
- [x] `showcase/scripts` vitest — 1780/1780 pass.
- [x] `validate-parity` MUST checks — 19/19 pass, 0 fail.
- [x] `validate-pins` ratchet — exits 0 against the updated baseline
(FAIL-set hash matched).
- [x] D6 `gen-ui-custom` cells verified live (google-adk + agno).
- [x] CI fully green — per-integration Depot Docker `build-check`,
Python unit tests (3.10 + 3.12), commitlint, format, oxlint, Validate
Showcase, production-pinning lint.

## Follow-ups

Consolidated follow-up ledger (bucket c/d items):
https://www.notion.so/copilotkit/37b3aa38185281e5b871d0b907aaef71
2026-06-10 07:28:35 -07:00
Jordan Ritter 3c84d93fbd test(showcase): fixture-collision ceiling ratchets (duplicates 288->290, shadows 151->123) + honest overflow diagnostics 2026-06-10 07:11:26 -07:00
Jordan Ritter 2dbecf5f93 test(showcase): reasoning error-path, history, and parity coverage + forwarded-props agent stub 2026-06-10 07:11:25 -07:00
Jordan Ritter 8c31efcfa7 docs(showcase): accurate reasoning/demo comments and docstrings 2026-06-10 07:11:25 -07:00
Jordan Ritter bb3dc3e48f fix(showcase): empty-run diagnostics and content-coercion hardening in reasoning agents 2026-06-10 07:11:25 -07:00
Jordan Ritter 687c749779 fix(showcase): thread full conversation history through reasoning agents 2026-06-10 07:11:25 -07:00
Jordan Ritter c98dcf28d5 fix(showcase): spring-ai robustness — aimock header hop, bounded executor, parse-fail logging, threadId fallback 2026-06-10 07:11:25 -07:00
Jordan Ritter e46913d557 fix(showcase): protocol-correct reasoning error paths (frame close, generic RUN_ERROR, no RUN_FINISHED after RUN_ERROR) 2026-06-10 07:07:26 -07:00
Jordan Ritter 82cfa4ccd6 fix(showcase): guard reasoning user-input extraction against non-string content
_extract_user_input in the ag2, crewai-crews, and langroid reasoning agents
documented a str return but passed AG-UI message content straight through.
Multimodal content can be a list of parts, which would flow unmodified into
the single caller in each file (_run_reasoning_agent ->
messages=[{"role": "user", "content": user_input}]) sent to the OpenAI
chat-completions API. Coerce: str passes through, a list joins its text
parts (dict or attr form), anything else falls back to str().

Callers (one per file):
- ag2/src/agents/reasoning_agent.py:104 -> chat message at :119
- crewai-crews/src/agents/reasoning_agent.py:108 -> chat message at :123
- langroid/src/agents/reasoning_agent.py:103 -> chat message at :118
2026-06-10 07:07:11 -07:00
Jordan Ritter 84f7299024 fix(showcase): harden copilotkit route error responses (no message/stack leak, errorId correlation) 2026-06-10 07:07:03 -07:00
Jordan Ritter f60b68b925 chore(showcase): keep google-adk interrupt-headless unsupported (needs aimock customEvents + ADK interrupt route)
(cherry picked from commit aac16af8ee32172e6886a83bf66931c54091f180)
2026-06-10 07:06:55 -07:00
Jordan Ritter 0015075b51 feat(showcase): port gen-ui-interrupt demo to google-adk (Strategy-B parity)
(cherry picked from commit 6cf8fb2f3f7008be82fc2dd012456907b2572265)
2026-06-10 07:06:55 -07:00
Jordan Ritter 4dbd520c2d feat(showcase): port threadid-frontend-tool-roundtrip demo to google-adk (parity)
(cherry picked from commit 269cce835e9e92743b565ffd6ace91f73419c6dc)
2026-06-10 07:06:55 -07:00
Jordan Ritter a4fc8c5aac feat(showcase): port hitl demo to google-adk (parity with gold)
(cherry picked from commit 9527ec763f9093d96fe614c79c00769ceb20a1a8)
2026-06-10 07:06:55 -07:00
Jordan Ritter 96c2438fab fix(showcase): port reasoning emission + bump @ag-ui to 0.0.55 (spring-ai)
Add a dedicated Spring/Java ReasoningController (/reasoning/) that
reimplements the ag2 reasoning_agent.py BEHAVIOR: it makes a direct
streaming chat-completions call, reads the native delta.reasoning_content
channel (with a <reasoning>...</reasoning> regex fallback), and emits
RUN_STARTED -> REASONING_MESSAGE_START/CONTENT/END -> TEXT_MESSAGE_*
-> RUN_FINISHED so the CopilotKit reasoning slot mounts
[data-testid="reasoning-block"].

Spring AI's ChatClient drops delta.reasoning_content and the AG-UI Java
SDK has no REASONING_MESSAGE_* event types (only THINKING_*, which
@ag-ui/client drops), so the controller manages its own SseEmitter and
writes the reasoning frames as raw JSON matching the @ag-ui/client 0.0.55
wire schema. Header forwarding (x-aimock-context) rides the existing
WebClientConfig exchange filter. Wire route reasoning-custom/-default (plus
legacy aliases) to /reasoning/, mirroring ag2's reasoningAgentNames, and
bump @ag-ui/client ^0.0.43 -> 0.0.55 for REASONING_MESSAGE_* decode support.

(cherry picked from commit 4d183371c489013f0a7bcce1a447078164974aef)
2026-06-10 07:06:54 -07:00
Jordan Ritter 373310e8f3 fix(showcase): port reasoning emission + bump @ag-ui to 0.0.55 (langroid)
(cherry picked from commit 7c3d02d92e55d1e9a5cb89d182658620c0eed99f)
2026-06-10 07:06:54 -07:00
Jordan Ritter fdf3b29999 fix(showcase): port reasoning emission + bump @ag-ui to 0.0.55 (crewai-crews)
(cherry picked from commit 0286f413d2ceb52744c917eb9fe9fdc5f28011f2)
2026-06-10 07:06:54 -07:00
Jordan Ritter 0838059bc8 fix(showcase): port reasoning emission + bump @ag-ui to 0.0.55 (ag2)
(cherry picked from commit 4b27db5d781f68e1658955bcd23f667e63d400b3)
2026-06-10 07:06:54 -07:00
Jordan Ritter cd0156cc61 fix(showcase): dependency pin hygiene — exact overrides, phantom deps, validate-pins ratchet 2026-06-10 07:06:42 -07:00
Jordan Ritter 207f24035a fix(showcase): bump google-adk @ag-ui/* to 0.0.55
(cherry picked from commit 09c22e070ac60b936a4a91a0fd3f34879604de77)
2026-06-10 07:06:30 -07:00
Jordan Ritter 54d1d81394 fix(showcase): bump ms-agent-dotnet @ag-ui/* to 0.0.55
(cherry picked from commit aa9f7e5e99085383ce60859c9c37692211ba8279)
2026-06-10 07:06:30 -07:00
Jordan Ritter 914db1288b fix(showcase): bump ms-agent-python @ag-ui/* to 0.0.55
(cherry picked from commit 6254bdad0b54a7930334129c715a1f18176fb094)
2026-06-10 07:06:30 -07:00
Jordan Ritter c99e240bbc fix(showcase): bump claude-sdk-typescript @ag-ui/* to 0.0.55
(cherry picked from commit 9996a01c17432d2a5fc508884156ffad11269c4b)
2026-06-10 07:06:30 -07:00
Jordan Ritter 7e0adf1b44 fix(showcase): bump strands @ag-ui/* to 0.0.55
(cherry picked from commit 30540f6c7ae1bef7b1ada449e9d0c83782b438d0)
2026-06-10 07:06:29 -07:00
Jordan Ritter 860f9fd19e fix(showcase): bump pydantic-ai @ag-ui/* to 0.0.55
(cherry picked from commit 4aa62ca567d8150c50309ca5290dcdc91f2f6c71)
2026-06-10 07:06:29 -07:00
Jordan Ritter 589bf36310 fix(showcase): bump claude-sdk-python @ag-ui/* to 0.0.55
(cherry picked from commit 776e91445a7ca47f584f9ca75a22754b1302c91b)
2026-06-10 07:06:29 -07:00
Jordan Ritter 37d4a35b69 fix(showcase): bump agno @ag-ui/* to 0.0.55
(cherry picked from commit befb301b086e87b1d3317f78301fac1232945240)
2026-06-10 07:06:29 -07:00
Jordan Ritter a628a166d3 fix(showcase): bump llamaindex @ag-ui/* to 0.0.55
(cherry picked from commit d22f7e6ab585fa77c31d7f6b567a296b3e815a29)
2026-06-10 07:06:29 -07:00
Ran Shemtov 0a76aa1ff7 Merge branch 'main' into ran/oss-248-a2ui-shared-params 2026-06-10 09:29:57 +02:00
Jordan Ritter 8eefbcc0c6 docs(showcase): document XDG isolate state paths and real --keep semantics
- isolate state paths updated for the XDG migration
- --keep now documented as leaving the isolated stack standing
2026-06-09 18:18:10 -07:00
Jordan Ritter b57b2162c1 test(showcase): pin isolate invariants in bats — trap wiring, races, anti-vacuity
- real-trap-path tests for --keep (no simulated trap shortcuts)
- sweep/lock/tombstone race pins (heartbeat resurrection, lock
  takeover, duplicate-name TOCTOU)
- sentinel anti-vacuity discipline so trap tests cannot pass vacuously
- reap-order probe pinning live-slot protection during sweeps
- root/PID-reuse/DST guards for liveness and age checks
2026-06-09 18:18:03 -07:00
Jordan Ritter f2ed2b28ef fix(showcase): harden isolate state machine — --keep wiring, default-stack guards, registry concurrency
- ISOLATE_KEEP promoted to a global so --keep survives cmd_test return
  into the trap scope
- early-die and default-stack protection: failed --isolate setup no
  longer tears down the default stack; half-initialized state is
  cleaned up on the way out
- liveness/PID/age reaping signals with a sweep lock: heartbeat
  updates, own-pid lock release, tombstones, and a claim-then-verify
  duplicate-name guard close slot-registry races (TOCTOU, lock
  takeover, reap order)
- teardown robustness: --volumes on every compose down, failed-down
  runs preserve state for diagnosis, reap remnants get a compose-down,
  path-traversal guard, uniform rm guards under set -e
- name validation: --isolate names must start with a lowercase letter
  or digit; reserved name 'showcase' rejected (it aliases the default
  stack)
- fail-loud warning before pre-down of an existing stack; help text
  updated

Result of an 8-round, 7-agent code-review loop with red-green
verified fixes.
2026-06-09 18:17:55 -07:00
Jordan Ritter 0b097d0c89 fix(showcase): fail loud on unnormalized env in expandImageConsumers; prune redundant assertions 2026-06-09 18:17:21 -07:00
Jordan Ritter 26a43a794f fix(showcase): reject prototype-key service names; sync redeploy scope docs 2026-06-09 18:09:01 -07:00
Jordan Ritter 6f3f9ef736 fix(showcase): validate imageOf env overlap, prototype-safe target lookup; clarify expansion docs 2026-06-09 18:00:46 -07:00
Jordan Ritter 58c655beee docs(showcase): correct stale redeploy default-scope claims; stub summary env in tests 2026-06-09 17:52:16 -07:00
Jordan Ritter 2af5d691ef fix(showcase): redeploy image consumers (harness-workers) when their shared image is rebuilt
harness-workers runs the SAME showcase-harness GHCR image as the harness
scheduler but has ciBuilt:false (it has no build slot of its own), and the
CI staging redeploy scope was derived purely from ciBuilt — so a main-merge
rebuild of showcase-harness:latest only bounced the scheduler while the
workers silently kept running the stale image (PR #5352's worker-side fixes
never reached staging).

Model image consumption explicitly in the SSOT instead:

- railway-envs.ts: new optional `imageOf` field on ServiceEntry — the SSOT
  key of the ciBuilt service whose image this entry runs. Set
  `imageOf: "harness"` on harness-workers. New module-load invariant
  `assertImageConsumersValid` (fail-loud, same style as
  assertDispatchNamesUnique): imageOf must name an existing SSOT key, the
  target must be ciBuilt, and the consumer itself must not be ciBuilt.
- redeploy-env.ts: new `expandImageConsumers(names, env)` applied inside
  runRedeploy — the redeploy set becomes the resolved scope PLUS any
  service whose imageOf points at a service already in scope. Env-aware:
  a consumer only joins envs it declares, so the staging-only worker
  never enters a prod redeploy (prod behavior unchanged).

No workflow change needed: showcase_build.yml keeps passing the
built-and-successful dispatch_names; the script expands them. Gate
behavior (gateIgnore / image-ref gate), the generated JSON
(emit --check passes byte-identical), and the promote dropdown are all
untouched. harness-legacy deliberately gets no imageOf (pinned pre-fleet
digest; must not follow rebuilds).
2026-06-09 17:40:41 -07:00
Mike Ryan a0840c0a15 fix(shell-docs): keep docs headings block-level (#5334)
## Summary
- fix shell-docs heading anchors so docs headings remain block-level
rows
- prevent adjacent headings, like `## What the runtime provides`
followed by `### Authentication & security`, from rendering on the same
line
- add a `globals.css` regression guard against bringing back
`inline-flex` on `.docs-heading`

## Verification
- `npm run test -- src/app/__tests__/globals-css.test.ts` in
`showcase/shell-docs`
- `npm run lint` in `showcase/shell-docs`
- `npm run test` in `showcase/shell-docs`
- `npm run typecheck` in `showcase/shell-docs`
- `npm run build` in `showcase/shell-docs`
- local visual check at `http://localhost:3003/backend/copilot-runtime`
with Playwright: the H2 and H3 render as separate full-width rows
- lefthook pre-commit passed `check-binaries`, `lint-fix`,
`test-and-check-packages`, and commitlint

## Notes
- `npm run lint` still reports existing warnings unrelated to this
change.
- `npm run build` still emits the existing Turbopack/NFT warning from
`next.config.ts` / `llms-mdx`, but completes successfully.
2026-06-09 17:25:34 -07:00
Jordan Ritter e90afdd223 feat(harness): GC long-dead worker rows, fail loud on gc/stale misconfig, demote restartsAttempted 2026-06-09 16:52:11 -07:00
Jordan Ritter 194d6219bc test(harness): cover fleet-health GC of dead roster rows 2026-06-09 16:52:10 -07:00
Jordan Ritter 6c30f523d7 fix(harness): graceful worker drain — abort in-flight runs, abandon partials, deregister with a latched write chain 2026-06-09 16:52:04 -07:00
Jordan Ritter 57b36f2e0f test(harness): cover graceful drain, deregistration races, and drain/timeout quadrants 2026-06-09 16:52:04 -07:00
Jordan Ritter cdc1e90e4a fix(harness): per-run-unique X-Test-Id to stop aimock fixture-sequence desync 2026-06-09 16:51:58 -07:00
Jordan Ritter 3684bd42cf test(harness): pin per-run-unique X-Test-Id across d6/d5/d4 drivers 2026-06-09 16:51:57 -07:00