Mint a run_id per D5/D6 feature run, inject x-diag-run-id + seed x-diag-hops=harness alongside x-aimock-context, and do a post-run aimock-journal join (timeout-bounded) that records a cv-verdict row (tagged harness-d5/harness-d6) localizing whether the context header reached aimock. Add claim/worker_id/snapshot/pool-condition breadcrumbs and surface previously-swallowed catches.
Shared single-line CVDIAG log format (redacted 12-char header prefix), x-diag-run-id/x-diag-hops correlation header constants, and a best-effort PocketBase diag_events sink (anonymously HTTP-readable) so the CV/x-aimock-context propagation chain can be traced mid-incident without Railway log access.
## Summary
The Threads locked-state card in every integration example tells users
to add a CopilotKit Intelligence license with `copilotkit
add-intelligence`. That's the wrong command:
1. `add-intelligence` is intentionally not wired into the CLI dispatch
yet.
2. Even when implemented, it only drops the Intelligence overlay — it
does not issue a license (licensed behavior is deferred to ENT-725).
The command that actually issues a license key and writes it to `.env`
is **`copilotkit license`**.
## Changes
- `locked-state.tsx` in all 18 examples: `copilotkit add-intelligence` →
`copilotkit license` (oxfmt collapsed the now-shorter `<code>` element
onto one line)
- `threads-drawer.module.css` descriptive comments (11 examples):
"add-intelligence command" → "license command"
- `examples/integrations/_intelligence/` references to `copilotkit
add-intelligence` are untouched — those describe the overlay scaffold
command, not licensing.
## Verification
- `git grep add-intelligence -- 'examples/**'` only matches
`_intelligence/` scaffold docs
- `pnpm parity:check` — no new findings (threads-drawer/** is
allowedDivergence; only pre-existing next-env.d.ts warns)
Closes ENT-804
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The locked-state card told users to add an Intelligence license with
`copilotkit add-intelligence`, but that command only drops the
Intelligence overlay and does not issue a license (and is not yet wired
into the CLI dispatch). The command that issues a license key is
`copilotkit license`.
The pool-fleet worker's Railway service is named `harness-workers` (PLURAL),
but the SSOT keyed it `showcase-harness-worker` (singular). The image-ref gate
matches SSOT keys to Railway service names verbatim, so the gate reported
`harness-workers` as an untracked Railway service AND the stale singular key
matched nothing. Rename the SSOT key (and every test/fixture reference) to the
exact Railway name `harness-workers`.
It stays the staging-only, domainless, probe-disabled worker that runs the
shared `showcase-harness` image: serviceId c2aa8a0b-…, staging instance
362c1e37-…, ciBuilt:false, gateIgnore:true, no build slot (so no dispatchName),
single `staging` env with no domain. Add a focused test pinning that shape.
Counts are unchanged (29 services / 26 CI_BUILT) — this is a rename, not an
addition; both harness workers already existed on main.
Verified LOCALLY against Railway: verify-railway-image-refs reports
`54 env-scoped instances verified (2 skipped)` — 0 violations, 0 missing, 0
untracked (harness-workers reconciled, harness-workers + harness-legacy the 2
gateIgnore'd skips). emit --check zero drift, Ruby parity green (borrowed
.up.railway.app host is parity-excluded), full scripts suite + typecheck green.
Replace ServiceEntry's parallel prodInstanceId/stagingInstanceId/domains/probe/
repoNameOverride fields with a single environments: Record<string, {instanceId,
domain?, probe?, repoName?}> map plus a hoisted env-independent probeDriver.
EnvName becomes an open string backed by an ENV_ID_BY_NAME registry so
accessors resolve arbitrary env names; a single-env service (the staging-only
showcase-harness-worker) now simply omits the absent env instead of carrying a
placeholder ID/borrowed host.
Accessors instanceIdFor/domainFor/repoNameFor index environments[env]
(domainFor still throws on missing/scheme); add envsFor(name),
serviceEnvPairs(), and probeEnabled(name, env). Generalize the image-ref gate
(iterate each service's declared environments, resolve env-id via the
registry, sum/iterate missingByEnv over registry env names) and verify-deploy's
host->env reverse-map (envForTarget iterates environments).
Pure TS-internal: emit-railway-envs-json.ts projects the env-map back onto the
FROZEN legacy JSON shape (prodInstanceId/stagingInstanceId/domains/probe/
repoNameOverride) via a documented legacyJsonCompat shim for the two domainless
harness workers, so railway-envs.generated.json stays byte-identical and Ruby
(bin/railway) + workflow jq + the parity test are untouched. Verified: emit
--check zero drift, Ruby test_expected_domains_parity green, golden snapshot
toEqual proves byte-identical resolution for every real (service, env) pair,
full scripts vitest suite green, showcase scripts typecheck clean.
Serializes the fully-resolved {service -> env -> {instanceId, domain, probe,
driver, repoName}} projection for all 29 services x 2 envs via the public
accessors (instanceIdFor/domainFor/repoNameFor) + per-entry probe config,
frozen as a fixture. This is the behavior-preservation guard for the
forthcoming env-map (Option C) refactor: resolved values must stay
byte-identical before and after.
Wire the resource-snapshot writer back into the fleet worker path with
per-replica attribution, add the worker_id column migration, and stamp
the legacy single-process path's worker_id as "" (empty string) to match
PocketBase storage and the surrounding codebase convention.
Move gen-ui-interrupt + interrupt-headless from features: to
not_supported_features: across affected integration manifests, and align
the generate-registry/generate-catalog scripts tests to the resulting
wired-feature counts (derive expected lengths from the parsed manifest
rather than hardcoding pre-quarantine numbers).
Make the showcase dev tool faithful to staging by construction. Two changes:
1. `showcase test --d5/--d6` now drives the fleet CONTROL-PLANE (producer ->
probe_jobs queue -> worker -> result-aggregator) instead of the legacy
in-process runLevel() driver. The new cli/control-plane-run.ts replicates
the deep/full producer tick exactly as runControlPlane wires it
(createE2eDeepServiceEnumerator / createServiceEnumerator over
createJobProducer + createFleetQueueClient), enqueues one operator-triggered
tick, and polls local PocketBase for the run's terminal cells. The running
worker fleet claims + runs the driver + the aggregator writes the d5/d6
status cells, so the dev tool exercises the IDENTICAL wiring + concurrency
as staging. The old in-process path stays available behind `--direct`.
2. `showcase up --dev` adds a docker-compose.dev.yml overlay that bind-mounts
each integration's source and overrides the run command with a stack-aware
hot-reload entrypoint (shared/dev/dev-entrypoint.sh: uvicorn --reload for
FastAPI agents, langgraph dev for graphs, next dev for the frontend). Edit a
source file and the component reloads in place with no image rebuild. The
built-image mode remains the faithful/staging-equivalent default.
## Problem
D5 dashboard cells were systemically RED while D6 was green. The D5
assistant-no-response failure traced to: the backend LLM call to aimock
arrived without `x-aimock-context`, so aimock's strict mode returned a
503, which surfaced as an `agent_run_error_event`, which led to a 30s
timeout and a red cell.
## Root cause
NOT fixtures / keying / aimock — all verified good; D6 greens the same
cells on the same V1 route. The actual cause: D5 ran as a **separate
driver** (`e2e_deep` / `d5-single-pill`) whose code path lost
`x-aimock-context` on the shared fleet, whereas D6's path keeps it. This
operational divergence was exposed when the pool migration put the
duplicate D5 path onto the fleet.
## Fix
**D5 = D6-take-one.** D5 now runs the `d6-all-pills` driver scoped to
`D5_REPRESENTATIVES` via `representativeOnly` + `rowPrefix: "d5"`:
- enumerator supplies `extraDriverInputs`; CLI builds inputs through
`buildDeepInputs`
- worker reads `rowPrefix` to filter the `d5:<slug>` aggregate
- deleted the `e2e_deep` kind and `d5-single-pill.ts`
## Verification
- Full harness typecheck clean (except the pre-existing
`resource-snapshot-writer.test.ts` `toReversed` lib-target error already
on main)
- ~2097 harness tests green
- 7-agent CR + confirmation round converged (zero bucket-a findings)
## NOTE
Merge auto-deploys to staging (mutable `:latest`, no rollback). This fix
is operational — real validation is watching the d5 dashboard cells go
green post-deploy. A clean local single-run cannot reproduce the
concurrency-driven staging failure.
## Deferred follow-ups (non-blocking)
- gate worker `rowPrefix`-override on `driverKind`
- `X-Test-Id` `d5-` for D5
- deploy-churn `incapable[]` (pre-existing)
- tests for `representativeOnly` + deploy-churn and the demos-path
The e2e_deep kind was removed (D5 now runs the D6 driver), leaving three
browser driver families: e2e_d6, e2e_demos, e2e_smoke. Update the stale
"four browser driver families" docstring and its test comment.
Asserts the real D5 invocation shape buildDeepInputs stamps (both knobs
together): only D5_REPRESENTATIVES featureTypes run AND every emitted key
(per-cell d5:<slug>/<ft> + aggregate d5:<slug>) uses the d5: prefix. The
existing tests cover the knobs in isolation only.
buildDeepInputs now forwards manifest.not_supported_features (matching D6's
buildFullInputs) so local CLI D5 runs don't false-red architecturally-
unsupported features. The D5 thrown-error terminal key changes from
d5-single-pill-e2e:<slug> to d5:<slug> (the driver's own emitAggregate key
shape) so a hard driver throw surfaces as a RED D5 cell, not a blank row.
Also corrects D5-scope docs: representativeOnly keeps the representative
featureTypes per D5_REPRESENTATIVES, not "one pill per category".
Removes E2E_DEEP_DRIVER_KIND and "e2e_deep" from the worker-internal
closed driver-kind set (D5 runs the e2e_d6 driver now), updates the
x-test-id-headers guard to assert on the surviving d6-all-pills driver
(d5-single-pill.ts was deleted), and repoints a stale doc comment.
Repoints the D5 probe at the unified D6 driver. The fleet D5
enumerator now stamps driverKind=e2e_d6 with representativeOnly + a
"d5" rowPrefix; the CLI's buildDeepInputs carries the same inputs;
config/probes/e2e-deep.yml declares kind=e2e_d6. Drops the e2e_deep
driver registration (orchestrator + worker registry + BROWSER_KINDS +
the worker-internal kind set), keeping the D5 producer schedule/cadence
intact. The worker now honors driverInputs.rowPrefix when filtering the
aggregate side-row out of captured cells so a "d5:<slug>" aggregate
doesn't leak into the D6 entry's "d6:<slug>" cell capture. This
eliminates the separate D5 launcher path whose own launcher instance +
cadence systematically dropped x-aimock-context against the shared fleet
pool (aimock strict 503 -> red).
Adds two input knobs to the d6-all-pills driver so it can run as "D5
take-one": `representativeOnly` filters the feature matrix to the
D5_REPRESENTATIVES set, and `rowPrefix` ("d5" | "d6", default "d6")
threads the dashboard key prefix through every emitted per-cell and
aggregate PB row. The representatives map is injectable for testing.
Everything else (route, headers, conversation, pooled launcher) is
unchanged. Red-green unit tests cover both knobs.
## Summary
The fleet **control-plane** runs the 8 in-process HTTP probe families
(smoke, starter_smoke, image_drift, qa, aimock_wiring, version_drift,
pin_drift, redirect_decommission) on its scheduler, but the on-demand
trigger endpoint `POST /api/probes/:id/trigger` was only mounted on the
legacy `boot()` path. There was no way to fire a family immediately on
the control-plane — operators had to wait on the slow cron.
This wires the **same** `registerProbesRoutes` onto the control-plane's
`buildServer` call in `runControlPlane`, using the control-plane's
already-built `httpProbeRegistry` / `httpProbeConfigs` / `scheduler` /
`httpRunWriter` and `OPS_TRIGGER_TOKEN`. Enables instant verification
instead of waiting on crons.
- **Triggerable ids** (control-plane): the prefixed in-process HTTP
probe scheduler ids — `probe:smoke`, `probe:image_drift`, `probe:qa`,
etc. (whatever HTTP families are loaded). The router's `isProbeId` guard
resolves against `httpProbeConfigs` (keyed `probe:<cfg.id>`).
- **404'd**: browser-only families (`probe:e2e_smoke`,
`probe:e2e_demos`, `probe:e2e_deep`, `probe:e2e_d6`) — they are
worker-routed, never run in-process — plus unknown ids and the
producer's own scheduler entries.
- **Fail-safe token handling mirrors `boot()` exactly**:
`OPS_TRIGGER_TOKEN` unset → router omitted (route 404s); set-but-empty /
whitespace-only → fail-loud at boot (refuses to mount an insecure
route); missing bearer token → 401.
## Scope
Minimal: only the `buildServer({ ..., probes })` wiring + the
boot()-equivalent token resolution in `runControlPlane`. No changes to
`probes.ts`, the in-process runner, or the worker/boot paths.
## Test plan
- [x] RED→GREEN: with the impl reverted, the GET-list / trigger-runs /
no-token-401 / empty-token-fail-loud tests fail (route absent → 404);
with the impl they pass.
- [x] New tests in `orchestrator.test.ts` (real `runControlPlane` boot
on a live port): GET `/api/probes` lists the 3 HTTP families (not the
browser kind); POST `probe:smoke/trigger` with bearer runs it; 401
without token; 404 for `probe:e2e_smoke` and unknown ids; router omitted
when token unset; fail-loud on empty token.
- [x] Full harness suite green: 120 files / 2135 tests.
- [x] `tsc -p tsconfig.build.json` clean.
- [x] oxlint / oxfmt clean on touched files.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The control-plane runs the 8 in-process HTTP probe families but the
on-demand trigger endpoint (POST /api/probes/:id/trigger) was only mounted
on the legacy boot() path, so operators had no way to fire a family
immediately and had to wait on the slow cron. Wire the same
registerProbesRoutes onto the control-plane's buildServer using its own
httpProbeRegistry/httpProbeConfigs/scheduler/httpRunWriter and
OPS_TRIGGER_TOKEN. Only the prefixed in-process probe ids (probe:<id>) are
triggerable; browser-only and unknown ids 404. Token handling mirrors
boot() (unset -> router omitted; set-but-empty -> fail-loud).
## Summary
Phase 2 of the harness pool-fleet migration: make the three remaining
BROWSER probe families (`e2e_smoke`, `e2e_demos`, `e2e_deep`) actually
run on the fleet by wiring producers into the control-plane. The worker
`DriverRegistry` for all four browser kinds already landed in #5283;
this PR is purely the PRODUCER side. Builds on #5283/#5284/#5285/#5286
(all merged).
### Unit A — three catalog-enumerator factories
`createE2eSmokeServiceEnumerator` / `createE2eDemosServiceEnumerator` /
`createE2eDeepServiceEnumerator`, each a thin specialization of the
generic `createServiceEnumerator` (the #5285 seam) with its own
`driverKind` + dashboard `probeKey` prefix and the shared
`D6_DISCOVERY_FILTER`:
| family | driverKind | probeKey prefix | verified against |
|---|---|---|---|
| smoke | `e2e_smoke` | `d4:<slug>` | `src/cli/targets.ts` `d4:${slug}`
|
| demos | `e2e_demos` | `e2e-demos:<slug>` |
`config/probes/e2e-demos.yml` `id: e2e-demos` |
| deep | `e2e_deep` | `d5-single-pill-e2e:<slug>` | `src/cli/targets.ts`
`d5-single-pill-e2e:${slug}` |
### Unit B — multi-schedule wiring in `runControlPlane`
`runControlPlane` now builds four producers and passes a `schedules`
array to `createControlPlane` (the #5285 multi-schedule API) via a new
pure `buildProducerSchedules` helper. Crons are read **literally from
the config YAMLs** — the deliberate offsets stagger the four families'
Playwright fan-outs on the shared `BrowserPool`:
| scheduleId | cron | source |
|---|---|---|
| `fleet-job-producer` | `40 * * * *` | d6 (unchanged; still honors
`FLEET_PRODUCER_CRON`) |
| `fleet-producer-e2e-smoke` | `*/15 * * * *` | `e2e-smoke.yml` |
| `fleet-producer-e2e-demos` | `10 * * * *` | `e2e-demos.yml` |
| `fleet-producer-e2e-deep` | `5,20,35,50 * * * *` | `e2e-deep.yml` |
The in-process HTTP probe runner (#5284) and the d6 producer's REQ-B
sweep leg are left intact (additive). The worker registry is **not**
touched.
### R-timeout (demos) — addressed
The demos driver's 20-min outer cap is threaded in-process via the
legacy `E2E_DEMOS_TIMEOUT_MS` env, which the **fleet worker never sets**
— so without a fix the 38-demo service would blow the driver's 5-min
`DEFAULT_TIMEOUT_MS` and go all-red. Fix:
- New `E2E_DEMOS_TIMEOUT_MS` SSOT const (mirrors `e2e-demos.yml`
`timeout_ms`) in the enumerator module.
- `createE2eDemosServiceEnumerator` conveys the cap per-job in
`driverInputs.timeout_ms` (smoke/deep convey nothing — verified no
`timeout_ms` on their specs, and d6's spec shape is unchanged).
- The demos driver now reads `input.timeout_ms` as a resolution source:
`ctx.env.E2E_DEMOS_TIMEOUT_MS` (legacy) > `input.timeout_ms` (fleet) >
`deps.timeoutMs` > `DEFAULT_TIMEOUT_MS`. Schema gains `timeout_ms:
z.number().int().positive().optional()`.
### Risks honored
- **R3 pool-contention (deploy-time):** all four families' jobs are
claimed by the same pooled worker(s) drawing from ONE `BrowserPool`
under `BROWSER_POOL_MAX_CONTEXTS` (24). Per-family `max_concurrency`
governs producer ENQUEUE width, not worker execution. The guard is the
offset crons + the 24-cap — preserved faithfully here. **Verify at
deploy time by triggering each family's producer and confirming the
BrowserPool does not starve.**
- **R1 context-headers (verified, no change):** both
`createPooledE2eSmokeLauncher` and `createPooledE2eDeepLauncher` thread
`contextOpts.extraHTTPHeaders`, so smoke + deep set their per-slug
`X-AIMock-Context` themselves.
## Test plan
- [x] Unit A: `createE2eSmoke/Demos/Deep` factories stamp the right
driverKind + probeKey shape and carry driverInputs; demos conveys
`timeout_ms` (default + override); d6 equivalence (no `timeout_ms`)
asserted — red→green verified.
- [x] Demos driver reads `input.timeout_ms` when the env is absent
(fleet path), env still wins over input (precedence) — red→green
verified.
- [x] Unit B: `buildProducerSchedules` emits 4 schedules with exact ids
+ crons; `FLEET_PRODUCER_CRON` override applies to d6 only;
`runControlPlane` registers all 4 producer schedules on the live
scheduler alongside the `probe:*` HTTP entries — red→green verified.
- [x] Full harness suite green: **120 files / 2128 tests passed**.
- [x] `tsc --noEmit` clean except the lone pre-existing `toReversed`
error.
- [ ] Deploy-time: trigger each family's producer and verify BrowserPool
does not starve under co-firing (R3).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Phase 2 of the harness pool-fleet migration — wire the three remaining BROWSER
probe families into the control-plane PRODUCER side so they actually run on the
fleet (the worker DriverRegistry for all four kinds already landed in #5283).
Unit A — three catalog-enumerator factories mirroring createD6ServiceEnumerator,
each delegating to the generic createServiceEnumerator with its own driverKind +
dashboard probeKey prefix and the shared D6_DISCOVERY_FILTER:
- createE2eSmokeServiceEnumerator → e2e_smoke / d4:<slug>
- createE2eDemosServiceEnumerator → e2e_demos / e2e-demos:<slug>
- createE2eDeepServiceEnumerator → e2e_deep / d5-single-pill-e2e:<slug>
R-timeout (demos): the demos driver's 20-min outer cap is threaded in-process via
the legacy E2E_DEMOS_TIMEOUT_MS env, which the fleet worker never sets. The demos
enumerator now conveys the YAML timeout_ms per-job in driverInputs.timeout_ms
(new E2E_DEMOS_TIMEOUT_MS SSOT const mirroring e2e-demos.yml), and the demos
driver reads input.timeout_ms as a resolution source (env > input > deps >
default) so the 38-demo service no longer blows the 5-min default to all-red.
Unit B — runControlPlane now builds four producers and passes a multi-schedule
manifest (buildProducerSchedules) to createControlPlane, each family on its own
cron read literally from the config YAMLs (the deliberate offsets stagger the
families' Playwright fan-outs on the shared BrowserPool):
- fleet-job-producer 40 * * * * (d6, unchanged; honors FLEET_PRODUCER_CRON)
- fleet-producer-e2e-smoke */15 * * * *
- fleet-producer-e2e-demos 10 * * * *
- fleet-producer-e2e-deep 5,20,35,50 * * * *
The in-process HTTP probe runner and the d6 producer's REQ-B sweep leg are left
intact (additive). The worker registry is untouched.
R1 (verified, no change): both createPooledE2eSmokeLauncher and
createPooledE2eDeepLauncher thread contextOpts.extraHTTPHeaders, so smoke + deep
set their per-slug X-AIMock-Context themselves.
## Summary
Make the fleet control-plane ALSO run the 8 HTTP-only probe families
in-process (they don't need the BrowserPool/worker), by lifting the
legacy `boot()` probe-loader machinery into `runControlPlane`. Before
this, the control-plane only ran the d6 producer and the 8 HTTP families
were dark on the fleet.
HTTP-only families now run in-process: `smoke`, `starter_smoke`,
`image_drift`, `qa`, `aimock_wiring`, `version_drift`, `pin_drift`,
`redirect_decommission`. The browser families (`e2e_d6` / `e2e_smoke` /
`e2e_demos` / `e2e_deep`) are deliberately NOT run in-process — d6 goes
via the worker producer path; the rest need a BrowserPool the
control-plane does not own. This is independent of the worker-registry /
worker-loop work.
## Design
- **`BROWSER_KINDS`** (exported) = `{e2e_d6, e2e_smoke, e2e_demos,
e2e_deep}`. HTTP = every kind NOT in this set. Single source of truth
for the partition.
- **`registerHttpProbeDrivers`** (exported) registers only the 8 HTTP
drivers (no BrowserPool drivers) — kept separate from
`registerAllProbeDrivers` so the control-plane's `probeRegistry` is
HTTP-only.
- In **`runControlPlane`**: build an HTTP-only `probeRegistry` + a
discovery registry wiring the SAME sources `boot()` uses
(`railway-services` cached 24h with an auth tracker, `pnpm-packages`), a
`createProbeLoader` scoped to HTTP kinds, and the same
`diffProbeSchedules`/`buildProbeInvoker` loop `boot()` runs —
registering one `probe:<id>` scheduler entry per YAML config on the
control-plane's scheduler. **Crons are driven FROM the YAML**
(`cfg.schedule`), never hardcoded. Each tick flows through the same
`statusWriter` pipeline as the worker-result aggregator. Hot-reload via
`probeLoader.watch` mirrors `boot()`.
- **`includeKind` predicate** added to `createProbeLoader`: a browser
YAML on disk is SKIPPED (not rejected) against the HTTP-only registry,
so a present `e2e_*` YAML never surfaces a spurious
`probes.reload.failed`.
## /health handling
The control-plane now owns in-process probe rules, so `ruleCount`
reflects the real in-process HTTP probe count (was a hardcoded `0`). The
`role: "control-plane"` rules>0 gate-drop is retained (liveness is
governed by the scheduler signals already folded into `loopOk`), but the
real count means a silent probe-loader failure (zero HTTP probes loaded)
is now VISIBLE on `/health` rather than masked behind the role
short-circuit. `schedulerJobCount` already counts the new `probe:`
entries (covers BOTH the producer entry and the probe entries).
Teardown: the HTTP-probe file watcher is torn down on `stop()` and on
bind failure.
## Red→Green evidence
Wrote failing tests first, confirmed RED against baseline (stashed the
impl):
- scheduling test: `probe:smoke` / `probe:image_drift` not registered
(FAIL)
- /health test: `rules` < 2 (FAIL)
- `BROWSER_KINDS is not iterable` (FAIL)
- loader `includeKind` skip behavior
After implementing, all GREEN.
## Test plan
- [x] New `runControlPlane` in-process HTTP probe tests (scheduling
partition + /health rule count + BROWSER_KINDS) — green, RED-verified
- [x] New `createProbeLoader` `includeKind` skip-not-reject test — green
- [x] Full harness vitest suite: 119 files / 2076 tests passed
- [x] Production build `tsc -p tsconfig.build.json`: clean (the only
`--noEmit` error is the pre-existing `toReversed` lib-target issue in an
unrelated test file, excluded from the build config)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Address the CR gaps on the in-process HTTP-probe control-plane:
- Sweep orphaned `running` probe_runs at control-plane boot (boot()'s
sweepStaleRuns never ran in fleet mode → leaked rows forever). Best-effort;
a sweep failure does not abort boot.
- Add a fail-loud BROWSER_KINDS / HTTP-driver disjointness assert at boot plus
a drift-lock test mirroring registerAllProbeDrivers, so a mis-added kind
can't silently go dark.
- Assert the PR's headline guarantees that were unasserted: cron-from-YAML
(exact values), /health loader-failure boot-survival + probes.reload.failed
emit + rules=0 observability, hot-reload add/remove + watcher teardown, a
discovery-backed family (qa) in the in-process schedule, and tightened
/health rule/job counts to exact equality.
- Bump the includeKind skip log to info; make diffHttpProbeSchedules'
unregister-failure post-state explicit (keep config observable); correct the
/health "no longer masks" comment, the discovery-source comment
(image_drift uses railway-services; version_drift uses pnpm-packages), the
reload-failed surface comment (no bus subscriber on the control-plane), and
drop transitional-rot tags.
The fleet control-plane previously ran only the d6 producer, leaving the 8
HTTP-only probe families (smoke, starter_smoke, image_drift, qa, aimock_wiring,
version_drift, pin_drift, redirect_decommission) dark on the fleet. Lift the
legacy boot() probe-loader machinery into runControlPlane so those families run
in-process:
- Add BROWSER_KINDS = {e2e_d6, e2e_smoke, e2e_demos, e2e_deep}; HTTP = every
other kind. Add registerHttpProbeDrivers (HTTP-only driver set, no BrowserPool
drivers).
- In runControlPlane, build an HTTP-only probeRegistry + discovery registry
(railway-services cached + pnpm-packages, mirroring boot()), a probe-loader
scoped to HTTP kinds, and the same diffProbeSchedules/buildProbeInvoker loop
boot() uses — registering one probe:<id> scheduler entry per YAML config.
Crons are driven from the YAML schedule. Browser e2e_* YAMLs route to the
worker producer path and are NOT scheduled in-process.
- Add an includeKind predicate to createProbeLoader so a browser YAML on disk is
SKIPPED (not rejected) against the HTTP-only registry — no spurious
probes.reload.failed.
- /health: ruleCount now reflects the in-process HTTP probe count (was a
hardcoded 0). The control-plane role still drops the rules>0 gate, but the
real count means a silent probe-loader failure is visible on /health rather
than masked. schedulerJobCount already counts the new probe entries.
- Tear down the HTTP-probe file watcher on stop() and on bind failure.
## Summary
Producer-side foundation for fleet framework item 3b. Two
**behavior-preserving** generalizations that make the seam capable of
multiple browser families and multiple producer cadences, while keeping
the d6 case **byte-identical**. No wiring is flipped on yet (see Out of
Scope).
### 1. Parameterized service enumerator
`showcase/harness/src/fleet/control-plane/catalog-enumerator.ts`
- New generic `createServiceEnumerator(params)`
(catalog-enumerator.ts:215) takes the service-set `filter`, the
`driverKind`, and a `probeKeyPrefix` (string prefix → `<prefix>:<slug>`,
or a builder fn).
- `createD6ServiceEnumerator` (catalog-enumerator.ts:280) is
re-expressed as a thin call passing the d6 params: `D6_DRIVER_KIND`
(`e2e_d6`), prefix `"d6"` (→ `d6:<slug>`), and `D6_DISCOVERY_FILTER`. d6
output is unchanged — same services, same filter, same kind, same keys.
### 2. Control-plane accepts an array of producer schedules
`showcase/harness/src/fleet/control-plane/control-plane.ts`
- New `ProducerSchedule` type (`{ scheduleId, cron, producer }`) + a
`schedules?` dep on `ControlPlaneDeps`.
- `createControlPlane` normalizes to an array (control-plane.ts:~232);
omitting `schedules` degenerates to the single d6 schedule on
`FLEET_PRODUCER_SCHEDULE_ID` (`fleet-job-producer`) @ `40 * * * *` —
current behavior preserved exactly.
- `start()` / `stop()` iterate the array, registering/unregistering each
scheduler entry and starting/stopping each producer.
## Out of scope (deferred — gated on other in-flight PRs)
- **No `runControlPlane` wiring** to actually PASS multiple schedules —
that edit conflicts with in-flight **#5284** (which edits
`runControlPlane`) and is deferred. This PR only makes
`control-plane.ts` *capable* of N schedules + generalizes the enumerator
seam; the wiring lands later.
- No `e2e_smoke` / `e2e_demos` / `e2e_deep` enumerators or producers
(Phase 2).
- No changes to `worker-loop.ts` / `payload-mapper.ts` /
`probe-loader.ts` (other PRs own those).
- No driverKind constant / contract changes.
## Test plan
- [x] Red→green TDD: 3 new enumerator tests (generic
kind/keys/filter/fn-prefix) + 2 new control-plane tests (N entries
registered with distinct crons; stop tears all down) failed before impl,
pass after.
- [x] Equivalence: all pre-existing d6-enumerator + single-schedule
control-plane tests pass unchanged.
- [x] Full harness suite green: **2078 passed** (119 files).
- [x] `tsc -p tsconfig.build.json` clean (exit 0).
- [x] Only the 4 intended files changed; no lockfile drift.
Do not merge — producer-side foundation only; wiring follows after #5284
lands.
## Summary
The `showcase_build` workflow's pre-build `verify-image-refs` gate (SSOT
= `showcase/scripts/railway-envs.ts`) was **failing all staging
deploys** with `1 untracked Railway services`. The interim
`harness-legacy` staging service (id
`11279eba-97eb-417e-82a5-7cb4254eb147`, project `showcase`, env
`staging`) exists on Railway but had no SSOT entry, so the Railway→SSOT
drift check failed → build skipped → nothing deployed (this blocked PR
#5283's merge build).
`harness-legacy` is the **interim legacy all-probe harness**
(`HARNESS_ROLE` unset) stood up to keep non-d6 coverage live during the
fleet migration. It is **not CI-built** (runs a pinned pre-fleet image
digest set out-of-band) and will be torn down at migration end.
## Change
- Adds a `harness-legacy` entry to `SERVICES` mirroring the
`showcase-harness-worker` precedent (PR #5280):
- `ciBuilt: false` — not built by `showcase_build`; no dedicated build
slot.
- `gateIgnore: true` — deliberately-untracked for the image-ref gate.
`findUntrackedServices` treats any SSOT entry as known, so this clears
the "untracked" failure; `gateValidated: false` keeps
`findMissingServices` from flagging it.
- `repoNameOverride` → `showcase-harness` for both envs (same image-ref
shape as the control-plane harness).
- Real serviceInstance IDs recorded for both envs (resolved via Railway
GraphQL); probe disabled in both envs.
- Regenerates `railway-envs.generated.json`.
- Updates service-count and gate-ignored carve-out assertions across the
three affected test files (28 → 29 services; `harness-legacy` added to
the two `GATE_IGNORED` sets).
## Verification
- **Live gate (red→green):** without the entry,
`verify-railway-image-refs.ts` exits **1** with `harness-legacy is not
in the SSOT`; with the entry it exits **0** — `✓ 54 env-scoped instances
verified (2 skipped)`.
- **Full scripts test suite green:** 1771 passed (49 files), including
`railway-envs.test.ts`, `verify-railway-image-refs.test.ts`, and
`emit-railway-envs-json.test.ts`.
- **`emit --check`** confirms `railway-envs.generated.json` is in sync.
- **tsc** clean (`tsc --noEmit -p showcase/scripts/tsconfig.json`).
## Test plan
- [x] `npx tsx showcase/scripts/verify-railway-image-refs.ts` exits 0
against live Railway
- [x] `vitest run` in `showcase/scripts` fully green
- [x] `npx tsx showcase/scripts/emit-railway-envs-json.ts --check`
passes
- [x] `tsc --noEmit -p showcase/scripts/tsconfig.json` clean
- [ ] CI green on this PR
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The showcase_build verify-image-refs gate (SSOT = showcase/scripts/railway-envs.ts)
was failing with "1 untracked Railway services" because the interim
harness-legacy staging service (the legacy all-probe harness kept live during
the pool-fleet migration) exists on Railway but had no SSOT entry. That
Railway->SSOT drift check skips the build, so nothing deploys.
Adds a harness-legacy SERVICES entry mirroring the showcase-harness-worker
precedent (PR #5280): ciBuilt:false (not built by showcase_build, runs a pinned
out-of-band digest) and gateIgnore:true (deliberately-untracked for the image-ref
gate). findUntrackedServices treats any SSOT entry as known, so this clears the
untracked failure; gateValidated:false keeps findMissingServices from flagging
it. Real serviceInstance IDs for both envs recorded from Railway GraphQL.
Regenerates railway-envs.generated.json and updates the service-count /
gate-ignored carve-out assertions (28->29 services).
Verified: live verify-railway-image-refs.ts now exits 0 ("54 env-scoped
instances verified, 2 skipped"); without the entry it exits 1 with the
harness-legacy untracked failure. Full scripts test suite green (1771 passed).
## Summary
Generalizes the fleet pool-worker from a single hardwired d6 driver into
a driver **registry** keyed by `payload.driverKind`, so a single worker
can host all four browser driver families (`e2e_d6`, `e2e_deep`,
`e2e_demos`, `e2e_smoke`). Phase-1 framework, item 3a of the harness
fleet migration.
- **`worker-loop.ts`** now accepts `drivers: Map<string, { driver;
payloadToInput }>` and dispatches each claimed job by
`payload.driverKind`. An unknown kind returns a terminal
`worker-protocol-violation` result (the same failure shape an unmappable
payload uses) — the worker never crashes on an unhandled kind. The
legacy single `driver`+`payloadToInput` pair is retained as a fallback
(so the pre-registry behavior and all existing callers keep working),
and construction fails loud if neither a non-empty registry nor the
legacy pair is supplied.
- **`payload-mapper.ts`** generalizes `createD6PayloadToInput` into a
shared `createPayloadToInput` plus per-kind aliases
(`createDeepPayloadToInput`, `createDemosPayloadToInput`,
`createSmokePayloadToInput`) and the four `E2E_*_DRIVER_KIND` constants.
The re-hydration is identical across families; each driver's own zod
schema is the validation gate.
- **`orchestrator.ts` `runWorker`** builds all four pooled drivers on
the shared `BrowserPool` and registers them by kind (lifted from the
legacy `registerAllProbeDrivers` pooled construction), then injects the
registry into the fleet worker.
- **`fleet/orchestrator.ts` `runWorker`** threads an optional `drivers`
registry through to `startWorkerLoop` (registry takes precedence; legacy
single-driver and self-contained-boot paths preserved).
**d6 routing is unchanged** — equivalence is the gate.
## Test plan
- [x] Red→green: new tests fail against the reverted implementation,
pass with it (verified by stashing the impl files and re-running).
- [x] worker-loop routing tests: `e2e_smoke`→smoke, `e2e_deep`→deep,
`e2e_d6`→d6 (equivalence), unknown kind→`worker-protocol-violation`,
matched-kind uses its own mapper, `startWorkerLoop` dispatches by kind
end-to-end.
- [x] payload-mapper tests: four driver-kind constants, per-kind mapper
re-hydration + key defaulting.
- [x] Full harness suite green: **2085 tests** across 119 files.
- [x] `tsc --noEmit` clean except the known pre-existing `toReversed`
error (present on `main`).
Do NOT merge — pending 7-agent CR.
The multi-schedule seam (createServiceEnumerator + the schedules[] capability
in control-plane) didn't meet the file's own best-effort/fail-loud bar. Harden
it now since Phase 2 builds on it (no production caller yet):
- stop(): guard each producer.stop() per-entry so one rejection no longer aborts
teardown of later schedules (leaked cron handlers + running producers).
- start(): pre-validate every schedule's cron up-front before starting any
producer, throwing an aggregated error naming the offending scheduleId — no
more half-started plane with `started` latched true.
- normalization: throw on duplicate scheduleId (replace-semantics would silently
collapse two producers onto one entry) and on an explicitly-empty schedules:[]
(distinct from omitted, which keeps the d6 default).
- createServiceEnumerator: require a non-empty filter.namePrefix (an absent
prefix would enumerate ALL services) and reject an empty probeKey from a
function-form prefix, naming the slug.
Also: narrow the d6 "byte-identical" docstring (specs identical; the
catalog-enumerated log adds driverKind), pluralize the start()/stop() +
module-header producer comments, drop the Phase 2 marker, and extract the shared
ServiceSetFilter type.
Producer-side foundation for fleet item 3b — two behavior-preserving
generalizations, byte-identical for the d6 case:
1. Generalize the d6 service enumerator into a parameterized
`createServiceEnumerator(params)` carrying the service-set filter, the
driverKind, and the probeKey prefix builder. `createD6ServiceEnumerator`
is now a thin wrapper passing the d6 params (e2e_d6 kind, d6:<slug> keys,
D6_DISCOVERY_FILTER), so d6 behavior is unchanged.
2. Generalize createControlPlane to accept an array of
{ scheduleId, cron, producer } entries and register each on the scheduler.
The single-d6 case degenerates to a one-element array on
fleet-job-producer @ 40 * * * *, preserving current behavior.
Out of scope (gated on in-flight PRs): orchestrator runControlPlane wiring
to pass multiple schedules (conflicts with #5284), the e2e_smoke/demos/deep
families (Phase 2), and worker-loop/payload-mapper/probe-loader.
CR fixes for the fleet worker driverKind→driver registry (PR #5283):
- Fix default (self-contained) worker boot: build the default d6 as a
registry entry { driver, payloadToInput, aggregateSlugKey } instead of a
bare driver with no mapper, so startWorkerLoop's construction guard no
longer throws "Fleet worker has no drivers".
- Thread aggregate-key derivation through DriverRegistryEntry
(aggregateSlugKey?), defaulting to d6:<slug> so the d6 cell-capture filter
stays byte-identical while non-d6 kinds can supply their own scheme.
- Add construction-time fail-loud assert that each registry entry's factory
kind matches its key; raise unknown-driver-kind log to error to match the
sibling protocol-violation logs.
- Extract shared buildPooledBrowserDrivers consumed by both
registerAllProbeDrivers and the worker registry; collapse no-op per-kind
payload-mapper aliases to the single createPayloadToInput.
- Introduce typed DriverKind union (contained to worker/payload-mapper);
keep contracts.driverKind a string wire boundary.
- Tests: new fleet/orchestrator.test.ts (default-boot equivalence), driver
construction-guard + custom aggregateSlugKey coverage, and lock-step +
registry-wiring pins (factory kind == constant).
Generalize the fleet worker from a single hardwired d6 driver into a
driver REGISTRY keyed by payload.driverKind, so one worker can host all
four browser driver families (e2e_d6, e2e_deep, e2e_demos, e2e_smoke).
- worker-loop: accept a `drivers: Map<kind, { driver, payloadToInput }>`
and dispatch each claimed job by `payload.driverKind`. Unknown kind →
terminal `worker-protocol-violation` (same shape as an unmappable
payload), never a worker crash. Legacy single `driver`+`payloadToInput`
pair retained as a fallback for back-compat. Fail-loud at construction
when neither a registry nor the legacy pair is supplied.
- payload-mapper: generalize createD6PayloadToInput into a shared
createPayloadToInput plus per-kind aliases and the four driver-kind
constants (the input re-hydration is identical across families; each
driver's own zod schema is the validation gate).
- orchestrator runWorker: build all four pooled drivers on the shared
BrowserPool and register them by kind, lifted from the legacy
registerAllProbeDrivers pooled construction.
d6 routing is unchanged (equivalence gate). Red-green tests cover
routing e2e_smoke/e2e_deep/e2e_d6 to their drivers and unknown-kind →
protocol-violation. Full harness suite green (2085 tests); tsc clean
except the known pre-existing toReversed error.
## Summary
Removes `examples/integrations/langgraph-python-threads/` (87 files) and
its two entries in `.github/config-allowlist.txt`.
The standalone "LangGraph Python + durable threads" template is defunct:
- The ENT-679 threads rollout was fully reverted from main on 2026-06-04
(#5215/#5216/#5217); the settled model is base `langgraph-python` +
Intelligence activation overlay, not a duplicate `-threads` template.
- The Intelligence CLI no longer scaffolds from it — verified zero
references on Intelligence `main` (a806a7e0) **and** in the published
`copilotkit@3.0.2` npm tarball.
## Verification
- `git grep langgraph-python-threads` across the tracked tree returns
nothing after this change (no hits in `pnpm-workspace.yaml`,
`pnpm-lock.yaml`, `_parity` manifest, CI workflows, or READMEs).
- The example was standalone (npm-based, not a pnpm workspace member,
not an Nx project) — zero workspace projects structurally affected.
- Lefthook commit gates green: check-binaries, sync-lockfile, lint-fix,
`nx run-many -t test --projects=packages/**` (25/25).
- `nx run-many -t lint,build --projects=packages/**`: all green except
pre-existing `@copilotkit/vue:lint` failures on main (files untouched by
this diff, introduced in 913c36b8b5).
Linear: ENT-800
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The first merge of main accidentally committed local lint/format
auto-fixes to 28 files unrelated to the example removal. Restore
them byte-for-byte from origin/main.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>