mirror of
https://github.com/vercel/workflow.git
synced 2026-09-14 19:59:43 +08:00
nathanc/shared-queue-http-handler
28 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2668e3325b |
Durable hook resume: write, then wake (#3841)
* test(core): reproduce lazy resume disposal race * Fix durable hook resume race * Fail closed on unknown hook wakes * Improve unsupported hook wake diagnostics * Address durable hook resume review feedback * Harden producer-committed wake handling * Serialize durable hook resume: write, then wake resumeHook() now dispatches strictly serially: the hook_received event is made durable first, and the workflow wake is published only after the write is acknowledged. The wake is a plain runId message (the shape the sequential path always published), so the producer-committed wake barrier, its queue-message field, and the HOOK_RESUME_INPUT_VERSION bump are all removed — no consumer or backend coordination is needed, and either side rolls back independently to today's behavior. The pre-write ops flush now partitions serialization ops: producer-push uploads are awaited before the event commits (the payload must not point at bytes still in flight), while consumer-settled reader ops — a dehydrated WritableStream, e.g. a manual webhook's responseWritable — are backgrounded. Awaiting those deadlocked the resume against its own wake (webhookWorkflow failing across the whole e2e matrix). Also: wake retries stop on definitive 4xx errors instead of burning the retry budget; WORKFLOW_DISABLE_LAZY_HOOK_RESUME no longer gates anything and is ignored; the internal resumeHookDurable alias is removed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Address review: retry classification, wake dedup, 409 passthrough - Wake retry classification now actually fires against @vercel/queue: its errors carry no status field, so classify by the World's deployment-unavailable hook, then numeric status, then the queue client's definitive-4xx error names. - The wake publish carries idempotencyKey `hook-<resumeId>` on the claim path, so a retried publish whose response was lost dedups instead of costing a duplicate full replay. - EntityConflictError (HTTP 409) from the durable write is no longer re-keyed to HookNotFoundError: every 409 the backend emits on this write today is transient (slot conflict past the server's retry budget, claim race) and committed nothing, so it surfaces retryable instead of presenting as a permanent 404. - Stamp workflow.hook.resume_committed / wake_published span attributes after each leg resolves, making stranded resumes (committed event, no wake) queryable from traces. - Document on the public resumeHook signature that passing the token (not a cached Hook) is what makes the write idempotent-on-retry. - Changeset/changelog: note the ended-run behavior change (late webhook deliveries to finished runs now 404 instead of 202) and the 409 passthrough. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Pranay Prakash <pranay.gp@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ffc58078d0 |
Stop logging on healthy workflow execution (#3878)
A successful run printed several lines that described the runtime working correctly. Most of it was fallout from defaulting the events transport to WebSockets (#3702): three breadcrumbs written while the transport was opt-in became default-path output, because each one reported a choice the caller no longer makes. - `world-vercel: using ws events transport (…)` ran once per cold start on every deployment, naming the transport it was always going to use. - The `projectConfig` proxy fallback warned once per process. That World cannot hold a socket, so with WS on by default every CLI command and the observability app warned about a fallback nobody asked for and nobody can act on. Debug-gated and reworded from "requested but" to "unavailable for". - The `max_duration` / `auth_expiry` drain notice is routine: the transport reconnects from the close that follows and no write is lost. Swept for the same shape elsewhere: - `world-local`'s queue-concurrency notice fired per message once a fan-out exceeded the limit — the semaphore doing its job. - `@workflow/world`'s active-run recovery line printed on every dev-server restart with work in flight. The re-enqueue *failure* above it stays unconditional; that one leaves a run unresumed. - The port-detection diagnostics in `@workflow/utils` keyed off `NODE_ENV=development`, which is the only environment that reaches them, so the gate made them unconditional for their whole audience. All of it moves behind `DEBUG=workflow:*` via a new `debugLog` in `@workflow/utils`, joining world-vercel's existing `httpLog` and `logRetry` output under one selector. Warnings and errors are untouched, so a run that actually goes wrong is no quieter than before — the ws-transport tests that assert failures are never silent still pass unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com> |
||
|
|
7e48e7b4de |
Re-enable the sealed log by default (#3737)
* Revert "[world] Make the sealed log opt-in instead of default-on (#3735)" Reverts |
||
|
|
dc68611fbf |
Default the events transport to WebSockets (#3702)
* Default the events transport to WebSockets WORKFLOW_EVENTS_TRANSPORT=http is the opt-out. Only that exact value disables it, so a typo'd or empty value fails toward the default rather than quietly pinning a deployment to HTTP. The prerequisite the gate named for defaulting on is met: postEventFrameOverWs opens a client span per frame. What is still missing is Vercel's outgoing-requests view, which reads instrumented fetch calls rather than spans and so cannot show a transport that issues no request. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * docs: WORKFLOW_EVENTS_TRANSPORT defaults to ws Three places still documented http as the default. Each now states the opt-out is the exact value http, rather than leaving 'default: ws' to imply that anything non-ws disables it — the asymmetry is deliberate in the code and is the part a reader would otherwise get wrong. Also drops 'Experimental' from the Vercel World page: a setting that is on for everyone by default is not opt-in experimental, whatever else it is. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Fix the gate's own unit tests for the flipped default Five tests in ws-transport.test.ts still encoded the opt-in semantics. Three were the isWsEventsTransportEnabled table itself; the other two (openWsChannel 'does nothing when the gate is off', and the channel release equivalent) relied on the suite's ambient unset environment meaning 'off', which it no longer does. Both now set http explicitly. Two tests in ws-transport-spans.test.ts asserted HTTP-side span behaviour the same way. The write one would have kept passing by falling through resolveWsTransport's null rather than because the gate was off - passing for the wrong reason, which is what this file exists to catch. Also makes the opt-out case-insensitive and trimmed. The gate is deliberately asymmetric - unrecognized values take the default - but that asymmetry should not extend to swallowing HTTP or ' http '. Whoever reaches for the escape hatch is plausibly mid-incident, and silently ignoring their opt-out over a capital letter is the same class of silent-wrong-transport bug this flip is meant to stop shipping. 554 tests pass in packages/world-vercel. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * ci: add a required forced-HTTP e2e lane (#3703) Flipping the default makes e2e-vercel-prod a WebSocket lane: it sets no WORKFLOW_EVENTS_TRANSPORT, and unset now means ws. Nothing in the file would exercise the HTTP events transport against a real deployment any more, so this is not additive coverage — it replaces coverage the flip silently removed. Unconditional and required rather than label-gated like the WS lane. HTTP is now the fallback, and the fallback is silent: resolveWsTransport returning null costs a write nothing and logs nothing, which is the shape of the durabench bug this stack came out of. Two apps rather than the WS lane's four, since every row is a real vercel deploy charged to every PR. nextjs-turbopack is the only fixture emitting OTEL spans, so it is the one that can show which transport actually ran; express covers the non-Next server path. Also corrects the WS lane's docblock, which claimed every other job exercises HTTP only. That stopped being true one commit ago. Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Fail loudly when step_completed falls back to HTTP under a strict flag The WS e2e lane asserts that the transport is harmless, not that it is used: an event written over HTTP produces the same run outcome as one written over the socket, so the lane stayed green through the entire period the transport was silently demoted. WORKFLOW_INTERNAL_EVENTS_TRANSPORT_STRICT turns that one case into a failed run, and the WS lane now sets it. Scoped to step_completed alone, because most fallback is legitimate: run_created is written outside any invocation that opens a channel; run_started routinely lands before the channel is registered (34% HTTP on a healthy deployment); step_created and wait_created mostly fold into events.createBatch, which is not wired to the socket; and a write after the invocation released its claim falls back by design. step_completed is issued after a step body has run, and was 100% ws across every WS-enabled deployment measured on two SDK versions. The flag reads as off unless the value is exactly 1 or true - the opposite asymmetry from the transport gate, which treats an unrecognized value as on. That gate risks a deployment sitting quietly on the wrong transport; this one fails runs, and should not be acquired by a typo. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * ci: run the WS transport lane on every PR It was opt-in behind ws-transport-test because four real vercel deploys were too much to charge an unrelated PR for a transport that was off by default. Flipping the default expires that reasoning from both ends: the cost is no longer for someone else's feature, and this is now the only lane that asserts the socket carried the events. e2e-vercel-prod inherits the new default but checks nothing, so behind a label the average PR would move every deployment onto WebSockets with nothing verifying they were used. Drops WS_REQUIRED from the gate along with it. That existed only to let the lane be legitimately skipped on an unlabelled PR; with no label the lane is required unconditionally, like e2e-vercel-prod and the HTTP lane, and the skipped case is now a failure rather than a warning. Gate script extracted and run against the cases that matter: ws skipped fails on a standard PR, ws skipped fails under workflow-server-test, and all-green passes. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * ci: widen the HTTP transport lane to six server shapes Before the flip, HTTP was the default and all 28 e2e-vercel-prod lane-runs covered it. After the flip they cover WebSockets instead, and this lane is the entirety of the HTTP coverage - two apps was too thin for a transport that is still supported. Six, not the full 14, because every row is a real vercel deploy charged to every PR. Chosen by server shape rather than count: example (baseline), nextjs-turbopack (Next, and the only fixture emitting OTEL spans), vite (Vite SSR), express (Node req/res), nitro (h3, also covers nuxt) and hono (fetch-API Request/Response, a different mount shape from express). The rest duplicate a shape already covered; python is left out because it has no conformance gate and needs routes this suite does not serve. The first four match the WS lane's matrix on purpose, so the same fixture runs on both transports and a failure on one can be read against the other. Project ids and slugs are copied from e2e-vercel-prod and verified equal to it; both lanes already use the same team and token. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> --------- Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com> |
||
|
|
f771585486 |
fix(world-vercel,world-local): hold process-wide state on globalThis (#3728)
* fix(world-vercel,world-local): hold process-wide state on globalThis Both packages are bundled into the host application's server build, and a bundler keys module identity on (resource, layer) — Next.js alone builds `instrument`, app-route, `ssr` and `edge` layers, so one process holds one copy of each of these modules per layer. Every module-scope `const`/`let` in them was therefore per-copy state wearing the costume of a process singleton. vercel/workflow#3493 made `@workflow/world-vercel` bundled rather than external and the events WebSocket transport regressed to HTTP for exactly this reason: the queue consumer registered its channel in the `instrument` copy's `Map` and the write path looked it up in the route copy's empty one. A deterministic miss, for the life of the process. `@workflow/world-local` had the same exposure all along — including `runFileLocks`, where a duplicated mutex simply stops mutually excluding. Add `globalSingleton()` to `@workflow/utils` (the primitive `@workflow/core` already hand-rolls for its World cache) and route every mutable module-scope binding in both worlds through it. Regression cover, in three layers: - `global-singleton.test.ts` pins the primitive's semantics. - `ws-transport-module-copies.test.ts` imports the module twice in one process and asserts a transport registered by one copy is found by the other — it fails on a plain module-scope `Map`, which is the shipped bug. - `scripts/lint/module-scope-state.mjs` fails the class: an AST rule banning mutable module-scope state in these packages, with `// per-copy-ok: <why>` as the deliberate escape. Wired into both packages' `vitest run src`, with fixture self-tests so it cannot rot into a no-op. * test(world-postgres): pin the module-scope-state rule for the postgres world It is deduped today only because `getRuntimeRequire()` loads it — a property of how it is loaded, not how it is written, and exactly what changed for world-vercel in #3493. The package is already clean; this keeps it that way. * docs(worlds): codify "a world must not hold mutable module state" A world package is loaded one of two ways, and only one of them gives it a single module instance: a runtime `require()` (deduped by Node) or the host's bundler (one copy per layer). Which one you get is a property of how the world is loaded, not of how it is written, and it changed under `world-vercel` in #3493 — so the rule has to be "never rely on module scope", not "rely on it until someone flips a config". Written down in the four places someone can meet it: - `docs/content/worlds/{v4,v5}/building-a-world.mdx` — a "Process-wide state" section for custom-world authors, with the loading modes spelled out and a nudge to prefer World-instance state over a global. - `packages/world/README.md` — the same constraint on the contract package. - `CLAUDE.md` — so the next contributor working in these packages sees it. - `packages/core/src/runtime/world.ts` — at the two static imports, which is where the difference between a bundled world and a required one originates. The rule's own error message now teaches it too, rather than naming a helper. Consolidates the guard while here: `@workflow/utils` owns the rule and its fixture self-tests, and sweeps every *published* `packages/world-*` discovered at runtime, so a world package added later is covered without anyone remembering. Each world keeps a one-assertion mirror for locality. * style: drop prose em dashes from this branch's new text #3704 landed a repo-wide writing pass hours after this branch was written and took `world-vercel/src` from 406 em dashes to 130 (`ws-transport.ts` alone went 35 to 1). This branch's docs section, README, comments and lint messages were written before that and would have put 36 of them straight back into the files that were just cleaned. Rewritten sentence by sentence rather than by substitution: an em dash becomes a colon, a comma, a full stop or a parenthetical depending on what it was doing. Also fixes a real defect the sweep surfaced: `world-postgres`'s guard test was generated through a shell heredoc and had literal backslash-backticks in its doc comment. * Update .changeset/world-module-scope-state.md Co-authored-by: Peter Wielander <mittgfu@gmail.com> Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> * fix(core): build the entrypoint's queue handler from getWorld() Adopted from #3666 by @MintedKenny, which implements #3665 and could not run CI as a fork PR. One line of behavior: `workflowEntrypoint`'s lazy handler init calls `getWorld()` rather than `getWorldHandlers()`. `getWorldHandlers()` owns a second, build-time-safe cache, so calling it from the runtime route built a *second* World in the same process. That costs a stateful World duplicate resources on every instance — world-postgres eagerly constructs a `pg.Pool` (default `max: 10`) and a nested world-local World in `createWorld()`, so self-hosted users have been paying for two of each — and, for a bundled world package, the two Worlds are built by two different module copies, which is the mechanism behind the WS transport regression the rest of this branch contains. The public `getWorldHandlers()` and its separate build-time cache are unchanged; only the runtime route stops using it. Kept from the original: the regression test asserting the factory runs exactly once, and the api-reference wording (re-applied over #3704's list punctuation). Not taken: renaming the `workflow.route.get_world_handlers` span. It is a distinct span from the per-request `workflow.route.get_world` at the top of the flow route, and reusing that name would collide with it in traces and in `runtime-trace-mode.test.ts`; a comment records why the name outlived the call. Co-authored-by: Kenneth <kenneth@standardforensics.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address AI review on the module-scope work Two blocking findings, both real: - **Cross-version state sharing** (`ws-transport.ts`). A process can hold two *published versions* of `@workflow/world-vercel` (a transitive dependency pinning an older `@workflow/core`, which depends on this package by exact version). Both wrote to the same unversioned `Symbol.for` key, so one version's write path could be handed a `WsEventsTransport` built by the other's class and frame against a protocol it may not share — with no version negotiation on the socket to catch it. `shapeVersion` cannot express this: the container is stable, the hazard is its contents. The registry and the events dispatcher recycler are now keyed by package version. The plain connection pools stay unversioned; sharing those across copies is the point. - **The documented pattern failed the rule this PR adds.** The custom-world docs teach `store[StateKey] ??= …`, which the rule flagged as a field write. It now recognizes state rooted at `globalThis`, following one alias hop, which is also what `core/private.ts:23` and `next/src/index.ts:58` are already doing correctly (core drops 26 findings to 22, next 7 to 6). The docs also now say outright that `globalSingleton()` is the same thing, since AGENTS.md prescribes it and the page did not mention it. Rule precision, from the review's probes: - `.mts`/`.cts` are scanned. `@workflow/world-testing` is authored in `.mts`, so its entry in the sweep was passing vacuously — with the walk fixed it reports a real finding, now annotated (it is a standalone `serve()` entry). - Mutations in top-level statements no longer count. A table filled at module evaluation is identical in every copy; divergence needs a later write. - `static` class fields are collected, attributed to the class name. - An *exported* binding initialized to an empty collection is a finding on its own, which approximates the cross-file case the walk cannot resolve. Six fixtures pin the new behavior. The rule's header now states what it does not see, and AGENTS.md states where the sweep stops and why core is not gated yet. Also tags `resetGlobalSingletonForTest` `@internal`. * fix(lint): attribute a static-field write to the field, not the class The static-field support added in the previous commit keyed `declared` on the class name, so a class carrying more than one mutable static reported one finding instead of one per field, and labelled the survivor with whichever mutation was seen first. On a two-static fixture it reported `static Registry.latch (`.set()`)`: the name of one field, the reason belonging to the other, pointing the reader at the wrong line. Key static fields `Class.field` and resolve a write to the same shape, via a new `memberPath()` that takes the first two segments of a member chain and tries that key before the bare root identifier. Two follow-ons fall out of having the path: - `this.field` inside a `static` member resolves to the class, which is the ordinary way to write the mutation. `staticClassOf()` returns nothing for an instance member, where `this` is an instance and the state is per-instance rather than per-copy, and nothing inside a nested `function`, which rebinds `this`. - `state.count++` is now a finding, like the `state.count += 1` that `assignment()` already reported. Fixtures pin all four, including the instance-field case that must stay clean. The four world packages still report zero, and the extracted `recordMutation()` keeps the file at its previous two Biome complexity warnings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: make module duplication inert across every bundled package `@workflow/core` is bundled into the host server build the same way the worlds are, and always has been — the original repro measured three live copies in every arm, including the pre-#3493 external one. One instance is not reachable: layers cannot share a module, and core cannot be external because it *is* workflow code (`runtime/start.ts:253` and nine methods in `runtime/run.ts` are `'use step'`), so it must go through the SWC loader. The Next integration already encodes that rule by removing workflow-bearing packages from `serverExternalPackages`. So the duplication stays and the hazard is removed instead, everywhere the duplication can happen. `@workflow/core` (22 findings to 0): warn-once latches in `constants.ts`, `start.ts` and `telemetry.ts`; the source-map tracer cache; the VM script cache; the QuickJS compiled-assets and baseline caches; the dev-server port cache (its own comment already said "per process"); the text codecs; the zstd browser decoder; and the `useStep` closure brand, where a function marked by one copy was invisible to another. The one with teeth was `step-single-flight.ts`: a per-copy map is not single-flight. Two invocations reaching it through different layers would each believe they were alone in the process and both run the step body, silently degrading in-process dedup to the cross-process residual its own doc scopes out to the ownership lease. Also `@workflow/world` (a warn-once set, hand-rolled onto `globalThis` to keep that package dependency-free), `@workflow/ai` (the lazy OTel API), and `@workflow/nest` (bootstrap config in a module-level `let` and two static class fields — configure one copy, read another, and the controller is unconfigured for the life of the process). Five sites are deliberately per-copy and now say why: state keyed on objects that never cross copies (the barrier safety-net `WeakSet`, the QuickJS pending byte `WeakMap`), the synchronously-scoped guest-code sink, and the OTel diagnostic that reports what *this* copy sees. The sweep now covers all of it. Packages with a single module graph stay out (build-time code, the CLI, the o11y UI, the test runner) and AGENTS.md records which and why. Found while doing this: two static fields on one class collapsed into a single entry in the rule, so `WorkflowModule.options` was invisible behind `WorkflowModule.outDir`. Statics are now keyed `Class.field`. * fix(world): suppress noAssignInExpressions on the globalThis idiom The hand-rolled form trips Biome, as it does in `packages/core/src/private.ts`, which carries the same suppression. Restructuring it into a helper function instead would hide the state behind a call the module-scope rule cannot follow, so the binding would stop being recognized as off-module and the package would report a finding for correct code. * fix: sweep every bundled package, and mark utils side-effect free @shalabhc asked on review whether `@workflow/utils` needs this too. It does, and so do three others: `utils`, `errors`, `serde` and `workflow` all end up in the host application's server build and none were in the sweep. All four report zero today, which is exactly the state `world-testing` appeared to be in before the `.mts` walk was fixed and it turned out to have a real finding. Being clean and being *checked* are different properties, and only the second one survives the next contributor. `sideEffects: false` on `@workflow/utils`: verified that every module in the package only declares (no import-time work), so a bundler can now drop the unused parts of the barrel instead of keeping all ~64 KB of it because three packages import one 476-byte function. --------- Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> Co-authored-by: Peter Wielander <mittgfu@gmail.com> Co-authored-by: Kenneth <kenneth@standardforensics.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Peter Wielander <peter.wielander@vercel.com> |
||
|
|
b2cac623d3 | [world] Make the sealed log opt-in instead of default-on (#3735) | ||
|
|
e1e64e3de3 |
docs: apply Vercel technical writing standards (#3704)
* docs: apply Vercel technical writing standards Audit the complete documentation corpus, package READMEs, skills, and source TSDoc/comments against the vercel-technical-writing skill and style-rules.md. Normalize sentence-case headings without changing published anchors, remove prose em dashes and filler wording, improve active voice and self-contained phrasing, standardize product/brand capitalization, American English, list punctuation, units, and code fence languages, and preserve exact runtime strings/table placeholders. All executable code is unchanged. Modified skills have their metadata versions bumped. * docs: extend writing audit to repository Markdown Apply the same technical-writing rules to design documents, compiler specifications, workbench guides, package changelogs, and the remaining tracked Markdown outside the deployed docs corpus. Preserve historical meaning, commands, output literals, table placeholders, and heading anchors. * docs: exclude generated package changelogs from audit |
||
|
|
7b79ba37cc |
Add support for 'noop' event type - spec version 7 (#3634)
Co-authored-by: Peter Wielander <peter.wielander@vercel.com> |
||
|
|
b3dbc6d264 | [docs] v5 changes docs: what's new, world upgrade guide, migration skills (#3100) | ||
|
|
9454d51db0 |
feat(core): resolve run.returnValue via a World long poll instead of a 1s poll (#3570)
Co-authored-by: Peter Wielander <mittgfu@gmail.com> Co-authored-by: Peter Wielander <peter.wielander@vercel.com> |
||
|
|
0b2797bbac | [next] Bundle the Vercel world into the Next.js server output (#3493) | ||
|
|
04e060a0ec | [world] Add WORKFLOW_NODE_HTTP to run the HTTP Worlds on node:http (#3461) | ||
|
|
de2a86c61c | [world] Make spec version 6 the current version (#3542) | ||
|
|
dc85865718 | [core] Drop pre-slot event ID support and preconditionGuard capability (#3519) | ||
|
|
01991edeeb |
feat(world-vercel): synthesize per-event client spans on the WS transport (#3452)
* feat(world-vercel): synthesize per-event client spans on the WS transport PR #3084 added the opt-in `WORKFLOW_EVENTS_TRANSPORT=ws` path and listed "no client-side span on the WS path" as a known limitation. Because event writes become multiplexed frames on one long-lived socket rather than individual `fetch` calls, the per-event `http POST` CLIENT span that the HTTP transport produced simply disappeared — traces went from one span per event to nothing between the invocation and the server. Restore it by synthesizing a request-shaped span around each frame, and give the upgrade its own span: - Extract `withHttpClientSpan` / `recordClientSpanStatus` from `instrumentedFetch` in `http-core.ts` so the synthetic span is emitted by the same envelope as the real one and cannot drift from it. `InstrumentedFetchOptions` now extends `HttpClientSpanOptions`. - `postEventFrameOverWs` opens `http POST` with `url.full` pointing at the v4 REST endpoint the frame is forwarded into, so per-event traces and latency dashboards keep working across the flag. Extract `eventsV4Url` so that URL cannot drift from the one the HTTP path actually requests. - Tag both transports with `workflow.events.transport` (`http` | `ws`) and `workflow.event.type`; the WS path additionally sets `network.protocol.name=websocket`, `workflow.events.ws.url` (the real wire destination) and `workflow.events.ws.req_id` (join key to the server's log line for the frame), so the span is never mistaken for a real HTTP request. - Add a `workflow.events.ws.connect` span around the upgrade — the one genuinely-HTTP request here, previously the invisible half of every WS write's latency — carrying `workflow.events.ws.reconnect_attempt`. This also puts `resolveUpgradeHeaders`' trace-context injection inside a client span, as AGENTS.md requires. - Fix `parseServer` to treat `wss:` as TLS (port 443, not 80). Out of scope, deliberately: per-frame `traceparent` (needs a frame-meta field plus a server change) and Vercel's outgoing-requests view (that instruments global `fetch`, so a frame structurally cannot appear there). Covered by `ws-transport-spans.test.ts`, which drives the real selection + transport + adapter stack over a fake socket and asserts span shape, failure reporting, retry behaviour and HTTP/WS parity. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhchaturvedi-7802 <shalabh.chaturvedi@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * chore: trim WS spans changeset to the user-facing summary Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhchaturvedi-7802 <shalabh.chaturvedi@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * fix(world-vercel): only tag event-write spans with transport Signed-off-by: Shalabh Chaturvedi <shalabh.chaturvedi@vercel.com> Co-Authored-By: shalabhchaturvedi-7802 <shalabh.chaturvedi@vercel.com> * fix(world-vercel): format WS transport span regression test Signed-off-by: Shalabh Chaturvedi <shalabh.chaturvedi@vercel.com> Co-Authored-By: Shalabh Chaturvedi <shalabh.chaturvedi@vercel.com> --------- Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com> |
||
|
|
b589460ce8 | [core] Report the replay position on every event write (#3479) | ||
|
|
6786db9953 | World-side incrementing event ID (specVersion 6) (#3389) | ||
|
|
264ddff67b |
Add WebSocket transport for step-execution event writes (opt-in) (#3084)
* sdk side for workflow server websockets Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * hardcoded workflow server Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * debug info Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * more debug Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * remove unnecessary debug Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * default on websockets, and override url Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * fix for missing funcs Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * make websockets opt outo Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * enable ws again Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * [revert later] reduce test to single test, test both http and ws at the same time Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * empty Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Run full suite with and without ws Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * empty Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * improve e2e test Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * minimize tests Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * fallback to http when proxy present Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * default to websockets, remove matrix Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * remove smoke test Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * update to new protocol Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * adjust for new protocol (runid in path) Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * fix ws transport error Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * blank Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * fix ws dep Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * fix ws external: only accelerators, not ws itself Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * add dedicated WS-transport e2e job; flip WS default back to opt-in Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * build all packages before local vercel build (needs workflow/nitro on disk) Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * force NITRO_PRESET=vercel for the local vercel build step Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * install vercel CLI once instead of npx-ing it per command Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * add changeset for WS events transport Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Harden the WS events transport and gate its e2e jobs Follow-ups from review of the WS transport. CI: `e2e-vercel-ws-transport` was wired into the `summary` job but not into `e2e-required-check`, so all three WS jobs could fail while the required check stayed green. Added to both branches of the status validation — including the `workflow-server-test` label branch, where the job runs under the same gating as `e2e-vercel-prod`. Transport: - `reqId` and the pending-reply map are now per connection rather than per transport. The protocol defines `reqId` as a per-connection counter, so a reconnected socket restarts at 1; with one shared map that collided with the previous socket's still-registered waiters. It also makes the superseded-socket guard structural instead of something the close path has to remember. - Post-open socket errors are no longer silent. The only `'error'` listener closed over the connect promise's `reject`, already settled once `'open'` fired, so every broken pipe / 1009 / protocol fault was swallowed and its requests hung with no per-request timeout to save them. Now logged, and the connection is torn down. - An unexpected close reconnects eagerly instead of waiting for the next write, since a socket breaking mid-run means more writes are coming. Bounded by exponential backoff, an attempt cap that falls back to lazy reconnect, a bail-out when a newer socket is already live, and an `unref()`ed timer so a backoff window can't delay handler exit. - `ws.send()` failures reject their request. `send()` doesn't throw on a non-OPEN socket — it reports through a callback we weren't passing — so the request just sat in `pending` forever. - The reserved `reqId: -1` malformed-frame reply and undecodable frames are logged loudly instead of dropped. - Auth headers resolve once per socket via a thunk, not once per event. The bearer only rides the upgrade, so the old code awaited `getVercelOidcToken()` on every write and discarded all but the first. Re-resolving on reconnect also means a new socket gets a fresh token. Adapter: a reply with no numeric status now fails closed. Defaulting to 200 reported a write as applied whenever the client met a frame it didn't understand — and the protocol is explicitly designed to grow new response variants. Tests: 24 new unit tests over the paths the e2e suite can't reach on demand (send failure mid-flight, error after open, late close from a superseded socket, reconnect backoff and give-up, sentinel/undecodable frame logging, one-token-per-socket) plus the adapter's fail-closed and typed-error mapping. Co-Authored-By: Shalabh Chaturvedi <shalabhc@users.noreply.github.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * ci: re-trigger to confirm the prior e2e failures were flake No code change. The 5 failures on 0bb21e7 clustered in a ~20s window across HTTP-path jobs (example/nuxt on the same test, sveltekit on a timeout) and one WS job (sleepingWorkflow's clock-skew assertion), which points at the environment rather than the transport changes. Re-running to confirm. Co-Authored-By: Shalabh Chaturvedi <shalabhc@users.noreply.github.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Add WS wire-contract conformance tests and pin the transport gate Three gaps in the existing coverage. **The HTTP path was already covered** — `events-v4.test.ts` has 22 tests, including five directly on `createWorkflowRunEventV4` over HTTP (alias URL, frame meta contents, response decoding, skipPreload/stateUpdatedAt forwarding). Those run with the gate unset, so they do confirm the two-branch refactor didn't disturb HTTP. No new tests needed there. **But nothing pinned the gate itself.** Every HTTP assertion stays green if the default flips to WS, because the transports are built to be indistinguishable at the result layer — and an earlier revision of this branch did flip the default deliberately, for benchmarking. Added tests for `isWsEventsTransportEnabled()` across values, and one that drives a real HTTP request through a MockAgent while asserting the WS transport is never constructed. **Nothing verified the bytes.** `ws-transport.test.ts` replies with whatever the test hands it, which proves the client's lifecycle but not that its frames are what workflow-server accepts. That's the drift the spec doc exists to prevent, and it already happened once: event meta flat on the frame where the server wanted it nested under `event`, with both sides' tests passing. `ws-protocol-conformance.test.ts` pairs the real client stack (through `createWorkflowRunEventV4`) with a fixture mirroring the server route's per-message handling: decode one frame, validate against a local copy of `WsRequestFrameSchema`, dispatch, encode the reply the way `replyMeta` does. `experimental_upgradeWebSocket` needs a real Vercel runtime, so the socket is faked — everything above it is genuine. Covers: the frame shape the server accepts (and that `reqId`/`type`/ `runId` don't leak into the event meta), payload passthrough, exactly one frame per message, 409 → the same typed error HTTP raises, fail-closed on an unknown reply variant, and reqId correlation across concurrent writes. Plus golden byte fixtures, since the schema copy is the one thing here that can silently drift. This is the "golden-frame interop test" the server spec lists as an open gap; the matching half still needs to land in workflow-server. Verified the conformance suite is not vacuous: flattening the client's frame meta fails 5 of its tests. Co-Authored-By: Shalabh Chaturvedi <shalabhc@users.noreply.github.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * match the HTTP RetryAgent's transient-failure policy on WS HTTP event writes go through an undici RetryAgent (RETRY_AGENT_OPTIONS): 5xx and transient connection errors are retried in-process, honoring Retry-After. The WS path never touches undici, so it shipped with no transient-failure handling at all — a single 503 or a mid-write reset surfaced straight to the step runtime and cost a whole step retry where HTTP would have absorbed it in milliseconds. That gap is invisible in a passing test run: writes still succeed, they just cost far more. So copy the policy rather than reinvent it — [500, 502, 503, 504] plus transport failures, undici's default backoff, Retry-After honored, and 429 deliberately excluded for the same reason RETRY_AGENT_OPTIONS excludes it (a firewall challenge this client cannot solve, which in-process retries only amplify). Adds WsTransportError so retryability is a typed property of the failure rather than something the adapter infers by string-matching. Splits resolveWsTransport()/wsReplyStatus() out of postEventFrameOverWs so the retry loop stays readable. The existing "fails closed on an error frame" test used status 500, which is now absorbed by the retry — switched to 403 so it keeps testing fail-closed rather than accidentally testing no-retry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * lazy-load `ws` so the default HTTP path never evaluates it events-v4.ts imports ws-transport.js unconditionally — the transport gate is a runtime branch, not a build-time one — so a top-level `import { WebSocket } from 'ws'` put `ws` and its optional native accelerators on the module-init path of every deployment, including the overwhelming majority that never opt in and never open a socket. Defer it to the first connect, memoized as a promise so concurrent first connects share one import. WebSocket.OPEN becomes an inlined constant so the readyState check doesn't pull the module in just to read it off the constructor. This does NOT remove the need for the bufferutil/utf-8-validate externals this branch also adds: webpack and Rollup both statically follow a dynamic import(), so the build-time story is unchanged. What it buys is that a deployment which never enables the transport never *evaluates* `ws`, so a mis-bundled accelerator can't break it. The test lives in its own file because vitest caches a vi.mock factory result for the life of the module registry — once any test in a file has connected, the factory never runs again and the counter can't distinguish "loaded lazily" from "loaded at import". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * release idle WS transports instead of renewing them forever The transports map was never pruned and WsEventsTransport had no way to close. Combined with eager reconnect that made a connection immortal by construction: the server drains at its own maxDuration and closes, the client immediately reopens, and the server pins a fresh invocation — for a run that finished long ago. A warm container ended up holding a live socket, and a live server invocation, for every runId it had ever served. workflow-server#683 already lists "one invocation stays resident per run rather than per write" as a known gap; this made it "per run, forever". Add close() plus a 60s idle release. There is no "run complete" signal to hang teardown off — the events adapter is a stateless per-write call — so idleness is the available proxy. 60s sits well below the server's ~680s drain deadline, so the client releases rather than the server reclaiming, and well above the gap between steps of an active run. scheduleReconnect() now bails when closed: close() closes the socket, which fires the same close handler an unexpected drop would, and without the guard the transport would instantly reconnect what it just released. request() revives an idle-closed transport rather than failing the write, re-registering itself only if nothing newer has claimed the map slot. Eviction therefore costs one handshake, not an error. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * refresh the bearer on an auth_expiry drain workflow-server#683 tags a drain frame with why it is closing: max_duration means the socket aged out and a plain reconnect is right, auth_expiry means the *bearer* ran out and reconnecting with the same one just earns a 401. This client logged the drain and ignored the reason, so against #683 an auth_expiry drain would burn all five reconnect attempts against a token the server had already rejected, then give up. Parse the reason (absent reads as max_duration, so this stays correct against the currently-deployed server) and thread forceRefresh through the getHeaders thunk, which triggers @vercel/oidc's refresh path via a wide expirationBufferMs. Worth being precise about when that can actually help. getVercelOidcToken resolves getContext().headers['x-vercel-oidc-token'] ?? env, and refreshToken() only writes the env var — the request-context header wins. So inside a deployed function there is genuinely no fresher token mid-invocation and the refresh is a no-op; outside one (CLI, local dev, a long-lived server) it works. That makes the guard the load-bearing half: if the re-resolved bearer is byte-identical, decline to reconnect, say so, and wait for the next write — which usually arrives on a new invocation carrying a new token. That failure is marked non-retryable so the retry loop doesn't spin on it either. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * document the WS path's instrumentation gap The HTTP branch goes through fetchV4 -> instrumentedFetch, which is not just a fetch wrapper: it opens the OTEL CLIENT span, injects trace context, sets the cache-bust header, emits the DEBUG logs, and routes through the global fetch that Vercel's observability "outgoing requests" view instruments. The comment on fetchV4 records why that matters — bypassing it via undici.request() is exactly what once made v4 event traffic disappear from the log viewer. The WS branch bypasses all of it. With the flag on, per-event writes have no client span, propagate no trace context to workflow-server, and don't appear in the outgoing-requests view; the server's own transport-tagged request metrics are the only remaining signal. That's acceptable for an opt-in POC behind a flag and unacceptable as a default, so write it down where someone deciding to flip the default will read it: instrumenting the transport is a prerequisite for that, not a follow-up nicety. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * ship the ws-accelerator externals instead of documenting a workaround `bufferutil` and `utf-8-validate` are optional native accelerators for `ws`, and neither is installed by default. Every bundler has to be told to leave them alone, for two different reasons: Rollup/Vite/Nitro fail the build outright (`Could not resolve "bufferutil" imported by "ws"`), while webpack bundles the JS wrapper without its native `.node` binding and throws `bufferUtil.mask is not a function` at runtime. The webpack half shipped in `@workflow/next`. The Rollup half only existed in `workbench/vite` and `workbench/tanstack-start` as `nitro.rollupConfig.external` — app configs, not shipped code. So a real user of `@workflow/vite`, `@workflow/nitro`, `@workflow/nuxt`, `@workflow/sveltekit` or `@workflow/astro` hit the same build failure the workbench had already worked around, and had to rediscover the fix. Fix it where it propagates: `workflowTransformPlugin` in `@workflow/rollup`, which all of those integrations already install. It is already the home of exactly this pattern for the optional `@opentelemetry/api` peer, so this sits next to its closest precedent. Note the treatment is deliberately the inverse of the OTEL one, which is externalized only when it *can't* be resolved. The OTEL API must load for tracing to work, so a self-contained output has to bundle it when present. These accelerators must specifically NOT load — they are a performance nicety with a correct try/catch fallback in `ws` — so unconditional external is both simpler and safer than risking a half-bundled native module. The two workbench configs drop their local copies, which is what proves the shipped fix actually works rather than being masked by them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Shalabh Chaturvedi <7066873+shalabhc@users.noreply.github.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * one retry policy for both transports, and no unanswerable waiters Two review findings on the WS events transport. **Retry belongs to `event-retry.ts`, not the adapter.** The WS path had its own retry loop, justified as mirroring undici's `RetryAgent`. That justification was wrong: `RetryHandler` defaults `methods` to GET/HEAD/ OPTIONS/PUT/DELETE/TRACE and nothing overrides it, so the `RetryAgent` never retried an event POST on either transport — which is precisely why `event-retry.ts` exists. Worse, that loop sat *inside* `withEventPostRetry`, so it defeated a compile-checked safety gate: `EVENT_RETRY_ELIGIBILITY` marks `step_started`, `step_retrying` and `hook_received` non-retryable (a replayed `step_started` double-increments `attempt`), and those frames were re-sent up to five times before the gate ever saw a failure. For eligible types the two loops multiplied: 3 outer attempts x 6 inner, with an inner backoff reaching 30s against an outer base deliberately set to 100ms. `postEventFrameOverWs` now makes one attempt and translates failures into the vocabulary that policy already speaks — a transport failure becomes a `WorkflowWorldError` with `code: 'TRANSPORT'`, exactly as `utils.ts` does for a failed `fetch`, and `isRetryableEventPostError` gains one clause keyed on that code. `WsTransportError` loses its `retryable` flag; its only consumer was the deleted loop. Two deliberate consequences. The code-keyed clause broadens HTTP in-process retry to `UND_ERR_CONNECT`, `UND_ERR_CLOSED` and `EAI_AGAIN`, which were in utils.ts's transient set but missing from event-retry.ts's — two hand-maintained lists collapsed into one semantic code. And the stale-token case (drain for auth expiry, refresh yields the same bearer) now gets two in-process attempts that cannot succeed, ~300ms before it falls through to queue redelivery; that is cheaper than keeping a WS-specific policy alive for one call site. `TIMEOUT` is deliberately not in the clause: utils.ts maps a caller-supplied `AbortError` onto it, and a cancelled write must not be re-issued. A status-less reply also stops being a bare `Error` — as one it failed `WorkflowWorldError.is()` and surfaced a protocol version skew as a USER_ERROR. It is now `code: 'PARSE_ERROR'`, the same code utils.ts uses for an unreadable HTTP body, and for the same reason: the write may or may not have landed. **No waiter is left unanswerable.** An undecodable frame, the server's malformed-frame sentinel (`reqId: -1`) and a non-numeric `reqId` were logged and dropped. None can be correlated by construction, so the request that provoked them stayed in `pending` with nothing in existence able to settle it — freed only by the server's own drain (~680s from connect), typically past the invocation's `maxDuration`. Each now fails the connection: every waiter learns why, and the socket is replaced. A reply for an id nobody is waiting on stays log-and-drop, deliberately — that request already settled, so nothing is orphaned, and failing the socket would punish healthy in-flight writes. A per-request deadline backs that up for whatever is left, including a server that accepts a frame and never answers it. Same knob as the HTTP path (`WORKFLOW_REQUEST_TIMEOUT_MS`, 60s), whose doc comment already describes this exact hang-to-SIGTERM pathology. One existing idle-teardown test needed the deadline raised: the idle window and the default deadline are both 60s, so a request could not outlive the former without also outliving the latter. The test is about `inFlight > 0` suppressing the teardown, so it now sets the deadline out of the way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Open the ws socket when the invocation starts, not on its first write Lazily connecting bills the whole handshake — an upgrade round-trip plus the OIDC token mint that rides it — to whichever event a fresh invocation writes first. When that is a `step_started` issued as the step body is already running, the event's server-recorded timestamp lands later than the work it describes: the step looks shorter than it was. That is the shape of the e2e timing failure on this branch, where a 9s step measured 6.5s from `getStepMetadata().stepStartedAt`. The queue handler is the earliest point that knows the run id, and a message delivered for a run means writes are coming, so `warmWsEventsTransport` starts the handshake there. By the first write it is done or in flight, and the write just uses it. Nothing about it is load-bearing: - It doesn't await, and can't fail the handler. A warm that fails logs and stops — a never-opened first connect is precisely the case `connect`'s close handler already declines to retry, so no backoff loop starts for a run that may never write. The first real write connects as it would have anyway, carrying the shared retry policy. - No-op unless `WORKFLOW_EVENTS_TRANSPORT=ws`, and no-op for the api-workflow proxy World, which can't serve an upgrade at all — the same fallback the write path takes. - Warming arms the idle timer as if a request had settled, so an invocation that warms and never writes (a health probe carrying the run id it is about to create) releases its socket on the usual 60s rather than stranding it. The socket is not `unref`'d, so a stranded one would hold this process and a server invocation open. Also closes a race that warming makes reachable: `close()` can only drop the connection it can see, so a release landing mid-handshake left the socket to install itself afterwards onto a transport already evicted from the cache, which nothing would then ever close. The `open` handler now declines to adopt a socket whose transport was released while it was connecting. This was already reachable via the eager reconnect path, just much harder to hit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * changeset: just the env var Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * inline the ws-accelerator predicate at its only call site Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * refactor(world-vercel): trim ws-transport comments Comments were 47% of the file. Cut the historical narration, the restatements of adjacent code, and the repeated rationale (the `unref` reasoning appeared four times, per-connection reqId three), keeping the non-obvious facts: `ws.send()` reports failure via callback instead of throwing, reqId is per-connection so `pending` must be too, the unknown-reqId case is deliberately non-fatal, the auth_expiry same-token bail-out, and why the idle timeout exists at all. No code changes. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * own transport selection in the transport module `events-v4.ts` was assembling the WS transport itself: reading the opt-in flag, resolving the URL, deciding which Worlds can use a socket, minting the per-connection header thunk, and holding the two once-per-process log latches. None of that is about turning an event into a frame, which is what the rest of that file does. Move it next to the socket it configures — `events-v4.ts` now consumes one seam (`resolveWsTransport`) plus the gate, and `queue.ts` gets `warmWsEventsTransport` from the module that owns the warm. `headersToRecord` now lives in `http-core.ts` because both callers need it and neither may import the other: `events-v4` already depends on the transport, so the reverse edge would be a cycle. Test fallout, and the reason the move is worth it: `events-v4-ws.test.ts` mocked `getWsEventsTransport` to observe the resolve step, which no longer intercepts anything now that the call is intra-module — an ESM mock replaces a module's exports, not its own call sites. That mock's tests were only ever about selection, so they move to `ws-transport.test.ts`, where the real selection code runs against the existing fake-socket harness instead of a stub. `resetWsEventsTransportsForTest` clears the log latches so the once-per-process assertions don't depend on test order. What stays behind mocks `resolveWsTransport` and covers what that file is actually for: reply frame in, `Response`-shaped result out — including the null-resolve fallback to HTTP, which nothing covered before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * import `ws` statically The lazy `import('ws')` was there to keep the package off the module-init path of deployments that never opt in — `events-v4.ts` imports this module unconditionally, since the transport gate is a runtime branch. Measured, that buys ~17ms: `require('ws')` is 16.5-18.0ms cold, 13 modules, and neither `bufferutil` nor `utf-8-validate` loads (optional peers, absent by default). Bundle size is identical either way — webpack and Rollup both statically follow a dynamic `import()`, which is why the externals in `@workflow/builders` are unaffected by this change. For 17ms it cost a memoized promise, an inlined `WS_READY_STATE_OPEN` (so a readyState check wouldn't force the module to load just to read a constant off the constructor), and a whole test file — `ws-transport-lazy.test.ts` had to live alone, because vitest caches a `vi.mock` factory result for the lifetime of a module registry, so only a file that connects exactly once can observe the laziness at all. It also skewed the thing this branch exists to measure. The import lands inside the first connect, so on a warm container it is billed to whichever event write opens the socket, inflating the timestamp of the step it labels — the same distortion the queue pre-warm was added to remove. Also drops `WS_READY_STATE_OPEN` in favour of `WebSocket.OPEN`, now that reading it is free. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * tighten the comments on the ws transport Comments only — no code changes in this commit. Cuts ~150 lines of prose across the WS additions. The rule applied: keep the design factors a future reader needs (why the connection is scoped to a run, why a bad reply takes the socket down, why the accelerators are externalized unconditionally, why `TIMEOUT` is excluded from the `TRANSPORT` classification) and drop the narrative of how the code got here — which revision did what, what an earlier attempt got wrong, what was measured on the way. That history lives in the PR and the git log, where it doesn't have to be re-read on every visit to the file. Biggest reductions: the retry essay above `postEventFrameOverWs` (30 lines to 11), the flag's OTEL-gap note (34 to 13), the OIDC refresh explainer (26 to 14), the accelerator rationale in `@workflow/builders` (26 to 14), and the conformance suite's header (28 to 17). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * inject W3C trace context on the ws upgrade Frames carry no headers, so the upgrade is the only place this transport can propagate context; the server parents a run's event spans to whichever invocation opened the socket. Covered in trace-propagation.test.ts, both with and without an active span. Splits the opt-in gate into an import-free ws-transport-enabled.ts so callers can answer it without loading this module (used by the next commit). Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * load the ws transport module only when it is enabled Both call sites read the gate from the import-free module and dynamically import ws-transport.js behind a true result, so a deployment on the HTTP default never pays ws's ~17ms of module init. The queue pre-warm absorbs it for one that opted in, keeping it off the first event write. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * document WORKFLOW_EVENTS_TRANSPORT as experimental Names the instrumentation gap (no client span per write) and the proxy path where the variable is ignored. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * correct why the ws accelerators are externalized No bundler fails the build on the unresolvable require — verified against Rollup 4.62. webpack half-bundles the native module and Vite substitutes a stub that makes the require succeed; both leave bufferUtil.mask undefined and throw only once a frame reaches the native masker at 48 bytes, which every CBOR event frame does. Same claim was repeated in the rollup plugin and its test. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * trim the WORKFLOW_EVENTS_TRANSPORT docs to user level Mirrors the other Vercel World env vars: same facts on both pages, each in its page's format. The instrumentation and socket-lifetime detail belongs in the code, not in a user-facing reference. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * cover the vite bundler in the ws transport lane Vite substitutes a stub for ws's absent native accelerators rather than failing the require, so nothing catches it until a masked frame reaches 48 bytes — and this job's three existing lanes are esbuild, turbopack and nitro. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * claim only what is measured about rollup and the ws accelerators The rationale asserted plain Rollup was "safe by accident" via a mechanism only ever observed in a minimal repro. Nitro traces and externalizes `ws` in a production build, so the bundled path is not reached there at all. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Give the events socket an explicit lifetime instead of an idle timer `openWsChannel` / `closeWsChannel` bracket one invocation of the flow route, and are the only calls anywhere that create a channel. Writes ask `resolveWsTransport` whether one is open — a lookup now, never a create — and take pooled HTTP when it says no. That removes the reason the idle timeout existed. A lazily-created socket has no owner, so a timer was the only thing able to end it, and the socket is not `unref`'d: the process could not exit, and a server invocation stayed pinned, for the full window past the last write. It also settles `run_created`. The trigger path opens no channel, so a lone write no longer pays for a handshake it cannot amortize — `start()` runs in an arbitrary request handler with no boundary the SDK can see. Refcounted rather than a flag: inline step executions ride the flow topic on per-step topics, so a run's steps can be concurrent invocations in one instance sharing the channel, and the first to finish must not cut the others short. A failed connect closes the channel so the invocation's writes fall back to HTTP instead of each paying its own doomed handshake. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Name the one reply header the WS path does not map The server copies six headers into an `event_ack`'s meta and this record maps five. The sixth, `X-API-Deprecated`, is inert today — the v4 route's middleware chain has no deprecation middleware to set it — but the record is the only header source a WS reply has, so an unmapped key is gone rather than merely unread, which is not true of the `Response` the HTTP path returns. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * docs: note that WORKFLOW_EVENTS_TRANSPORT=ws is ignored on the proxy path The api-workflow proxy is an HTTP-only REST gateway and does not forward a WebSocket upgrade, so a World configured with projectConfig keeps writing events over HTTP regardless of the setting. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * ci: gate the ws-transport e2e lanes on a label Three real `vercel deploy`s per run is too much to charge every unrelated PR in the repo for a transport that is off by default. PRs opt in with `ws-transport-test` (or `workflow-server-test`, which already exists to test the half of this the protocol lives in); main keeps the signal on every commit. The required aggregate has to allow the lane to be skipped in that case, so its status is asserted only when the lane was actually supposed to run. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * chore: regenerate pnpm-lock against current main main resolved `ws` to 8.20.0 as a transitive peer; this branch adds it as a direct dependency of world-vercel and floats it forward, which rewrites every `openai@x(ws@y)` peer key in the lockfile. Merging main textually combined the two, leaving those keys pointing at a `ws` entry the merged file no longer had — `--frozen-lockfile` then failed with ERR_PNPM_LOCKFILE_MISSING_DEPENDENCY on the PR's merge ref. Regenerated from main's lockfile so ours is a minimal delta on top of it. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * fix(world-vercel): align ws on the version main already resolves The lockfile broke on the PR's merge ref, not on this branch's head: main resolves ws@8.20.0 as a transitive peer, and a `^8.21.1` direct dep here floated it forward, rewriting all 73 `(ws@8.20.0)` peer keys. Git merged the two lockfiles without a conflict but left main-side keys pointing at a ws entry the merged file no longer had, so `--frozen-lockfile` failed with ERR_PNPM_LOCKFILE_MISSING_DEPENDENCY. `^8.20.0` resolves to the copy main already has, so the lockfile delta is the two importer entries instead of a repo-wide rewrite that re-breaks every time main moves. Also keeps one ws in the store rather than two. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Bind the channel release to the instance it claimed closeWsChannel resolved the transport by URL, but the refcount lives on the instance. A channel is evicted from the map as soon as it closes — a refused upgrade does that on the connect path — so the next opener for the same run registers a different instance under the same URL, and the first invocation's close then decremented that one instead. It dropped a socket a live invocation was still writing over, and for the event types EVENT_RETRY_ELIGIBILITY marks non-retryable there is no second attempt to carry the in-flight write over HTTP. openWsChannel now returns an idempotent release closed over the transport it incremented, and queue.ts holds that instead of re-resolving the run. The close awaits the open's own promise, so it also can no longer land ahead of the claim it releases. Also names the scope of the connect-failure de-opt: it covers the handshake only, so a channel that connects and then fails every write keeps taking the WS path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Decode a transport result, not a Response Main extracted the v4 POST decode into a helper typed `Response` while this branch narrowed the POST result to `FrameResponseLike`, because the WS branch synthesizes its result rather than holding a real `Response`. The two merge without a textual conflict and then fail to typecheck. Widen the helper: it reads only the two members `FrameResponseLike` declares, and a `Response` still satisfies them, so the HTTP call sites are unchanged. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Re-run CI Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Re-run CI Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Reconcile the WS transport with main's v4 POST rework main moved the materialized POST result off the `x-wf-*` response headers and onto a typed CBOR body, and added a second response shape: two callers now POST with `Accept: application/vnd.workflow.v4-frames` and read back a sentinel-terminated sequence of frames. A frame stream has no representation in a protocol that pairs one reply frame with one request frame, so the WS switch moves off the shared poster and onto `createWorkflowRunEventV4` alone — the materialized write, which is the hot per-step path this branch exists to shorten. `run_started` and the `hook_received` preload stay on HTTP. `decodeCreateEventResponse` takes `FrameResponseLike` rather than `Response` because the WS branch has none to hand over; a real `Response` satisfies the interface, so the HTTP callers are unchanged. The ids now come out of the CBOR body, so `replyMetaToHeaderRecord` no longer maps any `x-wf-*` name — only the two headers `errorFromV4Response` reads. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Re-run CI Resample the WS-arm sleepingWorkflow failure: it has now recurred on a second axis (nextjs-turbopack, 7709ms; previously vite, 7570ms), so the arm needs more samples before the skew can be called WS-specific or repo-wide flake. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Re-run CI Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Re-run CI Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * blank * blank --------- Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com> |
||
|
|
4bb86d3054 |
feat(world-vercel): support Hook minimum retention (#3286)
* feat(world-vercel): support Hook minimum retention * fix(core): fail deterministic Hook validation |
||
|
|
99f4aeb03d |
feat(world-postgres): support Hook minimum retention (#3276)
* feat(world-postgres): retain hook tokens after runs end
* refactor(world-postgres): reuse terminal run statuses
* docs: note Postgres Hook retention support
* fix(world-postgres): expose hook retention deadline
* Fix: Exhaustive `Record<AttributeKey, ...>` in `attribute-panel.tsx` is missing the `tokenRetentionUntil` key that was added to `HookSchema`, causing TS2741 and breaking every Vercel build.
This commit fixes the issue reported at packages/web-shared/src/components/sidebar/attribute-panel.tsx:426
## Bug
Commit `ad58321` added `tokenRetentionUntil: z.coerce.date().optional()` to `HookSchema` in `packages/world/src/hooks.ts:106`. This adds `tokenRetentionUntil` to the inferred `Hook` type.
In `packages/web-shared/src/components/sidebar/attribute-panel.tsx`, `AttributeKey` is a union that includes `keyof Hook`, so `tokenRetentionUntil` becomes a required member of the **exhaustive** `Record<AttributeKey, (value: unknown, context?: DisplayContext) => ...>` object literal `attributeToDisplayFn` (starting at line ~426).
Because the literal had no `tokenRetentionUntil` entry, `tsc` fails:
```
src/components/sidebar/attribute-panel.tsx(426,7): error TS2741:
Property 'tokenRetentionUntil' is missing in type '{ ... }' but required in type
'Record<AttributeKey, (value: unknown, context?: DisplayContext | undefined) => ReactNode>'.
```
This breaks `@workflow/web-shared#build` and therefore every Vercel deployment (17 failing deployments observed, all with this identical error).
## Fix
Added a `tokenRetentionUntil` entry to `attributeToDisplayFn`, placed alongside the other Hook date fields (`lastReceivedAt`, `disposedAt`):
```ts
tokenRetentionUntil: timestampWithTooltipOrNull,
```
`tokenRetentionUntil` is a `Date` field, and `timestampWithTooltipOrNull` (defined at line 402) is the display helper used by all the other surfaced date fields (`createdAt`, `startedAt`, `completedAt`, `retryAfter`, `resumeAt`, `occurredAt`). Given the intent of `ad58321` was to expose the hook retention deadline, surfacing it as a tooltip-annotated timestamp is the consistent choice.
Only `attributeToDisplayFn` is a fully exhaustive `Record<AttributeKey, ...>`; the other maps are `Partial<...>` / `Set`, so no other edits are required.
## Verification
`node_modules` are not installed in this sandbox, so `tsc` could not be executed directly. Verified structurally instead: the newly added `tokenRetentionUntil` entry (line 449) references `timestampWithTooltipOrNull`, which is defined in-file at line 402 and already used by the sibling date entries, so the fix satisfies the missing-key requirement without introducing new type errors.
Co-authored-by: Vercel <vercel[bot]@users.noreply.github.com>
Co-authored-by: VaguelySerious <mittgfu@gmail.com>
* docs(world-postgres): clarify expired hook rows
* feat(world-postgres): enforce Hook retention limit
* fix(world): remove duplicate Hook retention field
* fix(web-shared): remove duplicate retention renderer
* test(world): remove redundant retention coercion case
---------
Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
Co-authored-by: Vercel <vercel[bot]@users.noreply.github.com>
Co-authored-by: VaguelySerious <mittgfu@gmail.com>
|
||
|
|
e6f1b6f548 |
feat(world-local): support Hook minimum retention (#2866)
* feat(core): add hook token retention contract * refactor(core): constrain hook retention options * fix(core): preserve boolean hook visibility options * revert(core): preserve HookOptions interface * docs(core): clarify retained conflict ownership * docs(core): retain newest-wins conflict pattern * docs(core): simplify hook retention guidance * docs(core): explain retained token cleanup * docs(core): simplify idempotency guidance * docs(core): clarify retained token results * refactor(core): rename hook token expiration option * chore(core): name hook expiration changeset * docs(core): simplify Hook expiration language * docs(core): clarify Hook expiration deadline * docs(core): remove Hook deadline caveat * refactor(core): align Hook expiration field names * docs(core): narrow Hook expiration documentation * docs(core): clarify hook expiration availability * Update packages/core/src/workflow/hook.ts Co-authored-by: Peter Wielander <mittgfu@gmail.com> Signed-off-by: Nathan Colosimo <110621881+NathanColosimo@users.noreply.github.com> * docs(core): clarify Hook token expiration behavior * docs(core): explain active Hook expiration behavior * feat(world): advertise hook ttl capability * fix(core): validate hook ttl capability after main merge * refactor(core): rename hook expiry to minimum retention * docs: keep hook retention guidance on v5 * docs: define retained run availability * fix(core): validate Hook retention at creation * feat(core): define retained Hook lookup semantics * refactor(core): simplify hook retention checks * feat(world-local): support Hook token expiration * fix(world-local): make hook recovery atomic * refactor(world-local): align Hook minimum retention * fix(world-local): preserve Hook creation order * fix(world-local): expose retained Hooks consistently * refactor(world-local): simplify retained hook storage * fix(world-local): allow stale lock recovery * refactor(world-local): simplify hook retention storage Signed-off-by: Nathan Colosimo <110621881+NathanColosimo@users.noreply.github.com> * fix(world-local): serialize expired hook token handoff Signed-off-by: Nathan Colosimo <110621881+NathanColosimo@users.noreply.github.com> * fix(world-local): preserve hook creation order Signed-off-by: Nathan Colosimo <110621881+NathanColosimo@users.noreply.github.com> * refactor(world-local): clarify hook availability cleanup * docs: note Local World Hook retention support * fix(world-local): harden hook retention persistence * fix(web-shared): render hook retention deadline * fix(world-postgres): exclude unsupported hook retention * feat(world-local): enforce Hook retention limit * docs(world-local): clarify retention limit error * docs(world): clarify Hook retention deadline * docs(hooks): link retention configuration --------- Signed-off-by: Nathan Colosimo <110621881+NathanColosimo@users.noreply.github.com> Co-authored-by: Peter Wielander <mittgfu@gmail.com> |
||
|
|
1471f252fa | [core] Gate event creation on the loaded event count and restart replays in-process (#3145) | ||
|
|
2677653759 |
fix(world-local): bound stalled queue deliveries (#3255)
Signed-off-by: Andrew Barba <barba@hey.com> |
||
|
|
62d570ed4b | Remove retired v1 step route plumbing (#3061) | ||
|
|
fc81f4502f |
perf(core): immediate leading-edge dispatch for idle streams (flush window default 0) (#3088)
* perf(core): immediate leading-edge dispatch for idle streams (flush window default 0) Production producer-rate data (24h of client flush spans): most agents average 1.03-1.21 chunks per flush with 87-98% single-chunk flushes and >70% of chunks arriving more than 10ms after the previous request had already settled — a fixed 10ms leading window batches almost nothing for them while adding ~20% to isolated-chunk publish latency (~50ms median RTT). The one bursty producer (avg ~4-8 chunks/flush) gets its batching from in-flight accumulation, which does not depend on the window at all. The leading chunk of an idle sink now dispatches immediately by default (window 0): first chunk goes out at once, chunks arriving during its request coalesce into the next group, and each settle dispatches the accumulated group immediately — path-independent batching with no fixed tax on slow producers. A positive WORKFLOW_STREAM_FLUSH_INTERVAL_MS (or world.streamFlushIntervalMs, applying from the second group) opts into a windowed leading edge for slow-but-steady producers that prefer larger groups over first-chunk latency. Early-ack, the durability drain barrier, wire caps, and backpressure bounds are unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Update packages/world/src/interfaces.ts Co-authored-by: Peter Wielander <mittgfu@gmail.com> Signed-off-by: Karthik Kalyan <105607645+karthikscale3@users.noreply.github.com> * review: env var overrides world streamFlushIntervalMs; world option governs the leading edge too WORKFLOW_STREAM_FLUSH_INTERVAL_MS, when set, now takes precedence over world.streamFlushIntervalMs; otherwise the world option applies from the very first chunk (no more second-group lazy quirk). Deciding waits for the world when needed, which adds no latency: sendGroup awaits the same promise before any request can leave. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Signed-off-by: Karthik Kalyan <105607645+karthikscale3@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Peter Wielander <mittgfu@gmail.com> |
||
|
|
cdb3db4049 |
fix(world-postgres): abort stalled HTTP delivery on shutdown (#3064)
Signed-off-by: Joey Hotz <joeyhotz1@gmail.com> |
||
|
|
a5e6f1167a |
feat(core): add experimental Hook minimum retention (#2865)
* feat(core): add hook token retention contract * refactor(core): constrain hook retention options * fix(core): preserve boolean hook visibility options * revert(core): preserve HookOptions interface * docs(core): clarify retained conflict ownership * docs(core): retain newest-wins conflict pattern * docs(core): simplify hook retention guidance * docs(core): explain retained token cleanup * docs(core): simplify idempotency guidance * docs(core): clarify retained token results * refactor(core): rename hook token expiration option * chore(core): name hook expiration changeset * docs(core): simplify Hook expiration language * docs(core): clarify Hook expiration deadline * docs(core): remove Hook deadline caveat * refactor(core): align Hook expiration field names * docs(core): narrow Hook expiration documentation * docs(core): clarify hook expiration availability * Update packages/core/src/workflow/hook.ts Co-authored-by: Peter Wielander <mittgfu@gmail.com> Signed-off-by: Nathan Colosimo <110621881+NathanColosimo@users.noreply.github.com> * docs(core): clarify Hook token expiration behavior * docs(core): explain active Hook expiration behavior * feat(world): advertise hook ttl capability * fix(core): validate hook ttl capability after main merge * refactor(core): rename hook expiry to minimum retention * docs: keep hook retention guidance on v5 * docs: define retained run availability * fix(core): validate Hook retention at creation * feat(core): define retained Hook lookup semantics * refactor(core): simplify hook retention checks * docs(core): simplify retained conflict example * docs(core): flatten forward-to-owner example --------- Signed-off-by: Nathan Colosimo <110621881+NathanColosimo@users.noreply.github.com> Co-authored-by: Peter Wielander <mittgfu@gmail.com> |
||
|
|
8a872529fe |
docs: make /worlds the canonical home for World docs (#2934)
* docs: make /worlds the canonical home for World docs The world pages (Local/Postgres/Vercel) and Building a World were duplicated inside the v4 and v5 docs trees while /worlds/[id] rendered the v4 copy — hiding v5-only content like multi-region and leaving two diverging sources of truth. - Move world docs to an unversioned docs/content/worlds/ collection (based on the v5 copies, with inline 4.x callouts for factory naming and 5.x-only env vars), rendered at /worlds/* - Add /worlds/building-a-world; flatten the docs Deploying section to a single intro page and drop its Rocket icon - Point every link, frontmatter ref, and worlds-manifest docs field at /worlds/*; add redirects for the removed v5 and building-a-world URLs - Keep world docs on agent-facing surfaces: search, llms.txt, sitemap.md/.xml, and .md exports now serve the worlds collection - Extend the docs link linter to validate worlds pages (with heading anchors) and their outgoing links Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> * docs: version the world docs like the docs trees (v4/v5 switcher) Instead of a single unversioned copy, world docs now follow the same versioning strategy as the docs pages: content/worlds/v4 is served at /worlds/* (current) and content/worlds/v5 at /v5/worlds/*, restoring the original per-version content. Each world detail page (and Building a World) renders the docs version switcher — the worlds listing page has no natural home for it, so it lives on the world pages themselves. - Render-time href rewriting on v5 pages now covers /worlds/... links (shared rewriteHrefForVersion helper, also used by the v5 docs and cookbook routes), and the markdown-export rewrite does the same - v5 world pages are noindexed with a canonical to /worlds/<id>; community worlds stay unversioned (/v5/worlds/<id> redirects) - /v5/docs/deploying/world/* redirects now land on /v5/worlds/*; /v5/worlds and /v5/worlds/compare redirect to the unversioned pages - Link linter models the versioned worlds URL spaces (v5 pages resolve /worlds hrefs against the v5 collection); sitemap.md and the .md export routes cover /v5/worlds/* Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> * docs: fix v4 multi-region anchor and tighten version-prefix matching Address PR review: - The v4 Deploying page linked /worlds/vercel#multi-region, but the Multi-region section only exists on the v5 world page; use the explicit cross-version /v5/worlds/vercel#multi-region link (this was the Docs Links CI failure) - rewriteHrefForVersion now uses the boundary-checked hasPathPrefix (shared leaf module lib/geistdocs/path-prefix.ts, also used by source.ts) instead of bare startsWith - buildVersionUrl's shared-route fast path is segment-based rather than substring includes() Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> --------- Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |