* test: e2e coverage for run-idempotency conflict-handling strategies (#2387)
* test: e2e coverage for run-idempotency conflict-handling strategies
Covers the patterns documented in foundations/idempotency:
- claim-only hook mutex: token claimed and held with no payload data,
duplicate identifies the owner, token released after completion
- adopt the owner's result via conflict.returnValue
- signal the owner: duplicate forwards its payload via resumeHook
- supersede: duplicate cancels the owner and reclaims the token
- route-side resume-or-start retry pattern reaching the started run
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: fix adopt-owner-result race — gate owner completion on observed conflict
On slow runtimes the duplicate's first invocation could land after the
owner completed and released the token, making the duplicate a fresh
owner that waits forever for a payload (90s timeout across CI matrices).
Poll the duplicate's event log for hook_conflict before resuming the
owner, and widen the test timeout for the added gate budget.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* review: assert superseded owner's returnValue rejection; empty changeset
- Await run1.returnValue and assert WorkflowRunCancelledError so the
cancellation is verified end-to-end and no rejection leaks from the
supersede test.
- Test-only PR: use an empty changeset.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: retrigger preview deployments (turbopack deployment for 2e9d000 wedged in esbuild hang)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci: bust poisoned turbo cache entry for nextjs-turbopack build
The 2e9d000 deployment's next build crashed in an esbuild hang but its
task (70724907c9dd3a29) was recorded into the turbo remote cache anyway,
so every subsequent build with the same input hash replays the broken
artifact (missing routes-manifest). Change a build input to force a
fresh execution.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
* fix(backport): adapt conflict-handling tests to stable's getConflict/getRun API
The backport of #2387 used APIs that only exist on `main`, breaking the
nextjs-turbopack/webpack builds and the e2e suite on `stable`:
- `hookAdoptOwnerResultWorkflow`/`hookSupersedeOwnerWorkflow` read
`conflict.returnValue`/`conflict.cancel()`, but on `stable`
`getConflict()` resolves with `{ runId }`. Resolve the owning run via
`getRun(conflict.runId)` inside a step (the documented stable pattern)
to await its result / cancel it.
- Import `resumeHook` from `workflow/api` in 99_e2e.ts (was used by
`forwardPayloadToOwner` but never imported).
- Convert the backported `waitForHook(token, { runId })` call sites to
`waitForHookState(token, predicate)`; `waitForHook` does not exist on
`stable` (#2405 standardized on `waitForHookState`).
Both workbench builds and `biome check` pass locally.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(backport): import HookNotFoundError in e2e test
The backported conflict tests call `HookNotFoundError.is()` in
`hookClaimOnlyMutexWorkflow` (token-release wait) and the resume-or-start
route test, but the import was never carried into the stable test file —
causing a runtime `ReferenceError: HookNotFoundError is not defined`.
Import it from `@workflow/errors` (matches `main`).
Verified locally against nextjs-turbopack: the two previously-failing
tests plus the adopt/signal/supersede rewrites all pass (5/5).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Pranay Prakash <pranay.gp@gmail.com>
* feat: replace hook.hasConflict with hook.getConflict (Promise<Run | null>)
hasConflict's boolean didn't expose WHICH run owns the token, so the
duplicate run couldn't act on the conflict. getConflict resolves with
null once registration commits, or with a Run handle for the conflicting
run — letting the workflow return/log the owner's runId, inspect its
status, await its result, or cancel it and continue, all in code.
The workflow-mode create-hook module exposes the bundle's compiled Run
class (durable step-proxy methods) on a well-known symbol so the host-
side hook consumer can construct the conflicting run inside the VM.
Contexts without the class (plain unit tests) fall back to a { runId }
object, which is also the documented v4 shape (no native Run
serialization in v4).
* fix: never resolve getConflict with a non-Run fallback shape
getConflict's contract is Promise<Run | null>. In the degenerate cases
where a real Run cannot be constructed — a hook_conflict event persisted
by an old world without conflictingRunId, or a context that never loaded
the workflow-mode create-hook module — reject with HookConflictError
instead of resolving with a { runId }-shaped impostor.
Test harnesses now register the Run class on the (VM) globalThis like
real bundles do.
* refactor: make getConflict a method — hook.getConflict()
A property getter that triggers registration/suspension reads as passive
state; a method makes the side effect explicit. Update implementation,
types, tests, e2e workflows, docs, and changeset.
* review: guard Run class registration, fix anchors, clarify changeset
- Only register WORKFLOW_RUN_CLASS when the workflow runtime is present
(WORKFLOW_CREATE_HOOK installed on globalThis), so host imports of the
workflow-mode module neither mutate the host global nor expose the
non-step-proxy host Run.
- Drop #run-idempotency link fragments — that section lands in the
stacked docs PR (#2011), which restores the anchored links.
- Note in docs that getConflict() rejects with HookConflictError for
legacy hook_conflict events lacking the owner's run ID.
- Changeset now calls out the hasConflict -> getConflict() replacement.
* refactor: resolve the conflicting Run through the serialization class registry
Replace the bespoke WORKFLOW_RUN_CLASS global with the registry the
serialization pipeline already uses to revive Run instances:
- The SWC plugin already auto-registers the workflow bundle's compiled
Run in globalThis[workflow-class-registry], but under a path-derived
classId the host cannot know statically. The workflow-mode create-hook
module now aliases it under a stable id (class//workflow//Run) via a
new aliasSerializationClass() helper (a plain registry entry —
registerSerializationClass cannot be reused since the plugin's IIFE
already defined the non-configurable classId property).
- createConflictingRun() looks the class up with
getSerializationClass(RUN_CLASS_ID, ctx.globalThis) and constructs
through its WORKFLOW_DESERIALIZE hook, exactly as the Instance reviver
would for a serialized Run crossing from a step into the workflow.
- Because the registry is keyed per-global, no environment guard is
needed: a stray host-side import registers the host Run on the host
registry, which is the correct class for that context. The
WORKFLOW_CREATE_HOOK guard, the ??=, and the WORKFLOW_RUN_CLASS symbol
are all deleted.
Verified: 1156 core unit tests; compiled workbench bundle contains the
stable alias alongside the plugin's path-derived registration with zero
WORKFLOW_RUN_CLASS references; all 5 hookGetConflict e2e tests pass
against a local nextjs-turbopack dev server, including conflict
resolution reading conflict.status through a durable step.
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Nathan Rajlich <n@n8.io>
* feat: add hook ready promise
* test: cover hook ready continuation scheduling
* feat: replace hook.ready with hook.hasConflict (Promise<boolean>)
- hook.hasConflict resolves true when the token is owned by another
active hook, false once registration is committed — no throw, so
workflows can branch on conflicts early. Awaiting it suspends the
workflow to commit the hook registration (createHook alone does not).
- Chain the already-created fast-path through promiseQueue so
resolution order matches event-log order (review feedback).
- Skip inline step execution when a suspension has an awaited hook
creation so the hasConflict continuation can advance independently
of step execution (review feedback).
- Update unit tests, e2e tests, workbench workflows, and v4/v5 docs.
* docs: fix inconsistent hasConflict bullet in create-webhook reference
State both resolution values explicitly (true = token already owned,
false = registered) instead of a parenthetical that only described the
false case.
* docs: require docs preview links in PR descriptions for docs changes
* docs: restore SWC Plugin heading in AGENTS.md
---------
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Nathan Rajlich <n@n8.io>
* [world-vercel] Validate ref resolve responses before use
When workflow-server returns a ref body to the SDK, the bytes are
fed into the workflow runtime's event log and deserialized via
`decodeFormatPrefix`. The SDK always writes ref payloads with at
least a 4-byte format prefix (see `encodeWithFormatPrefix` in
`@workflow/core`), so a zero-byte response — or one whose length
disagrees with `Content-Length` — is never a valid stored value.
Before this change, `resolveRefDescriptor` had no validation: a
200 with an empty body would be passed downstream as a zero-length
Uint8Array, which then failed deep inside replay with:
Data too short to contain format prefix: expected at least 4 bytes, got 0
By that point the workflow's in-memory event snapshot is already
poisoned with the empty payload, so every subsequent replay
deterministically reproduces the same failure, downstream
`resumeHook()` calls surface as `Hook not found`, and the run
only unsticks when stale-run cleanup terminates the sandbox.
This catches the failure at the transport boundary instead, where
it can be retried as a `WorkflowWorldError`. Both an empty body
and a length mismatch (truncated streaming response) are rejected.
This is the SDK-side companion to vercel/workflow-server#432, which
adds the same validation on the server side.
* Address review: reject <4-byte bodies, handle malformed Content-Length
Three review changes:
1. Reject any body shorter than the 4-byte format-prefix length, not
just zero-byte bodies. The SDK guarantees every stored ref payload
starts with a 4-byte format prefix (FORMAT_PREFIX_LENGTH in
@workflow/core), so a 1-3 byte body would also fail downstream
replay with the same 'Data too short to contain format prefix'
error this PR exists to prevent.
2. Parse Content-Length safely with parseInt + Number.isFinite +
non-negative checks instead of bare Number(). A non-numeric value
like 'abc' would otherwise produce NaN and silently surface as a
'truncated' error, masking the real cause. Malformed values are
treated as absent; the minimum-length check still defends against
actual truncation in that case.
3. Add tests for the truncated-body-without-Content-Length case
(chunked transfer where Content-Length validation can't see the
truncation), and for a malformed Content-Length header that should
be ignored rather than misreported as truncation.
The validation logic also moves into a small assertValidRefBody
helper to keep the inner trace function under the noExcessiveCognitiveComplexity limit.
* Address review: scope 4-byte minimum to binary refs, strict Content-Length parsing
- Only apply the 4-byte format-prefix minimum to application/octet-stream
payloads; CBOR refs can legitimately be 1-byte primitives (true/0/null).
- Require Content-Length to be a plain run of digits before comparing;
parseInt would otherwise accept numeric-prefixed garbage ('12junk' -> 12).
- Make the changeset succinct.
* Address review: skip Content-Length check for compressed responses
fetch/undici transparently decompresses gzip/br bodies but leaves
Content-Length describing the encoded (compressed) size, so comparing it
against the decompressed byteLength would reject valid compressed refs as
a phantom 'ref-body-length-mismatch'. Skip the comparison when a
non-identity Content-Encoding is present; an absent or 'identity' encoding
is still validated. Adds regression tests for both cases.
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
The nitro-native-build changelog sample calls useStorage() (a Nitro
server-side auto-import) from a step, which the docs type-checker
couldn't resolve and failed with TS2552. Add liberal global
declarations for useStorage/useDatabase/useRuntimeConfig so Nitro
auto-imports type-check in docs samples.
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(next): always apply turbopack content condition regardless of builder mode
When lazy discovery is enabled (deferred builder), shouldApplyTurboCondition
was false, so turbopack.rules were added with no content filter — causing the
workflow loader to run on every JS/TS file. Apply the content condition
unconditionally so the loader only fires on files with workflow directives.
* add changeset
---------
Signed-off-by: Will Binns-Smith <wbinnssmith@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: JJ Kasper <jj@jjsweb.site>
* fix(world-vercel): retry transient response-body parse failures in the HTTP client
A sporadic failure reading/decoding a 2xx response body (truncated or
terminated stream, connection reset mid-body, or a gateway returning a
non-CBOR/JSON body) was surfaced immediately as a PARSE_ERROR. The
shared RetryAgent only retries connection/5xx failures — body
consumption happens after it returns the response, so these escape its
retry logic.
Retry such failures inside `makeRequest` with bounded exponential
backoff, scoped to idempotent methods (GET/HEAD) so writes are never
replayed. This fixes the reported `events.list` parse failure at the
adapter layer.
* fix(core): propagate exhausted transient world errors to the queue
Pairs with the world-vercel in-adapter retry: when a response-body parse
failure survives the adapter's retries (or comes from a non-idempotent
write that is never retried in-process), it must not fail the run.
Re-throw such transient world errors from the replay loop so they
propagate to the queue handler, which replays the whole run — safe
because replay is idempotent. Schema-validation contract errors stay
fatal.
* Revert "fix(core): propagate exhausted transient world errors to the queue"
This reverts commit 7bb62e9f81.
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Nathan Rajlich <n@n8.io>
* test(e2e): cover WritableStream passed as start() argument
Adds an e2e workflow + test where a parent workflow gets a WritableStream
via getWritable(), forwards it through start() to a child workflow, and
the child step writes raw bytes to it. Asserts the external reader on
the parent's stream observes the exact bytes the child wrote.
* fix(core): avoid double-framing when WritableStream is forwarded via start()
When a workflow's getWritable() handle is passed across start() to a
child workflow, the parent step's reviver wraps it in a serialize
transform that pipes into a workflow server stream. Until now,
getExternalReducers.WritableStream then installed a second serialize
transform on top of that — so every chunk the child step wrote got
devalue-framed twice but only deframed once on the reader side, and
external consumers saw the inner frame instead of the original bytes.
Fix: tag every user-visible writable that's already backed by a
workflow server stream with its (runId, name). When the external
reducer recognizes those tags during dehydration, it bridges bytes
straight from the new child-side server stream to the original server
stream instead of piping through the user's writable. That leaves the
producer-side serialize transform (installed once by the child's step
reviver) as the only framing layer in the chain.
* fix(core): forward (runId, name) when a tagged WritableStream crosses start()
Replaces the previous in-process bridge with first-class writable
forwarding at the descriptor level. When a parent workflow's
getWritable() handle is passed as an argument to a child workflow,
the dehydrated descriptor now carries the original (runId, name).
The child run's step-side reviver opens the writable against the
parent's server stream directly and resolves the parent run's
encryption key (encrypt-only) via getEncryptionKeyForRun.
This removes the architectural limitation that the bridge could
only stay alive for the duration of the parent step process — on
Vercel that capped forwarding at ~15 minutes regardless of the
child run's lifetime, dropping any writes the child made after the
parent step process exited.
importKey() now accepts a usages parameter, defaulting to
['encrypt', 'decrypt']. The cross-run forwarding path imports with
['encrypt'] only so a compromised child run cannot decrypt any
existing data on the parent's stream — only contribute new writes.
* test: rename writable-forwarded workflows and cover step-context getWritable()
Addresses PR review:
- Rename writableForwardedToChildChildWorkflow → writableForwardedChildWorkflow
(drops the duplicated 'Child' segment).
- Split writableForwardedToChildWorkflow into two variants covered by a
test.each: writableForwardedFromWorkflowWorkflow (workflow-context
getWritable, the original test) and writableForwardedFromStepWorkflow
(step-context getWritable passed directly into start() from the same
step that called getWritable()).
- Terser changeset description.
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Nathan Rajlich <n@n8.io>