* [world-vercel] Validate ref resolve responses before use
When workflow-server returns a ref body to the SDK, the bytes are
fed into the workflow runtime's event log and deserialized via
`decodeFormatPrefix`. The SDK always writes ref payloads with at
least a 4-byte format prefix (see `encodeWithFormatPrefix` in
`@workflow/core`), so a zero-byte response — or one whose length
disagrees with `Content-Length` — is never a valid stored value.
Before this change, `resolveRefDescriptor` had no validation: a
200 with an empty body would be passed downstream as a zero-length
Uint8Array, which then failed deep inside replay with:
Data too short to contain format prefix: expected at least 4 bytes, got 0
By that point the workflow's in-memory event snapshot is already
poisoned with the empty payload, so every subsequent replay
deterministically reproduces the same failure, downstream
`resumeHook()` calls surface as `Hook not found`, and the run
only unsticks when stale-run cleanup terminates the sandbox.
This catches the failure at the transport boundary instead, where
it can be retried as a `WorkflowWorldError`. Both an empty body
and a length mismatch (truncated streaming response) are rejected.
This is the SDK-side companion to vercel/workflow-server#432, which
adds the same validation on the server side.
* Address review: reject <4-byte bodies, handle malformed Content-Length
Three review changes:
1. Reject any body shorter than the 4-byte format-prefix length, not
just zero-byte bodies. The SDK guarantees every stored ref payload
starts with a 4-byte format prefix (FORMAT_PREFIX_LENGTH in
@workflow/core), so a 1-3 byte body would also fail downstream
replay with the same 'Data too short to contain format prefix'
error this PR exists to prevent.
2. Parse Content-Length safely with parseInt + Number.isFinite +
non-negative checks instead of bare Number(). A non-numeric value
like 'abc' would otherwise produce NaN and silently surface as a
'truncated' error, masking the real cause. Malformed values are
treated as absent; the minimum-length check still defends against
actual truncation in that case.
3. Add tests for the truncated-body-without-Content-Length case
(chunked transfer where Content-Length validation can't see the
truncation), and for a malformed Content-Length header that should
be ignored rather than misreported as truncation.
The validation logic also moves into a small assertValidRefBody
helper to keep the inner trace function under the noExcessiveCognitiveComplexity limit.
* Address review: scope 4-byte minimum to binary refs, strict Content-Length parsing
- Only apply the 4-byte format-prefix minimum to application/octet-stream
payloads; CBOR refs can legitimately be 1-byte primitives (true/0/null).
- Require Content-Length to be a plain run of digits before comparing;
parseInt would otherwise accept numeric-prefixed garbage ('12junk' -> 12).
- Make the changeset succinct.
* Address review: skip Content-Length check for compressed responses
fetch/undici transparently decompresses gzip/br bodies but leaves
Content-Length describing the encoded (compressed) size, so comparing it
against the decompressed byteLength would reject valid compressed refs as
a phantom 'ref-body-length-mismatch'. Skip the comparison when a
non-identity Content-Encoding is present; an absent or 'identity' encoding
is still validated. Adds regression tests for both cases.
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
The nitro-native-build changelog sample calls useStorage() (a Nitro
server-side auto-import) from a step, which the docs type-checker
couldn't resolve and failed with TS2552. Add liberal global
declarations for useStorage/useDatabase/useRuntimeConfig so Nitro
auto-imports type-check in docs samples.
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(next): always apply turbopack content condition regardless of builder mode
When lazy discovery is enabled (deferred builder), shouldApplyTurboCondition
was false, so turbopack.rules were added with no content filter — causing the
workflow loader to run on every JS/TS file. Apply the content condition
unconditionally so the loader only fires on files with workflow directives.
* add changeset
---------
Signed-off-by: Will Binns-Smith <wbinnssmith@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: JJ Kasper <jj@jjsweb.site>
* fix(world-vercel): retry transient response-body parse failures in the HTTP client
A sporadic failure reading/decoding a 2xx response body (truncated or
terminated stream, connection reset mid-body, or a gateway returning a
non-CBOR/JSON body) was surfaced immediately as a PARSE_ERROR. The
shared RetryAgent only retries connection/5xx failures — body
consumption happens after it returns the response, so these escape its
retry logic.
Retry such failures inside `makeRequest` with bounded exponential
backoff, scoped to idempotent methods (GET/HEAD) so writes are never
replayed. This fixes the reported `events.list` parse failure at the
adapter layer.
* fix(core): propagate exhausted transient world errors to the queue
Pairs with the world-vercel in-adapter retry: when a response-body parse
failure survives the adapter's retries (or comes from a non-idempotent
write that is never retried in-process), it must not fail the run.
Re-throw such transient world errors from the replay loop so they
propagate to the queue handler, which replays the whole run — safe
because replay is idempotent. Schema-validation contract errors stay
fatal.
* Revert "fix(core): propagate exhausted transient world errors to the queue"
This reverts commit 7bb62e9f81.
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Nathan Rajlich <n@n8.io>
* test(e2e): cover WritableStream passed as start() argument
Adds an e2e workflow + test where a parent workflow gets a WritableStream
via getWritable(), forwards it through start() to a child workflow, and
the child step writes raw bytes to it. Asserts the external reader on
the parent's stream observes the exact bytes the child wrote.
* fix(core): avoid double-framing when WritableStream is forwarded via start()
When a workflow's getWritable() handle is passed across start() to a
child workflow, the parent step's reviver wraps it in a serialize
transform that pipes into a workflow server stream. Until now,
getExternalReducers.WritableStream then installed a second serialize
transform on top of that — so every chunk the child step wrote got
devalue-framed twice but only deframed once on the reader side, and
external consumers saw the inner frame instead of the original bytes.
Fix: tag every user-visible writable that's already backed by a
workflow server stream with its (runId, name). When the external
reducer recognizes those tags during dehydration, it bridges bytes
straight from the new child-side server stream to the original server
stream instead of piping through the user's writable. That leaves the
producer-side serialize transform (installed once by the child's step
reviver) as the only framing layer in the chain.
* fix(core): forward (runId, name) when a tagged WritableStream crosses start()
Replaces the previous in-process bridge with first-class writable
forwarding at the descriptor level. When a parent workflow's
getWritable() handle is passed as an argument to a child workflow,
the dehydrated descriptor now carries the original (runId, name).
The child run's step-side reviver opens the writable against the
parent's server stream directly and resolves the parent run's
encryption key (encrypt-only) via getEncryptionKeyForRun.
This removes the architectural limitation that the bridge could
only stay alive for the duration of the parent step process — on
Vercel that capped forwarding at ~15 minutes regardless of the
child run's lifetime, dropping any writes the child made after the
parent step process exited.
importKey() now accepts a usages parameter, defaulting to
['encrypt', 'decrypt']. The cross-run forwarding path imports with
['encrypt'] only so a compromised child run cannot decrypt any
existing data on the parent's stream — only contribute new writes.
* test: rename writable-forwarded workflows and cover step-context getWritable()
Addresses PR review:
- Rename writableForwardedToChildChildWorkflow → writableForwardedChildWorkflow
(drops the duplicated 'Child' segment).
- Split writableForwardedToChildWorkflow into two variants covered by a
test.each: writableForwardedFromWorkflowWorkflow (workflow-context
getWritable, the original test) and writableForwardedFromStepWorkflow
(step-context getWritable passed directly into start() from the same
step that called getWritable()).
- Terser changeset description.
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Nathan Rajlich <n@n8.io>
* Add server-backed exact ID search to the Events tab.
Replace client-side substring filtering with API lookups for full correlation and event IDs so searches work beyond the first loaded page.
* Fix exact ID search dimming and support wrun_ correlation IDs.
Disable group dimming for server search results and accept run IDs in the exact ID parser so run-level correlation search works.
* Fix dimmed row when searching by event ID for run-level events.
Map selectedGroupKey to __run__ for run-level search results so the matched row is treated as related instead of dimmed.
* Remove run ID search from Events tab exact ID lookup.
Workflow-server only accepts step, wait, and hook correlation IDs — not wrun_. Update the search placeholder and validation toast accordingly.
* Harden exact ID search UX and correlation fetch limits.
Normalize lowercase ULIDs, scope Enter toasts to ID-like input, abort stale searches, disable search when unavailable, expand parser tests, and cap correlation pagination in workflow web.
* Fix search clear race and surface truncated correlation results.
Guard successful exact-ID search against aborted requests, invalidate in-flight work when the input clears, and return truncation metadata from correlation pagination.
* Differentiate exact ID search errors from not-found results.
Return a discriminated union from onExactIdSearch and show search errors in the Events tab instead of mislabeling them as missing IDs.
* Apply suggestion from @VaguelySerious
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Signed-off-by: Karthik Kalyan <105607645+karthikscale3@users.noreply.github.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
* [world-vercel] Add /run-id sub-export with tagged ULID encode/decode
Encodes a tag bit, 5-bit version, and 6-bit Vercel region ID into a
ULID-shaped string used for workflow run IDs. Tagged values remain
valid 26-char Crockford-Base32 ULIDs so they still sort and round-trip
through any system that accepts ULIDs.
* [world-vercel] Add string-value assertions to run-id tests
Add exact-string expectations for encoded outputs at known inputs,
covering the default region/version pair, numeric region IDs, version
overrides, boundary values (all-zero, all-max), the dirty-input
overwrite case, and the lexicographic-order checks. Also adds an
explicit byte-array expectation for the canonical ULID-spec example
string and an additional first-char-range coverage test for isTagged.
* [world-vercel] Remove internal-repo reference from regions doc comment
* [world-vercel] Address PR review feedback on run-id sub-export
- isTaggedString now fully validates the input as a 26-char Crockford
Base32 ULID (delegating to ulidToBytes) instead of only inspecting
the first character. This fixes false positives on inputs like
'4UUUU...' that have a valid tag-bit position but invalid chars
later in the string.
- isTagged() now accepts `unknown` to match its documented behavior
of safely rejecting non-string inputs without requiring callers to
cast.
- Introduce `RegionKey` for the full set of keys including 'unknown',
and narrow `RegionCode` to `Exclude<RegionKey, 'unknown'>` so the
return type of `lookupRegion` and the `DecodedRunId.region` field
accurately reflect that 'unknown' is never produced. Updates
`encode` to reject 'unknown' as a region code string at runtime
(callers wanting the unknown sentinel should pass numeric 0).
* [world-vercel] Move tagged-ULID metadata to the top of randomness
Address review feedback on #1978:
1. **Metadata at top of randomness, not bottom.** Place `regionId` (6
bits) in the high bits of byte[6] and `version` (5 bits) straddling
bytes 6 and 7, leaving the bottom 69 bits of randomness untouched by
`encode`. This means a `monotonicFactory()`-style ULID generator's
intra-millisecond bottom-bit increments survive encoding intact, so
consecutive `encode(ulid(), region, { version })` calls with the
same metadata produce strictly increasing strings. Previously the
metadata sat in the bottom 11 bits — exactly the bits the monotonic
factory uses — causing same-ms collisions/inversions.
2. **DecodedRunId is now a discriminated union.** When `tagged: false`,
the `regionId`, `version`, and `region` fields are typed as
`null` instead of being populated with garbage bits from arbitrary
ULIDs. This forces callers to discriminate on `tagged` before
reading metadata.
3. **regionIdFor: keep runtime backstop, mark as ignored for coverage.**
The unreachable-in-TS branch stays as a defensive runtime check for
callers crossing a JS/TS boundary; an istanbul/c8 ignore comment
keeps coverage tools quiet.
Doc strings and tests updated accordingly. The new layout adds a test
verifying that a sequence of incrementing-bottom-bit ULIDs (simulating
`monotonicFactory()`) round-trips through `encode` as a strictly
increasing sequence.
108/108 world-vercel tests pass; typecheck clean.
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
The World service omits input fields from run/step snapshot responses when
payloads are externalized as RemoteRef blobs. Mirrors existing
WorkflowRunWithoutData / StepWithoutData types and extends #1939's treatment
of output/error/completedAt to input.
Fixes WorkflowWorldError schema validation failures on every event
acknowledgement after the 2026-05-12 service-side RemoteRef rollout.
Fixes#1977.
Signed-off-by: Nathan Rajlich <n@n8.io>
Signed-off-by: adamiBs <bsadambs@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Nathan Rajlich <n@n8.io>
Translates the diagrammed pattern into 5 prefix-replay tests that run the
same workflow against progressively longer prefixes of a 5-event log
(hook_created, wait_created, hook_received A, wait_completed,
hook_received B). Each test asserts the consumer takes the same
deterministic path: suspending at the right intermediate point with the
right invocationsQueue state, or completing with race winners
[hookA, sleep]. The full-log test verifies that the trailing
hook_received B is consumed by the dangling race-2 hook awaiter without
producing an unconsumed-event error. Runs in both sync and async
deserialization modes via the existing defineTests harness.
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(world-local): prevent path traversal via request-supplied IDs (#1829)
* fix(world-local): prevent path traversal via request-supplied IDs
Request-supplied identifiers (runId, eventId, stepId, hookId, correlationId,
stream names, and tags) flowed directly into path.join() calls, allowing a
client to send values like '../../../package' and cause the backend to read
or write files outside the workflow data directory.
Add a centralized validator (assertSafeEntityId) that rejects IDs which are
empty, start with '.', or contain path separators or NUL bytes. Apply it at
each storage-layer entry point that composes IDs into filesystem paths:
fs.taggedPath / readJSONWithFallback / paginatedFileSystemQuery, the runs /
steps / events / hooks storage methods, and the streamer.
* address review feedback
- UnsafeEntityIdError now extends WorkflowWorldError for consistency with
other storage-layer errors and the platform error-to-HTTP mapping.
- Add resolveWithinBase(basedir, ...segments) containment helper and
apply it at every taggedPath / readJSONWithFallback / .locks path
construction site in events-storage and legacy, so a forgotten
assertSafeEntityId at a future call site can't silently regress.
- Truncate attacker-controlled values in the error message.
- Drop unused assertSafeEntityIds helper and the unreachable typeof
check under the TS signature.
- Fix docstrings on assertSafeEntityId / taggedPath JSDoc example /
filePrefix validation comment to match what the code actually does.
- handleLegacyEvent now re-asserts runId locally so the invariant is
documented at the call site instead of implicitly inherited from
events.create.
---------
Co-authored-by: JJ Kasper <jj@jjsweb.site>
* fix(world-local): tighten ID validation and add streamer regression tests
Addresses code review feedback on the path-traversal backport:
- Reject dots inside entity IDs in `assertSafeEntityId`. Internal IDs
(ULIDs, step_N, etc.) never contain dots, but `stripTag()` /
`getObjectCreatedAt()` strip a trailing `.[tag]` suffix from filenames,
so a request-supplied runId like `wrun_123.foo` would be silently
mangled during listing/pagination.
- Reject empty `correlationId` on events that include one. The event
schemas only require `z.string()`, so without this check a
step_created / hook_created / wait_created request with
`correlationId: ''` would silently be written under a malformed
composite key like `${runId}-`.
- Add streamer regression tests covering writeToStream, closeStream,
listStreamsByRunId, and getStreamChunks (the v4-shape surface that
this backport touches independently of main).
---------
Co-authored-by: JJ Kasper <jj@jjsweb.site>
Turborepo replays nextjs-turbopack:build from cache without restoring the
Vercel diagnostics manifest (.vercel/output/diagnostics/workflows-manifest.json),
which causes the Vercel deployment to fail post-build. Add .vercel/output/**
to the workbench's Turbo outputs so it is persisted and replayed. Applies to
both nextjs-turbopack and nextjs-webpack (whose turbo.json is a symlink).
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>