* [world-vercel] Validate ref resolve responses before use
When workflow-server returns a ref body to the SDK, the bytes are
fed into the workflow runtime's event log and deserialized via
`decodeFormatPrefix`. The SDK always writes ref payloads with at
least a 4-byte format prefix (see `encodeWithFormatPrefix` in
`@workflow/core`), so a zero-byte response — or one whose length
disagrees with `Content-Length` — is never a valid stored value.
Before this change, `resolveRefDescriptor` had no validation: a
200 with an empty body would be passed downstream as a zero-length
Uint8Array, which then failed deep inside replay with:
Data too short to contain format prefix: expected at least 4 bytes, got 0
By that point the workflow's in-memory event snapshot is already
poisoned with the empty payload, so every subsequent replay
deterministically reproduces the same failure, downstream
`resumeHook()` calls surface as `Hook not found`, and the run
only unsticks when stale-run cleanup terminates the sandbox.
This catches the failure at the transport boundary instead, where
it can be retried as a `WorkflowWorldError`. Both an empty body
and a length mismatch (truncated streaming response) are rejected.
This is the SDK-side companion to vercel/workflow-server#432, which
adds the same validation on the server side.
* Address review: reject <4-byte bodies, handle malformed Content-Length
Three review changes:
1. Reject any body shorter than the 4-byte format-prefix length, not
just zero-byte bodies. The SDK guarantees every stored ref payload
starts with a 4-byte format prefix (FORMAT_PREFIX_LENGTH in
@workflow/core), so a 1-3 byte body would also fail downstream
replay with the same 'Data too short to contain format prefix'
error this PR exists to prevent.
2. Parse Content-Length safely with parseInt + Number.isFinite +
non-negative checks instead of bare Number(). A non-numeric value
like 'abc' would otherwise produce NaN and silently surface as a
'truncated' error, masking the real cause. Malformed values are
treated as absent; the minimum-length check still defends against
actual truncation in that case.
3. Add tests for the truncated-body-without-Content-Length case
(chunked transfer where Content-Length validation can't see the
truncation), and for a malformed Content-Length header that should
be ignored rather than misreported as truncation.
The validation logic also moves into a small assertValidRefBody
helper to keep the inner trace function under the noExcessiveCognitiveComplexity limit.
* Address review: scope 4-byte minimum to binary refs, strict Content-Length parsing
- Only apply the 4-byte format-prefix minimum to application/octet-stream
payloads; CBOR refs can legitimately be 1-byte primitives (true/0/null).
- Require Content-Length to be a plain run of digits before comparing;
parseInt would otherwise accept numeric-prefixed garbage ('12junk' -> 12).
- Make the changeset succinct.
* Address review: skip Content-Length check for compressed responses
fetch/undici transparently decompresses gzip/br bodies but leaves
Content-Length describing the encoded (compressed) size, so comparing it
against the decompressed byteLength would reject valid compressed refs as
a phantom 'ref-body-length-mismatch'. Skip the comparison when a
non-identity Content-Encoding is present; an absent or 'identity' encoding
is still validated. Adds regression tests for both cases.
The nitro-native-build changelog sample calls useStorage() (a Nitro
server-side auto-import) from a step, which the docs type-checker
couldn't resolve and failed with TS2552. Add liberal global
declarations for useStorage/useDatabase/useRuntimeConfig so Nitro
auto-imports type-check in docs samples.
* fix(next): always apply turbopack content condition regardless of builder mode
When lazy discovery is enabled (deferred builder), shouldApplyTurboCondition
was false, so turbopack.rules were added with no content filter — causing the
workflow loader to run on every JS/TS file. Apply the content condition
unconditionally so the loader only fires on files with workflow directives.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* add changeset
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: JJ Kasper <jj@jjsweb.site>
* feat(world-vercel): allow a custom dispatcher
Add a `dispatcher` option to `APIConfig`/`createVercelWorld` so callers can
supply a custom undici dispatcher. It is threaded through every request path
(HTTP and the queue) and defaults to the shared undici RetryAgent.
* docs(world-vercel): explain why dispatcher is typed unknown
* Reduce unnecessary CI runtime
* Fix shared E2E artifact extraction path
* Stabilize getWorkflowPort timeout test on Windows
* Preserve UI unit coverage on CI fast path
* fix(world-vercel): retry transient response-body parse failures in the HTTP client
A sporadic failure reading/decoding a 2xx response body (truncated or
terminated stream, connection reset mid-body, or a gateway returning a
non-CBOR/JSON body) was surfaced immediately as a PARSE_ERROR. The
shared RetryAgent only retries connection/5xx failures — body
consumption happens after it returns the response, so these escape its
retry logic.
Retry such failures inside `makeRequest` with bounded exponential
backoff, scoped to idempotent methods (GET/HEAD) so writes are never
replayed. This fixes the reported `events.list` parse failure at the
adapter layer.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(core): propagate exhausted transient world errors to the queue
Pairs with the world-vercel in-adapter retry: when a response-body parse
failure survives the adapter's retries (or comes from a non-idempotent
write that is never retried in-process), it must not fail the run.
Re-throw such transient world errors from the replay loop so they
propagate to the queue handler, which replays the whole run — safe
because replay is idempotent. Schema-validation contract errors stay
fatal.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Revert "fix(core): propagate exhausted transient world errors to the queue"
This reverts commit 7bb62e9f81.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(world-local): skip Nov 2025 ghost versions on npm
`@workflow/world-local` versions 5.0.0-beta.8, beta.9, and beta.10 are
published on npm from an abandoned November 2025 release train (see "About
this version" timestamps on npmjs.com). They use the old `createEmbeddedWorld`
export and 4.x dependencies, while current code uses `createLocalWorld` and
5.x deps.
The last "Version Packages (beta)" PR (#2147) bumped world-local's
package.json from beta.7 → beta.8. The actual publish to npm was either
silently skipped or 409'd, so the npm slot still holds the November
content while `@workflow/core@5.0.0-beta.9` was published with a hard
dependency pin on `@workflow/world-local@5.0.0-beta.8` — pointing
downstream consumers (e.g. Ash) at incompatible code:
npm ERR! MISSING_EXPORT: "createLocalWorld" is not exported by
".../@workflow/world-local@5.0.0-beta.8/.../dist/index.js"
Bump packages/world-local/package.json to 5.0.0-beta.10 (the highest
contaminated slot) so the next Version Packages PR computes
5.0.0-beta.11, which is a free slot on npm. That cascades a patch bump
to core (workspace:*) → core@5.0.0-beta.10 with a clean pin to
world-local@5.0.0-beta.11.
No functional code changes.
* Tighten changeset to one sentence
Signed-off-by: Nathan Rajlich <n@n8.io>
---------
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: Nathan Rajlich <n@n8.io>
* fix(core,errors): classify SDK encryption failures as RUNTIME_ERROR
SDK-level AES-GCM encrypt/decrypt failures are never the user's fault,
but the run-failure classifier was tagging them as USER_ERROR because
the native Web Crypto OperationError (most commonly raised by
AESCipherJob.onDone on GCM auth-tag mismatch) does not match any
RUNTIME_ERROR_CHECKS entry.
Introduce a new RuntimeDecryptionError (subclass of WorkflowRuntimeError)
that the encryption module throws when subtle.encrypt/subtle.decrypt
fails, with the original DOMException as cause plus diagnostic context
(operation, byteLength, printable/hex format prefix of the input
header). classifyRunError now picks it up via RUNTIME_ERROR_CHECKS, so
these failures surface as RUNTIME_ERROR with a proper named class for
dashboards and triage.
* Trim changeset description to one sentence
* Trim historical-context comments
* docs: add runtime-decryption-failed troubleshooting page (v4 + v5)
* fix(core): round-trip RuntimeDecryptionError context, fix formatPrefix, propagate through serialization wrappers
Addresses review feedback on #2145:
- Add a RuntimeDecryptionError reducer/reviver (+ SerializableSpecial
entry + globalThis registration) so its `context` (operation,
byteLength, formatPrefix) survives the dehydrate/hydrate run-error
round trip instead of being dropped by the generic Error reducer.
- Stop capturing `formatPrefix` in the low-level encryption layer, which
only sees the stripped AES payload (nonce bytes), not the outer `encr`
marker. The serialization layer now attaches the real envelope prefix.
- Rethrow RuntimeDecryptionError unchanged from the serialize/dehydrate
catch blocks instead of reframing it as a SerializationError, so an
encryption failure during dehydration stays a RUNTIME_ERROR rather than
being misclassified as USER_ERROR.
* fix(core): enrich stream decrypt errors with envelope prefix + fix lint
- Mirror the catch/enrich/rethrow block from serialization/encryption.ts
around the stream-path aesGcmDecrypt() call so auth-tag failures on
encrypted stream frames also carry context.formatPrefix = 'encr'
(addresses review feedback). Add a tampered-frame test.
- Fix all auto-fixable Biome lint findings in the touched files
(template literals, useless try/catch wrappers, optional chaining,
non-null assertions).
* Add server-backed exact ID search to the Events tab.
Replace client-side substring filtering with API lookups for full correlation and event IDs so searches work beyond the first loaded page.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix exact ID search dimming and support wrun_ correlation IDs.
Disable group dimming for server search results and accept run IDs in the exact ID parser so run-level correlation search works.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix dimmed row when searching by event ID for run-level events.
Map selectedGroupKey to __run__ for run-level search results so the matched row is treated as related instead of dimmed.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Remove run ID search from Events tab exact ID lookup.
Workflow-server only accepts step, wait, and hook correlation IDs — not wrun_. Update the search placeholder and validation toast accordingly.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Harden exact ID search UX and correlation fetch limits.
Normalize lowercase ULIDs, scope Enter toasts to ID-like input, abort stale searches, disable search when unavailable, expand parser tests, and cap correlation pagination in workflow web.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix search clear race and surface truncated correlation results.
Guard successful exact-ID search against aborted requests, invalidate in-flight work when the input clears, and return truncation metadata from correlation pagination.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Differentiate exact ID search errors from not-found results.
Return a discriminated union from onExactIdSearch and show search errors in the Events tab instead of mislabeling them as missing IDs.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Apply suggestion from @VaguelySerious
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>