* Reduce unnecessary CI runtime
* Fix shared E2E artifact extraction path
* Stabilize getWorkflowPort timeout test on Windows
* Preserve UI unit coverage on CI fast path
* fix(world-vercel): retry transient response-body parse failures in the HTTP client
A sporadic failure reading/decoding a 2xx response body (truncated or
terminated stream, connection reset mid-body, or a gateway returning a
non-CBOR/JSON body) was surfaced immediately as a PARSE_ERROR. The
shared RetryAgent only retries connection/5xx failures — body
consumption happens after it returns the response, so these escape its
retry logic.
Retry such failures inside `makeRequest` with bounded exponential
backoff, scoped to idempotent methods (GET/HEAD) so writes are never
replayed. This fixes the reported `events.list` parse failure at the
adapter layer.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(core): propagate exhausted transient world errors to the queue
Pairs with the world-vercel in-adapter retry: when a response-body parse
failure survives the adapter's retries (or comes from a non-idempotent
write that is never retried in-process), it must not fail the run.
Re-throw such transient world errors from the replay loop so they
propagate to the queue handler, which replays the whole run — safe
because replay is idempotent. Schema-validation contract errors stay
fatal.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Revert "fix(core): propagate exhausted transient world errors to the queue"
This reverts commit 7bb62e9f81.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(world-local): skip Nov 2025 ghost versions on npm
`@workflow/world-local` versions 5.0.0-beta.8, beta.9, and beta.10 are
published on npm from an abandoned November 2025 release train (see "About
this version" timestamps on npmjs.com). They use the old `createEmbeddedWorld`
export and 4.x dependencies, while current code uses `createLocalWorld` and
5.x deps.
The last "Version Packages (beta)" PR (#2147) bumped world-local's
package.json from beta.7 → beta.8. The actual publish to npm was either
silently skipped or 409'd, so the npm slot still holds the November
content while `@workflow/core@5.0.0-beta.9` was published with a hard
dependency pin on `@workflow/world-local@5.0.0-beta.8` — pointing
downstream consumers (e.g. Ash) at incompatible code:
npm ERR! MISSING_EXPORT: "createLocalWorld" is not exported by
".../@workflow/world-local@5.0.0-beta.8/.../dist/index.js"
Bump packages/world-local/package.json to 5.0.0-beta.10 (the highest
contaminated slot) so the next Version Packages PR computes
5.0.0-beta.11, which is a free slot on npm. That cascades a patch bump
to core (workspace:*) → core@5.0.0-beta.10 with a clean pin to
world-local@5.0.0-beta.11.
No functional code changes.
* Tighten changeset to one sentence
Signed-off-by: Nathan Rajlich <n@n8.io>
---------
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: Nathan Rajlich <n@n8.io>
* fix(core,errors): classify SDK encryption failures as RUNTIME_ERROR
SDK-level AES-GCM encrypt/decrypt failures are never the user's fault,
but the run-failure classifier was tagging them as USER_ERROR because
the native Web Crypto OperationError (most commonly raised by
AESCipherJob.onDone on GCM auth-tag mismatch) does not match any
RUNTIME_ERROR_CHECKS entry.
Introduce a new RuntimeDecryptionError (subclass of WorkflowRuntimeError)
that the encryption module throws when subtle.encrypt/subtle.decrypt
fails, with the original DOMException as cause plus diagnostic context
(operation, byteLength, printable/hex format prefix of the input
header). classifyRunError now picks it up via RUNTIME_ERROR_CHECKS, so
these failures surface as RUNTIME_ERROR with a proper named class for
dashboards and triage.
* Trim changeset description to one sentence
* Trim historical-context comments
* docs: add runtime-decryption-failed troubleshooting page (v4 + v5)
* fix(core): round-trip RuntimeDecryptionError context, fix formatPrefix, propagate through serialization wrappers
Addresses review feedback on #2145:
- Add a RuntimeDecryptionError reducer/reviver (+ SerializableSpecial
entry + globalThis registration) so its `context` (operation,
byteLength, formatPrefix) survives the dehydrate/hydrate run-error
round trip instead of being dropped by the generic Error reducer.
- Stop capturing `formatPrefix` in the low-level encryption layer, which
only sees the stripped AES payload (nonce bytes), not the outer `encr`
marker. The serialization layer now attaches the real envelope prefix.
- Rethrow RuntimeDecryptionError unchanged from the serialize/dehydrate
catch blocks instead of reframing it as a SerializationError, so an
encryption failure during dehydration stays a RUNTIME_ERROR rather than
being misclassified as USER_ERROR.
* fix(core): enrich stream decrypt errors with envelope prefix + fix lint
- Mirror the catch/enrich/rethrow block from serialization/encryption.ts
around the stream-path aesGcmDecrypt() call so auth-tag failures on
encrypted stream frames also carry context.formatPrefix = 'encr'
(addresses review feedback). Add a tampered-frame test.
- Fix all auto-fixable Biome lint findings in the touched files
(template literals, useless try/catch wrappers, optional chaining,
non-null assertions).
* Add server-backed exact ID search to the Events tab.
Replace client-side substring filtering with API lookups for full correlation and event IDs so searches work beyond the first loaded page.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix exact ID search dimming and support wrun_ correlation IDs.
Disable group dimming for server search results and accept run IDs in the exact ID parser so run-level correlation search works.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix dimmed row when searching by event ID for run-level events.
Map selectedGroupKey to __run__ for run-level search results so the matched row is treated as related instead of dimmed.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Remove run ID search from Events tab exact ID lookup.
Workflow-server only accepts step, wait, and hook correlation IDs — not wrun_. Update the search placeholder and validation toast accordingly.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Harden exact ID search UX and correlation fetch limits.
Normalize lowercase ULIDs, scope Enter toasts to ID-like input, abort stale searches, disable search when unavailable, expand parser tests, and cap correlation pagination in workflow web.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix search clear race and surface truncated correlation results.
Guard successful exact-ID search against aborted requests, invalidate in-flight work when the input clears, and return truncation metadata from correlation pagination.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Differentiate exact ID search errors from not-found results.
Return a discriminated union from onExactIdSearch and show search errors in the Events tab instead of mislabeling them as missing IDs.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Apply suggestion from @VaguelySerious
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
The shared Button's dark-mode hover relied on an unregistered `dark-theme:`
Tailwind variant, so the inverted (default) button lost its hover style — and
the background resolved to transparent when consumed by apps that supply their
own Geist tokens (e.g. vercel/front). Use Geist's literal hover fallbacks
driven by arbitrary ancestor-theme variants instead, render the previously
missing focus-visible ring, and apply Geist's 4px tiny radius to the xs size.
Authored to compile under both Tailwind v3 and v4.
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
The World service omits input fields from run/step snapshot responses when
payloads are externalized as RemoteRef blobs. Mirrors existing
WorkflowRunWithoutData / StepWithoutData types and extends #1939's treatment
of output/error/completedAt to input.
Fixes WorkflowWorldError schema validation failures on every event
acknowledgement after the 2026-05-12 service-side RemoteRef rollout.
Fixes#1977.
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: Nathan Rajlich <n@n8.io>
Translates the diagrammed pattern into 5 prefix-replay tests that run the
same workflow against progressively longer prefixes of a 5-event log
(hook_created, wait_created, hook_received A, wait_completed,
hook_received B). Each test asserts the consumer takes the same
deterministic path: suspending at the right intermediate point with the
right invocationsQueue state, or completing with race winners
[hookA, sleep]. The full-log test verifies that the trailing
hook_received B is consumed by the dangling race-2 hook awaiter without
producing an unconsumed-event error. Runs in both sync and async
deserialization modes via the existing defineTests harness.
* [world-vercel] Add /run-id sub-export with tagged ULID encode/decode
Encodes a tag bit, 5-bit version, and 6-bit Vercel region ID into a
ULID-shaped string used for workflow run IDs. Tagged values remain
valid 26-char Crockford-Base32 ULIDs so they still sort and round-trip
through any system that accepts ULIDs.
* [world-vercel] Add string-value assertions to run-id tests
Add exact-string expectations for encoded outputs at known inputs,
covering the default region/version pair, numeric region IDs, version
overrides, boundary values (all-zero, all-max), the dirty-input
overwrite case, and the lexicographic-order checks. Also adds an
explicit byte-array expectation for the canonical ULID-spec example
string and an additional first-char-range coverage test for isTagged.
* [world-vercel] Remove internal-repo reference from regions doc comment
* [world-vercel] Address PR review feedback on run-id sub-export
- isTaggedString now fully validates the input as a 26-char Crockford
Base32 ULID (delegating to ulidToBytes) instead of only inspecting
the first character. This fixes false positives on inputs like
'4UUUU...' that have a valid tag-bit position but invalid chars
later in the string.
- isTagged() now accepts `unknown` to match its documented behavior
of safely rejecting non-string inputs without requiring callers to
cast.
- Introduce `RegionKey` for the full set of keys including 'unknown',
and narrow `RegionCode` to `Exclude<RegionKey, 'unknown'>` so the
return type of `lookupRegion` and the `DecodedRunId.region` field
accurately reflect that 'unknown' is never produced. Updates
`encode` to reject 'unknown' as a region code string at runtime
(callers wanting the unknown sentinel should pass numeric 0).
* [world-vercel] Move tagged-ULID metadata to the top of randomness
Address review feedback on #1978:
1. **Metadata at top of randomness, not bottom.** Place `regionId` (6
bits) in the high bits of byte[6] and `version` (5 bits) straddling
bytes 6 and 7, leaving the bottom 69 bits of randomness untouched by
`encode`. This means a `monotonicFactory()`-style ULID generator's
intra-millisecond bottom-bit increments survive encoding intact, so
consecutive `encode(ulid(), region, { version })` calls with the
same metadata produce strictly increasing strings. Previously the
metadata sat in the bottom 11 bits — exactly the bits the monotonic
factory uses — causing same-ms collisions/inversions.
2. **DecodedRunId is now a discriminated union.** When `tagged: false`,
the `regionId`, `version`, and `region` fields are typed as
`null` instead of being populated with garbage bits from arbitrary
ULIDs. This forces callers to discriminate on `tagged` before
reading metadata.
3. **regionIdFor: keep runtime backstop, mark as ignored for coverage.**
The unreachable-in-TS branch stays as a defensive runtime check for
callers crossing a JS/TS boundary; an istanbul/c8 ignore comment
keeps coverage tools quiet.
Doc strings and tests updated accordingly. The new layout adds a test
verifying that a sequence of incrementing-bottom-bit ULIDs (simulating
`monotonicFactory()`) round-trips through `encode` as a strictly
increasing sequence.
108/108 world-vercel tests pass; typecheck clean.
* fix(web-shared): use inline blur style for tailwind v3 compatibility
Co-authored-by: Mitul Shah <mitulxshah@gmail.com>
* fix(web-shared): use blur-[4px] arbitrary value for tailwind v3/v4 compat
Co-authored-by: Mitul Shah <mitulxshah@gmail.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* docs(cookbook): replace child workflow polling with hook resume pattern
Recommend startAndWait() with withChildCompletionHook() for v4 and v5 child
workflow orchestration instead of getRun().status polling loops.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix child-workflows cookbook review feedback
Tighten resumeParentCompletion to a discriminated union so hook.resume
typechecks, add zod to the vitest workbench, remove unused resumeHook
import, and add an empty changeset per AGENTS.md.
Co-authored-by: Cursor <cursoragent@cursor.com>
* docs(cookbook): trim child-workflows hook resume guide
Remove redundant polling comparison copy, the getRun() alternative section, and v5-only start() tips to keep the cookbook focused on the hook pattern.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix cookbook pattern for AI SDK
The AI SDK cookbook entry presented `streamText` inside a `"use step"`
turn function with tools also marked `"use step"`. That implies tools
are individually durable, but the `"use step"` directive is a no-op
when called from another step — so tools run as plain inline functions
inside `runTurn`, and the durability boundary is the entire turn.
Changes:
- Remove the `"use step"` directive from tool implementations in the
workflow code sample and add an explanatory comment.
- Update the frontmatter summary and intro paragraph to drop the
inaccurate "tools remain durable steps" claim.
- Add a "Tools are not individually durable" entry to Pitfalls with
consequences and mitigations (idempotency or `DurableAgent`).
- Add a `runTurn` durability-boundary bullet to "How it works".
- Add a "Tool call durability" row to the `streamText` vs `DurableAgent`
comparison table.
- Fix two misleading Key APIs bullets that claimed tools wrap
`"use step"` functions and that `"use step"` makes tool executions
durable.
Applied identically to both v4 and v5 cookbook entries.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Apply suggestion from @VaguelySerious
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Signed-off-by: Karthik Kalyan <105607645+karthikscale3@users.noreply.github.com>
* Correct outdated DurableAgent guidance in AI SDK cookbook.
The callout and comparison table incorrectly claimed DurableAgent lacks stopWhen, structured output, and onStepFinish — update them to reflect the actual implementation and clarify when raw streamText() is still appropriate.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Reword tool comment to describe current behavior, not a changelog.
Address review feedback: the inline comment should explain how tools run inside runTurn without referencing removed "use step" directives.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Signed-off-by: Karthik Kalyan <105607645+karthikscale3@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
* fix(build): use inline sourcemaps across all workspace packages
Prevents a Turbopack bug on Windows that caused the SWC worker to
crash when reading external .js.map files in workspace-linked packages
(paths were concatenated with mixed separators like
'packages/serde/dist\\index.js.map'). Once the worker crashed, the
module graph entered a broken state where all subsequent requests
returned 500, manifesting as flaky E2E Windows tests where the
'should rebuild on imported step dependency change' test would time
out waiting for /api/workflows/start to succeed.
Extends the original fix from #352 (previously applied only to
@workflow/core and workflow) to the shared base tsconfig.json so that
all packages producing a dist/ output benefit.
* Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Signed-off-by: Nathan Rajlich <n@n8.io>
* Tighten changeset description
---------
Signed-off-by: Nathan Rajlich <n@n8.io>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* fix(world-local): tighten ID validation
Forward-port of code review fixes from #1920 (backport of #1829):
- Reject dots inside entity IDs in `assertSafeEntityId`. Internal IDs
(ULIDs, step_N, etc.) never contain dots, but `stripTag()` /
`getObjectCreatedAt()` strip a trailing `.[tag]` suffix from filenames,
so a request-supplied runId like `wrun_123.foo` would be silently
mangled during listing/pagination.
- Reject empty `correlationId` on events that include one. The event
schemas only require `z.string()`, so without this check a
step_created / hook_created / wait_created request with
`correlationId: ''` would silently be written under a malformed
composite key like `${runId}-`.
The streamer regression tests from #1920 are not forward-ported because
main's streamer surface differs (renamed methods, separate test
coverage) from stable's v4 shape.
* simplify changeset
Turborepo replays nextjs-turbopack:build from cache without restoring the
Vercel diagnostics manifest (.vercel/output/diagnostics/workflows-manifest.json),
which causes the Vercel deployment to fail post-build. Add .vercel/output/**
to the workbench's Turbo outputs so it is persisted and replayed. Applies to
both nextjs-turbopack and nextjs-webpack (whose turbo.json is a symlink).
* [ci] Attribute backport changelog entries to original PR author
`@changesets/changelog-github` resolves a changeset commit to its
associated PR via the GitHub GraphQL `associatedPullRequests` field. For
commits landed on `stable` via our backport workflow, that resolves to
the backport PR authored by `github-actions[bot]` (since the backport
workflow uses `createCommitOnBranch` to produce signed commits), so the
generated changelog ends up with "Thanks @github-actions!" instead of
the original contributor.
This adds a small `.changeset/changelog.mjs` wrapper around
`@changesets/changelog-github` that detects backport PRs by matching the
title (`Backport #N: ...`) or body (`Automated backport of #N to
` + '`stable`' + `...`) produced by `.github/workflows/backport.yml`, resolves the
original PR number, and injects `pr:`/`commit:` directives into the
changeset summary before delegating to the upstream generator. The
result is that the rendered changelog entry attributes the change to the
original PR and author, while the commit link still points at the
backport commit on the release branch.
* Address review feedback
- Use `Bearer` instead of `Token` for the GitHub GraphQL auth header
for consistency with the rest of the repo (Copilot review on PR #2091).
- Add a defensive `formatError` helper so `console.warn` in the catch
blocks doesn't itself throw when a non-Error value is thrown (Copilot
review on PR #2091).
* cleanup items
* Update events-list.tsx
* nice
* Update attribute-panel.tsx
* woo
* remove toast
* Fix: Unused variable `selectedResource` causes TypeScript build failure (TS6133) due to `noUnusedLocals: true` in tsconfig.
This commit fixes the issue reported at packages/web-shared/src/components/new-trace-viewer/trace-viewer.tsx:462
**Bug explanation:**
The variable `selectedResource` is declared on line 462 of `trace-viewer.tsx`:
```ts
const selectedResource = selectedSpan?.resource as string | undefined;
```
However, all references to this variable were removed in the PR (the colored resource badge in the panel header was removed), leaving behind the unused declaration. The project's TypeScript configuration has `noUnusedLocals: true`, which causes TypeScript to emit error TS6133 for any declared-but-unused local variables. This is confirmed directly in the Vercel build logs:
```
@workflow/web-shared:build: src/components/new-trace-viewer/trace-viewer.tsx(462,9): error TS6133: 'selectedResource' is declared but its value is never read.
```
This caused the `@workflow/web-shared#build` task to fail with exit code 2, which in turn caused the entire Vercel deployment to fail.
**Fix explanation:**
Removed the unused `const selectedResource = selectedSpan?.resource as string | undefined;` declaration on line 462. The nearby `selectedResourceId` variable (which was NOT removed) remains in place and is still actively used in the JSX below. This is a minimal one-line deletion that resolves the build failure.
Co-authored-by: Vercel <vercel[bot]@users.noreply.github.com>
Co-authored-by: mitul-s <mitulxshah@gmail.com>
* tweak
* Update events-list.tsx
* Update events-list.tsx
* cleanup
* Update attribute-panel.tsx
* copy button
* Create button.tsx
* cleanup
* rename `this` to Context
* cleanup
* Update button.tsx
* avoid decrypt flashign
* cleanup
* Update attribute-panel.tsx
* nav on detail card
* polish
* Update detail-card.tsx
* Update trace-viewer.tsx
* colors
* Update attribute-panel.tsx
* Update detail-card.tsx
* Update copyable-data-block.tsx
* Detail Pane + Other cleanup items (#2020)
* polish
* Update attribute-panel.tsx
* Create wise-frogs-thank.md
Signed-off-by: Mitul Shah <mitulxshah@gmail.com>
---------
Signed-off-by: Mitul Shah <mitulxshah@gmail.com>
---------
Signed-off-by: Mitul Shah <mitulxshah@gmail.com>
Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
Co-authored-by: Vercel <vercel[bot]@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* [next] make lazyDiscovery the default in withWorkflow
Flips the default for `workflows.lazyDiscovery` from `false` to `true`
so new projects get deferred workflow discovery automatically on Next.js
versions that support deferred entries (>= 16.2.0-canary.48). Older
versions continue to fall back to eager discovery.
Users can still opt back into eager discovery explicitly by passing
`workflows: { lazyDiscovery: false }`.
Also:
- Remove the now-redundant `lazyDiscovery: true` from the Next.js
workbench apps.
- Reword the fallback warning for clarity when lazy is the default.
- Update the local-build e2e assertion to match the new warning text.
- Update the withWorkflow docs with the new default.
* [workbench] remove commented 'export default nextConfig' lines
* [core] Exclude inline step execution from replay timeout
The v5 combined workflow+step handler wraps inline step bodies in the
same setTimeout(..., REPLAY_TIMEOUT_MS) guard that previously only
bounded the v4 'workflows' function's fast deterministic replay. As a
result, any workflow with a single step exceeding 240s hard-fails with
FatalError: Workflow replay exceeded maximum duration (240s) after 4
attempts — even though the step could legitimately run for the full
function maxDuration (up to 800s on Pro Fluid).
Replace the setTimeout guard with a per-invocation budget that only
accumulates non-step time. pauseReplayBudget() / resumeReplayBudget()
bracket each executeStep() call, and the loop checks the budget at
iteration boundaries. The retry-then-fail semantics from #1567 are
preserved verbatim for the pure-replay case.
Also adds a WORKFLOW_REPLAY_TIMEOUT_MS env var override (clamped to
30s..780s) so operators can adjust the bare-replay ceiling without
patching @workflow/core.
Fixes#2009.
* Address PR review feedback
- Extract budget bookkeeping into ReplayBudget class (replay-budget.ts)
with sentinel-protected idempotent pause()/resume() to avoid
double-counting in future refactors that nest step execution
- Restore VERCEL_URL gate around process.exit(1) so a long pure-replay
in local dev/non-Vercel runtimes can't hard-kill the host process
- Warn (once per distinct raw value) when WORKFLOW_REPLAY_TIMEOUT_MS is
clamped or rejected, so misconfiguration is observable
- Correct Hobby maxDuration comment (60s standard / 300s Fluid)
- Document budget-check responsiveness trade-off vs. old setTimeout
- Tighten describe-error test assertions to match the full new hint
- Shorten changeset description
- Add ReplayBudget unit tests (9) including 8-minute step regression
- Add warn-once tests for getReplayTimeoutMs (extended)
* Replace VERCEL_URL gate with World capability
Per review feedback, gating runtime behavior on process.env.VERCEL_URL
leaks deployment-environment concerns into @workflow/core. Replace the
check with a new optional capability on the World interface:
processExitTriggersQueueRedelivery (default false).
- @workflow/world: declare the new optional capability on World
- @workflow/world-vercel: set it to true (Vercel fails the function
invocation on non-zero exit and VQS redelivers via fresh invocation)
- @workflow/core: handleReplayBudgetExhausted reads world.processExitTriggersQueueRedelivery
instead of process.env.VERCEL_URL; behavior is otherwise unchanged
- Add 4 unit tests for handleReplayBudgetExhausted covering both branches
(exit-for-redelivery and best-effort run_failed)
---------
Co-authored-by: Peter Wielander <mittgfu@gmail.com>