* [core] Exclude inline step execution from replay timeout
The v5 combined workflow+step handler wraps inline step bodies in the
same setTimeout(..., REPLAY_TIMEOUT_MS) guard that previously only
bounded the v4 'workflows' function's fast deterministic replay. As a
result, any workflow with a single step exceeding 240s hard-fails with
FatalError: Workflow replay exceeded maximum duration (240s) after 4
attempts — even though the step could legitimately run for the full
function maxDuration (up to 800s on Pro Fluid).
Replace the setTimeout guard with a per-invocation budget that only
accumulates non-step time. pauseReplayBudget() / resumeReplayBudget()
bracket each executeStep() call, and the loop checks the budget at
iteration boundaries. The retry-then-fail semantics from #1567 are
preserved verbatim for the pure-replay case.
Also adds a WORKFLOW_REPLAY_TIMEOUT_MS env var override (clamped to
30s..780s) so operators can adjust the bare-replay ceiling without
patching @workflow/core.
Fixes#2009.
* Address PR review feedback
- Extract budget bookkeeping into ReplayBudget class (replay-budget.ts)
with sentinel-protected idempotent pause()/resume() to avoid
double-counting in future refactors that nest step execution
- Restore VERCEL_URL gate around process.exit(1) so a long pure-replay
in local dev/non-Vercel runtimes can't hard-kill the host process
- Warn (once per distinct raw value) when WORKFLOW_REPLAY_TIMEOUT_MS is
clamped or rejected, so misconfiguration is observable
- Correct Hobby maxDuration comment (60s standard / 300s Fluid)
- Document budget-check responsiveness trade-off vs. old setTimeout
- Tighten describe-error test assertions to match the full new hint
- Shorten changeset description
- Add ReplayBudget unit tests (9) including 8-minute step regression
- Add warn-once tests for getReplayTimeoutMs (extended)
* Replace VERCEL_URL gate with World capability
Per review feedback, gating runtime behavior on process.env.VERCEL_URL
leaks deployment-environment concerns into @workflow/core. Replace the
check with a new optional capability on the World interface:
processExitTriggersQueueRedelivery (default false).
- @workflow/world: declare the new optional capability on World
- @workflow/world-vercel: set it to true (Vercel fails the function
invocation on non-zero exit and VQS redelivers via fresh invocation)
- @workflow/core: handleReplayBudgetExhausted reads world.processExitTriggersQueueRedelivery
instead of process.env.VERCEL_URL; behavior is otherwise unchanged
- Add 4 unit tests for handleReplayBudgetExhausted covering both branches
(exit-for-redelivery and best-effort run_failed)
---------
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
* ci: extract wait-for-vercel-project to vercel/wait-for-deployment-action
The action's logic was duplicated between this repo and
vercel/workflow-server, which is annoying to keep in sync. Move it to
a standalone repository so both can consume the same pinned build.
Changes:
- Delete .github/actions/wait-for-vercel-project entirely.
- Replace all five `uses: ./.github/actions/wait-for-vercel-project`
references with `uses: vercel/wait-for-deployment-action@<sha>` in:
benchmarks.yml, dispatch-front-workflow-release-pr.yml,
docs-checks.yml, tarballs-checks.yml, tests.yml
- All `with:` inputs (project-slug, environment, timeout,
check-interval, github-token) are unchanged — the new action's
input contract is backwards-compatible.
The new action is ESM-only, targets Node 24, ships a ~12KB bundle
(down from ~830KB in the old in-repo version) by dropping
@actions/core and its transitive undici dependency, and is
unit-tested. See https://github.com/vercel/wait-for-deployment-action.
* ci: bump wait-for-deployment-action to fix/status-context-auto for verification
Repinning to vercel/wait-for-deployment-action#fix/status-context-auto
(SHA 04d46ef) which fixes the broken 'opt-out' heuristic that made
status-context resolution silently disabled for every consumer.
Reproduced in this repo's E2E logs:
Looking for GitHub deployment in environment "Preview – example-workflow"
Deployment ID resolution disabled (status-context is empty)
Deployment ready: https://example-workflow-...labs.vercel.dev
Run E2E Tests: VERCEL_DEPLOYMENT_ID= <-- empty
Will repin to the post-merge main SHA once CI is green.
* ci: bump wait-for-deployment-action pin to merged main SHA
Repinning from the fix/status-context-auto branch (04d46ef) to the
post-merge main SHA (0e2b0c5, vercel/wait-for-deployment-action#4).
The deployment-id resolution fix verified against the prior fix-branch
pin (E2E tests now read VERCEL_DEPLOYMENT_ID=dpl_... correctly across
the matrix; only flaky/unrelated Vercel deployment failures remain).
* ci: grant statuses:read alongside deployments:read
The wait-for-deployment-action also reads the 'Vercel – <slug>'
combined commit status to resolve the dpl_xxx ID. The official
permissions table lists statuses:read for
GET /repos/{owner}/{repo}/commits/{ref}/status.
GitHub Actions doesn't currently support prefilling workflow_dispatch
inputs via URL query params (community/community#51159), so the
"override via workflow_dispatch with this commit SHA" instruction in
the no-backport comment required the reader to go look up the SHA
themselves. Paste the full 40-char SHA into the comment in a fenced
code block so it's one click to copy into the "Commit SHA" input on
the workflow run page.
* test(e2e): cover WritableStream passed as start() argument
Adds an e2e workflow + test where a parent workflow gets a WritableStream
via getWritable(), forwards it through start() to a child workflow, and
the child step writes raw bytes to it. Asserts the external reader on
the parent's stream observes the exact bytes the child wrote.
* fix(core): avoid double-framing when WritableStream is forwarded via start()
When a workflow's getWritable() handle is passed across start() to a
child workflow, the parent step's reviver wraps it in a serialize
transform that pipes into a workflow server stream. Until now,
getExternalReducers.WritableStream then installed a second serialize
transform on top of that — so every chunk the child step wrote got
devalue-framed twice but only deframed once on the reader side, and
external consumers saw the inner frame instead of the original bytes.
Fix: tag every user-visible writable that's already backed by a
workflow server stream with its (runId, name). When the external
reducer recognizes those tags during dehydration, it bridges bytes
straight from the new child-side server stream to the original server
stream instead of piping through the user's writable. That leaves the
producer-side serialize transform (installed once by the child's step
reviver) as the only framing layer in the chain.
* fix(core): forward (runId, name) when a tagged WritableStream crosses start()
Replaces the previous in-process bridge with first-class writable
forwarding at the descriptor level. When a parent workflow's
getWritable() handle is passed as an argument to a child workflow,
the dehydrated descriptor now carries the original (runId, name).
The child run's step-side reviver opens the writable against the
parent's server stream directly and resolves the parent run's
encryption key (encrypt-only) via getEncryptionKeyForRun.
This removes the architectural limitation that the bridge could
only stay alive for the duration of the parent step process — on
Vercel that capped forwarding at ~15 minutes regardless of the
child run's lifetime, dropping any writes the child made after the
parent step process exited.
importKey() now accepts a usages parameter, defaulting to
['encrypt', 'decrypt']. The cross-run forwarding path imports with
['encrypt'] only so a compromised child run cannot decrypt any
existing data on the parent's stream — only contribute new writes.
* test: rename writable-forwarded workflows and cover step-context getWritable()
Addresses PR review:
- Rename writableForwardedToChildChildWorkflow → writableForwardedChildWorkflow
(drops the duplicated 'Child' segment).
- Split writableForwardedToChildWorkflow into two variants covered by a
test.each: writableForwardedFromWorkflowWorkflow (workflow-context
getWritable, the original test) and writableForwardedFromStepWorkflow
(step-context getWritable passed directly into start() from the same
step that called getWritable()).
- Terser changeset description.
Major-version refs like `@v2`/`@v5` resolve to mutable refs on the
upstream repos — sometimes a tag, sometimes a branch (e.g. marocchino
keeps `v1`/`v2`/`v3` as branches), and dawidd6 force-pushes the bare
`v6` tag forward outside of releases. A compromised maintainer account
could push new code that our CI picks up on the next run with
GITHUB_TOKEN (or, for changesets/action, NPM_TOKEN) in hand.
Pin all third-party `uses:` references to full commit SHAs with a
trailing version comment so the upstream release is still visible to
reviewers. Dependabot/Renovate can keep these fresh going forward.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(backport): mark not-maintained-on-stable list as exhaustive and require git verification
The backport workflow's AI decision prompt only listed `docs/` (outside
`docs/content/`) and `skills/` as paths not maintained on `stable`, but
didn't flag that list as exhaustive. On at least one commit, the AI
generalized the pattern and incorrectly claimed `tarballs/` was also
not maintained on `stable` (it is) — likely because the commit subject
mentioned "preview tarball" and the changed files included a docs
preview smoke check workflow.
Tighten the prompt to:
- explicitly mark the not-maintained list as exhaustive;
- instruct the AI to run `git ls-tree origin/stable -- <path>` to
verify any other path it wants to cite as main-only;
- warn it not to infer main-only-ness from suggestive names like
"docs", "preview", "tarball", or "workflow".
* ci(backport): link to the backport job run in posted comments and PR body
Each of the comments the backport workflow posts (no-backport,
backport-created, conflict-failure) and the body of the backport PR
it opens now includes a link to the GitHub Actions run that produced
the decision. Makes it easy to jump from the comment straight to the
opencode output, prompt, and AI reasoning when the AI gets a call wrong.
The repo's enterprise `~ALL` required-signatures ruleset rejects the
unsigned commits produced by `peaceiris/actions-gh-pages@v4`, breaking
the benchmark and E2E result publishing jobs on `main`. Replace those
steps with a shared composite action that uses the GraphQL
`createCommitOnBranch` mutation — same pattern already used by
`backport.yml` — so commits are signed automatically by GitHub and
satisfy the rule.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The test server's flow handler awaited flowPOST before incrementing the
invocation counter, which races with the test polling getRun() to see
the completed status. When the workflow completed inside flowPOST and
flushed the run to the DB, the test could observe the completed state
and immediately query /_flow-invocations before the counter was bumped,
yielding a flaky 'expected 0 to be 1' assertion.
Increment the counter before awaiting flowPOST so the count is
observable as soon as the run transitions to completed.
* Forward port stale wait replay fix
* Guard V5 replay writes against stale events
* Revert "Guard V5 replay writes against stale events"
This reverts commit 22e74d3558.
* Update wait replay comments
Sets pnpm's `minimumReleaseAge` to 2 days (company-wide standard) and
excludes internal scopes from the gate.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Shiki wraps each highlighted token in its own <span>, which breaks Geist
Mono programming ligatures like `===`, `!==`, `=>`. The ligature glyph is
rendered at the advance width of a single character (~8.4px) instead of
three (~25.2px), causing it to visually overlap the preceding token.
Disable `font-variant-ligatures` inside `pre code` so each character
renders at its true monospace width.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(world-postgres): bootstrap graphile-worker schema in setup CLI
`workflow-postgres-setup` now installs the `graphile_worker` schema in
addition to the drizzle migrations so that by the time any consumer
calls `world.start()`, both schemas already exist. This eliminates the
inter-process race on graphile-worker's `installSchema` where
concurrent `CREATE SCHEMA IF NOT EXISTS` calls could both pass the
MVCC-snapshotted existence check and one would fail with
`duplicate key value violates unique constraint "pg_namespace_nspname_index"`.
Reproduced locally against a fresh postgres:18-alpine with 8 parallel
`makeWorkerUtils().migrate()` calls — 7/8 fail without the pre-bootstrap,
0/8 fail after running `workflow-postgres-setup` first.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Apply suggestions from code review
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
---------
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Nathan Rajlich <n@n8.io>
* Expose conflicting run id on hook conflicts
* Mark hook conflict run id as future required
* Address hook conflict docs review
* Address hook conflict review comments
* Fix hook conflict docs typecheck
Use Tailwind's typed arbitrary duration syntax for trace viewer zoom controls so Tailwind v3 consumers do not emit ambiguous utility warnings when scanning the package.
* drop setup-command input from reusable community-world workflows
The community-world matrix is produced by running
scripts/create-community-worlds-matrix.mjs in the fork PR's checkout,
so any field on it is attacker-controlled. Forwarding
matrix.world.setup-command into the reusable workflow and eval-ing it
let a malicious fork PR execute arbitrary shell on the runner.
Replace the pass-through with a hardcoded per-world-id case in the
reusable workflows (only turso currently needs a setup step) and drop
the setup field from the matrix generator.
* rename step to "Per-world setup"
Addresses Copilot review feedback: the step no longer executes an
arbitrary command, so the old name was misleading.
Register the Fantastic Four community world packages in the worlds manifest and show the Redis variants in the Embedded docs section.
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
Co-authored-by: Pranay Prakash <pranay.gp@gmail.com>
The AI SDK renamed `Experimental_Agent` to `ToolLoopAgent`. Update the
"Building Durable AI Agents" page's API route snippet (v4 and v5) so it
matches the current AI SDK API.
Co-authored-by: Cursor <cursoragent@cursor.com>
Drops the label-based backport override in favor of workflow_dispatch.
The pull_request_target trigger has security concerns (it runs with
write permissions on PR-controlled events), and we already have a
manual dispatch path that covers the same use case.
* Update workflow-trace-view.tsx
* Update trace viewer layout to be in a row
Signed-off-by: Mitul Shah <mitulxshah@gmail.com>
---------
Signed-off-by: Mitul Shah <mitulxshah@gmail.com>
Large backport prompts (commit message + diff capped at 200KB) can
exceed Linux's `ARG_MAX` and cause `opencode run` to fail with
"Argument list too long" (exit 126). Redirect the prompt files into
`opencode run`'s stdin instead of passing them on the command line.
* docs: split v4/v5 content, fix version switcher end-to-end
## Content restructuring
- Split `docs/content/docs/` into `docs/content/docs/v4/` and
`docs/content/docs/v5/` so each version is a fully independent
content tree with no shared-file coupling
- v4 excludes the four pages that are v5-only (AbortController
cancellation docs and the serializable-abort-controller internal page)
- v5 retains all pages; `preRelease` frontmatter field removed (no
longer needed now that each version is its own folder)
- Removed `AbortController` / `AbortSignal` from v4 serialization page
(section moved to v5 only)
## Fumadocs source
- Added `v4docs` and `v5docs` as separate `defineDocs()` collections in
`source.config.ts`; shared `docsSchema` (no more `preRelease` field)
- `source.ts` exports both `source` (v4, `baseUrl: /docs`) and
`v5Source` (v5, same base URL)
## Version routing
- `version-source.ts` simplified: `filterPreReleaseFromNodes` and
`isPreReleaseUrl` logic removed; v4 tree uses `source`, v5 tree uses
`v5Source` + `rewriteNodeUrls`
- v4 `page.tsx`: removed `preRelease` guard (v4Source has no such pages)
- v5 `page.tsx`: uses `v5Source` for `getPage` / `generateStaticParams`
/ `generateMetadata`; `v5Link` wrapper rewrites `/docs/…` hrefs to
`/v5/docs/…` so inline MDX links stay in the v5 context
## Versioned cookbook
- Added `app/[lang]/v5/cookbook/` layout + page (mirrors v4 but uses
`v5Source`, `rewriteCookbookUrlForVersion`, and `V5CookbookLink`)
- `getCookbookTree` accepts a `versionPrefix` parameter; sidebar URLs
are prefixed accordingly (`/v5/cookbook/…`)
- `cookbook-tree.ts`: added `skipVersions?: string[]` per-recipe field
for version-specific exclusions; `distributed-abort-controller` is
marked `skipVersions: ['v5']`
## Version switcher — state & navigation
- New `VersionProvider` context (`hooks/geistdocs/use-version.tsx`)
backed by `localStorage`: URL is source of truth on versioned pages,
`localStorage` carries the preference across non-versioned pages
(cookbook overview, worlds, etc.)
- `VersionSwitcher` uses `useVersion()` context instead of URL-only
detection; now visible on all pages including cookbook
- `DesktopMenu` and `MobileMenu` use `activeVersion` from context so
the "Docs" and "Cookbook" navbar links resolve to the correct version
prefix on every page
- `buildVersionUrl` expanded to handle `/cookbook/…` paths alongside
`/docs/…`; non-versioned routes (worlds, api) return unchanged
- `switchVersion` does a `HEAD` probe before navigating; falls back to
the versioned cookbook or docs home if the target page doesn't exist
in that version (handles v4-only → v5 and v5-only → v4 cases)
## Cookbook content (v5)
- Rewrote `agent-cancellation` recipe using a single `AbortController`
pattern; removed Hard Cancellation vs Stop Signal two-approach
comparison
- Deleted `distributed-abort-controller` recipe from v5 (native
`AbortController` serialization makes it unnecessary)
- Removed references to distributed-abort-controller from
`cookbook/index.mdx` and `common-patterns/timeouts.mdx`
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(docs): use abortSignal (not signal) in DurableAgent.stream() options
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(docs): update prepack scripts to use versioned content paths
Content moved from docs/content/docs/ to docs/content/docs/v5/ on main
(pre-release channel). Stable branch will use v4/ after backport.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>