* Reduce unnecessary CI runtime
* Fix shared E2E artifact extraction path
* Stabilize getWorkflowPort timeout test on Windows
* Preserve UI unit coverage on CI fast path
* ci: extract wait-for-vercel-project to vercel/wait-for-deployment-action
The action's logic was duplicated between this repo and
vercel/workflow-server, which is annoying to keep in sync. Move it to
a standalone repository so both can consume the same pinned build.
Changes:
- Delete .github/actions/wait-for-vercel-project entirely.
- Replace all five `uses: ./.github/actions/wait-for-vercel-project`
references with `uses: vercel/wait-for-deployment-action@<sha>` in:
benchmarks.yml, dispatch-front-workflow-release-pr.yml,
docs-checks.yml, tarballs-checks.yml, tests.yml
- All `with:` inputs (project-slug, environment, timeout,
check-interval, github-token) are unchanged — the new action's
input contract is backwards-compatible.
The new action is ESM-only, targets Node 24, ships a ~12KB bundle
(down from ~830KB in the old in-repo version) by dropping
@actions/core and its transitive undici dependency, and is
unit-tested. See https://github.com/vercel/wait-for-deployment-action.
* ci: bump wait-for-deployment-action to fix/status-context-auto for verification
Repinning to vercel/wait-for-deployment-action#fix/status-context-auto
(SHA 04d46ef) which fixes the broken 'opt-out' heuristic that made
status-context resolution silently disabled for every consumer.
Reproduced in this repo's E2E logs:
Looking for GitHub deployment in environment "Preview – example-workflow"
Deployment ID resolution disabled (status-context is empty)
Deployment ready: https://example-workflow-...labs.vercel.dev
Run E2E Tests: VERCEL_DEPLOYMENT_ID= <-- empty
Will repin to the post-merge main SHA once CI is green.
* ci: bump wait-for-deployment-action pin to merged main SHA
Repinning from the fix/status-context-auto branch (04d46ef) to the
post-merge main SHA (0e2b0c5, vercel/wait-for-deployment-action#4).
The deployment-id resolution fix verified against the prior fix-branch
pin (E2E tests now read VERCEL_DEPLOYMENT_ID=dpl_... correctly across
the matrix; only flaky/unrelated Vercel deployment failures remain).
* ci: grant statuses:read alongside deployments:read
The wait-for-deployment-action also reads the 'Vercel – <slug>'
combined commit status to resolve the dpl_xxx ID. The official
permissions table lists statuses:read for
GET /repos/{owner}/{repo}/commits/{ref}/status.
Major-version refs like `@v2`/`@v5` resolve to mutable refs on the
upstream repos — sometimes a tag, sometimes a branch (e.g. marocchino
keeps `v1`/`v2`/`v3` as branches), and dawidd6 force-pushes the bare
`v6` tag forward outside of releases. A compromised maintainer account
could push new code that our CI picks up on the next run with
GITHUB_TOKEN (or, for changesets/action, NPM_TOKEN) in hand.
Pin all third-party `uses:` references to full commit SHAs with a
trailing version comment so the upstream release is still visible to
reviewers. Dependabot/Renovate can keep these fresh going forward.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The repo's enterprise `~ALL` required-signatures ruleset rejects the
unsigned commits produced by `peaceiris/actions-gh-pages@v4`, breaking
the benchmark and E2E result publishing jobs on `main`. Replace those
steps with a shared composite action that uses the GraphQL
`createCommitOnBranch` mutation — same pattern already used by
`backport.yml` — so commits are signed automatically by GitHub and
satisfy the rule.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* drop setup-command input from reusable community-world workflows
The community-world matrix is produced by running
scripts/create-community-worlds-matrix.mjs in the fork PR's checkout,
so any field on it is attacker-controlled. Forwarding
matrix.world.setup-command into the reusable workflow and eval-ing it
let a malicious fork PR execute arbitrary shell on the runner.
Replace the pass-through with a hardcoded per-world-id case in the
reusable workflows (only turso currently needs a setup step) and drop
the setup field from the matrix generator.
* rename step to "Per-world setup"
Addresses Copilot review feedback: the step no longer executes an
arbitrary command, so the old name was misleading.
* Add stable Next.js eager and lazy test coverage
* Address PR review feedback
* Fix eager Next step route builds
* Fix eager Next manifest refreshes
* Fix eager Next e2e stack assertions
* Externalize native step bundle bindings
* Lazy load Vercel world runtime
* Fix Next dev step sourcemap assertions
* Consolidate eager build changesets
* Fix Vercel world tracing in Next deployments
* Externalize Vercel world in Next builds
* Fix webpack tracing for Vercel world deps
* Fix eager workflow route bundling
* Rely on Next server externals
* Pass stale-banner via path: to sticky-pull-request-comment instead of message:
The 'Update existing test comment with stale warning' step inlined the
previous comment body via ${{ steps.get-comment.outputs.previous-results }}
into the action's `message:` input. As the test matrix grows, the
resulting argv can exceed ARG_MAX and the action fails with
'Argument list too long' — observed on a feature branch where the
matrix doubled.
Write the rendered stale-banner message to
$RUNNER_TEMP/stale-comment.md in the github-script step and pass the
path to sticky-pull-request-comment via its `path:` input instead.
This is robust to any future matrix size.
* Apply same fix to benchmarks.yml
Same ARG_MAX hazard exists in the benchmark workflow's stale-warning
step. Apply the identical `path:`-instead-of-`message:` refactor:
- The github-script step now writes the rendered stale-banner to
$RUNNER_TEMP/stale-comment.md and exposes the path as a step output.
- The sticky-pull-request-comment 'Update existing benchmark comment
with stale warning' step uses `path:` instead of inlining
${{ steps.get-comment.outputs.previous-results }} via `message:`.
The final 'Update PR comment with results' step in this workflow
already used `path: benchmark-summary.md`; only the stale-banner
update was inlined.
* Use `github.run_started_at` for stale-comment timestamps
The 'Started at:' label was sourced from `github.event.pull_request.updated_at`,
which is the PR metadata-update timestamp — not the workflow run start
time. That made the displayed timestamp:
- coupled to PR edits (label changes, description edits, etc.) rather
than to the actual CI run, and
- stale on workflow re-runs (an empty re-run would still show the
original PR-update time).
Switch all six occurrences across `tests.yml` and `benchmarks.yml` to
`github.run_started_at`, the canonical "this CI run started at"
timestamp.
* ci: switch Vercel deployment-protection bypass to OIDC Trusted Sources
The e2e, benchmark, and docs-smoke CI jobs previously used the static
`VERCEL_AUTOMATION_BYPASS_SECRET` deployment-protection bypass token
to reach protected Vercel deployments. Switch them over to the new OIDC
Trusted Sources flow: the GitHub Actions runner mints a short-lived
OIDC token via `core.getIDToken()` and forwards it on requests in the
`x-vercel-trusted-oidc-idp-token` header.
Each workbench project (and `workflow-docs`) has been configured with a
matching trusted-source rule:
aud=https://github.com/vercel, repository=vercel/workflow
The shared header helper now lives at `scripts/trusted-sources-headers.mjs`
and is imported by both the e2e/bench tests and the docs smoke script,
removing the previous duplication.
* rename to VERCEL_OIDC_TOKEN and wire through world-vercel
- Rename the env var from VERCEL_TRUSTED_OIDC_TOKEN to VERCEL_OIDC_TOKEN
to match Vercel's convention (also read by @vercel/oidc's
getVercelOidcToken()).
- In @workflow/world-vercel, replace the legacy
VERCEL_WORKFLOW_SERVER_PROTECTION_BYPASS / x-vercel-protection-bypass
flow with VERCEL_OIDC_TOKEN / x-vercel-trusted-oidc-idp-token. The
trusted-source header is attached on every outbound workflow-server
request (both proxied through api.vercel.com and direct).
- Drop the bypass header from the encryption-key and
resolve-latest-deployment fetches: those go to api.vercel.com which
is public.
- Drop VERCEL_WORKFLOW_SERVER_PROTECTION_BYPASS plumbing from tests.yml.
- Update the pending world-vercel changeset to describe the final
trusted-sources flow.
* .
* .
* ci: add statuses:read permission for wait-for-vercel-project action
The action queries /commits/{sha}/status (Commit Statuses API) in addition
to the Deployments API, in order to extract the Vercel `dpl_...` ID. With
an explicit permissions block in place, GITHUB_TOKEN now needs
`statuses: read` or the action 403s when resolving the deployment ID.
Reported by Copilot review on #1882.
* ci(docs): log status code and body when waitForServer times out
Helps diagnose deployment-protection / OIDC-trusted-source bypass
failures (e.g. SSO redirects) on the workflow-docs preview.
* ci(docs): log OIDC token claims (aud, repository, etc.) for diagnostics
Helps determine whether the bypass is failing because of missing
trusted-source config, claim mismatch, or audience mismatch.
* ci(docs): add curl debug step to verify OIDC header reaches Vercel
* .
* ci: remove debug logging now that trusted-sources config is correct
The fetch-failure root cause was the trusted-sources rule format: the
labs workbench projects had been PATCHed with just `to.slugs` (no
`preset`), but Vercel's edge requires the dashboard-form-style
`to.preset: 'all-custom'` field plus `development` in the slug list to
match incoming requests. After re-PATCHing all projects with the
correct format, the bypass works end-to-end.
* ci(docs): debug — test trusted-sources bypass against docs and labs deployments
Trying repository_owner claim added to one labs project to see if that
fixes the bypass.
* ci(docs): revert curl debug step
The GitHub Actions OIDC trusted-sources bypass returns 401 on all tested
projects regardless of claim configuration (including workflow-docs which
was set up via the dashboard). This is not a per-project config issue.
Need to investigate with Vercel team before continuing.
* ci(docs): probe trusted-sources bypass and surface x-vercel-id
Adds a debug step that does two HEAD requests against the docs preview
deployment (with and without the OIDC trusted-sources header) and prints
the response status line plus `x-vercel-id` for each. The proxy-side
trusted-sources changes for GitHub Actions OIDC tokens are rolling out
gradually (~12+ hours), so the edge-node identifier in `x-vercel-id`
helps explain why a request might succeed or fail during the rollout
window.
Also includes `x-vercel-id` in the `waitForServer` timeout error so
post-mortem analysis of failing runs has the same edge-node info.
* ci(docs): drop trusted-sources curl probe — bypass works once proxy fix reaches the serving edge node
The probe served its purpose: confirmed the bypass is functional once
the request lands on a region that has the proxy-side trusted-sources
fix rolled out. The waitForServer error message still surfaces
x-vercel-id for any future rollout-window debugging.
* .
* world-vercel: log outbound OIDC token claims once per process
Adds a one-shot diagnostic that prints the non-sensitive claims of the
OIDC token (`iss`, `aud`, `owner_id`, `project_id`, `environment`,
`sub`, `scope`, `exp`) on the first request that uses bearer auth.
This is invaluable for debugging Vercel deployment-protection
trusted-source rule mismatches: a 401 from the edge tells you nothing
about why the rule didn't match, and the token's claims are the only
thing that determines that. The signature is never logged.
Gated to once per process — Vercel-issued tokens are process-stable for
the lambda's lifetime so further log lines would just be redundant
spam.
* world-vercel: route trusted-sources header through getVercelOidcToken()
The Authorization bearer correctly preferred config.token (a static
Vercel auth token from CLI / Actions runner) and fell back to
getVercelOidcToken() inside a Vercel function. But the trusted-sources
bypass header (x-vercel-trusted-oidc-idp-token) was being read directly
from process.env.VERCEL_OIDC_TOKEN inside getHeaders(). That env var is
the bake-time token, frozen at deployment-creation time — on a project
that has been redeployed after a settings change, it carries stale
claims (e.g. an iss from when the project was briefly in 'global' mode)
that no longer match the workflow-server's trusted-sources rule.
Move trusted-sources header attachment from getHeaders() (sync) to
getHttpConfig() (async) and source it from getVercelOidcToken(). That
function reads getContext().headers['x-vercel-oidc-token'] first — a
freshly minted per-request token that always reflects current project
settings — and only falls back to the env var when that header is
missing.
Bearer auth source remains config.token-first.
Also expand the diagnostic to log claims from BOTH the per-request OIDC
token AND the bake-time env var so the divergence is visible in logs
when debugging future trusted-source mismatches.
Removes the now-misleading getProtectionBypassHeader() helper (its
'read env var directly' semantics were exactly the bug).
* world-vercel: skip OIDC trusted-sources header on proxied path
The two outbound flows have different auth requirements:
1. Proxied (usingProxy=true) — calls api.vercel.com/v1/workflow.
Public endpoint, authenticated with a static Vercel auth token via
config.token. The api-workflow proxy mints its own OIDC token
before forwarding to workflow-server, so the trusted-sources
bypass header on the SDK→proxy hop is meaningless. CLI, GitHub
Actions, and other API-client callers take this path.
2. Direct (usingProxy=false) — runs inside a Vercel deployment
talking straight to workflow-server. workflow-server validates a
Vercel OIDC bearer; Vercel's edge validates the trusted-sources
header. Both must come from getVercelOidcToken() (the per-request
fresh token), not process.env.VERCEL_OIDC_TOKEN (the bake-time
token that can be stale after a project config change).
Previously getHttpConfig attached x-vercel-trusted-oidc-idp-token on
both paths whenever getVercelOidcToken() resolved. That accidentally
forwarded the GitHub Actions OIDC token (when wired into
VERCEL_OIDC_TOKEN by the test runner) onto every SDK→proxy request,
which is harmless but wrong-by-design — the proxy is public, doesn't
look at that header on its inbound side, and the GHA token isn't its
intended audience.
Bearer auth source rules:
- Proxied: only config.token. (No fallback to OIDC; that auth
pathway doesn't go through the proxy's auth checks.)
- Direct: config.token (for tests / local dev), falling back to
getVercelOidcToken() (for Vercel-runtime calls).
* world-vercel: throw if proxied path is hit without a Vercel auth token
The api-workflow proxy authenticates the caller with a regular Vercel
auth token (not OIDC), so reaching the proxied path with no
config.token is always wrong: the proxy will reject the request and
the SDK caller would see an opaque 401 with no actionable hint.
Throw at config-resolution time with a clear message that points to
the WORKFLOW_VERCEL_AUTH_TOKEN env var the SDK reads from. Adds tests
covering both the no-token-throws case and the with-token-attaches-
bearer-and-skips-trusted-sources case.
* test(e2e): include x-vercel-id in startWorkflowViaHttp error message
When the trusted-sources bypass returns 401, the error message now
surfaces the response's x-vercel-id header so we can identify which
edge node served the failure. Helps distinguish proxy-rollout
incompleteness from actual config errors during incremental
rollouts of edge-side changes.
* ci: mint GHA OIDC tokens on demand to survive 5-minute expiry
GitHub Actions OIDC tokens have a hard 5-minute lifetime that cannot be
extended (no API to ask for a longer TTL — exp is always iat + ~300s).
Pre-minting once at the start of the job and shipping the result down
to the test runner via env var means tests that run late in the suite
hit an expired token and 401 on /api/trigger-pages (and any other
trusted-sources protected endpoint).
Move minting into scripts/trusted-sources-headers.mjs:
- getTrustedSourcesHeaders() is now async.
- It calls the runner's ACTIONS_ID_TOKEN_REQUEST_URL endpoint directly
(the env vars GHA exposes when permissions: id-token: write is on)
and re-mints 60s before the cached token's exp.
- Falls back to process.env.VERCEL_OIDC_TOKEN for non-GHA contexts
(Vercel runtime, local dev).
Workflow files drop the now-redundant 'Mint OIDC token' step and the
VERCEL_OIDC_TOKEN env-var passthrough on the test step. The runner env
vars propagate to subsequent steps automatically.
Updates all 17 callers in e2e.test.ts / bench.bench.ts / utils.ts /
docs/scripts/check-docs-smoke.mjs to await the now-async call.
* address PR #1882 code review
- Drop `statuses: read` from the three workflow permission blocks (the
wait-for-vercel-project action works without it on a public repo).
- Revert the `x-vercel-id` debug logging in `startWorkflowViaHttp`.
- Delete `packages/world-vercel/src/jwt-claims.ts` (debug-only helper).
- Drop the JWT claims diagnostic logging from `getHttpConfig`.
- Tighten the auth-flow comment in `getHttpConfig` and remove the
historical 'no longer attaches' note from `getHeaders`/its test.
- Restore `.changeset/world-vercel-protection-bypass.md` (already
shipped in a beta release per .changeset/pre.json).
- Trim the `.changeset/world-vercel-trusted-sources.md` description to
one short paragraph.
* docs(AGENTS): document local VERCEL_OIDC_TOKEN via vercel env pull
Configured trustedSources.projects on all 11 workbench app projects so
each one accepts a Vercel-issued OIDC token from any of the others. A
developer running e2e locally can now do `vercel env pull` from any
workbench app's directory and use the resulting VERCEL_OIDC_TOKEN to
bypass Deployment Protection on any of the workbench preview/prod
deployments — no need to disable protection on the project just to run
the suite locally.
A recurring Turbopack-on-Windows bug causes the dev server to enter a
'MODULE_UNPARSABLE' state during HMR in dev.test.ts, after which every
request returns 500. The remaining e2e suite then polls stuck workflows
for 60s each, burning the full 30-minute job window before getting
cancelled (~50% of recent main runs).
Bail out of the Windows e2e job as soon as dev.test.ts fails, and
health-check the dev server before kicking off test:e2e so any other
silent breakage is surfaced quickly instead of via a 30-minute timeout.
* ci: refactor wait-for-vercel-project to use GitHub Deployments API
Replaces the Vercel SDK / Vercel API token-based implementation with one
that resolves the deployment URL via the GitHub Deployments API:
- Find the GitHub Deployment for (target SHA, environment) where
environment matches the Vercel-app-created "Preview \u2013 <slug>" or
"Production \u2013 <slug>" naming pattern.
- Wait for the latest deployment status to be `success` (or `inactive`
when Vercel skips a duplicate build, in which case its environment_url
still points at the live deployment).
- Probe the URL to confirm the edge can route to it (any non-5xx
response counts as live, including 401/403 from Deployment Protection
and 404/405 from the app). Manual redirect handling treats redirects
to vercel.com as "still building".
- Resolve the dpl_xxx deployment ID from the matching commit status
(Vercel posts `Vercel \u2013 <slug>` statuses where target_url's last
path segment is the inspector ID == deployment ID without the prefix).
Inputs change: project-slug + bypass-secret + github-token (with
GITHUB_TOKEN default) replace team-id + project-id + vercel-token.
Removes the @vercel/sdk dependency, shrinking the bundled dist from
5.4MB to 829KB. The VERCEL_DOCS_TOKEN secret is no longer referenced
anywhere in the repo and can be deleted from GH after this lands.
* ci(wait-for-vercel-project): drop URL probe and bypass-secret input
The GitHub Deployment status transitions to `success` only after the
Vercel app finishes building and routing is live, so an extra HTTP
liveness probe of the deployment URL was redundant. Removing it lets
us also drop the bypass-secret input \u2014 protected deployments don't
need a workaround anymore because we never make the request.
Reduces the action surface area and eliminates a runtime fetch.
* ci(wait-for-vercel-project): address PR review
- Fail loudly when the dpl_xxx deployment ID can't be resolved instead
of returning an empty string. Consumers wire this into
VERCEL_DEPLOYMENT_ID, which world-target uses to pick between the
vercel and local worlds (packages/utils/src/world-target.ts), so an
empty value would silently flip execution mode.
- Pass the GitHub App token to wait-for-vercel-project in the dispatch
release workflow. The job sets `permissions: contents: read`, which
blocks the default GITHUB_TOKEN from reading the Deployments API.
The App token (already generated for workflow,front) has the
necessary scopes.
* ci: fix VERCEL_WORKFLOW_SERVER_* ternary so main actually unsets them
In GitHub Actions expressions, '' is falsy, so the original
`cond && '' || secrets.X` pattern always fell through to the secret
regardless of branch. The result was that pushes to main were sending
preview workflow-server values to production, causing 'invalid_url'
errors on `x-vercel-workflow-api-url` across all e2e jobs.
Flip the condition so the secret sits in the truthy branch and ||
correctly selects '' on main.
* ci: skip pnpm cache in matrix-generation jobs
The Get Test Matrix and Get Community Worlds Matrix jobs only run a
small Node script to emit a JSON matrix; they never run `pnpm install`.
With `cache: 'pnpm'` set on actions/setup-node, the post-job cache save
step fails with 'Path Validation Error' because the pnpm store path was
never created, marking the whole job as failed.
Add a cache-pnpm input to setup-workflow-dev (default true) and opt out
in the two matrix-generation jobs.
* CI script improvements
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* address codex feedback: harden remaining workflows
- add allowlist regex in prepare-workbench-path to block path traversal
- move matrix/input values to env vars across e2e-vercel-prod,
benchmarks (local/postgres/vercel), and the reusable community-world
workflows
- validate app-name/world-id/world-package inputs in the reusable
community-world workflows
- pipe getCommunityWorldsMatrix script output through jq -c to prevent
\$GITHUB_OUTPUT injection
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* feat(world-vercel): support WORKFLOW_VERCEL_PROTECTION_BYPASS env var
Allows sending a Vercel Deployment Protection bypass secret via the
`x-vercel-protection-bypass` header on all outbound requests made by
the Vercel world, enabling use against protected deployments (e.g.
previews, or workflow-server once protection is enabled).
* feat(world-vercel): support VERCEL_WORKFLOW_SERVER_URL env var
Replace hard-coded WORKFLOW_SERVER_URL_OVERRIDE constant with a function
that reads from the VERCEL_WORKFLOW_SERVER_URL env var. Allows configuring
the workflow-server URL per-deployment (e.g. workbench Preview envs
pointing to a branch deployment) without editing source.
* fix(world-vercel): preserve inline WORKFLOW_SERVER_URL_OVERRIDE const
Keep the inline const as an empty-string literal so external CI rewrite
tooling continues to work unmodified; the env var is a fallback when the
inline value is empty.
* refactor(world-vercel): rename to VERCEL_WORKFLOW_SERVER_PROTECTION_BYPASS
Align env var naming with VERCEL_WORKFLOW_SERVER_URL.
* ci: expose workflow-server protection bypass env vars to e2e-vercel-prod
Set VERCEL_WORKFLOW_SERVER_URL and VERCEL_WORKFLOW_SERVER_PROTECTION_BYPASS
on PR runs so e2e tests hit the protected workflow-server preview; leave
unset on main so production runs use the public default URL.
* refactor(world-vercel): address PR review comments
- Consolidate bypass header logic in getHeaders() to reuse
getProtectionBypassHeader() instead of duplicating env lookup.
- Use consistent 'Authorization' casing in direct fetch() calls.
- Add unit tests for getProtectionBypassHeader, getHttpUrl, and getHeaders
covering env var toggling and proxy/override combinations.
* ci: upgrade pnpm/action-setup to v6 and read version from package.json
Removes hardcoded pnpm version (10.14.0) from all workflows and instead
reads the version from the packageManager field in package.json, so CI
stays in sync with the version used locally.
* ci: update setup-workflow-dev composite action to use pnpm/action-setup@v6
Also removes the pnpm-version input since the action now reads the
version from package.json#packageManager.
* ci: downgrade pnpm/action-setup to v5
v6 installs pnpm 11 RC/beta, which has a regression
(pnpm/pnpm#11264, pnpm/action-setup#225/#227/#228) that causes
'ERR_PNPM_BROKEN_LOCKFILE: expected a single document in the stream'
when the project's packageManager pins a 10.x pnpm version. v5 is the
latest stable release before v6 and supports reading the version from
package.json#packageManager.
* test: improve e2e test failure diagnostics with run context and GitHub annotations
When e2e tests fail, automatically dump workflow run diagnostics (status,
input/output, error details, event timeline, dashboard link) to the CI
logs. Emit GitHub Actions annotations that surface on PR file diffs.
Fix collectedRunIds which was declared but never populated, enabling
observability links in the PR comment. Enrich the aggregation script
to include run IDs and dashboard URLs for failed tests.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: increase diagnostics hook timeout and fix flaky Vercel Prod tests
- Increase onTestFailed hook timeout to 30s (default was 10s) so
diagnostics can fetch run data even after slow test timeouts
- parallelSleepWorkflow: increase elapsed threshold from 10s to 25s to
accommodate Vercel cold start latency
- webhookWorkflow: increase hook polling deadline from 30s to 60s and
test timeout from 60s to 120s for slow Vercel webhook registration
- readableStreamWorkflow: stop reading once expected content is received
instead of waiting for stream close (which can hang on Vercel), and
increase test timeout to 120s
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: emit ::error annotations via process.stdout.write to bypass vitest ANSI prefix
Vitest's console interceptor prepends ANSI escape codes to console.log
output, which prevents GitHub Actions from parsing ::error workflow
commands. Use process.stdout.write() directly to ensure clean output.
Also enhance the custom reporter to emit annotations in onFinished
(which runs after vitest output is complete) as a reliable fallback,
and enrich failure data from the diagnostics sidecar.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: only show observability links for vercel-prod test failures
Community world and local tests don't run on Vercel's backend, so
dashboard links are meaningless for those categories. Previously,
test name collisions across sidecar files could cause community
test failures to show Vercel dashboard URLs from vercel-prod runs.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: link annotations to test files instead of symlinked workflow sources
The workflow source files in workbench/ are symlinks that GitHub can't
resolve, causing annotations to show raw paths like #L0 instead of
linking to code. Now:
- utils.ts: omit file= from onTestFailed annotations (just show title)
- github-reporter.ts: use the actual test file path (e.g.
packages/core/e2e/e2e.test.ts) which GitHub can resolve
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: use types.isNativeError() for cross-VM Error serialization
FatalError was not properly serialized when passed from workflow code into a step function because the Error reducer checked `value instanceof global.Error` where `global` is the VM's globalThis. Errors created in the host context (like FatalError from @workflow/errors) have a different Error prototype than the VM context, so the instanceof check returned false and the error was silently dropped.
Replaced with `types.isNativeError()` from `node:util` which uses V8's internal type tag and works across VM context boundaries.
* ci: don't cancel in-progress CI runs on main branch
## Summary
Refactors the E2E tests to call `start()` from `workflow/api` directly instead of going through the `/api/trigger` HTTP endpoint in each workbench app. This removes a layer of indirection — the tests now use the same API that users would use to start workflows programmatically.
### Before
```ts
const run = await triggerWorkflow('addTenWorkflow', [123]);
const returnValue = await getWorkflowReturnValue(run.runId);
```
- `triggerWorkflow()` sent an HTTP POST to `/api/trigger` on the workbench app
- The workbench app looked up the workflow function, called `start()`, and returned the run ID
- `getWorkflowReturnValue()` polled `GET /api/trigger?runId=...` until the workflow completed
### After
```ts
const run = await start(await e2e('addTenWorkflow'), [123]);
const returnValue = await run.returnValue;
```
- `e2e()` / `getWorkflowMetadata()` fetches the manifest from `/.well-known/workflow/v1/manifest.json` to look up the correct `workflowId`
- `start()` is called directly from the test process via the configured World
- `run.returnValue` polls for completion via the World (no HTTP polling endpoint needed)
### Changes
**`packages/core/e2e/e2e.test.ts`**
- Removed `triggerWorkflow()` and `getWorkflowReturnValue()` helpers
- Added `fetchManifest()` to fetch and cache the workflow manifest from the deployment
- Added `getWorkflowMetadata(file, fn)` to look up `{ workflowId }` from the manifest
- Added `e2e(fn)` shorthand for the common case of `workflows/99_e2e.ts`
- All tests call `start()` and `run.returnValue` directly
- Error tests use `.catch()` to inspect `WorkflowRunFailedError`
- Output stream tests use `run.getReadable()` directly (skipped on local world where cross-process streaming isn't supported)
- `beforeAll` configures the local World with the correct data directory and base URL
- Pages Router tests use `startWorkflowViaHttp()` to specifically validate the HTTP trigger path
**Workbench apps (hono, express, fastify, nest)**
- Removed `/api/trigger` route handlers
- Kept `/api/hook`, `/api/test-direct-step-call`, `/api/test-health-check` endpoints
- Re-added `_workflows.js` side-effect import for hono/express/fastify to maintain Nitro's HMR dependency graph
**Deleted trigger-only route files** from: nextjs-turbopack, nextjs-webpack, vite, sveltekit, astro, nuxt, nitro-v2, nitro-v3, example
**`.github/workflows/tests.yml`**
- Added `WORKFLOW_PUBLIC_MANIFEST: '1'` to all E2E test jobs
### Dependencies
Stacked on #963 which adds `WORKFLOW_PUBLIC_MANIFEST` support to all framework builders.
Added `WORKFLOW_SERVER_URL_OVERRIDE` configuration to the Vercel world adapter and removed the deprecated `WORKFLOW_VERCEL_SKIP_PROXY` and `WORKFLOW_VERCEL_BACKEND_URL` environment variables.
### What changed?
- Added a changeset for a patch release across multiple packages
- Removed `WORKFLOW_VERCEL_SKIP_PROXY` environment variable from GitHub workflow tests
- Removed `WORKFLOW_VERCEL_BACKEND_URL` from environment variables in CLI and core packages
- Simplified the URL resolution logic in the Vercel world adapter
- Added support for a `WORKFLOW_SERVER_URL_OVERRIDE` constant for testing against different workflow-server versions
- Added the `x-vercel-workflow-api-url` header when the URL override is set
### How to test?
1. Verify that Vercel deployments continue to work without the removed environment variables
2. Test with a custom workflow server URL by setting the `WORKFLOW_SERVER_URL_OVERRIDE` constant in the world-vercel package
### Why make this change?
This change simplifies the configuration for the Vercel world adapter by removing deprecated environment variables and standardizing on a cleaner approach for specifying the workflow API URL. The new implementation automatically determines whether to use the proxy based on project configuration, making it more intuitive and reducing the need for explicit configuration.
* Refactor e2e tests for errors
* Improve step error tests to check workflow return value
Step error workflows now catch the error and return message/stack,
making assertions cleaner. Tests verify both:
- Workflow return value (caught error message)
- CLI step result (original stack with function names)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Add assertion for stack trace in caught step error
With the fix in step.ts that propagates original stack traces,
we can now verify that caught step errors include function names
directly in the workflow return value (not just via CLI).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* logging
* Fix: When a step handler is re-invoked after max retries are exhausted, the step_failed event doesn't include a stack trace, breaking stack trace propagation for this edge case.
This commit fixes the issue reported at packages/core/src/runtime/step-handler.ts:127-146
## Stack trace loss in step_failed event when max retries are exceeded
**What fails:** Step handler does not include stack property in step_failed event when step is re-invoked after max retries exhausted, causing FatalError in workflow to lose original stack trace
**How to reproduce:**
1. Create a step that always fails with an error that includes a stack trace
2. Exhaust all retry attempts (e.g., maxRetries = 3, so 4 total attempts)
3. The step handler is re-invoked (attempt > maxRetries + 1)
4. This triggers the edge case at lines 127-146 in step-handler.ts
5. A step_failed event is created with fatal: true but without stack property
6. When step.ts processes this event (lines 103-104), the stack remains unset since event.eventData.stack is undefined
**Result:** FatalError created in workflow has default stack (from step.ts) instead of original error stack from step execution
**Expected:** Stack trace should be propagated consistently across all error paths. All other step_failed events with fatal: true include stack property (lines 311-313 and 351-353), but edge case was missing it.
**Root cause:** When step handler is re-invoked after max retries exhausted (attempt > maxRetries + 1), the code creates a step_failed event without including step.error?.stack, which contains the previous error information from the last failed attempt. Other code paths capture this correctly.
Co-authored-by: Vercel <vercel[bot]@users.noreply.github.com>
* add withData
* Enable source maps for step bundles and validate in e2e tests
- Move sourcemap generation from intermediate workflow bundle to steps bundle
- Enhance step error tests to validate function names and source files in stack traces
- Remove isLocalDeployment() checks since source maps now work in dev mode
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Add changeset for step bundle source maps
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Add changeset for core package e2e test improvements
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Enable source maps in CI e2e tests
Add NODE_OPTIONS="--enable-source-maps" to all e2e test jobs to ensure
stack traces show original source file paths.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Fix step error source map checks for local prod builds
Add hasStepSourceMaps() helper that correctly identifies when source maps
are expected to work:
- Vercel prod: works (production builds have proper source maps)
- Local dev: works (DEV_TEST_CONFIG is set, uses step bundle with inline source maps)
- Local prod: doesn't work (nitro/bundler output doesn't preserve source maps)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Add hasWorkflowSourceMaps() utility for vite-based framework exception
- Add isViteBasedFramework() helper to detect vite, sveltekit, astro apps
- Add hasWorkflowSourceMaps() to check if workflow errors have source maps
(known issue: vite-based frameworks in local deployments don't preserve them)
- Refactor e2e.test.ts to use the new utility instead of inline check
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* prevent sourcemap checks in nextjs and sveltekit (known offenders)
* improve source maps matrix check
* fix matrix again
* ugh more matrix ignoring
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: Vercel <vercel[bot]@users.noreply.github.com>
* perf: add 5s delay to benchmark steps to simulate real work
* perf: add realistic workloads to benchmark steps
Add realistic workloads to benchmark step functions:
- doWork() - 1 second delay to simulate real computation
- stressTestStep() - 1 second delay to simulate real computation
- genBenchStream() - generates ~5KB of data in 50 chunks
- transformStream() - uppercases stream content (renamed from doubleNumbers)
Add "Slurp Time" metric to stream benchmarks measuring time from
first byte to complete stream consumption, complementing TTFB.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(e2e): increase dev test timeouts for Windows
Windows file watching and rebuilding is slower than macOS/Linux,
causing the 10s timeouts to fail. Increased to 30s to accommodate.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix(ci): treat community world failures as warnings, not errors
Community world tests/benchmarks are now non-blocking:
- Removed from has_failures check in both tests.yml and benchmarks.yml
- Added separate has_warnings output for community failures
- Split PR comment notices: ❌ for failures, ⚠️ for community warnings
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: apply PR review suggestions
- Fix chunk size calculation: chunkSize - 11 for ~100 bytes per chunk
- Remove redundant metadata check (always truthy)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
The publish-results job was being skipped because it depends on the
summary job which uses `if: always()`. When a job uses `always()`,
downstream jobs need to also use `always()` in their condition,
otherwise GitHub Actions skips them when any upstream job fails.
This was causing e2e-results.json to never be published to gh-pages.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: compare benchmarks against PR base branch instead of main
- Use github.event.pull_request.base.ref instead of hardcoded main
- Remove search_artifacts: true to ensure most recent baseline is used
- For stacked PRs, this compares against the parent PR's baseline
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: group failed e2e tests by category and app in summary
Instead of listing each failed test as a separate item, group them by:
1. Category (world): e.g., "Community Worlds", "Vercel Production"
2. App (framework): e.g., "mongodb", "turso", "nextjs-turbopack"
This makes the summary much more readable when there are many failures.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: ensure local E2E tests always produce JSON output
- Add 'fastify' to app detection list in aggregate-e2e-results.js
- Change && to ; so e2e tests run even if dev.test.ts fails
- This ensures local-dev, local-prod, and local-postgres categories
appear in the E2E summary comment
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: ensure local E2E tests always produce JSON output
- Add 'fastify' to app detection list in aggregate-e2e-results.js
- Change && to ; so e2e tests run even if dev.test.ts fails
- This ensures local-dev, local-prod, and local-postgres categories
appear in the E2E summary comment
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* feat: publish CI results to GitHub Pages for docs
- Add generate-docs-data.js script to create JSON summaries from CI artifacts
- Add publish-results job to tests.yml and benchmarks.yml workflows
- Update docs/lib/worlds-data.ts to fetch from GitHub Pages URLs
- Results published to https://vercel.github.io/workflow/ci/
This allows the docs worlds page to display actual test/benchmark
results without requiring a GITHUB_TOKEN.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: correct outputFile path for local E2E test artifacts
The --outputFile path was using ../../ which placed files outside the
repo because pnpm run test:e2e executes from workspace root, not from
the cd'd workbench directory. This prevented local-dev, local-prod, and
local-postgres test results from being uploaded as artifacts.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: show green checkmark for skipped tests instead of warning
Skipped tests are intentional and shouldn't show as warnings in the
E2E test summary comments.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* feat: use collapsible sections in benchmark PR comment
Wrap each benchmark, stream benchmarks section, and summary tables in
<details> toggles to make the PR comment more compact and readable.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* feat: add Vercel observability links to benchmark PR comments
- Store runId in benchmark timing data
- Add project-slug to Vercel benchmark matrix
- Pass WORKFLOW_VERCEL_PROJECT_SLUG env var to benchmarks
- Store Vercel metadata (teamSlug, projectSlug, environment) in timing files
- Generate observability deep links for each Vercel world benchmark
- Show observability links below Production (Vercel) tables
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
* fix: use correct Vercel project slugs for observability links
- nextjs-turbopack → example-nextjs-workflow-turbopack
- nitro-v3 → workbench-nitro-workflow
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>