mirror of
https://github.com/vercel/workflow.git
synced 2026-09-14 19:59:43 +08:00
codex/atomic-start-postgres
27 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
27d0ce7904 |
Route preview benchmarks through the e2e server (#3274)
* Expose Workflow web server override Signed-off-by: Karthik Kalyanaraman <karthik.kalyanaraman@vercel.com> * Use neutral workflow server test URL Signed-off-by: Karthik Kalyanaraman <karthik.kalyanaraman@vercel.com> * Route preview benchmarks through e2e server Signed-off-by: Karthik Kalyanaraman <karthik.kalyanaraman@vercel.com> * Use an empty changeset Signed-off-by: Karthik Kalyanaraman <karthik.kalyanaraman@vercel.com> --------- Signed-off-by: Karthik Kalyanaraman <karthik.kalyanaraman@vercel.com> |
||
|
|
11dc036854 |
ci: stop deploying changeset-release/main, run its e2e against production (#3243)
* ci: stop deploying changeset-release/main, run its e2e against production The changesets action force-pushes `changeset-release/main`, and it can point at exactly main's HEAD SHA. Vercel keeps one commit status per project per SHA, so when both a production deployment (from main) and a preview deployment (from changeset-release/main) are built for the same commit, whichever finishes last owns the status. On 2026-07-30 the preview finished last, so `vercel/wait-for-deployment-action` — which reads the deployment ID out of that status — handed production e2e runs a preview deployment ID and forked runs across environments. Disable git deployments for that branch in every Vercel project rooted in this repo, and give the changeset PR's Vercel e2e lanes a deployment to test that actually exists: main's production deployment for the PR's base SHA, resolved by SHA so a mid-flight production build is waited out rather than silently replaced by an older one. Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> * ci: resolve changeset-release e2e deployments with the wait action, tokenless Per review: with changeset-release/main no longer deployed, main SHAs can never again be deployed to a second environment of these projects, so the per-SHA commit status the action reads is unambiguous for exactly this lane. Reuse vercel/wait-for-deployment-action with environment: production and sha pinned to the PR base SHA instead of the Vercel-API polling script, drop the script and its VERCEL_TOKEN usage, and inherit the action's inactive/skipped-build handling. Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> --------- Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> |
||
|
|
d53b055a2b | [ci] Run benchmarks in-deployment to avoid proxy overhead (#2967) | ||
|
|
8977666479 | [ci] Benchmark comment: show avg-latency deltas vs main (#2842) | ||
|
|
da4e0995b0 | [ci] Overhaul performance benchmarks: focused metrics + sticky PR comment (#2820) | ||
|
|
2bb4164c9c |
Add Platformatic World to worlds-manifest.json (#1450)
* Add Platformatic World to worlds-manifest.json Signed-off-by: marcopiraccini <marco.piraccini@gmail.com> * ci fixup Signed-off-by: marcopiraccini <marco.piraccini@gmail.com> * ci: pin platformatic world image to 0.8.1 and harden community-world runner Signed-off-by: marcopiraccini <marco.piraccini@gmail.com> * platforamtic-world version Signed-off-by: marcopiraccini <marco.piraccini@gmail.com> * ci: wire generic docker service-type into community benchmark workflow The shared community-worlds matrix now emits service-type "docker" for any world with non-builtin or multiple services (e.g. Platformatic, which needs postgres + the platformatic/workflow image). tests.yml's e2e-community path already handles it, but benchmarks.yml's benchmark-community path (label-gated, non-blocking) did not — so a "community-benchmarks" run would start no services and fail. Mirror the e2e "Start Docker services" step, package-version pin, and docker cleanup into benchmark-community-world.yml, and pass `services`/`version` through from benchmarks.yml. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Signed-off-by: marcopiraccini <marco.piraccini@gmail.com> Co-authored-by: Pranay Prakash <pranay.gp@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3867270be8 |
Reduce unnecessary CI runtime (#2151)
* Reduce unnecessary CI runtime * Fix shared E2E artifact extraction path * Stabilize getWorkflowPort timeout test on Windows * Preserve UI unit coverage on CI fast path |
||
|
|
3c50f8c77b |
ci: extract wait-for-vercel-project to vercel/wait-for-deployment-action (#2065)
* ci: extract wait-for-vercel-project to vercel/wait-for-deployment-action
The action's logic was duplicated between this repo and
vercel/workflow-server, which is annoying to keep in sync. Move it to
a standalone repository so both can consume the same pinned build.
Changes:
- Delete .github/actions/wait-for-vercel-project entirely.
- Replace all five `uses: ./.github/actions/wait-for-vercel-project`
references with `uses: vercel/wait-for-deployment-action@<sha>` in:
benchmarks.yml, dispatch-front-workflow-release-pr.yml,
docs-checks.yml, tarballs-checks.yml, tests.yml
- All `with:` inputs (project-slug, environment, timeout,
check-interval, github-token) are unchanged — the new action's
input contract is backwards-compatible.
The new action is ESM-only, targets Node 24, ships a ~12KB bundle
(down from ~830KB in the old in-repo version) by dropping
@actions/core and its transitive undici dependency, and is
unit-tested. See https://github.com/vercel/wait-for-deployment-action.
* ci: bump wait-for-deployment-action to fix/status-context-auto for verification
Repinning to vercel/wait-for-deployment-action#fix/status-context-auto
(SHA 04d46ef) which fixes the broken 'opt-out' heuristic that made
status-context resolution silently disabled for every consumer.
Reproduced in this repo's E2E logs:
Looking for GitHub deployment in environment "Preview – example-workflow"
Deployment ID resolution disabled (status-context is empty)
Deployment ready: https://example-workflow-...labs.vercel.dev
Run E2E Tests: VERCEL_DEPLOYMENT_ID= <-- empty
Will repin to the post-merge main SHA once CI is green.
* ci: bump wait-for-deployment-action pin to merged main SHA
Repinning from the fix/status-context-auto branch (04d46ef) to the
post-merge main SHA (0e2b0c5, vercel/wait-for-deployment-action#4).
The deployment-id resolution fix verified against the prior fix-branch
pin (E2E tests now read VERCEL_DEPLOYMENT_ID=dpl_... correctly across
the matrix; only flaky/unrelated Vercel deployment failures remain).
* ci: grant statuses:read alongside deployments:read
The wait-for-deployment-action also reads the 'Vercel – <slug>'
combined commit status to resolve the dpl_xxx ID. The official
permissions table lists statuses:read for
GET /repos/{owner}/{repo}/commits/{ref}/status.
|
||
|
|
ee61817865 |
ci: pin third-party GitHub Actions to commit SHAs (#2050)
Major-version refs like `@v2`/`@v5` resolve to mutable refs on the upstream repos — sometimes a tag, sometimes a branch (e.g. marocchino keeps `v1`/`v2`/`v3` as branches), and dawidd6 force-pushes the bare `v6` tag forward outside of releases. A compromised maintainer account could push new code that our CI picks up on the next run with GITHUB_TOKEN (or, for changesets/action, NPM_TOKEN) in hand. Pin all third-party `uses:` references to full commit SHAs with a trailing version comment so the upstream release is still visible to reviewers. Dependabot/Renovate can keep these fresh going forward. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
94572d741f |
ci: publish gh-pages updates via signed GraphQL commits (#2047)
The repo's enterprise `~ALL` required-signatures ruleset rejects the unsigned commits produced by `peaceiris/actions-gh-pages@v4`, breaking the benchmark and E2E result publishing jobs on `main`. Replace those steps with a shared composite action that uses the GraphQL `createCommitOnBranch` mutation — same pattern already used by `backport.yml` — so commits are signed automatically by GitHub and satisfy the rule. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
4708a77a35 |
CI: drop setup-command input from reusable community-world workflows (#1828)
* drop setup-command input from reusable community-world workflows The community-world matrix is produced by running scripts/create-community-worlds-matrix.mjs in the fork PR's checkout, so any field on it is attacker-controlled. Forwarding matrix.world.setup-command into the reusable workflow and eval-ing it let a malicious fork PR execute arbitrary shell on the runner. Replace the pass-through with a hardcoded per-world-id case in the reusable workflows (only turso currently needs a setup step) and drop the setup field from the matrix generator. * rename step to "Per-world setup" Addresses Copilot review feedback: the step no longer executes an arbitrary command, so the old name was misleading. |
||
|
|
1203dae70c |
Friendlier workflow errors (consolidated) (#1849)
* Introduce structured context-violation errors + Ansi renderer Phase 1: Add Ansi rendering helpers (frame, hint, note, help, code, inline) to @workflow/errors, and a chalk mock for readable snapshot tests. Phase 2: Add four context-violation error classes to @workflow/core (NotInWorkflowContextError, NotInStepContextError, NotInWorkflowOrStepContextError, UnavailableInWorkflowContextError) and apply them to all twelve user-facing throw sites so errors now include docs links and a structured "what/why/fix" frame. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * Address review: tighten changeset, implement ansifyName, harden Ansi - Tighten phase 1 changeset to a single sentence (per pranaygp review) and switch to double-quoted frontmatter (per Copilot + repo convention). - Implement `ansifyName` to actually apply dim styling to workflow/ / step/ prefixes; add an `Ansi.dim` helper to `@workflow/errors` so callers don't need to import chalk directly. - Remove the `void getWorkflowMetadata;` workaround in context-errors.ts by dropping the unused value import (we only needed the type and symbol). - Render the plain-Error throw in `workflow/get-workflow-metadata.ts` with `Ansi.frame` + docs link so the VM path matches the structured-class styling from the sibling step path (still uses a plain Error to avoid the module-init cycle). - Guard `buildUnderline` against zero-length markers so a stray empty token can't produce a negative `String.repeat` count. * Structured runtime logger metadata + fold in replay-timeout logging Adds a `.child()` and `.forRun(runId, workflowName)` child-logger API to the structured logger so runtime/step code doesn't have to repeat `workflowRunId`/`workflowName`/`stepId` on every call. Normalizes error metadata to structured `errorName` / `errorMessage` / `errorStack` fields instead of ad-hoc `error: err.message` strings, and adds comments to silent catches that swallow expected idempotency conflicts. Also folds in the pending changes from #1812 so that PR can be closed: - Standardize the console prefix to `[workflow-sdk]`. - Split the replay-timeout log into a warn-while-retrying vs. error-when-giving-up, and surface the underlying error when we can't mark a timed-out run as failed. - Include the error stack in the "Fatal runtime error during workflow setup" log and in the top-level user-code workflow error log so the stack surfaces in flattened log drains. - Drop the `[Workflows] "<runId>" - ` prefix from `buildWorkflowSuspensionMessage` — the structured logger now attaches run context. Supersedes #1812. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * Use double-quoted changeset frontmatter per repo convention * Add SerializationError + apply to user-facing serialization sites Phase 4 of friendlier errors: introduce a `SerializationError` class with an optional `hint` and a docs link (workflow-sdk.dev/err/serialization-failed), and adopt it at every user-facing serialization boundary in @workflow/core: - Locked ReadableStream at a workflow boundary - Unregistered class / missing `classId` / missing `WORKFLOW_DESERIALIZE` - Attempting to return step functions to clients or call workflow functions directly - Webhook `respondWith()` called outside a step - `dehydrate*` / `getSerializeStream` failures (workflow args/return, step args/return, stream chunks) Internal invariants (format prefix length checks, unknown format bytes, missing `STREAM_NAME_SYMBOL`, encryption key/size guards, etc.) now throw `WorkflowRuntimeError` instead of plain `Error` so the classifier and logger treat them consistently. `formatSerializationError` now returns `{ message, hint }` so the hint fragment can be rendered with the standard SerializationError framing instead of being baked into the message string. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * Use double-quoted changeset frontmatter per repo convention * Presentation-only user vs SDK error attribution Add describeError() that derives attribution and class-aware hints from existing error classes + RUN_ERROR_CODES — no event data changes. Wire into step failures, max-delivery exhaustion, run failures, and fatal setup errors so terminal logs include errorAttribution and a hint for known error types. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * Address review: describeError accepts precomputed errorCode + instanceof - `describeError(err, errorCode?)` now accepts an optional precomputed `RunErrorCode`. `classifyRunError(err)` only narrows to USER_ERROR / RUNTIME_ERROR, so the REPLAY_TIMEOUT and MAX_DELIVERIES_EXCEEDED branches were previously unreachable from the step / run failure log sites. Callers that know the failure category (runtime.ts for replay timeout and max-deliveries exhaustion) now pass the code in. - Context-violation checks use `instanceof` against the actual classes from context-errors.ts instead of a name-string set. Type-safe + survives class renames. - Wire the new hints through to the REPLAY_TIMEOUT and MAX_DELIVERIES_EXCEEDED log sites so those branches actually render a hint now. - 3 new tests cover the reachable code paths + precomputed-code override. - Changeset frontmatter switched to double quotes per repo convention. * Cosmetic consistency pass on remaining bare throws Internal invariants now use WorkflowRuntimeError so describeError attributes them to the SDK: missing startedAt, VM generateKey, closure-vars outside step context, ENOTSUP. defineHook().resume() formats schema validation failures as a readable list instead of a JSON blob. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * Use double-quoted changeset frontmatter per repo convention * Data-driven describeRunError + expose via @workflow/core/describe-error Observability renderers read persisted run_failed / step_failed event data, not live Error instances. describeRunError takes { errorCode, errorName } and returns the same { attribution, hint } shape as describeError, so the CLI and web UI can derive user-vs-SDK framing from the event log directly. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * Friendlier build-time errors: WorkflowBuildError class + applications Add `WorkflowBuildError` class in `@workflow/errors` with optional `hint` for an actionable next step, and apply it in `@workflow/builders` at user-facing sites: failed esbuild phases, unresolved built-in steps, and empty esbuild output now throw `WorkflowBuildError` with a hint pointing at the likely fix. Runtime invariants remain plain `Error`. * Polish friendlier-errors rendering: drop functionName leak, simplify docs link, redirect stack - Drop the readonly `functionName` param-property on context-error classes so util.inspect no longer prints a trailing `{ functionName: 'foo()' }` block. - Replace the `DocLink` ("label: https://…") shape with a plain `DocsUrl` template-literal type. Error output now renders a single clean line: `docs: https://…` (new `Ansi.docs` helper) instead of the noisier "note: Read more about foo(): https://…". - Add throw helpers (`throwNotInWorkflowContext`, etc.) that call `Error.captureStackTrace(err, stackStartFn)` on V8 engines so the top frame of the thrown error points at the user's call site instead of at the gate function inside the framework. Callers pass themselves as the boundary. - Refactor `defineHook()` (both root and `/workflow`) to use named function closures rather than `this.create`/`this.resume`, since the stack redirect relies on a stable function identity that survives destructuring. - Update context-errors.test.ts to snapshot the new `docs:` framing and to add a regression test asserting the top stack frame is the user call site. * Consolidate friendlier-errors stack: fix ANSI leak + non-retry semantics Addresses PR review feedback across the 8-phase friendlier-errors stack and fixes issues surfaced by manual testing (createHook() inside a step): - ANSI no longer leaks into .message / .stack. Context-violation errors now store plain text on .message and render the colored framed form lazily via [util.inspect.custom] / toString(). Structured logs, log drains, CBOR-serialized events, and JSON payloads no longer contain raw \x1B[...m bytes. - Context violations are now fatal. ContextViolationError sets fatal = true; FatalError.is(err) recognizes any error with a fatal: true own property. Calling createHook() from a step no longer burns three retry attempts on a guaranteed-to-fail context violation. - Ansi helpers moved to @workflow/errors/ansi subpath so imports from @workflow/errors no longer pull chalk into consumers that only want error classes (addresses reviewer VaguelySerious). - Shared redirectStackToCaller helper in packages/core/src/capture-stack.ts, used by both context-errors.ts and workflow/get-workflow-metadata.ts (addresses Copilot review on #1849). - Structured framed content: ContextViolationError now takes a structured FramedContent (title segments + detail branches) and renders plain/pretty from the same source of truth. Tightens the eight existing phase changesets to 1-2 sentences each and adds four new scoped changesets (errors-ansi-subpath, context-errors-plain-message, context-errors-fatal, capture-stack-shared) for the followup fixes, so the final changelog history stays readable. * test: update step-handler mocks for scoped forRun() logger The runtime logger now uses .forRun(runId, name, {stepId, stepName}) to attach scope context, so 409-handling log calls no longer repeat {workflowRunId, stepId} in every metadata bag — those live on the scoped logger instance. Update the mock to return itself from forRun() and tighten assertions to check both the log args (errorName/errorMessage) and the forRun() scope. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * Mark SerializationError fatal + route dehydration through step-failure path SerializationError now carries readonly fatal = true. Step-return dehydration is wrapped inside the user-code try/catch so that the resulting error flows through userCodeFailed → step_failed → FatalError.is() short-circuit instead of bubbling up as HTTP 500 and triggering a queue retry loop. Retrying a step that returned a non-POJO is guaranteed to fail the same way, so this saves ~20s and 3 near- identical error blocks per serialization failure. * Add logging snapshot tests + manual-test artifacts Snapshot tests lock in the exact shape of: - describeError() payloads (attribution, errorCode, hint) for every classification — plain Error, SerializationError, context-violation, WorkflowRuntimeError, REPLAY_TIMEOUT, MAX_DELIVERIES_EXCEEDED. - The scoped-logger call signature for the two canonical runtime failure paths (fatal-bubble and hit-max-retries), so refactors of forRun() / child() metadata merging can't silently change what users see in their log drains. SerializationError now also has a direct test for readonly fatal=true + FatalError.is() recognition. pr-artifacts/ contains real log-output snapshots from running the nextjs-turbopack workbench against five error scenarios. These are reference material for reviewers and are flagged to be removed before merge. * Readable step-fatal logs: inline stack + friendly step/workflow names The step-level fatal-error log used to embed the full stack trace inside an `errorStack` string field in the metadata object, so util.inspect rendered it as a quote-escaped, line-continuation blob when the log hit the terminal — unreadable in practice. Move framing + stack into the log *message* (matching the workflow-level log in runtime.ts) and keep the metadata object compact with only the indexable structured fields (`errorAttribution`, `errorName`, `errorMessage`, `hint`, IDs). Log drains still get the same keys; humans now see a readable stack trace. Also introduce `formatStepName` / `formatWorkflowName` in `@workflow/utils` that render machine names (`step//./workflows/1_simple//add`) as `add (./workflows/1_simple)` in log framings, using the existing `parseStepName` / `parseWorkflowName` parsers. Applied to step-fatal, hit-max-retries, exceeded-max-retries, and workflow-threw log sites. Artifacts in pr-artifacts/ updated to show the new output shape, and renamed .log → .md since they're Markdown and IDE previews are nicer that way. * Opinionated pretty formatter for runtime structured-log metadata Replace util.inspect's default object dump (which quote-escapes multi-line stacks and paragraph hints into a single-line JSON-y blob) with a workflow-aware formatter that composes the entire log line into a single string passed to console.error / console.warn. Highlights of the new output: - Per-run / per-step IDs render with their parsed friendly names so users see `wrun_… · simple (./workflows/1_simple)` instead of just the raw `workflowName: 'workflow//./workflows/1_simple//simple'`. - Color-coded attribution badge (user error red / sdk error magenta) paired with the error class in bold. - Hints render as a paragraph under `hint:` rather than a backslash- `\n`-escaped string. - Drops redundant fields (errorStack always; errorMessage when it's already in the parent message) to avoid double-printing. - Unknown fields fall through as a sorted `key value` tail so we never silently drop log information. @workflow/errors/ansi gains bold/red/magenta helpers used by the formatter. The web / web-shared packages don't consume stderr — they read structured event payloads from the World event log — so this is presentation-only at the runtime layer. * ci(benchmarks): disable pnpm cache for getCommunityWorldsMatrix The job never runs `pnpm install` (it just calls `node` against a checked-in script), so the pnpm store path never exists. The post-job `actions/setup-node@v4` cache-save then fails with `Path Validation Error: Path(s) specified in the action for caching do(es) not exist` and red-X's the entire job even though the matrix step succeeded. The setup-workflow-dev composite already has a `cache-pnpm` opt-out input for this exact case — wire it through here. * Address PR review comments: inspect dedup, cause leak, retry-loop tests - ContextViolationError: util.inspect(err) duplicated every framed detail line because the stack-tail strip only sliced the first message line. V8's Error.stack reads `Name: messageLine1\n messageLine2\n at ...`, so for our multi-line `title\n╰▶ docs: …` messages every detail line was getting prepended twice (once in the pretty form, once via the unsliced message tail). Count the actual message lines and slice past all of them. Repro test asserts `╰▶ docs:` appears exactly once. - WorkflowError: stop assigning `cause: undefined` as an enumerable own property when no cause is provided. Subclasses (every error in this PR) inherit the parent constructor; the unconditional assignment polluted `util.inspect(err)` output with `{ cause: undefined, … }` on every no-cause instance. The `super(...)` call already conditionally sets `.cause` non-enumerably when `options.cause` is provided. - step-handler.test.ts: add a regression-gate suite that exercises the fatal-vs-retryable retry-loop wiring directly. Asserts that an error with `fatal: true` produces exactly one `step_failed` event with no `step_retrying`, and that a non-fatal `Error` retries via `step_retrying` on early attempts and emits `step_failed` once the retry budget is exhausted. Catches the silent-regression case where `fatal = true` is removed from a context-violation error class but the `FatalError.is()` unit tests stay green. * Consolidate changesets + remove pr-artifacts Address review feedback to drastically shorten the changesets — fold the 15 file-by-file entries into a single user-facing changeset for @workflow/core / errors / builders / utils. Also drop the pr-artifacts/ folder (reviewer-only log captures, no longer needed). * Polish runtime error logging: layout, stack trim, hint consolidation Five user-driven fixes from manual smoke-testing of #1849: 1. Logger layout. composeLogLine() now puts the structured-fields block (attribution badge, run/step IDs, error code) **between** the framing line and the stack body, instead of after it where 30+ lines of stack buried the most useful information. The framing stays at the top, stack at the bottom, structured info readable at a glance. 2. Stack trim. Drops framework-internal frames (`node_modules/.pnpm/`, `node:internal/`, Turbopack-bundled `node_modules__pnpm_*` chunks, `_next_dist_*` chunks) and caps the surviving frame count at 6 so the stack stays compact even on heavy async wrappers. Suppressed runs emit one summary line so users know the trim happened. 3. Wrapper-route noise. The nextjs-turbopack workbench's start route was catching `WorkflowRunFailedError` rejection on `Promise.race([readLoop(), run.returnValue])` and re-logging it via `console.error('Error in workflow stream:', error)` plus `controller.error(error)` — which then triggered Next.js's `⨯ failed to pipe response` overlay. The SDK already logs the failure cleanly upstream and the runId is on the response header, so the wrapper now closes the SSE stream cleanly on WorkflowRunFailedError. 4. Consistent framed `╰▶ hint:` / `╰▶ docs:` layout for all errors that carry a hint or docs slug. WorkflowError, SerializationError, and WorkflowBuildError now share one `appendFramedDetails` helper matching the box-drawing structure that ContextViolationError already used. Was: blank-line-separated `Learn more: <url>`. Now: one tree, indistinguishable from context-violation rendering. 5. Drop the duplicate logger-side `hint` field. Hints now live on the error message only — actionable hints get serialized into the event log, rehydrated on the workflow side, and shown in observability automatically. The previous logger-only hint duplicated stderr but never made it past the step boundary. Updated SerializationError hint to point at the foundations doc ("Ensure you're returning workflow serializable types. Check the serialization docs to see what's serializable: https://workflow-sdk.dev/docs/foundations/serialization") instead of the hardcoded `(plain objects, arrays, primitives, …)` list, which drifted out of sync as the supported types grew. Same hint reuses for step args, workflow args/return, stream messages, and any other site that goes through `formatSerializationError`. Also retitled the retry summary `3 retries` → `3 max retries` since "3 retries" next to "4 attempts" was ambiguous (already-happened vs. budget). * Trim error-card title + drop machine step name from persisted error - ErrorStackBlock (web observability): show just the first non-empty trimmed line of the error message in the card title with single-line truncation. Multi-line messages (`Failed to serialize step return value\n╰▶ hint: …`) were rendering the entire framed body in the title, pushing the copy button off-screen and burying the scannability of the headline. Full message stays in the body via the stack (V8 prepends `Name: message` to `Error.stack`), so no information is lost; hover-tooltip exposes the full title text. - Persisted error message: drop the `Step "step//./.../foo"` machine name from `Step failed after N retries: …` and `Step exceeded max retries (…)` strings. Observability already attributes the event to a specific step via the UI tree, and the CLI logger emits the friendly `Step foo (./...) hit max retries` framing on its own line. Embedding the raw `step//./...` machine name in the persisted message text was duplicate noise. * Update .changeset/friendlier-errors.md Co-authored-by: Peter Wielander <mittgfu@gmail.com> Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> * Update .changeset/pretty-log-format.md Co-authored-by: Peter Wielander <mittgfu@gmail.com> Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> * Update SerializationError snapshot tests for slug-less message The class no longer attaches a slug-based `╰▶ docs:` line — the foundations URL is embedded directly in the hint via the `formatSerializationError` helper in @workflow/core. Update the test expectations accordingly: - bare-title case is now a single line (no docs link) - hint case renders one `╰▶ hint: …` branch (no second branch) * Update serialization.test.ts hint assertions for foundations URL Four `should throw error for an unsupported type` cases were still asserting on the old hardcoded type list. Update to the new hint phrasing that points at the foundations doc, matching the change in `formatSerializationError` (`packages/core/src/serialization/errors.ts`). --------- Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> Co-authored-by: Peter Wielander <mittgfu@gmail.com> |
||
|
|
059821cb39 |
ci: pass stale-banner via path: to sticky-pull-request-comment in tests + benchmarks workflows (#1887)
* Pass stale-banner via path: to sticky-pull-request-comment instead of message:
The 'Update existing test comment with stale warning' step inlined the
previous comment body via ${{ steps.get-comment.outputs.previous-results }}
into the action's `message:` input. As the test matrix grows, the
resulting argv can exceed ARG_MAX and the action fails with
'Argument list too long' — observed on a feature branch where the
matrix doubled.
Write the rendered stale-banner message to
$RUNNER_TEMP/stale-comment.md in the github-script step and pass the
path to sticky-pull-request-comment via its `path:` input instead.
This is robust to any future matrix size.
* Apply same fix to benchmarks.yml
Same ARG_MAX hazard exists in the benchmark workflow's stale-warning
step. Apply the identical `path:`-instead-of-`message:` refactor:
- The github-script step now writes the rendered stale-banner to
$RUNNER_TEMP/stale-comment.md and exposes the path as a step output.
- The sticky-pull-request-comment 'Update existing benchmark comment
with stale warning' step uses `path:` instead of inlining
${{ steps.get-comment.outputs.previous-results }} via `message:`.
The final 'Update PR comment with results' step in this workflow
already used `path: benchmark-summary.md`; only the stale-banner
update was inlined.
* Use `github.run_started_at` for stale-comment timestamps
The 'Started at:' label was sourced from `github.event.pull_request.updated_at`,
which is the PR metadata-update timestamp — not the workflow run start
time. That made the displayed timestamp:
- coupled to PR edits (label changes, description edits, etc.) rather
than to the actual CI run, and
- stale on workflow re-runs (an empty re-run would still show the
original PR-update time).
Switch all six occurrences across `tests.yml` and `benchmarks.yml` to
`github.run_started_at`, the canonical "this CI run started at"
timestamp.
|
||
|
|
cd50618d1f |
ci: switch Vercel deployment-protection bypass to OIDC Trusted Sources (#1882)
* ci: switch Vercel deployment-protection bypass to OIDC Trusted Sources The e2e, benchmark, and docs-smoke CI jobs previously used the static `VERCEL_AUTOMATION_BYPASS_SECRET` deployment-protection bypass token to reach protected Vercel deployments. Switch them over to the new OIDC Trusted Sources flow: the GitHub Actions runner mints a short-lived OIDC token via `core.getIDToken()` and forwards it on requests in the `x-vercel-trusted-oidc-idp-token` header. Each workbench project (and `workflow-docs`) has been configured with a matching trusted-source rule: aud=https://github.com/vercel, repository=vercel/workflow The shared header helper now lives at `scripts/trusted-sources-headers.mjs` and is imported by both the e2e/bench tests and the docs smoke script, removing the previous duplication. * rename to VERCEL_OIDC_TOKEN and wire through world-vercel - Rename the env var from VERCEL_TRUSTED_OIDC_TOKEN to VERCEL_OIDC_TOKEN to match Vercel's convention (also read by @vercel/oidc's getVercelOidcToken()). - In @workflow/world-vercel, replace the legacy VERCEL_WORKFLOW_SERVER_PROTECTION_BYPASS / x-vercel-protection-bypass flow with VERCEL_OIDC_TOKEN / x-vercel-trusted-oidc-idp-token. The trusted-source header is attached on every outbound workflow-server request (both proxied through api.vercel.com and direct). - Drop the bypass header from the encryption-key and resolve-latest-deployment fetches: those go to api.vercel.com which is public. - Drop VERCEL_WORKFLOW_SERVER_PROTECTION_BYPASS plumbing from tests.yml. - Update the pending world-vercel changeset to describe the final trusted-sources flow. * . * . * ci: add statuses:read permission for wait-for-vercel-project action The action queries /commits/{sha}/status (Commit Statuses API) in addition to the Deployments API, in order to extract the Vercel `dpl_...` ID. With an explicit permissions block in place, GITHUB_TOKEN now needs `statuses: read` or the action 403s when resolving the deployment ID. Reported by Copilot review on #1882. * ci(docs): log status code and body when waitForServer times out Helps diagnose deployment-protection / OIDC-trusted-source bypass failures (e.g. SSO redirects) on the workflow-docs preview. * ci(docs): log OIDC token claims (aud, repository, etc.) for diagnostics Helps determine whether the bypass is failing because of missing trusted-source config, claim mismatch, or audience mismatch. * ci(docs): add curl debug step to verify OIDC header reaches Vercel * . * ci: remove debug logging now that trusted-sources config is correct The fetch-failure root cause was the trusted-sources rule format: the labs workbench projects had been PATCHed with just `to.slugs` (no `preset`), but Vercel's edge requires the dashboard-form-style `to.preset: 'all-custom'` field plus `development` in the slug list to match incoming requests. After re-PATCHing all projects with the correct format, the bypass works end-to-end. * ci(docs): debug — test trusted-sources bypass against docs and labs deployments Trying repository_owner claim added to one labs project to see if that fixes the bypass. * ci(docs): revert curl debug step The GitHub Actions OIDC trusted-sources bypass returns 401 on all tested projects regardless of claim configuration (including workflow-docs which was set up via the dashboard). This is not a per-project config issue. Need to investigate with Vercel team before continuing. * ci(docs): probe trusted-sources bypass and surface x-vercel-id Adds a debug step that does two HEAD requests against the docs preview deployment (with and without the OIDC trusted-sources header) and prints the response status line plus `x-vercel-id` for each. The proxy-side trusted-sources changes for GitHub Actions OIDC tokens are rolling out gradually (~12+ hours), so the edge-node identifier in `x-vercel-id` helps explain why a request might succeed or fail during the rollout window. Also includes `x-vercel-id` in the `waitForServer` timeout error so post-mortem analysis of failing runs has the same edge-node info. * ci(docs): drop trusted-sources curl probe — bypass works once proxy fix reaches the serving edge node The probe served its purpose: confirmed the bypass is functional once the request lands on a region that has the proxy-side trusted-sources fix rolled out. The waitForServer error message still surfaces x-vercel-id for any future rollout-window debugging. * . * world-vercel: log outbound OIDC token claims once per process Adds a one-shot diagnostic that prints the non-sensitive claims of the OIDC token (`iss`, `aud`, `owner_id`, `project_id`, `environment`, `sub`, `scope`, `exp`) on the first request that uses bearer auth. This is invaluable for debugging Vercel deployment-protection trusted-source rule mismatches: a 401 from the edge tells you nothing about why the rule didn't match, and the token's claims are the only thing that determines that. The signature is never logged. Gated to once per process — Vercel-issued tokens are process-stable for the lambda's lifetime so further log lines would just be redundant spam. * world-vercel: route trusted-sources header through getVercelOidcToken() The Authorization bearer correctly preferred config.token (a static Vercel auth token from CLI / Actions runner) and fell back to getVercelOidcToken() inside a Vercel function. But the trusted-sources bypass header (x-vercel-trusted-oidc-idp-token) was being read directly from process.env.VERCEL_OIDC_TOKEN inside getHeaders(). That env var is the bake-time token, frozen at deployment-creation time — on a project that has been redeployed after a settings change, it carries stale claims (e.g. an iss from when the project was briefly in 'global' mode) that no longer match the workflow-server's trusted-sources rule. Move trusted-sources header attachment from getHeaders() (sync) to getHttpConfig() (async) and source it from getVercelOidcToken(). That function reads getContext().headers['x-vercel-oidc-token'] first — a freshly minted per-request token that always reflects current project settings — and only falls back to the env var when that header is missing. Bearer auth source remains config.token-first. Also expand the diagnostic to log claims from BOTH the per-request OIDC token AND the bake-time env var so the divergence is visible in logs when debugging future trusted-source mismatches. Removes the now-misleading getProtectionBypassHeader() helper (its 'read env var directly' semantics were exactly the bug). * world-vercel: skip OIDC trusted-sources header on proxied path The two outbound flows have different auth requirements: 1. Proxied (usingProxy=true) — calls api.vercel.com/v1/workflow. Public endpoint, authenticated with a static Vercel auth token via config.token. The api-workflow proxy mints its own OIDC token before forwarding to workflow-server, so the trusted-sources bypass header on the SDK→proxy hop is meaningless. CLI, GitHub Actions, and other API-client callers take this path. 2. Direct (usingProxy=false) — runs inside a Vercel deployment talking straight to workflow-server. workflow-server validates a Vercel OIDC bearer; Vercel's edge validates the trusted-sources header. Both must come from getVercelOidcToken() (the per-request fresh token), not process.env.VERCEL_OIDC_TOKEN (the bake-time token that can be stale after a project config change). Previously getHttpConfig attached x-vercel-trusted-oidc-idp-token on both paths whenever getVercelOidcToken() resolved. That accidentally forwarded the GitHub Actions OIDC token (when wired into VERCEL_OIDC_TOKEN by the test runner) onto every SDK→proxy request, which is harmless but wrong-by-design — the proxy is public, doesn't look at that header on its inbound side, and the GHA token isn't its intended audience. Bearer auth source rules: - Proxied: only config.token. (No fallback to OIDC; that auth pathway doesn't go through the proxy's auth checks.) - Direct: config.token (for tests / local dev), falling back to getVercelOidcToken() (for Vercel-runtime calls). * world-vercel: throw if proxied path is hit without a Vercel auth token The api-workflow proxy authenticates the caller with a regular Vercel auth token (not OIDC), so reaching the proxied path with no config.token is always wrong: the proxy will reject the request and the SDK caller would see an opaque 401 with no actionable hint. Throw at config-resolution time with a clear message that points to the WORKFLOW_VERCEL_AUTH_TOKEN env var the SDK reads from. Adds tests covering both the no-token-throws case and the with-token-attaches- bearer-and-skips-trusted-sources case. * test(e2e): include x-vercel-id in startWorkflowViaHttp error message When the trusted-sources bypass returns 401, the error message now surfaces the response's x-vercel-id header so we can identify which edge node served the failure. Helps distinguish proxy-rollout incompleteness from actual config errors during incremental rollouts of edge-side changes. * ci: mint GHA OIDC tokens on demand to survive 5-minute expiry GitHub Actions OIDC tokens have a hard 5-minute lifetime that cannot be extended (no API to ask for a longer TTL — exp is always iat + ~300s). Pre-minting once at the start of the job and shipping the result down to the test runner via env var means tests that run late in the suite hit an expired token and 401 on /api/trigger-pages (and any other trusted-sources protected endpoint). Move minting into scripts/trusted-sources-headers.mjs: - getTrustedSourcesHeaders() is now async. - It calls the runner's ACTIONS_ID_TOKEN_REQUEST_URL endpoint directly (the env vars GHA exposes when permissions: id-token: write is on) and re-mints 60s before the cached token's exp. - Falls back to process.env.VERCEL_OIDC_TOKEN for non-GHA contexts (Vercel runtime, local dev). Workflow files drop the now-redundant 'Mint OIDC token' step and the VERCEL_OIDC_TOKEN env-var passthrough on the test step. The runner env vars propagate to subsequent steps automatically. Updates all 17 callers in e2e.test.ts / bench.bench.ts / utils.ts / docs/scripts/check-docs-smoke.mjs to await the now-async call. * address PR #1882 code review - Drop `statuses: read` from the three workflow permission blocks (the wait-for-vercel-project action works without it on a public repo). - Revert the `x-vercel-id` debug logging in `startWorkflowViaHttp`. - Delete `packages/world-vercel/src/jwt-claims.ts` (debug-only helper). - Drop the JWT claims diagnostic logging from `getHttpConfig`. - Tighten the auth-flow comment in `getHttpConfig` and remove the historical 'no longer attaches' note from `getHeaders`/its test. - Restore `.changeset/world-vercel-protection-bypass.md` (already shipped in a beta release per .changeset/pre.json). - Trim the `.changeset/world-vercel-trusted-sources.md` description to one short paragraph. * docs(AGENTS): document local VERCEL_OIDC_TOKEN via vercel env pull Configured trustedSources.projects on all 11 workbench app projects so each one accepts a Vercel-issued OIDC token from any of the others. A developer running e2e locally can now do `vercel env pull` from any workbench app's directory and use the resulting VERCEL_OIDC_TOKEN to bypass Deployment Protection on any of the workbench preview/prod deployments — no need to disable protection on the project just to run the suite locally. |
||
|
|
3a08eaa0a1 |
ci: refactor wait-for-vercel-project to use GitHub Deployments API (#1861)
* ci: refactor wait-for-vercel-project to use GitHub Deployments API Replaces the Vercel SDK / Vercel API token-based implementation with one that resolves the deployment URL via the GitHub Deployments API: - Find the GitHub Deployment for (target SHA, environment) where environment matches the Vercel-app-created "Preview \u2013 <slug>" or "Production \u2013 <slug>" naming pattern. - Wait for the latest deployment status to be `success` (or `inactive` when Vercel skips a duplicate build, in which case its environment_url still points at the live deployment). - Probe the URL to confirm the edge can route to it (any non-5xx response counts as live, including 401/403 from Deployment Protection and 404/405 from the app). Manual redirect handling treats redirects to vercel.com as "still building". - Resolve the dpl_xxx deployment ID from the matching commit status (Vercel posts `Vercel \u2013 <slug>` statuses where target_url's last path segment is the inspector ID == deployment ID without the prefix). Inputs change: project-slug + bypass-secret + github-token (with GITHUB_TOKEN default) replace team-id + project-id + vercel-token. Removes the @vercel/sdk dependency, shrinking the bundled dist from 5.4MB to 829KB. The VERCEL_DOCS_TOKEN secret is no longer referenced anywhere in the repo and can be deleted from GH after this lands. * ci(wait-for-vercel-project): drop URL probe and bypass-secret input The GitHub Deployment status transitions to `success` only after the Vercel app finishes building and routing is live, so an extra HTTP liveness probe of the deployment URL was redundant. Removing it lets us also drop the bypass-secret input \u2014 protected deployments don't need a workaround anymore because we never make the request. Reduces the action surface area and eliminates a runtime fetch. * ci(wait-for-vercel-project): address PR review - Fail loudly when the dpl_xxx deployment ID can't be resolved instead of returning an empty string. Consumers wire this into VERCEL_DEPLOYMENT_ID, which world-target uses to pick between the vercel and local worlds (packages/utils/src/world-target.ts), so an empty value would silently flip execution mode. - Pass the GitHub App token to wait-for-vercel-project in the dispatch release workflow. The job sets `permissions: contents: read`, which blocks the default GITHUB_TOKEN from reading the Deployments API. The App token (already generated for workflow,front) has the necessary scopes. |
||
|
|
e2ef3568a5 |
CI script improvements (#1826)
* CI script improvements Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * address codex feedback: harden remaining workflows - add allowlist regex in prepare-workbench-path to block path traversal - move matrix/input values to env vars across e2e-vercel-prod, benchmarks (local/postgres/vercel), and the reusable community-world workflows - validate app-name/world-id/world-package inputs in the reusable community-world workflows - pipe getCommunityWorldsMatrix script output through jq -c to prevent \$GITHUB_OUTPUT injection Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
afa3931a59 | Add stress benchmarks: 1000-step, data payload, and stream tests (#1214) | ||
|
|
81a883bc9b |
ci: don't cancel in-progress CI runs on main branch (#1166)
* fix: use types.isNativeError() for cross-VM Error serialization FatalError was not properly serialized when passed from workflow code into a step function because the Error reducer checked `value instanceof global.Error` where `global` is the VM's globalThis. Errors created in the host context (like FatalError from @workflow/errors) have a different Error prototype than the VM context, so the instanceof check returned false and the error was silently dropped. Replaced with `types.isNativeError()` from `node:util` which uses V8's internal type tag and works across VM context boundaries. * ci: don't cancel in-progress CI runs on main branch |
||
|
|
0edcccf84d | [ci] Skip for CI checks when doing vercel backend testing only (#1063) | ||
|
|
86f62f2779 |
Refactor e2e tests to no longer use "trigger" endpoint (#958)
## Summary
Refactors the E2E tests to call `start()` from `workflow/api` directly instead of going through the `/api/trigger` HTTP endpoint in each workbench app. This removes a layer of indirection — the tests now use the same API that users would use to start workflows programmatically.
### Before
```ts
const run = await triggerWorkflow('addTenWorkflow', [123]);
const returnValue = await getWorkflowReturnValue(run.runId);
```
- `triggerWorkflow()` sent an HTTP POST to `/api/trigger` on the workbench app
- The workbench app looked up the workflow function, called `start()`, and returned the run ID
- `getWorkflowReturnValue()` polled `GET /api/trigger?runId=...` until the workflow completed
### After
```ts
const run = await start(await e2e('addTenWorkflow'), [123]);
const returnValue = await run.returnValue;
```
- `e2e()` / `getWorkflowMetadata()` fetches the manifest from `/.well-known/workflow/v1/manifest.json` to look up the correct `workflowId`
- `start()` is called directly from the test process via the configured World
- `run.returnValue` polls for completion via the World (no HTTP polling endpoint needed)
### Changes
**`packages/core/e2e/e2e.test.ts`**
- Removed `triggerWorkflow()` and `getWorkflowReturnValue()` helpers
- Added `fetchManifest()` to fetch and cache the workflow manifest from the deployment
- Added `getWorkflowMetadata(file, fn)` to look up `{ workflowId }` from the manifest
- Added `e2e(fn)` shorthand for the common case of `workflows/99_e2e.ts`
- All tests call `start()` and `run.returnValue` directly
- Error tests use `.catch()` to inspect `WorkflowRunFailedError`
- Output stream tests use `run.getReadable()` directly (skipped on local world where cross-process streaming isn't supported)
- `beforeAll` configures the local World with the correct data directory and base URL
- Pages Router tests use `startWorkflowViaHttp()` to specifically validate the HTTP trigger path
**Workbench apps (hono, express, fastify, nest)**
- Removed `/api/trigger` route handlers
- Kept `/api/hook`, `/api/test-direct-step-call`, `/api/test-health-check` endpoints
- Re-added `_workflows.js` side-effect import for hono/express/fastify to maintain Nitro's HMR dependency graph
**Deleted trigger-only route files** from: nextjs-turbopack, nextjs-webpack, vite, sveltekit, astro, nuxt, nitro-v2, nitro-v3, example
**`.github/workflows/tests.yml`**
- Added `WORKFLOW_PUBLIC_MANIFEST: '1'` to all E2E test jobs
### Dependencies
Stacked on #963 which adds `WORKFLOW_PUBLIC_MANIFEST` support to all framework builders.
|
||
|
|
6208616d06 |
Add sequential step benchmarks and fix broken benchmark infrastructure (#845)
* Add 50, 100, 500 concurrent step benchmarks Enable Promise.all and Promise.race benchmarks for 50, 100, and 500 concurrent steps (previously 100+ were skipped). Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * Add 50, 100, 500 sequential step benchmarks Extends the sequential step benchmarks to test workflows with 50, 100, and 500 sequential steps in addition to the existing 10 step test. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * Add full/quick benchmark suite toggle for CI - Add BENCHMARK_FULL_SUITE env var to control which benchmarks run - Quick suite (default for PRs): 10, 25, 50 step benchmarks - Full suite (main branch, manual dispatch): adds 100, 500 step benchmarks - Add workflow_dispatch input to manually trigger full suite from GitHub UI - Skip 100+ sequential and concurrent step benchmarks by default This keeps PR benchmarks fast while allowing full stress testing on demand. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * Add 25 sequential steps benchmark Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * Fix benchmark API calls to use binary format PR #853 changed the workflow trigger API to expect binary data (application/octet-stream) instead of JSON. The e2e tests were updated but the benchmark file was missed, causing all benchmarks to fail silently since Jan 28. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * Only run full benchmark suite on manual dispatch Remove automatic full suite on main branch pushes - only run full suite when manually triggered with full_suite=true. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> |
||
|
|
8ff4b3a508 |
Enable Vercel Deployment Protection bypass secret for E2E tests (#580)
This allows us to enable the Deployment Protection feature for the Vercel test projects. |
||
|
|
ee4fff6814 |
benchmarking: make steps simulate real work (+ misc improvements) (#565)
* perf: add 5s delay to benchmark steps to simulate real work * perf: add realistic workloads to benchmark steps Add realistic workloads to benchmark step functions: - doWork() - 1 second delay to simulate real computation - stressTestStep() - 1 second delay to simulate real computation - genBenchStream() - generates ~5KB of data in 50 chunks - transformStream() - uppercases stream content (renamed from doubleNumbers) Add "Slurp Time" metric to stream benchmarks measuring time from first byte to complete stream consumption, complementing TTFB. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * fix(e2e): increase dev test timeouts for Windows Windows file watching and rebuilding is slower than macOS/Linux, causing the 10s timeouts to fail. Increased to 30s to accommodate. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * fix(ci): treat community world failures as warnings, not errors Community world tests/benchmarks are now non-blocking: - Removed from has_failures check in both tests.yml and benchmarks.yml - Added separate has_warnings output for community failures - Split PR comment notices: ❌ for failures, ⚠️ for community warnings 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: apply PR review suggestions - Fix chunk size calculation: chunkSize - 11 for ~100 bytes per chunk - Remove redundant metadata check (always truthy) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
54ba1888cd |
fix: compare benchmarks against PR base branch instead of main (#560)
* fix: compare benchmarks against PR base branch instead of main - Use github.event.pull_request.base.ref instead of hardcoded main - Remove search_artifacts: true to ensure most recent baseline is used - For stacked PRs, this compares against the parent PR's baseline 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: group failed e2e tests by category and app in summary Instead of listing each failed test as a separate item, group them by: 1. Category (world): e.g., "Community Worlds", "Vercel Production" 2. App (framework): e.g., "mongodb", "turso", "nextjs-turbopack" This makes the summary much more readable when there are many failures. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: ensure local E2E tests always produce JSON output - Add 'fastify' to app detection list in aggregate-e2e-results.js - Change && to ; so e2e tests run even if dev.test.ts fails - This ensures local-dev, local-prod, and local-postgres categories appear in the E2E summary comment 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: ensure local E2E tests always produce JSON output - Add 'fastify' to app detection list in aggregate-e2e-results.js - Change && to ; so e2e tests run even if dev.test.ts fails - This ensures local-dev, local-prod, and local-postgres categories appear in the E2E summary comment 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * feat: publish CI results to GitHub Pages for docs - Add generate-docs-data.js script to create JSON summaries from CI artifacts - Add publish-results job to tests.yml and benchmarks.yml workflows - Update docs/lib/worlds-data.ts to fetch from GitHub Pages URLs - Results published to https://vercel.github.io/workflow/ci/ This allows the docs worlds page to display actual test/benchmark results without requiring a GITHUB_TOKEN. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: correct outputFile path for local E2E test artifacts The --outputFile path was using ../../ which placed files outside the repo because pnpm run test:e2e executes from workspace root, not from the cd'd workbench directory. This prevented local-dev, local-prod, and local-postgres test results from being uploaded as artifacts. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: show green checkmark for skipped tests instead of warning Skipped tests are intentional and shouldn't show as warnings in the E2E test summary comments. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * feat: use collapsible sections in benchmark PR comment Wrap each benchmark, stream benchmarks section, and summary tables in <details> toggles to make the PR comment more compact and readable. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * feat: add Vercel observability links to benchmark PR comments - Store runId in benchmark timing data - Add project-slug to Vercel benchmark matrix - Pass WORKFLOW_VERCEL_PROJECT_SLUG env var to benchmarks - Store Vercel metadata (teamSlug, projectSlug, environment) in timing files - Generate observability deep links for each Vercel world benchmark - Show observability links below Production (Vercel) tables 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * fix: use correct Vercel project slugs for observability links - nextjs-turbopack → example-nextjs-workflow-turbopack - nitro-v3 → workbench-nitro-workflow 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
5dd15452cd |
Test and benchmark community worlds against e2e tests (#482)
* Add new github workflow * Enable pull_request trigger for community worlds workflow 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Add community worlds manifest and generation scripts - Add community-worlds.json manifest as single source of truth - Add scripts/generate-community-worlds-workflow.mjs to generate CI workflow - Add scripts/generate-community-worlds-docs.mjs to generate docs section - Update aggregate-benchmarks.js to load community worlds dynamically - Add pnpm generate:community-worlds script - Update docs/deploying/world/index.mdx with community worlds The manifest-based approach allows: - E2E tests to be auto-generated from the manifest - Benchmark aggregation to include community worlds - Docs to stay in sync with tested worlds 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Fix YAML syntax error - quote strings starting with @ The @ symbol has special meaning in YAML, so package names like @workflow-worlds/turso need to be quoted. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Fix Redis health-cmd quoting, remove unpublished starter world - Quote health-cmd when it contains spaces (fixes Docker arg parsing) - Remove @workflow-worlds/starter as it's not published to npm 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Add benchmarks and summary job for community worlds - Add build job to share artifacts between benchmark jobs - Add benchmark jobs for Turso, MongoDB, and Redis worlds - Update summary job to show both E2E and benchmark status matrix - Add left border/indent to sidebar child items for visual hierarchy - Update workflow generator to support benchmark generation 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Reuse build artifacts for E2E tests E2E jobs now depend on the shared build job and download artifacts instead of rebuilding packages from scratch. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Add Worlds Ecosystem dashboard to docs - Create worlds-manifest.json with official and community worlds - Add aggregate-worlds-data.mjs script for processing E2E and benchmark results - Create WorldsDashboard, WorldCard, and BenchmarkChart components - Add /docs/worlds page showing compatibility status and performance - Include sample data for development The dashboard shows: - E2E test progress per world (pass/fail/skip counts) - Benchmark performance comparison across all worlds - Filter by official vs community worlds 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Fix CI to output JSON test results and add Jazz world - Update workflow generator to output JSON test results from vitest - Upload E2E results as artifacts for parsing in summary job - Summary job now shows actual pass/fail/skip counts per world - Add Jazz world to worlds-manifest.json (requires external credentials) - Add update-worlds-status.yml workflow to auto-update dashboard data - Update TypeScript types to support null lastRun and metrics 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Refactor community worlds to use reusable workflows Instead of a generated workflow file, integrate community world testing directly into tests.yml and benchmarks.yml using reusable workflows. - Add reusable workflows for E2E tests: e2e-community-world.yml (no services), e2e-community-world-mongodb.yml, e2e-community-world-redis.yml - Add reusable workflows for benchmarks: benchmark-community-world.yml, benchmark-community-world-mongodb.yml, benchmark-community-world-redis.yml - Update tests.yml to call reusable workflows for Turso, MongoDB, Redis - Update benchmarks.yml to include community world benchmarks in summary - Delete generated community-worlds.yml and generator script This approach: - Inherits proper Rust/SWC setup from the main workflows - Keeps all CI in the established patterns - Makes adding new community worlds straightforward 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Fix benchmark timing file naming for community worlds Add WORKFLOW_BENCH_BACKEND env var support to bench.bench.ts so community world benchmarks generate timing files with the correct backend suffix (e.g., bench-timings-nextjs-turbopack-turso.json instead of -local.json). 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Add @workflow-worlds/starter to community worlds test matrix 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Unify worlds manifest and add dynamic GitHub API fetching - Merge community-worlds.json into worlds-manifest.json with type field - Add server-side data fetching from GitHub API for worlds dashboard - Remove static worlds-status.json, fetch CI artifacts dynamically - Update all references to use unified manifest format - Remove obsolete community-worlds.yml and update-worlds-status.yml workflows 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Skip community worlds for non-nextjs-turbopack in benchmark summary Community worlds only run against nextjs-turbopack, so hide the "missing" rows for Express and Nitro frameworks in the benchmark comparison tables. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Fix stream benchmark detection to check for actual TTFB data The previous check `!== null` incorrectly returned true for undefined, causing all benchmarks to show TTFB columns. Now explicitly checks for a number type. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Hide Worlds Ecosystem page from sidebar The page is still accessible via direct link at /docs/worlds but won't appear in the navigation until it's been further iterated on. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Fix review comments: trailing newline and division by zero - Add trailing newline when replacing Community Worlds section in docs - Fix division by zero in WorldCard benchmark calculation when metrics is empty 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Consolidate community world workflows with service-type parameter - Create setup-workflow-dev composite action for common setup steps - Add service-type input to benchmark-community-world.yml and e2e-community-world.yml - Use conditional job execution (if: inputs.service-type == 'mongodb') to handle different services - Update benchmarks.yml and tests.yml to pass service-type parameter - Delete redundant workflow files: - benchmark-community-world-mongodb.yml - benchmark-community-world-redis.yml - e2e-community-world-mongodb.yml - e2e-community-world-redis.yml Reduces workflow files from 11 to 7 and eliminates ~500 lines of duplicated YAML. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Generate community world test matrix from worlds-manifest.json Replace hardcoded community world jobs with dynamic matrix generation using scripts/create-community-worlds-matrix.mjs. This allows adding/removing community worlds by editing the manifest instead of multiple workflow files. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Apply suggestion from @vercel[bot] Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com> * Add Samples column and separate local/production benchmarks - Add Samples column to all benchmark tables showing iteration count - Separate benchmark results into Local Development and Production sections - Add explanatory context for each section (localhost vs Vercel deployment) - Add GitHub action step summaries to e2e community world tests - Create aggregate-e2e-results.js script for parsing vitest JSON output 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Remove obsolete generate-community-worlds-docs script The Worlds Ecosystem page now fetches from worlds-manifest.json at runtime, making this script unnecessary. The npm script also referenced a non-existent workflow generator script. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Maximize composite action usage and reorganize benchmark output - Update setup-workflow-dev composite action with optional Rust, install-dependencies, and install-args inputs - Update tests.yml to use composite action in unit, e2e-vercel-prod, getTestMatrix, e2e-local-*, and getCommunityWorldsMatrix jobs - Update benchmarks.yml to use composite action in build, benchmark-local, benchmark-postgres, benchmark-vercel, and getCommunityWorldsMatrix jobs - Reorganize benchmark output to group by benchmark test with local/production tables within each benchmark - Remove invalid $schema reference from worlds-manifest.json 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Add beads stealth mode stuff (for personal claude memory - will remove stealth if people want) * Add E2E test results PR comment summary - Add pr-comment-start job to create/update PR comment when tests start - Add artifact uploads to all e2e test jobs (vercel-prod, local-dev, local-prod, local-postgres, windows) - Update e2e-community-world.yml with consistent artifact naming (e2e-community-*) - Add summary job to aggregate all e2e results and update PR comment - Extend aggregate-e2e-results.js with --mode aggregate for multi-job PR summary - Group results by category (Vercel Production, Local Development, etc.) 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Add step summaries to all e2e test jobs Add "Generate E2E summary" step to each individual e2e job: - e2e-vercel-prod - e2e-local-dev - e2e-local-prod - e2e-local-postgres - e2e-windows Each job now outputs pass/fail/skip counts to GITHUB_STEP_SUMMARY. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Consolidate community world workflows to single job Replace 3 mutually exclusive jobs (e2e/e2e-mongodb/e2e-redis) with a single job that starts services via docker run when needed. This eliminates the skipped job entries that appear in the GitHub Actions UI. - Use conditional docker run steps instead of services: block - Add health check loops to wait for service readiness - Add cleanup step to stop containers 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> * Fix extractWorldId to handle community world artifact naming Add handling for `e2e-results-community-{world}` pattern so community world test results are properly extracted (e.g., `e2e-results-community-turso` now correctly extracts `turso` instead of `community-turso`). 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com> |
||
|
|
6e8e828252 |
Add stream benchmarks and slightly improve the benchmarking code (#470)
* Add stream benchmarks and cnealup for benchmark code * 10m delay for local world race * potential fix for first bute" * Improve sticky comment stuff * local world: ignore controller close errors * show missing data in benchmark comment * nitrpicks * Improve leaderboard * Add benchmark comparisons against main * changeset |
||
|
|
a8f48c5a08 | add benchmarking (#460) |