mirror of
https://github.com/vercel/workflow.git
synced 2026-09-14 19:59:43 +08:00
main
427 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fb9e27589d |
perf(core): commit pre-claimed inline pairs in their own batch chunk (#4098)
* perf(core): commit pre-claimed inline pairs in their own batch chunk The batched fan-out fold sorted the pre-claimed inline [step_created, step_started] pairs to the front and then filled the same chunk with plain step/wait creates up to MAX_BATCH_FANOUT_EVENTS. Only that chunk gates the inline bodies, so the bodies waited on a 32-row commit (a 61-item DynamoDB transaction on the Vercel backend) when all they needed was the pairs' own rows. Chunk the pair rows and the plain creates separately: the pairs get their own leading chunk(s), still adjacent and never split across a boundary, and the plain creates fill the subsequent chunks of 32, committing concurrently beside the pair chunk and gating only their own queue publishes. Eligibility, the batch-of-one rule, pairCommits gating, per-chunk publishes and the foreign-interleaving diagnostic are unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * perf(core): fold pairs only with two inline steps; guard a lone plain create With the pairs in a chunk of their own, plain creates alongside them share no round trip with the pairs, so "company" no longer makes a lone pair worth folding. `inlinePairFoldEligible` now requires two or more inline steps: a lone inline step keeps the lazy `step_started` (one row, optimistic-start capable, bump-and-report) while the eager creates beside it still batch. A plain partition of exactly one entry is the batch-of-one case again, so it takes the guarded single create (slot-snapshot params, bump-and-report, same conflict tolerance and `createdStepCorrelationIds` bookkeeping) instead of a one-row createBatch, and its queue publish still waits for that create. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
c29200fac5 |
docs(ai): clean up WorkflowAgent docs and examples (#3891)
Co-authored-by: Peter Wielander <peter.wielander@vercel.com> |
||
|
|
357aa7c38a | [core] QuickJS engine reports its log position and consumes returned event pages; document who names a position (#4106) | ||
|
|
e00b1a57ee |
perf(world-vercel): batch a fan-out's step-execution queue publishes (#3838)
* perf(world-vercel): batch a fan-out's step-execution queue publishes A `Promise.all` fan-out dispatched one queue message per branch. Those publishes ride the shared default undici agent (8 connections, HTTP/1.1, `pipelining: 1` — see `getQueueDispatcher`), and `handleSuspension` is awaited in full before the first inline step body runs, so an N-branch fan-out paid ~N/8 serialized round trips straight onto time-to-first-step. The `step_created` writes were already batched and HTTP/2-multiplexed; the publishes were the remaining per-branch round trip. Adds an optional `Queue.queueBatch`, implemented on `@vercel/queue`'s `experimental_sendBatch` (0.5.1), and uses it for the batched fan-out fold's publishes. Each commit chunk now publishes in one request instead of up to 32. `queueBatch` reports per-entry outcomes rather than throwing, because a batch can partially fail. `queueMessages` in core keeps the previous all-or-nothing behavior for this call site: it rejects if any entry failed, so the delivery is redelivered and republishes the set, deduped by the per-step `idempotencyKey` the caller already passed. Worlds without `queueBatch` fall back to concurrent single sends. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(core): reject a short queueBatch result set instead of reading it as success `queueMessages` only inspected `error`, so a World whose `queueBatch` returned fewer results than it was given messages reported success for the whole batch. The omitted entries were never published and nothing raised: `handleSuspension` resolved, the delivery was acked, and those steps were never dispatched, so the run stalls with no error recorded anywhere. Reproduced at 64 branches against a World returning half its results: 32 of 63 steps silently lost. world-vercel guards this internally and `@vercel/queue` length-checks its own response, so it was not reachable through the world added here. It is reachable through the interface `building-a-world` opens to third-party worlds, which is where the check belongs. Documented on the interface and in the guide alongside it. Also notes that the batch grouping degenerates to one request per message under WORKFLOW_SEQUENTIAL_REPLAYS=1 (per-step physical topics are one of the routing dimensions groups split on), and corrects the comment claiming the error's `retryable` flag is consumed downstream: nothing reads it yet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(world-vercel): carry trace context on each batched queue message `experimental_sendBatch` injects the active trace context into the multipart REQUEST headers, and the per-part headers it builds never see it. VQS stores headers per message and re-emits a stored `traceparent` at delivery as `x-vercel-queue-traceparent`, which is what lets a consumer attach a span link back to its producer, so a batched message arrived with no producer context and its `vqs.process` span got no link. `send()` is unaffected: for a single message the request headers ARE that message's headers. At 64 branches that was 63 of 64 step dispatches losing the transport-level producer link. The run's own step tracing was never affected: that carrier travels in the message payload (`WorkflowInvokePayload.traceCarrier`), which is what the consumer builds its trace context from, not a header. Injects the active context into each entry's headers in `queueBatch` — last, so it wins over caller-supplied `opts.headers` exactly as the SDK's own injection does — and honors VERCEL_QUEUE_TRACE_PROPAGATION so that kill switch still covers both paths. `getTraceContextHeaders()` is factored out of `injectTraceContextIntoHeaders` so the two share one source. Verified on the wire against a stub VQS speaking the real batch endpoint: `traceparent` carrying the producer's traceId/spanId lands on all 64 multipart parts through the real SDK, with the per-message idempotency keys still alongside it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Karthik Kalyanaraman <karthik.kalyanaraman@vercel.com> |
||
|
|
0392d69fcf |
workflow docs: add a bunch of missing material (#4068)
* cancellable steps and use of asyncio.timeout * typed streams * hook return values * share_sandboxes * deterministic helpers |
||
|
|
c09c1bb6ea |
fix(core): drain step stream writes before completion (#3941)
## Summary & Motivation ### Situation - Workflow stream writers are expected to call `releaseLock()` when a step finishes writing so another step can acquire the stream. - The runtime observes that release and drains the server sink, but `step_completed` currently races the overall stream operation against 500ms. - A slow PUT can therefore continue under `waitUntil` while the next step starts and reads a stale tail. - Release can also happen while native `writer.write()` promises remain unsettled, leaving frames upstream of the server sink when a naive drain runs. - Writers intentionally kept locked must remain non-blocking so producer and consumer steps can overlap. ### Fix - Treat a writer released before step return as an implicit durable handoff boundary. - At step end, acquire the unlocked stream with a temporary writer and enqueue an internal checkpoint behind all writes queued by the released writer. - Once the checkpoint crosses serialization, wait for those frames to reach the server sink and drain the group-commit PUT before `step_completed`. - If the writer remains locked, do not wait for durability; preserve the existing 500ms inline-loop heuristic and background `waitUntil` lifecycle. - Drain failures or the 30-second safety timeout fail/retry the step. Client disconnect errors remain non-fatal. ## Test Plan - Covers released and held locks, release with unsettled writes, delayed first writer acquisition, forwarded writable arguments, drain timeout/failure, and multiple streams. - `pnpm --filter @workflow/core build` - `pnpm --filter @workflow/core typecheck` - `pnpm --filter @workflow/core test` — 2,405 passed, 3 expected failures, 1 skipped |
||
|
|
ec57aff3be | [core] Log pending consumers in divergence diagnostics (#4021) | ||
|
|
45a3072948 | [core] Fix the python e2e conformance suite after the retention merge (#4022) | ||
|
|
d4817ce548 |
feat(streams): add WebSocket writer lifecycle (#3833)
## Summary & Motivation Implements the client half of `workflow-stream-ws/v1` behind the existing default-off `WORKFLOW_STREAMS_TRANSPORT=ws` gate, populating the `createWriteSession` seam only when opted in. - Writes and closes are serialized over one socket per writer lifetime; groups above the v1 per-request chunk cap are split without resetting writer-local sequence. - Any failure before the upgrade is accepted (declined upgrade, proxy, load error, a dispatch that beats the handshake) falls back to HTTP for the rest of the writer's life. - Once a write is on the socket, a missing or uncorrelatable reply poisons the session rather than replaying over HTTP, since a duplicate append cannot be ruled out. - Idle clean closes reconnect with the same writer id, capped at three attempts so a draining server can't hot-loop. - The handshake gets a `workflow.stream.ws.connect` span and each frame a synthesized `http POST` span, so per-event tracing survives the non-`fetch` transport. ## Test Plan New unit tests cover the lifecycle, fallback, and poisoning paths; 598 `@workflow/world-vercel` tests plus package build/typecheck and workspace lint/format pass. Root build/typecheck couldn't run locally (missing Rust toolchain for the unrelated `@workflow/swc-plugin`). |
||
|
|
efbdc213a0 |
[core] Make hook.metadata a lazy Promise getter (#3988)
* [core] Make `hook.metadata` a lazy Promise getter Hydrating a hook's metadata is a decrypting READ: it needs the owning run's payload keys, and resolving those costs a run fetch plus a `run-key` API round trip (~350ms). `getHookByToken()` did that work eagerly on every lookup that found a metadata-bearing hook, so callers that only wanted `runId`/`token` — and hook resumption, which never reads metadata at all — paid for it anyway. `metadata` is now a getter returning a memoized Promise, the same shape as `run.returnValue`. The lookup is one read again; hydration and the key resolution behind it happen on first access, or never. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com> * docs: surface the lazy hook.metadata change in What's new, the migration skill, and the resumeHook reference Adds the breaking-change row to the v5 What's new page and puts that page in the sidebar as the first visible entry (the /v5/docs redirect to getting-started is unchanged). Teaches the migrating-workflow-v4-to-v5 skill the `await hook.metadata` rewrite and bumps its version. Points the resumeHook reference at HookWithLazyMetadata, and notes on the World storage page that world.hooks.getByToken() returns raw serialized metadata. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * [core] Export the lazy-metadata hook type as `Hook` from `workflow/api` `getHookByToken()` and `resumeHook()` return `Hook`, not a separate `HookWithLazyMetadata`: one public hook type whose `metadata` is a lazy Promise, mirroring `Run` for runs. The World-level record from `@workflow/world` is unchanged and is referenced as `WorldHook` inside the runtime. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * [core] Define the lazy `metadata` getter in place; tighten changeset and docs wording Review feedback: the hook record a World returns is a fresh object per lookup and the eager path mutated it anyway, so define the getter on it directly instead of copying it with Object.create(). The changeset is one sentence, and the docs describe hydration as extra network round trips rather than decryption, since not every World encrypts. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com> Co-authored-by: Pranay Prakash <1797812+pranaygp@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
2354301f39 | [world-vercel] Bound the queue client's requests (#4049) | ||
|
|
51a181af91 | docs: document WORKFLOW_NODE_HTTP in the v4 World docs (#4050) | ||
|
|
f83e8367f4 |
[world-vercel] Honor WORKFLOW_NODE_HTTP on the queue transport (#4044)
* fix(world-vercel): honor WORKFLOW_NODE_HTTP on the queue transport getQueueDispatcher was the one dispatcher getter that ignored the flag. The reasoning was that `undefined` cannot move the queue client onto node:http (QueueClient takes a dispatcher and no fetch override), so returning it would only drop this package's pool tuning and fall back to undici's global agent. That misses what the flag is actually for. The deployments that need it are the ones where the undici copy *this package bundles* is unusable, and `undefined` does move the request off that copy: global fetch dispatches on the runtime's own undici instead. On such a deployment every other request survives while the queue client keeps dispatching through the broken copy, and an acknowledgeMessage that never resolves means the message is redelivered for as long as the platform keeps killing the invocation holding it. Losing pool tuning is the correct trade under a flag whose premise is that the bundled undici is not usable here. An explicit config.dispatcher still wins. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: correct the queue-send note for WORKFLOW_NODE_HTTP The queue client is a partial exception to the flag, not a full one: it cannot move to node:http, but it does honor the flag by dispatching through the runtime's own copy of the HTTP client library instead of the copy the World bundles. That distinction is the whole point when the bundled copy is what does not work. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
fdeb642270 |
feat(streams): add WebSocket capability gate (#3764)
## Summary & Motivation `WORKFLOW_STREAMS_TRANSPORT=ws` advertises client support for `workflow-stream-ws/v1` on stream writes; anything else keeps HTTP. It's a capability signal only — the server decides each upgrade, so there's no version or tenant-policy heuristic on the client side. ## Test Plan Unit tests for the gate's accepted values; typecheck, build, and lint pass locally. |
||
|
|
c340820411 |
Add Run#getWritable() for appending to another run's stream (#3972)
## Summary & Motivation Lets a long-lived run own a shared stream while independent runs append through `run.writable` or `run.getWritable()` using only the owner's run ID. The handle seals to the owner's public key when it has one, so it grants append access without read capability, and carries the existing forwarding symbols so passing it through `start()` and into a step keeps the owner's identity. ## Test Plan Tests added |
||
|
|
61fb1f93bd | [core] Add a retention option to start() (#3787) | ||
|
|
c129332923 |
[docs] Fix prose typos across v4/v5 docs and the SWC plugin README (#3948)
Copy-edit only, no behavior or API changes: - "it's a only a short step" -> "it's only a short step" (ai/index) - "When you tool needs" -> "When your tool needs" (ai/defining-tools) - "Workflow operation that suspend" -> "operations that suspend" (ai/sleep-and-delays) - "extend out ... to use emit" -> "extend our ... to emit", and "other tools calls ... inject out own" -> "other tool calls ... inject our own" (ai/streaming-updates-from-tools) - "non-yet-standard" -> "not-yet-standard" (how-it-works/understanding-directives) - "apps ... and needs no special configuration" -> "and need no special configuration" on the five getting-started pages that disagreed with the other five (express, fastify, hono, nuxt, vite) - drop the orphan "needed." line after "No separate command is required." and fix "the local installed version" -> "the locally installed version" (observability/index) - "Time between emissions of a chunk" -> "emission of a chunk" (observability/tracing) - "determine that is safe" -> "determine that it is safe" (whats-new) - "three rules bind an implementation" -> "four rules": the list has four bullets, and skills/migrating-world-v4-to-v5 already says four (worlds/upgrading-to-v5) - drop the stray duplicate "Workflow" before the Workflow SDK link in @workflow/swc-plugin's README, and the duplicated horizontal rule before "## Detect mode" in its spec Each fix is applied to both the v4 and v5 copies wherever the same text exists in both. Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com> |
||
|
|
f9073d0739 |
Add attribute inspection to the CLI (#3950)
* Add attribute inspection to the CLI and probe the cancel window once
`wf inspect attributes` lists the distinct attribute keys on a project's
runs with their run counts and first/last seen times, and
`wf inspect runs --attribute key=value` filters by them. Between them
they turn attributes from something you can only write into something
you can discover and query. Both are analytics-only — storage has no
cross-run attribute index — so the listing says so rather than falling
back, and the filter warns and is ignored the way --since/--until
already do.
The flag is parsed and bounded in lib/inspect so the error names
--attribute rather than the parameter it becomes, and so it is testable
next to the other inspect flag helpers. It splits on the first `=` only,
since a value may contain one, and keeps an empty value, which matches
runs whose attribute was set to the empty string.
`wf cancel` also probed the plan's listing window inside its per-status
fan-out, so a four-status cancel issued four identical probes. The
window is a property of the plan rather than of a status, so the probe
is hoisted above the fan-out: eight requests become five. The harness
only ever modelled the storage path, so that probe logic had no
coverage; the new test fails with two probes before the change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Do not depend on an unreleased world export for the flag cap
The --attribute cap was imported from @workflow/world, where the
constant is added by a different branch, so on main it resolved to
undefined and `values.length > undefined` was always false: the flag
accepted any number of pairs and the test for it never threw.
Declare the cap in the CLI instead. The World and the backend enforce
the same bound independently, and this copy exists only so the error can
name the flag the user typed rather than the parameter it becomes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Degrade sleeps to the event log and bound the inspect flags
`wf inspect sleeps` was the only list path that could not degrade: it
branched on analytics being present and either returned or exited, so on
any backend providing analytics the storage branch below it was
unreachable and an analytics failure ended the command. It now warns and
falls through, like the run, step, and event listings. An argument the
World rejected is not retried — the same argument fails either path, so
falling back would trade a precise message for a slower failure.
handleApiError also only recognised errors carrying an HTTP status.
A client-side argument rejection has none, because no request was made,
so it fell past every branch and was rethrown as an unhandled error. It
is now reported as given: the message already names the method, the
parameter, and what it received.
--limit and --runId are checked before any backend setup so a mistyped
value names the flag and costs no round trip. The limit bound is
deliberately looser than the per-endpoint caps, which differ by resource
and stay with the World; this one catches a typo'd digit or a negative.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Scope --attribute to inspect and document the inspect flags
--attribute was added to the shared cliFlags, which cancel, health,
start, and web all spread — so `workflow health --attribute k=v` parsed
and was silently ignored. It belongs with the other inspect-only
filters in the command's own flags, next to --runId and --since.
The configuration reference documented every shared flag but none of
the inspect-only ones, so --runId, --stepId, --hookId, --since/--until,
--withData and --decrypt had no entries at all. They now do, in an
Inspect filtering section, alongside --attribute. --status and
--workflowName were documented under bulk cancel only; both also filter
inspect listings, which is now noted where they are.
--limit's entry described a default with no bound and is now rejected
outside 1 to 1000, so it says so, and points out that individual
listings cap lower.
The attributes guide claimed filtering was available "through the
Analytics API", which is no longer the whole story: the CLI can now
discover keys and filter by them, so that section splits into a CLI half
and an API half.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Stop dropping inspect flags silently
Three flags the caller typed were being discarded without saying so —
the same failure the World argument guards were added to remove,
reintroduced one layer up.
--attribute and --since/--until warned that the backend has no
analytics read path, but that condition is also false when --withData
asks for payloads, which only storage carries. Blaming the backend for
the caller's own flag sends them looking in the wrong place, so the
warning now names whichever applies.
inspect attributes dropped --sort entirely, explained only by a code
comment. It is forwarded now, and still left unset when absent so the
backend's alphabetical key order stands rather than the `desc` the
time-ordered listings impose.
A repeated --attribute key silently kept the last value, and a test
asserted that as if it were intended. Matching is per-key, so resolving
it means discarding a filter the caller typed: it is rejected instead.
The shared --limit entry also stated the 1-to-1000 bound that only
inspect enforces, which is wrong for cancel's own 1-to-500. The bound
moves to an inspect entry and the shared one points at both.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Reject --attribute on listings that cannot use it
Only the runs listing filters by attributes, but the flag was parsed for
every inspect resource: `inspect steps --attribute tenant=acme` returned
a normal, unfiltered step list with no warning, as did events, hooks,
attributes, and `inspect run <id>`, which already names one run. That is
the silent drop the preceding commit set out to remove, missed one layer
up in the command itself.
Validated alongside the other flag bounds, before any backend setup, so
a flag on the wrong subcommand costs no round trip.
Covered at the command level as well as in the unit, since the defect
was not in the validator but in nothing calling it: the tests drive
`Inspect.run` with a mocked setup module and assert the backend is never
reached. Five of them fail without this change.
Reported in review; verified against a real project rather than found
by the suite, which is why the command-level coverage goes in with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Resolve the test's oclif root without a URL pathname
`new URL('../..', import.meta.url).pathname` yields `/D:/a/...` on
Windows — a leading slash before the drive letter — so `Config.load`
could not find package.json and every command-level test failed there
while passing on Linux. `fileURLToPath` handles both.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Address review on the attribute flag
Attribute keys naming an Object.prototype member were rejected as
duplicates before anything was stored, because the duplicate check used
`in`, which walks the prototype. `--attribute toString=v` failed on
first sight, and `__proto__=v` would have set the prototype rather than
stored a value had it got that far. The map is null-prototype now and
the check uses Object.hasOwn.
--url and --web return before the filter is parsed, and neither
forwards it, so `inspect runs --attribute k=v --url` opened an
unfiltered view and a malformed pair skipped validation entirely. Both
are rejected: the dashboard takes no attribute filter.
--sort carried an oclif default of desc, so the "forward only when
asked" check in the attribute listing was always true and overrode the
backend's alphabetical key order. Every time-ordered listing already
falls back to desc itself, so the parser-level default is gone and the
flag now means what it says.
The docs claimed --since and --until must be supplied together, but the
CLI resolves the pair before the World sees it: --since alone is valid
and --until defaults to now. Only --until alone is rejected.
The vercel[bot] comment about ANALYTICS_MAX_ATTRIBUTE_FILTERS not being
exported was already addressed in
|
||
|
|
7cc5c88a8b |
[core] Settle a hook's awaiter in-process instead of re-invoking, on creation and on conflict (#3938)
* [core] Settle a hook's awaiter in-process instead of re-invoking, on creation and on conflict * [core] Address review: deterministic hook signal tests, split changesets, document the boundary - hook.test.ts: drive the idle poll with explicit macrotask turns instead of a fixed 20ms sleep (Copilot) - Split the changeset so each package's entry says only what changed in it - runtime-tuning docs: hook-only suspensions no longer always park; the hook write continuation is the one exception Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
fbfb9fe869 |
Validate world.analytics arguments up front (#3943)
* Validate world.analytics arguments up front
Every analytics method now checks its arguments before making a request
and throws a RangeError naming the limit it broke: the ids, the
pagination limit against the cap for that listing, and the attribute
filter's pair count, key length and value size. Because analytics is an
optional capability, callers wrap it in a catch, which turned an invalid
argument into what looked like an empty result rather than an error.
Two arguments that used to be dropped silently now fail too. A limit of 0
fell back to the default page size, and a startTime without a matching
endTime turned a listing you meant to bound into an unbounded one that
looked like a normal answer.
Exports ANALYTICS_RUN_SCOPED_PAGE_LIMIT, ANALYTICS_PAGE_LIMIT and
ANALYTICS_MAX_ATTRIBUTE_FILTERS so callers can check the bounds
themselves.
Deprecates analytics.events.listByCorrelationId() in favour of
analytics.events.list({ runId, correlationId }), which issues the same
request and also accepts an eventType filter. It keeps its own
implementation rather than delegating: list() treats correlationId as
optional and skips an empty one, so a delegation would turn an empty id
into an unfiltered listing of the run.
Documents every analytics method in the reference. events.getMany() was
missing entirely, seven methods shared one code block with no parameters
or return shapes, and none of the limits were written down.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Parse attribute key timestamps as UTC
firstSeenAt and lastSeenAt were the only analytics timestamps still on a
plain date coercion. The values arrive without a timezone designator, so
that read them in the process's local zone and every other field in the
namespace read them as UTC — a seven-hour skew on those two fields alone
for anyone running outside UTC.
The added test fails without the fix under TZ=America/Los_Angeles.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs: drop the deprecated correlation-id listing from the reference
A reference page should describe the API you should reach for. The
deprecation notice lives on the method itself, so editors surface it
where it matters without the page advertising a method nobody should
start using. Also drops it from the page-limit table.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Close two gaps in the analytics argument guards
Run ids were validated with workflowRunIdSchema while every other id used
a regex mirroring the backend. Those disagree: z.ulid() accepts a
lowercase body and a first character above 7, and the backend accepts
neither, so the most-used parameter had the leakiest guard and still
produced the 400 this is meant to prevent. Run ids now use the same
pattern as the rest.
A supplied-but-empty filter value was also still being dropped —
correlationId, the optional runId scope on hooks.get, and workflowName
all tested truthiness. Dropping one widens the result set rather than
narrowing it, so an empty correlationId listed the whole run and an
empty workflowName listed every workflow. That is the same failure the
limit and time-window guards were added to prevent, and the comment on
listByCorrelationId already described the hazard. They now compare
against undefined, so an empty id throws and an empty name is forwarded
for the backend to match.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Raise argument rejections as a typed, non-retryable error
The guards threw bare RangeErrors, which left a caller — or an agent
driving this API — parsing English to decide whether to fix the call or
retry it. They now raise WorkflowWorldError with
code: 'INVALID_ARGUMENT', the code the rest of this client already uses
for its transport and throttle failures, so the retry decision is a
field lookup. normalizeEventIds moves with them rather than staying the
one guard that throws a different type.
Also sharpens the four messages that made a caller do the work:
a half-open window now names the bound that is missing rather than
restating the rule, an inverted window prints both ends, and the
attribute-value and event-id batch errors report the size measured
rather than only the bound they broke.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Name the method and the field on an argument rejection
Two things a caller could not get without reading prose. The same guard
runs behind several methods, so `runId must be a workflow run id` was
identical whether it came from events.list or steps.get — fine with a
stack, lossy once the error has crossed a log line. And the offending
argument was only available as the first token of the message, which is
the part most likely to be reworded.
Messages now open with the method, and WorkflowWorldError carries an
optional `field`:
analytics.runs.list: pagination.limit must be an integer between 1
and 100 (received 9999)
→ code: 'INVALID_ARGUMENT', field: 'pagination.limit'
`field` is additive on the error class and set only by these guards, so
nothing that reads the existing properties changes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
aba7c8d5d0 | docs: expand Python workflow guide (#3923) | ||
|
|
3c08778905 |
[core] Retain workflow VMs across waits (#3892)
* perf(core): retain workflow VMs across waits * test(core): cover retained wait wake races Signed-off-by: Nathan Colosimo <110621881+NathanColosimo@users.noreply.github.com> --------- Signed-off-by: Nathan Colosimo <110621881+NathanColosimo@users.noreply.github.com> |
||
|
|
40bed1777a |
python docs: add some missing material (#3875)
This merges a bunch of stuff currently on the vercel.com docs. I'm going to go remove those next to centralize things for now. |
||
|
|
2668e3325b |
Durable hook resume: write, then wake (#3841)
* test(core): reproduce lazy resume disposal race * Fix durable hook resume race * Fail closed on unknown hook wakes * Improve unsupported hook wake diagnostics * Address durable hook resume review feedback * Harden producer-committed wake handling * Serialize durable hook resume: write, then wake resumeHook() now dispatches strictly serially: the hook_received event is made durable first, and the workflow wake is published only after the write is acknowledged. The wake is a plain runId message (the shape the sequential path always published), so the producer-committed wake barrier, its queue-message field, and the HOOK_RESUME_INPUT_VERSION bump are all removed — no consumer or backend coordination is needed, and either side rolls back independently to today's behavior. The pre-write ops flush now partitions serialization ops: producer-push uploads are awaited before the event commits (the payload must not point at bytes still in flight), while consumer-settled reader ops — a dehydrated WritableStream, e.g. a manual webhook's responseWritable — are backgrounded. Awaiting those deadlocked the resume against its own wake (webhookWorkflow failing across the whole e2e matrix). Also: wake retries stop on definitive 4xx errors instead of burning the retry budget; WORKFLOW_DISABLE_LAZY_HOOK_RESUME no longer gates anything and is ignored; the internal resumeHookDurable alias is removed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Address review: retry classification, wake dedup, 409 passthrough - Wake retry classification now actually fires against @vercel/queue: its errors carry no status field, so classify by the World's deployment-unavailable hook, then numeric status, then the queue client's definitive-4xx error names. - The wake publish carries idempotencyKey `hook-<resumeId>` on the claim path, so a retried publish whose response was lost dedups instead of costing a duplicate full replay. - EntityConflictError (HTTP 409) from the durable write is no longer re-keyed to HookNotFoundError: every 409 the backend emits on this write today is transient (slot conflict past the server's retry budget, claim race) and committed nothing, so it surfaces retryable instead of presenting as a permanent 404. - Stamp workflow.hook.resume_committed / wake_published span attributes after each leg resolves, making stranded resumes (committed event, no wake) queryable from traces. - Document on the public resumeHook signature that passing the token (not a cached Hook) is what makes the write idempotent-on-retry. - Changeset/changelog: note the ended-run behavior change (late webhook deliveries to finished runs now 404 instead of 202) and the 409 passthrough. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Pranay Prakash <pranay.gp@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ffc58078d0 |
Stop logging on healthy workflow execution (#3878)
A successful run printed several lines that described the runtime working correctly. Most of it was fallout from defaulting the events transport to WebSockets (#3702): three breadcrumbs written while the transport was opt-in became default-path output, because each one reported a choice the caller no longer makes. - `world-vercel: using ws events transport (…)` ran once per cold start on every deployment, naming the transport it was always going to use. - The `projectConfig` proxy fallback warned once per process. That World cannot hold a socket, so with WS on by default every CLI command and the observability app warned about a fallback nobody asked for and nobody can act on. Debug-gated and reworded from "requested but" to "unavailable for". - The `max_duration` / `auth_expiry` drain notice is routine: the transport reconnects from the close that follows and no write is lost. Swept for the same shape elsewhere: - `world-local`'s queue-concurrency notice fired per message once a fan-out exceeded the limit — the semaphore doing its job. - `@workflow/world`'s active-run recovery line printed on every dev-server restart with work in flight. The re-enqueue *failure* above it stays unconditional; that one leaves a run unresumed. - The port-detection diagnostics in `@workflow/utils` keyed off `NODE_ENV=development`, which is the only environment that reaches them, so the gate made them unconditional for their whole audience. All of it moves behind `DEBUG=workflow:*` via a new `debugLog` in `@workflow/utils`, joining world-vercel's existing `httpLog` and `logRetry` output under one selector. Warnings and errors are untouched, so a run that actually goes wrong is no quieter than before — the ws-transport tests that assert failures are never silent still pass unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Pranay Prakash <1797812+pranaygp@users.noreply.github.com> |
||
|
|
1c44cc8c3f |
[world-vercel] Fail the run on a lost event payload instead of retrying forever (#3742)
* fix(world-vercel): fail the run on a lost event payload instead of retrying
A frame stream that dies mid-body reaches us as a truncated response, which
is exactly what a dropped socket looks like. So an event whose stored payload
is permanently gone was indistinguishable from a transient blip, and the
runtime kept redelivering a replay that could never succeed: one run re-read
a single missing payload 12,932 times in 26 minutes, and the backend query
behind each attempt throttled its table.
The World now sends a terminal `{_error: 1, code}` frame for failures that a
retry cannot fix. Handle it:
- `payload-missing` raises `CorruptedEventLogError`, so the run fails with
`CORRUPTED_EVENT_LOG` rather than looping. The log does reference a payload
nothing can produce.
- An unknown code raises a `WorkflowWorldError` with no retryable code and no
status, which is also terminal. A future code stays safe without needing a
client release first.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Revert the world-vercel URL override to empty
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(world-vercel): classify terminal stream errors
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
---------
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Alex Langenfeld <alex.langenfeld@vercel.com>
|
||
|
|
cc6eb7e837 | [world-vercel] Fix deploymentId "latest" resolving against the wrong team (#3844) | ||
|
|
d9e0777eb8 | [core] Never write hook_received eagerly on the lazy resume path (#3794) | ||
|
|
556f3f080a |
[core] Retain workflow VMs across hooks (#3604)
* Retain workflow VMs across hooks * Refine retained VM decisions and diagnostics * Fix hook suspension assertion * Harden retained hook race coverage * Simplify retention blocker log metadata * Preserve workflow suspension compatibility * Bound retained VM serialization diagnostics * Clarify bounded serialization diagnostics |
||
|
|
91eb1ae924 |
chore(docs): use geistdocs 1.23.1 (#3777)
Co-authored-by: Peter Wielander <mittgfu@gmail.com> Co-authored-by: Peter Wielander <peter.wielander@vercel.com> |
||
|
|
d62b44473b |
[core] Prune schema modules from workflow bundles (#3550)
* [core] Prune schema modules from workflow bundles * [world] Inline one-off validation options * refactor(world): simplify event schema boundaries * refactor(world): simplify event schema boundaries * fix(world): keep noop metadata schema-free * refactor(world): drop zod 4.4 compatibility * test(builders): cover workflow API bundle boundary |
||
|
|
7e48e7b4de |
Re-enable the sealed log by default (#3737)
* Revert "[world] Make the sealed log opt-in instead of default-on (#3735)" Reverts |
||
|
|
dc68611fbf |
Default the events transport to WebSockets (#3702)
* Default the events transport to WebSockets WORKFLOW_EVENTS_TRANSPORT=http is the opt-out. Only that exact value disables it, so a typo'd or empty value fails toward the default rather than quietly pinning a deployment to HTTP. The prerequisite the gate named for defaulting on is met: postEventFrameOverWs opens a client span per frame. What is still missing is Vercel's outgoing-requests view, which reads instrumented fetch calls rather than spans and so cannot show a transport that issues no request. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * docs: WORKFLOW_EVENTS_TRANSPORT defaults to ws Three places still documented http as the default. Each now states the opt-out is the exact value http, rather than leaving 'default: ws' to imply that anything non-ws disables it — the asymmetry is deliberate in the code and is the part a reader would otherwise get wrong. Also drops 'Experimental' from the Vercel World page: a setting that is on for everyone by default is not opt-in experimental, whatever else it is. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Fix the gate's own unit tests for the flipped default Five tests in ws-transport.test.ts still encoded the opt-in semantics. Three were the isWsEventsTransportEnabled table itself; the other two (openWsChannel 'does nothing when the gate is off', and the channel release equivalent) relied on the suite's ambient unset environment meaning 'off', which it no longer does. Both now set http explicitly. Two tests in ws-transport-spans.test.ts asserted HTTP-side span behaviour the same way. The write one would have kept passing by falling through resolveWsTransport's null rather than because the gate was off - passing for the wrong reason, which is what this file exists to catch. Also makes the opt-out case-insensitive and trimmed. The gate is deliberately asymmetric - unrecognized values take the default - but that asymmetry should not extend to swallowing HTTP or ' http '. Whoever reaches for the escape hatch is plausibly mid-incident, and silently ignoring their opt-out over a capital letter is the same class of silent-wrong-transport bug this flip is meant to stop shipping. 554 tests pass in packages/world-vercel. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * ci: add a required forced-HTTP e2e lane (#3703) Flipping the default makes e2e-vercel-prod a WebSocket lane: it sets no WORKFLOW_EVENTS_TRANSPORT, and unset now means ws. Nothing in the file would exercise the HTTP events transport against a real deployment any more, so this is not additive coverage — it replaces coverage the flip silently removed. Unconditional and required rather than label-gated like the WS lane. HTTP is now the fallback, and the fallback is silent: resolveWsTransport returning null costs a write nothing and logs nothing, which is the shape of the durabench bug this stack came out of. Two apps rather than the WS lane's four, since every row is a real vercel deploy charged to every PR. nextjs-turbopack is the only fixture emitting OTEL spans, so it is the one that can show which transport actually ran; express covers the non-Next server path. Also corrects the WS lane's docblock, which claimed every other job exercises HTTP only. That stopped being true one commit ago. Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Fail loudly when step_completed falls back to HTTP under a strict flag The WS e2e lane asserts that the transport is harmless, not that it is used: an event written over HTTP produces the same run outcome as one written over the socket, so the lane stayed green through the entire period the transport was silently demoted. WORKFLOW_INTERNAL_EVENTS_TRANSPORT_STRICT turns that one case into a failed run, and the WS lane now sets it. Scoped to step_completed alone, because most fallback is legitimate: run_created is written outside any invocation that opens a channel; run_started routinely lands before the channel is registered (34% HTTP on a healthy deployment); step_created and wait_created mostly fold into events.createBatch, which is not wired to the socket; and a write after the invocation released its claim falls back by design. step_completed is issued after a step body has run, and was 100% ws across every WS-enabled deployment measured on two SDK versions. The flag reads as off unless the value is exactly 1 or true - the opposite asymmetry from the transport gate, which treats an unrecognized value as on. That gate risks a deployment sitting quietly on the wrong transport; this one fails runs, and should not be acquired by a typo. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * ci: run the WS transport lane on every PR It was opt-in behind ws-transport-test because four real vercel deploys were too much to charge an unrelated PR for a transport that was off by default. Flipping the default expires that reasoning from both ends: the cost is no longer for someone else's feature, and this is now the only lane that asserts the socket carried the events. e2e-vercel-prod inherits the new default but checks nothing, so behind a label the average PR would move every deployment onto WebSockets with nothing verifying they were used. Drops WS_REQUIRED from the gate along with it. That existed only to let the lane be legitimately skipped on an unlabelled PR; with no label the lane is required unconditionally, like e2e-vercel-prod and the HTTP lane, and the skipped case is now a failure rather than a warning. Gate script extracted and run against the cases that matter: ws skipped fails on a standard PR, ws skipped fails under workflow-server-test, and all-green passes. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * ci: widen the HTTP transport lane to six server shapes Before the flip, HTTP was the default and all 28 e2e-vercel-prod lane-runs covered it. After the flip they cover WebSockets instead, and this lane is the entirety of the HTTP coverage - two apps was too thin for a transport that is still supported. Six, not the full 14, because every row is a real vercel deploy charged to every PR. Chosen by server shape rather than count: example (baseline), nextjs-turbopack (Next, and the only fixture emitting OTEL spans), vite (Vite SSR), express (Node req/res), nitro (h3, also covers nuxt) and hono (fetch-API Request/Response, a different mount shape from express). The rest duplicate a shape already covered; python is left out because it has no conformance gate and needs routes this suite does not serve. The first four match the WS lane's matrix on purpose, so the same fixture runs on both transports and a failure on one can be read against the other. Project ids and slugs are copied from e2e-vercel-prod and verified equal to it; both lanes already use the same team and token. Co-Authored-By: opencode <opencode@vercel.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> --------- Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com> |
||
|
|
f771585486 |
fix(world-vercel,world-local): hold process-wide state on globalThis (#3728)
* fix(world-vercel,world-local): hold process-wide state on globalThis Both packages are bundled into the host application's server build, and a bundler keys module identity on (resource, layer) — Next.js alone builds `instrument`, app-route, `ssr` and `edge` layers, so one process holds one copy of each of these modules per layer. Every module-scope `const`/`let` in them was therefore per-copy state wearing the costume of a process singleton. vercel/workflow#3493 made `@workflow/world-vercel` bundled rather than external and the events WebSocket transport regressed to HTTP for exactly this reason: the queue consumer registered its channel in the `instrument` copy's `Map` and the write path looked it up in the route copy's empty one. A deterministic miss, for the life of the process. `@workflow/world-local` had the same exposure all along — including `runFileLocks`, where a duplicated mutex simply stops mutually excluding. Add `globalSingleton()` to `@workflow/utils` (the primitive `@workflow/core` already hand-rolls for its World cache) and route every mutable module-scope binding in both worlds through it. Regression cover, in three layers: - `global-singleton.test.ts` pins the primitive's semantics. - `ws-transport-module-copies.test.ts` imports the module twice in one process and asserts a transport registered by one copy is found by the other — it fails on a plain module-scope `Map`, which is the shipped bug. - `scripts/lint/module-scope-state.mjs` fails the class: an AST rule banning mutable module-scope state in these packages, with `// per-copy-ok: <why>` as the deliberate escape. Wired into both packages' `vitest run src`, with fixture self-tests so it cannot rot into a no-op. * test(world-postgres): pin the module-scope-state rule for the postgres world It is deduped today only because `getRuntimeRequire()` loads it — a property of how it is loaded, not how it is written, and exactly what changed for world-vercel in #3493. The package is already clean; this keeps it that way. * docs(worlds): codify "a world must not hold mutable module state" A world package is loaded one of two ways, and only one of them gives it a single module instance: a runtime `require()` (deduped by Node) or the host's bundler (one copy per layer). Which one you get is a property of how the world is loaded, not of how it is written, and it changed under `world-vercel` in #3493 — so the rule has to be "never rely on module scope", not "rely on it until someone flips a config". Written down in the four places someone can meet it: - `docs/content/worlds/{v4,v5}/building-a-world.mdx` — a "Process-wide state" section for custom-world authors, with the loading modes spelled out and a nudge to prefer World-instance state over a global. - `packages/world/README.md` — the same constraint on the contract package. - `CLAUDE.md` — so the next contributor working in these packages sees it. - `packages/core/src/runtime/world.ts` — at the two static imports, which is where the difference between a bundled world and a required one originates. The rule's own error message now teaches it too, rather than naming a helper. Consolidates the guard while here: `@workflow/utils` owns the rule and its fixture self-tests, and sweeps every *published* `packages/world-*` discovered at runtime, so a world package added later is covered without anyone remembering. Each world keeps a one-assertion mirror for locality. * style: drop prose em dashes from this branch's new text #3704 landed a repo-wide writing pass hours after this branch was written and took `world-vercel/src` from 406 em dashes to 130 (`ws-transport.ts` alone went 35 to 1). This branch's docs section, README, comments and lint messages were written before that and would have put 36 of them straight back into the files that were just cleaned. Rewritten sentence by sentence rather than by substitution: an em dash becomes a colon, a comma, a full stop or a parenthetical depending on what it was doing. Also fixes a real defect the sweep surfaced: `world-postgres`'s guard test was generated through a shell heredoc and had literal backslash-backticks in its doc comment. * Update .changeset/world-module-scope-state.md Co-authored-by: Peter Wielander <mittgfu@gmail.com> Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> * fix(core): build the entrypoint's queue handler from getWorld() Adopted from #3666 by @MintedKenny, which implements #3665 and could not run CI as a fork PR. One line of behavior: `workflowEntrypoint`'s lazy handler init calls `getWorld()` rather than `getWorldHandlers()`. `getWorldHandlers()` owns a second, build-time-safe cache, so calling it from the runtime route built a *second* World in the same process. That costs a stateful World duplicate resources on every instance — world-postgres eagerly constructs a `pg.Pool` (default `max: 10`) and a nested world-local World in `createWorld()`, so self-hosted users have been paying for two of each — and, for a bundled world package, the two Worlds are built by two different module copies, which is the mechanism behind the WS transport regression the rest of this branch contains. The public `getWorldHandlers()` and its separate build-time cache are unchanged; only the runtime route stops using it. Kept from the original: the regression test asserting the factory runs exactly once, and the api-reference wording (re-applied over #3704's list punctuation). Not taken: renaming the `workflow.route.get_world_handlers` span. It is a distinct span from the per-request `workflow.route.get_world` at the top of the flow route, and reusing that name would collide with it in traces and in `runtime-trace-mode.test.ts`; a comment records why the name outlived the call. Co-authored-by: Kenneth <kenneth@standardforensics.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address AI review on the module-scope work Two blocking findings, both real: - **Cross-version state sharing** (`ws-transport.ts`). A process can hold two *published versions* of `@workflow/world-vercel` (a transitive dependency pinning an older `@workflow/core`, which depends on this package by exact version). Both wrote to the same unversioned `Symbol.for` key, so one version's write path could be handed a `WsEventsTransport` built by the other's class and frame against a protocol it may not share — with no version negotiation on the socket to catch it. `shapeVersion` cannot express this: the container is stable, the hazard is its contents. The registry and the events dispatcher recycler are now keyed by package version. The plain connection pools stay unversioned; sharing those across copies is the point. - **The documented pattern failed the rule this PR adds.** The custom-world docs teach `store[StateKey] ??= …`, which the rule flagged as a field write. It now recognizes state rooted at `globalThis`, following one alias hop, which is also what `core/private.ts:23` and `next/src/index.ts:58` are already doing correctly (core drops 26 findings to 22, next 7 to 6). The docs also now say outright that `globalSingleton()` is the same thing, since AGENTS.md prescribes it and the page did not mention it. Rule precision, from the review's probes: - `.mts`/`.cts` are scanned. `@workflow/world-testing` is authored in `.mts`, so its entry in the sweep was passing vacuously — with the walk fixed it reports a real finding, now annotated (it is a standalone `serve()` entry). - Mutations in top-level statements no longer count. A table filled at module evaluation is identical in every copy; divergence needs a later write. - `static` class fields are collected, attributed to the class name. - An *exported* binding initialized to an empty collection is a finding on its own, which approximates the cross-file case the walk cannot resolve. Six fixtures pin the new behavior. The rule's header now states what it does not see, and AGENTS.md states where the sweep stops and why core is not gated yet. Also tags `resetGlobalSingletonForTest` `@internal`. * fix(lint): attribute a static-field write to the field, not the class The static-field support added in the previous commit keyed `declared` on the class name, so a class carrying more than one mutable static reported one finding instead of one per field, and labelled the survivor with whichever mutation was seen first. On a two-static fixture it reported `static Registry.latch (`.set()`)`: the name of one field, the reason belonging to the other, pointing the reader at the wrong line. Key static fields `Class.field` and resolve a write to the same shape, via a new `memberPath()` that takes the first two segments of a member chain and tries that key before the bare root identifier. Two follow-ons fall out of having the path: - `this.field` inside a `static` member resolves to the class, which is the ordinary way to write the mutation. `staticClassOf()` returns nothing for an instance member, where `this` is an instance and the state is per-instance rather than per-copy, and nothing inside a nested `function`, which rebinds `this`. - `state.count++` is now a finding, like the `state.count += 1` that `assignment()` already reported. Fixtures pin all four, including the instance-field case that must stay clean. The four world packages still report zero, and the extracted `recordMutation()` keeps the file at its previous two Biome complexity warnings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: make module duplication inert across every bundled package `@workflow/core` is bundled into the host server build the same way the worlds are, and always has been — the original repro measured three live copies in every arm, including the pre-#3493 external one. One instance is not reachable: layers cannot share a module, and core cannot be external because it *is* workflow code (`runtime/start.ts:253` and nine methods in `runtime/run.ts` are `'use step'`), so it must go through the SWC loader. The Next integration already encodes that rule by removing workflow-bearing packages from `serverExternalPackages`. So the duplication stays and the hazard is removed instead, everywhere the duplication can happen. `@workflow/core` (22 findings to 0): warn-once latches in `constants.ts`, `start.ts` and `telemetry.ts`; the source-map tracer cache; the VM script cache; the QuickJS compiled-assets and baseline caches; the dev-server port cache (its own comment already said "per process"); the text codecs; the zstd browser decoder; and the `useStep` closure brand, where a function marked by one copy was invisible to another. The one with teeth was `step-single-flight.ts`: a per-copy map is not single-flight. Two invocations reaching it through different layers would each believe they were alone in the process and both run the step body, silently degrading in-process dedup to the cross-process residual its own doc scopes out to the ownership lease. Also `@workflow/world` (a warn-once set, hand-rolled onto `globalThis` to keep that package dependency-free), `@workflow/ai` (the lazy OTel API), and `@workflow/nest` (bootstrap config in a module-level `let` and two static class fields — configure one copy, read another, and the controller is unconfigured for the life of the process). Five sites are deliberately per-copy and now say why: state keyed on objects that never cross copies (the barrier safety-net `WeakSet`, the QuickJS pending byte `WeakMap`), the synchronously-scoped guest-code sink, and the OTel diagnostic that reports what *this* copy sees. The sweep now covers all of it. Packages with a single module graph stay out (build-time code, the CLI, the o11y UI, the test runner) and AGENTS.md records which and why. Found while doing this: two static fields on one class collapsed into a single entry in the rule, so `WorkflowModule.options` was invisible behind `WorkflowModule.outDir`. Statics are now keyed `Class.field`. * fix(world): suppress noAssignInExpressions on the globalThis idiom The hand-rolled form trips Biome, as it does in `packages/core/src/private.ts`, which carries the same suppression. Restructuring it into a helper function instead would hide the state behind a call the module-scope rule cannot follow, so the binding would stop being recognized as off-module and the package would report a finding for correct code. * fix: sweep every bundled package, and mark utils side-effect free @shalabhc asked on review whether `@workflow/utils` needs this too. It does, and so do three others: `utils`, `errors`, `serde` and `workflow` all end up in the host application's server build and none were in the sweep. All four report zero today, which is exactly the state `world-testing` appeared to be in before the `.mts` walk was fixed and it turned out to have a real finding. Being clean and being *checked* are different properties, and only the second one survives the next contributor. `sideEffects: false` on `@workflow/utils`: verified that every module in the package only declares (no import-time work), so a bundler can now drop the unused parts of the barrel instead of keeping all ~64 KB of it because three packages import one 476-byte function. --------- Signed-off-by: Pranay Prakash <pranay.gp@gmail.com> Co-authored-by: Peter Wielander <mittgfu@gmail.com> Co-authored-by: Kenneth <kenneth@standardforensics.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Peter Wielander <peter.wielander@vercel.com> |
||
|
|
b2cac623d3 | [world] Make the sealed log opt-in instead of default-on (#3735) | ||
|
|
e1e64e3de3 |
docs: apply Vercel technical writing standards (#3704)
* docs: apply Vercel technical writing standards Audit the complete documentation corpus, package READMEs, skills, and source TSDoc/comments against the vercel-technical-writing skill and style-rules.md. Normalize sentence-case headings without changing published anchors, remove prose em dashes and filler wording, improve active voice and self-contained phrasing, standardize product/brand capitalization, American English, list punctuation, units, and code fence languages, and preserve exact runtime strings/table placeholders. All executable code is unchanged. Modified skills have their metadata versions bumped. * docs: extend writing audit to repository Markdown Apply the same technical-writing rules to design documents, compiler specifications, workbench guides, package changelogs, and the remaining tracked Markdown outside the deployed docs corpus. Preserve historical meaning, commands, output literals, table placeholders, and heading anchors. * docs: exclude generated package changelogs from audit |
||
|
|
7b79ba37cc |
Add support for 'noop' event type - spec version 7 (#3634)
Co-authored-by: Peter Wielander <peter.wielander@vercel.com> |
||
|
|
b3dbc6d264 | [docs] v5 changes docs: what's new, world upgrade guide, migration skills (#3100) | ||
|
|
d5d19fc8d5 | docs: update Geistdocs to 1.20.4 (#3654) | ||
|
|
9b1b8c7111 | [core] Pin correlation-id draw order to event-log order (#3700) | ||
|
|
9454d51db0 |
feat(core): resolve run.returnValue via a World long poll instead of a 1s poll (#3570)
Co-authored-by: Peter Wielander <mittgfu@gmail.com> Co-authored-by: Peter Wielander <peter.wielander@vercel.com> |
||
|
|
5b5a926f88 |
fix(core): make step-argument serialization failures catchable in workflow code (#3675)
* fix(core): make step-argument serialization failures catchable in workflow code
A step whose arguments fail to serialize is now finalized by the
suspension handler as step_created + step_failed (mirroring a step-body
failure) instead of rejecting the whole suspension. The next replay —
forced in-process, since no step message is dispatched for the failed
step — rejects the step's promise with the SerializationError, so a
try/catch around the step call observes it. Uncaught, the error
propagates out of the workflow body and fails the run as a fatal
USER_ERROR immediately, instead of redelivering the orchestrator
message until max deliveries (49/48) as reported in production on v4.
* Serialize the step_failed error with the VM global; one-sentence changeset
Addresses review feedback: dehydrateStepError in
finalizeUnserializableStep now receives suspension.globalThis like every
other dehydration in this file. Error detection is realm-independent, so
the host-created SerializationError serializes identically, but VM-realm
values guest code threw into the cause chain are now detected by the
realm-sensitive reducers.
* Address review: QuickJS engine support, deferred-batch join, drain gate, placeholder marker, telemetry, docs
- QuickJS: dumpPendingOps now catches a step input's serialization
failure per-op, reframes it as a SerializationError with the same
framed message as dehydrateStepArguments, and surfaces it on the
pending op instead of failing the whole collection. The entrypoint's
dispatchPendingOps finalizes such steps as step_created (placeholder
input) + step_failed, excludes them from inline claims and queue
publishes, marks them handled, and raises the requeue signal so the
failure is observed even when the feed lags — mirroring the node:vm
engine, so both engines agree: catchable in workflow code, USER_ERROR
with the framed message when uncaught. Both step-argument e2e tests
now pass on WORKFLOW_VM=quickjs.
- runtime.ts: the failed-step replay path now joins
suspensionResult.deferredBatchWork before continuing, so a trailing
chunk commit or step-message publish rejection propagates instead of
being swallowed after ack; committed inline claims are documented as
deliberately handed to owned recovery.
- Terminal drain: finalization is gated on a stepDispatch target. The
drain caller has no replay to observe a finalization, so a completed
run no longer gains failed-step rows for an unawaited unserializable
step — the rethrown error is swallowed by the drain's catch,
preserving its pre-existing behavior.
- The placeholder input now carries a marker string ('[input
unavailable: step argument serialization failed]', shared via
runtime/unserializable-step.ts) so inspect/o11y don't render the
failed step as a genuine zero-argument call.
- New workflow.steps.failed_serialization span attribute on the
suspension span, so occurrence is measurable without log search.
- Docs: v5 serialization-failed error page documents where each
boundary's failure surfaces (catchable step failure vs run failure)
and the no-retry USER_ERROR semantics; foundations/errors-and-retries
gains a Serialization Failures section with the try/catch shape.
* Guard the finalization crash window; self-contained docs samples
- A crash or transient failure between finalization's two durable
writes leaves a lone placeholder step_created, and redelivery then
dispatches the step through normal crash recovery — previously
running user code with the placeholder arguments. The placeholder
now carries a structural flag on the input triple's top level (which
user code never controls, so no false positives), and the step
executor checks it after hydration: instead of running the body, it
throws the intended fatal SerializationError, completing the
interrupted finalization as step_failed. Applies to both engines
(they share the placeholder and the executor).
- Regression tests: executor fails a placeholder-input step without
running the body (and doesn't trip on a genuine argument equal to
the display marker); handleSuspension rejects for redelivery when
step_failed can't be written after step_created landed, leaving the
recoverable placeholder behind; mixed bad-step + large fan-out
returns the failure set alongside still-pending deferredBatchWork
whose rejection surfaces — the contract the runtime's failed-step
join (added previously) relies on.
- Docs: the two new code samples are now self-contained so the docs
code-sample typecheck passes.
|
||
|
|
0b2797bbac | [next] Bundle the Vercel world into the Next.js server output (#3493) | ||
|
|
37e1d9e5a9 |
Batch: pre-claim inline steps in the same batch (#3568)
* Pre-claim inline steps inside the suspension batch (born-running pairs) Restacked onto main after #3025's squash-merge; folds in the review-round changes to the flush loop (per-write requestId attribution on createBatch, and the seeded/advancing slot-bump expectation, now shared with the pre-claim ceiling). Fold each lazy-inline step's deferred writes into the batched fan-out as an adjacent [step_created, step_started] pair: the created row carries the input, the started row is a bare ownership-stamped claim the server folds into one born-running create. The whole scheduling turn commits as ONE durable write, inline bodies start straight off that commit (in parallel with the VQS publishes for backgrounded steps), and executeStep gains a pre-claimed mode that runs or skips the body off the batch's per-event verdict - a pair 409 is the same skipped outcome as losing the lazy claim. The lone-inline case keeps the optimistic lazy path (a pair-only batch buys nothing over the single claim). Also threads per-event computeInstanceId through the World batch request, and folds the batch's committed slot ceiling into the inline slot snapshot so terminal writes stop being answered with reports echoing the batch's own events. * Parallel chunk commits, per-chunk continuation, batch span attributes Production trace of a 67-event fan-out showed the three batch chunks POSTing back-to-back (~230ms each) with no bodies or queue messages until all three settled (~670ms). Three changes: - Chunks now POST concurrently. Slot assignment is the server's, so parallel chunks race for slot ranges exactly like the pre-fold path's parallel single writes did; entity conditions, not commit order, carry correctness. The foreign-interleaving diagnostic is computed once over the whole fold (committed span vs seed) instead of per chunk. - Per-chunk continuation: each chunk's step-execution queue messages publish the moment ITS creates are durable (in-flush, via stepDispatch, same message shape and idempotency key as the caller's dispatch pass - the affected steps are pre-reported in queuedStepCorrelationIds so the caller skips them). Only the chunk carrying the inline pairs gates handleSuspension's return (opt-in via allowDeferredBatchWork); trailing chunk commits + all publishes ride result.deferredBatchWork, which the runtime joins next to the dispatch join before it can ack - the every-create-durable-before-ack contract is unchanged, the bodies just start off the pair chunk instead of the slowest chunk. - OTel: batch identity attributes (workflow.batch.size, per-type workflow.batch.shape) now live on the world.events.createBatch span (instrumentObject) instead of the http POST span, which keeps only wire-level facts (transport, bytes) and no longer sets workflow.event.type - that attribute names a single event write and tagging a batch with its first event's type misclassifies traffic. * Address review: settle deferred fold on failure, drop pair-batch retry Three fixes from review of the deferred/parallel-chunk fold. 1. A pair-chunk rejection escaped `handleSuspension` while the trailing chunks' commits and publishes were still in flight. `deferredBatchWork` never reaches the caller once the handler throws, so nothing joined that work — exactly the state `settlePhase` exists to prevent: a sibling create landing after the rejection commits an event from the abandoned replay's seeded sequence and races the caller's restart reload. The failure path now settles `trailing` before rethrowing. 2. Every pair-carrying chunk gates the return, not just the first. Pairs sort to the front and two rows per inline step fit inside one chunk, so this is one commit today, but `findIndex` silently degraded if either cap moved: a pair in an unawaited chunk yields no `inlineClaims` entry, the caller falls back to a lazy `step_started`, and that races this same fold's in-flight pair for the same step. constants.test.ts now pins the cap relationship. 3. A batch carrying a `step_started` is no longer retried in-process. The born-running pair does converge to a 409, but the pre-claim caller reads a pair 409 as "a concurrent writer owns this step" and skips the body — and on a retry that is indistinguishable from "my own first attempt committed the pair". Skipping there stranded a running step stamped with this invocation's own message id until the ownership lease expired (860s), where the single-POST path deliberately fails the delivery and recovers through owned-recovery in seconds. Same reasoning `EVENT_RETRY_ELIGIBILITY` already applies to `step_started`. Also asserts `lazyStepInput` / `preclaimedStart` mutual exclusivity in executeStep instead of only documenting it, and adds the changeset. Tests: +1 suspension-handler (pair-chunk failure settles the trailing chunk before escaping — fails without fix 1), +1 constants (cap relationship), +1 world-vercel (a born-running pair batch is single-attempt), and the existing batch-retry test retargeted at an entity-conditioned batch. Full @workflow/core unit suite 2178 green, @workflow/world-vercel 514 green, typecheck green across core / world / world-vercel. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Guard inline bodies against unhandledRejection; review follow-ups The dispatch/deferred-batch joins now sit between the step promises' creation and the `Promise.all` that reads them, so a body rejecting in that window had no handler attached at the microtask checkpoint — an unhandledRejection, fatal under Node's default --unhandled-rejections=throw. A 412 fenced claim races exactly that window, and `deferredBatchWork` widens it by a trailing-chunk round trip. Attach a no-op catch at creation, the same way `dispatchesSettled` already does two lines up; the awaits below still decide the outcome. Review follow-ups: - `workflow.batch.shape` is sorted by event type. Map iteration is first-seen order, so a pre-claimed fold and a pure eager fold rendered the same composition as different strings, which is not groupable as a dimension. - A lost pre-claim reports StepSkipReason `running`, not `completed`. The pair's 409 says the step already exists and its claim winner is executing; the other skip site is a genuine terminal-state conflict, and tagging both `completed` left the attribute unable to separate the two. - `batchCommittedSlotCeiling`'s docstring now says the echo is only fully suppressed for a single-chunk fold: on a multi-chunk fan-out an inline terminal write issued before the trailing chunks land still names a position below them and still draws a report. - The defensive throw on a missing dehydrated input records where it lands — the pair is already durable, so it fails with the step claimed and its body unrun, recovered on redelivery via owned-recovery rather than failing cleanly. No regression test for the unhandledRejection: the existing inlineClaimRejectionScenario runs both steps inline, so `dispatches` is empty and the join resolves in a microtask — the window never opens and a test there passes with or without the fix. Reproducing it needs a scenario with a backgrounded step and a slow queue publish alongside the fenced claim. Full @workflow/core unit suite 2178 green, @workflow/world-vercel 514 green, typecheck and biome clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Pin per-event computeInstanceId on the batch wire Batch encoding is a separate path from the single-event POST, so the frame meta had no coverage: the only assertion was at the World-call boundary. Adds a wire-level test that a pre-claimed pair's step_started half carries computeInstanceId in its frame meta and the step_created half does not. Verified it fails when the threading in createWorkflowRunEventBatch is removed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Guard the pre-claim path as inert on Worlds without createBatch world-local and world-postgres do not implement createBatch, so the fold never engages there — but the runtime passes ownerMessageId and allowDeferredBatchWork unconditionally. The existing "keeps the single path when the World lacks createBatch" test passed neither, so it never covered the pre-claim path at all. Assert the inertness with the params the runtime actually sends: no claims, no deferred work, no slot ceiling, the lazy-inline step still carrying its input, and no step_started reaching the world. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Peter Wielander <peter.wielander@vercel.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
04e060a0ec | [world] Add WORKFLOW_NODE_HTTP to run the HTTP Worlds on node:http (#3461) | ||
|
|
df0103bc3c |
Document stream reader request cancellation (#3581)
## Summary & Motivation The earlier timeout-docs attempt in #534 was closed without merging. Without `supportsCancellation`, a browser disconnect leaves the stream reader reconnecting until the function hits `FUNCTION_INVOCATION_TIMEOUT`, so the streaming guide now documents the `vercel.json` opt-in, with a warning that it terminates everything matching the configured path, and the resumable-streams guide points at it. ## Test Plan Docs only; no tests. ## Docs Preview | Page | v4 | v5 | | --- | --- | --- | | Streaming | [Preview](https://workflow-docs-git-alangenfeld-timeout-docs.vercel.sh/docs/foundations/streaming#avoiding-function-timeouts-after-client-disconnects) | [Preview](https://workflow-docs-git-alangenfeld-timeout-docs.vercel.sh/v5/docs/foundations/streaming#avoiding-function-timeouts-after-client-disconnects) | | Resumable Streams | [Preview](https://workflow-docs-git-alangenfeld-timeout-docs.vercel.sh/docs/ai/resumable-streams) | [Preview](https://workflow-docs-git-alangenfeld-timeout-docs.vercel.sh/v5/docs/ai/resumable-streams) | Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com> |
||
|
|
b0adb50bce |
feat(world,world-vercel): createBatch — ordered batch event write with per-event results (#3025)
## createBatch: the client half of the v4 batch event write (per-event
results, no fence)
> **Note:** this PR was rebuilt from scratch. The previous revision
implemented the retired "v2 suspension fence" design
(`expectedRunVersion` / `batchId` / `logicalCreatedAt`, a
grammar-validated collect mode, world-postgres migration 0016,
world-local claim machinery). The server redesigned its endpoint in
place (vercel/workflow-server#646, merged and deployed) and this branch
now targets that contract on top of current `main` (specVersion 6 slot
identity). The old head is tagged `batch-client-v2-fence-design`; prior
review threads reference deleted code.
### The server contract this targets
`POST /api/v4/runs/:runId/events/batch` (workflow-server#646): an
**ordered** list of v4 frames — byte-identical to single-event POST
frames, **no batch-level meta** — committed in one DynamoDB transaction
per attempt, answered with HTTP 200 + `{ results }`: one entry per
frame, in request order. Each event reports what its own single POST
would have returned: `200` + the materialized entity, or the single-path
status/code (e.g. `409`/`conflict` for an event an earlier delivery
already applied). A transport retry of a committed batch converges to
all-409s with nothing written twice — idempotency comes from per-entity
conditions, not batch bookkeeping. Slot-identity runs only (specVersion
≥ 6 — what `world-vercel` stamps on every new run since #3389).
### What this revision ships
1. **`@workflow/world` — the spec addition.** `Storage['events']` gains
one optional method; **method presence is the capability declaration**
(no capability flag, no stub required):
```ts
createBatch?(
runId: string,
events: BatchEventRequest[],
params?: CreateEventBatchParams
): Promise<EventBatchResult>;
interface BatchEventRequest {
event: CreateEventRequest; // same discriminated union as the single create
occurredAt?: Date; // under slot identity: the source of the durable createdAt
}
type BatchEventItemResult = // one per submitted event, in request order
| { status: 200; event: Event; run?: WorkflowRun; step?: Step; wait?: Wait }
| { status: number; error: string; message: string };
interface EventBatchResult { results: BatchEventItemResult[] }
```
Contract: **ordered** (events land in the log in request order),
**per-event outcomes** (each event reports what its own single `create`
would have returned — success discriminated by `error === undefined`),
**idempotent on retry** (per-entity conditions make a retried committed
batch converge to per-event 409s). Worlds that don't implement it keep
the single-event path. `world-local` and `world-postgres` deliberately
do NOT implement it — batching a local/in-process write buys nothing
(this deletes the old revision's riskiest surface: the hand-written
postgres migration and the world-local claim machinery).
2. **`@workflow/world-vercel`** — the wire adapter: per-event frames
concatenated in order (reusing the single-frame encoder; each frame
carries its own `occurredAt`, which under slot identity is the source of
the durable `createdAt` — this natively closes the replay-clock question
the old `logicalCreatedAt` field existed for), CBOR `{ results }`
decoded against the **same per-type zod schemas as the single POST**,
loud `SCHEMA_VALIDATION` on any malformed response (wrong length,
invalid item), and the standard typed error mapping for request-level
failures.
3. **Retry policy** — a `batchIdempotent` override in the event-retry
eligibility machinery: the whole batch POST retries transient transport
failures/5xx (and waits out 429 `Retry-After` per #3504) regardless of
the contained event types, because per-event entity conditions make the
retry converge; the per-type non-retryability matrix guards *single*
posts (where e.g. a retried bare `step_started` would increment
`attempt`) and doesn't apply inside a batch.
Tests: 7 wire tests — frame encoding/ordering + **no fence fields on the
wire**, per-event result mapping (successes typed, failures passed
through), malformed-response failures (length mismatch, invalid item
body with index), typed request-level 400s, in-process 5xx retry,
empty-batch guard — plus 9 suspension-handler tests for the runtime
fold: ordering (steps then waits), per-event 409 tolerance, non-409
failure propagation, every gate exclusion (flag off / no `createBatch` /
pre-slot run / hook writes), 32-cap chunking, and lazy-inline exclusion.
Full `world-vercel` suite: 508 passed; full `@workflow/core` suite: 2126
passed.
### The runtime integration: batched suspension fan-out (ON by default)
The suspension handler folds a **clean fan-out** — the suspension's
eager `step_created` + `wait_created` writes — into `createBatch` calls
of at most **32 events**, and uses the batch endpoint **exactly when two
or more batchable eager events exist**: a lone eager event takes the
ordinary single write (same round trip, and it keeps the slot-snapshot +
bump-and-report the single path provides) (mirroring the server's
transaction budgets: 2 items/event against the 100-item cap, 768 KB
inline-byte budget; larger fan-outs commit in successive batches). The
gate requires: World implements `createBatch` ∧ run on slot identity
(specVersion ≥ 6) ∧ no attribute writes ∧ no hook writes ∧ no resilient
step dispatch. **Everything outside the gate keeps the single-event path
byte-for-byte**, and lazy-inline steps keep deferring their
`step_created` to the lazy start exactly as before.
Per-event semantics mirror the single path: a `409` is the same
already-exists tolerance as `EntityConflictError` (the conflicted step
is not marked owned); any other per-event failure fails the suspension
write the way a single-path rejection would. Slot bumps (the batch
endpoint has no bump-and-report) are tolerated and logged — the same
accepted exposure as a dropped truncated skipped-slot report on the
single path.
**On by default**, with the `WORKFLOW_TURBO`-shaped kill switch as the
operator escape hatch: **`WORKFLOW_BATCH_TRANSITIONS=0`** (or `false`)
disables batching and restores the exact prior one-write-per-event path.
Documented in the worlds configuration reference and the changelog
entry. Burn-in watch: the `event_batch`-tagged slot-conflict metrics and
DynamoDB throttle monitors on the server side.
### Docs
- New v5 changelog entry **`changelog/batched-event-writes`**
documenting the World spec addition (full `createBatch` signature +
contract — the signature block is compile-checked against
`@workflow/world` by the docs code-sample checker), the runtime fold,
and the follow-up.
- `configuration/worlds` gains the **`WORKFLOW_BATCH_TRANSITIONS`**
reference entry: default on, `=0`/`false` as the documented escape
hatch.
### Staged follow-up: the deferred sequential transition (the STSO win)
Hold `step_completed(N)` across the replay turn and commit
`[step_completed(N), step_created(N+1), step_started(N+1)]` as one batch
at the next lazy start (the server folds the pair born-running). This
needs the synthetic-completion replay machinery rebuilt against today's
runtime (parallel inline batches, turbo's run-ready barrier, optimistic
starts, slot bookkeeping) — it stays a separate PR so the SDK's most
sensitive replay path gets its own focused review. Its acceptance
criteria are already agreed: the runtime eligibility matrix as unit
tests, and an e2e that asserts ≥1 POST to `/events/batch` and **zero**
single-event POSTs for the batched transitions.
### Compatibility
- Old servers: no `/batch` route → 404/405 → callers fall back to
single-event posts (the runtime PRs will latch this per run).
- Pre-slot runs: request-level 400 (`batch-requires-slot-identity`) →
same fallback.
- No `WORKFLOW_SERVER_URL_OVERRIDE` pin this time — the server endpoint
is merged and deployed to production.
Refs: vercel/workflow-server#646 (endpoint), vercel/workflow-server#780
(unbatchable-types design space), #3389 (slot identity), #3504 (429
retry).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
1321570464 | [docs] Document duplicate-event handling, and describe webhook token generation accurately (#3497) | ||
|
|
de2a86c61c | [world] Make spec version 6 the current version (#3542) | ||
|
|
dc85865718 | [core] Drop pre-slot event ID support and preconditionGuard capability (#3519) |