Add isDecisionEvent and gate the stateEventCount fence on it in
world-local and world-postgres: facts (completions, receipts, non-lazy
step_started claims) pass unfenced and never bump the fence even when
the runtime attaches a snapshot. A stale fact is byte-identical to a
fresh one and takes its meaning from its commit-assigned log position,
so fencing it prevents nothing and converts steady traffic into
replay-restart churn (measured: 776 wait_completed + 1107 step_started
rejections in one 48-run storm).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The run_started preload and the step-terminal inline delta are
events-directory scans whose latency grows with the directory; holding
the per-run append serializer across them starved the run's other
writers and, under load, tripped the stale-lock breaker on a healthy
holder — minting duplicate positions. Attach both read pages after the
lock releases (their point-in-time semantics are unchanged), and err
the stale threshold far to the late side: breaking early corrupts,
breaking late merely stalls a crashed run's writers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A fenced create whose snapshot is stale was rejected at the publish
point — after entity writes, claim files, and the terminal marker had
already landed. Unlike postgres, those writes cannot roll back: a
rejected step_created stranded a step entity whose step executes with
no step_created in the log (an unreplayable completion), and a rejected
run_completed left the terminal marker behind, rejecting every later
hook resume. Admit or reject the session while the append is still a
pure no-op.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Assign every event a dense per-run seq and a tail-dominant event key at
its publish point, under an on-disk per-run append serializer (the
filesystem analogue of world-postgres's advisory transaction lock), and
enforce the stateEventCount decision fence under the same lock. Extend
the SDK's contiguity assertion to the pre-replay convergence point so
the inline write-response delta and run_started preload are covered too.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* origin/pgp/replay-engine-determinism-tests:
fix(core): reset the idle poll's budget on delivery progress
fix(core): do not start the abandon clock before a payload can be claimed
fix(core): never retire a committed delivery on a timer, and count parked deliveries as in-flight
fix(core): make delivery-barrier retirement deterministic
test(core): mark known-failing determinism repros as expected failures
test(core): reproduce delivery-barrier starvation under unrelated delivery traffic
test(core): characterize divergence from a below-watermark event
test(core): reproduce ReplayDivergenceError from a buffered hook payload
test(core): reproduce delivery-barrier idle collapse and premature retirement
* pgp/autoincrement-event-sort-ids:
Key the fence's sibling credit on a per-invocation writerId
Fence decisions on foreign decisions, not on facts
Assign commit-ordered event positions in world-postgres and fence stale writers
Derive the precondition watermark from the log's maximum, not its tail
Remove the temporary backend URL override
Temporarily point CI at the backend branch preview
Gate event creation on the loaded event count and restart replays in-process
# Conflicts:
# packages/core/src/runtime/helpers.test.ts
# packages/core/src/runtime/step-executor.ts
# packages/world-vercel/src/events.ts
* Fix Biome lint violations and add Biome CI check
Biome was not configured to respect .gitignore, so ~92% of the 13,355
reported diagnostics came from gitignored build artifacts. Enable VCS
integration (useIgnoreFile), apply safe auto-fixes across the repo, fix
the remaining mechanical errors by hand, downgrade judgment-call a11y /
dangerouslySetInnerHTML rules to warnings, and add a 'biome ci' job to
the Lint workflow so violations block PRs going forward.
* Use an empty changeset (no behavior change, no release needed)
* feat(world): persist the compute instance that ran each step attempt
Add CreateEventParams.computeInstanceId (ambient per-event identity, mirroring requestId) and a readable Event.computeInstanceId. Core stamps it on every step_started write; world-vercel forwards it in the v4 frame meta next to vercelId. Lets observability distinguish steps sharing a compute instance from those on different instances or invocations.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
* test: cover computeInstanceId threading from params to v4 frame meta
world-vercel: computeInstanceId reaches the v4 frame meta, rides alongside vercelId rather than replacing it, and is omitted when unset. core: step_started carries it without displacing the stateUpdatedAt precondition guard (both share one params object).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
* fix(web-shared): render computeInstanceId in the attribute panel
AttributeKey derives from keyof Event, so adding computeInstanceId to the event schema widened it and left the exhaustive attributeToDisplayFn map incomplete (TS2741). Renders it beside deploymentId as 'Compute Instance ID', copyable like the other opaque ids.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
* fix(world-postgres): exclude computeInstanceId from the events column contract
The events table asserts satisfies DrizzlishOfType<...Omit<Event, 'occurredAt'>...>, so adding computeInstanceId to the event schema broke the build (TS1360). This world does not persist it, matching how occurredAt is already handled.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
* refactor: move computeInstanceId to the analytics read contract
The server routes computeInstanceId into ClickHouse and returns it on AnalyticsEvent/AnalyticsStep, never on the event record — so Event.computeInstanceId was dead on read and zod would strip the field off the analytics wire. Move it to AnalyticsEventSchema/AnalyticsStepSchema (beside vercelId/requestId, the same class of ambient provenance), which also drops the world-postgres column-contract exclusion entirely.
Also: hoist the duplicated step_started params into one local, extract the repeated mock-agent harness in events.test.ts, and use vi.spyOn plus an identity assertion against COMPUTE_INSTANCE_ID.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
---------
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Counted flat, `IDLE_POLL_DEADLINE_ROUNDS` is not a deadline but a cap on
how many events one drain window may deliver.
Delivering a window in log order costs one macrotask per deferring
delivery, so a window of K events needs K poll rounds to clear. Past the
budget `scheduleWhenIdle` fires anyway, into the middle of the batch — and
for a hook consumer that callback raises a `WorkflowSuspension`, so the run
suspends carrying none of the work the remaining deliveries were about to
create. A new test in `delivery-barrier-idle-starvation.test.ts` puts 40
deliveries in one window and shows the callback landing after exactly 16 of
them.
The budget now restarts whenever a delivery reaches workflow code, tracked
as `deliveryProgress` and incremented by `markDelivered()` only. Excluding
the abandonment safety net's retirements is what keeps the escape hatch
working: a stream of pokes to a hook nobody reads retires barrier after
barrier and delivers nothing, so the budget still runs out and the
suspension still fires. That is the 3m19s stall the deadline was added for,
and its test still passes.
Ordering machinery that slows delivery down is only safe if the liveness
timers around it measure progress rather than elapsed rounds; this is the
one that did not.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
An unarmed delivery barrier is retired on the inference that no consumer
has claimed the payload. That inference is not available until a consumer
could have: `claim()` cannot hand anything over until the payload's own
hydration resolves, so while it is still hydrating, "nobody has claimed
it" carries no information.
Left unguarded, the two retirement conditions conspire during that window.
The payload's own hydration holds `pendingDeliveries` above zero, so the
idle route cannot fire, and the abandon deadline retires the barrier
because its hydration was slow. Against a production hydration (object
storage fetch plus decrypt, 10-500ms) and a deadline of a few dozen
`setTimeout(0)` ticks, that is the common case for a buffered payload
rather than an edge one.
`registerDeliveryBarrier` now takes `abandonableAfter` and does not start
the retirement poll until it settles; `workflow/hook.ts` passes the
buffered payload's own hydration. The waiting-consumer path is unaffected:
it registers armed, and an armed barrier leaves the poll on its first
check either way.
This is the same slow-hydration trap b58fdcf0d had on the armed side,
found by re-auditing the deadline against production latency after that
commit's storm run. The regression test holds hydration open well past the
deadline and asserts both that the barrier survives and that a later step
result still orders behind the claim.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
Follow-up to the previous commit, from review of #3196 and from an e2e
regression that review work exposed.
Retirement by the safety net is now UNARMED-only. The abandon deadline used to
fire for any barrier, which reintroduced the collapse at a longer timescale:
production hydration (object-storage fetch plus decrypt) runs 10-500ms, far
past any tick budget worth setting, so the slowest deliveries would have been
exactly the ones force-retired mid-flight. Only an unclaimed buffered hook
payload — the one delivery that can be abandoned at the root — is retirable by
the net; every other barrier leaves the registry from its own delivery chain.
All five call sites were audited against that invariant: step_completed,
step_failed, wait_completed, the waiting-consumer and claim() hook paths, and
the abort hook all attach their chain unconditionally.
Deliveries parked in awaitEarlierDeliveries are now counted, and the count
gates scheduleWhenIdle. A delivery whose only remaining gate is a
recentlyDeliveredBarriers entry is invisible to both pendingDeliveries (already
released in its hydration slot) and to a scan of the live registry (already
empty), so a suspension armed in that window preempted it and the run suspended
carrying none of the work the delivery was about to create. That is how the
second payload of the hookWithSleepWorkflow e2e went missing. Guarded by a unit
test that fails without the counter.
Review fixes:
- Starvation suite: the poke storm held pendingDeliveries above zero with a
single unmatched increment, so it modelled a stuck hydration rather than
overlapping traffic. It now overlaps genuinely, leaks nothing, and asserts
both properties. The header no longer attributes the observed 3m19s stall to
this mechanism, which the suite does not establish.
- Added the liveness guard asked for in review: 228 unclaimed payloads gating a
later delivery still drain in a handful of ticks.
- Added the mirrored-log control asked for in review: the interleaving where
the step result legitimately precedes the wait completion must keep replaying
cleanly, so the ordering assertions cannot be satisfied by flipping the bias.
- Below-watermark suite: the read shape is not a world-postgres class.
world-vercel paginates DynamoDB by event ULID with an `eid:` cursor over a
strictly-after range, and workflow-server documents its IDs as monotonic only
intra-instance. Recorded in the header; the second test is renamed to say it
characterizes current behavior rather than pinning a contract.
- Call-site comments that documented idle retirement of parked barriers as a
deliberate tradeoff now describe what the net actually does.
- Stale private.ts / step.ts line references replaced with symbol names.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
* origin/main:
[benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main (#3213)
feat(core): emit faas.instance span attribute for compute instance identity (#2989)
[world-local] Retry transient EPERM unlink failures on Windows (#3215)
[world-vercel] Raise H2 receive windows on the events agent (#3212)
Version Packages (beta) (#3185)
Align Streams UI with trace viewer (#3197)
fix(core): don't observe idle while a committed delivery is parked behind its deferral (#3198)
[world-vercel] Make HTTP/2 actually multiplex on the events path (#3190)
[RFC] feat(nitro): embed observability dashboard in-process at /_workflow (#2548)
* Split STSO by inline vs queue-hop steps, add distribution diffs vs main
The sequential-steps benchmark's STSO metric mixed two unrelated
phenomena: gaps between steps running back-to-back in the same warm
process, and gaps across an invocation boundary (queue dispatch, client
reinit, event-log replay), which cost ~10x more. The old step-index
windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured
neither cleanly: whether a boundary happened to land inside a window
moved that window's P99 by hundreds of percent, which is most of the
run-to-run variance the benchmark comment was reporting.
The workflow now tags each step with whether it was the first step body
executed in its process ('queue-hop') or a later one in the same warm
process ('inline') via a process-global, so the split is ground truth
rather than inferred from step index or trace timestamps. STSO is
reported as two rows over *every* gap in the run instead of three
sampled windows. No targets on the new rows — the old ones described the
index-bucketed grouping.
computeStats now keeps the full sorted sample array alongside the
percentiles, and the comment renders a histogram + cumulative-time diff
against `main` under the table, one per STSO kind. Percentiles alone
hide how many samples moved and by how much, which is exactly where the
variance lives. Inline rows use a fixed 50ms bin width (the adaptive
width is coarse enough to hide structure inside that cluster); queue-hop
rows keep the adaptive width. Negative gaps (clock skew between two step
bodies' clocks) get their own bucket rather than being counted with the
slow tail.
Raw samples are stripped from the comment's embedded data block — ~1000
per run would exceed GitHub's comment size limit within a couple of
history entries — so the histogram renders for the current run only,
while collapsed history keeps its tables. Until this lands on `main` no
baseline has raw samples, so the section renders this run's distribution
as a single series.
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
* Clarify what stripping raw samples from the data block does not affect
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
* Drop the bucket tables; fold counts and deltas into the histogram bars
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
* Collapse the STSO distribution section into a dropdown
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
* Fix footer assertion after the dropdown wording change
Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>
---------
Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
Three defects in the delivery-barrier registry could hand branch-deciding
deliveries to workflow code out of event-log order, surfacing as
CORRUPTED_EVENT_LOG on replay. All three are fixed here, and the twelve
expected-fail repros that documented them are flipped back to plain `it`.
Barriers no longer retire on a global idle tick. #3198 stopped the idle CHECK
from observing idle while a committed delivery is parked; this stops the same
predicate from retiring the barriers themselves. `pendingDeliveries` tracks
only the host-side hydration window — released inside the promiseQueue slot,
before the detached continuation that hands the value over, and never touched
at all by a wait_completed — so one idle tick used to retire every live
barrier while their deliveries were still parked in awaitEarlierDeliveries.
Retirement moves to its own poll, which retires only an UNARMED barrier: an
unclaimed buffered hook payload, the one delivery that can be abandoned at the
root. Every other kind is committed and retires from its own chain.
Both polls get a deadline. Without one neither has any escape from continuous
unrelated traffic: an abandoned barrier under a stream of deliveries to a
never-read hook starved indefinitely, observed live as a run making no
progress for 3m19s across 228 pokes. The barrier deadline is counted in raw
ticks, so the traffic keeping the system busy cannot stretch it;
scheduleWhenIdle's is counted in poll rounds, each of which waits out a full
promiseQueue drain, because that function also schedules suspensions, where
firing early preempts data delivery.
A retired delivery stays visible to the registry for one more macrotask.
markDelivered() resolved a barrier one statement before the resolve() that
wakes the branch, so anything reading the registry in between — a delivery
consumed in a later drain window, or a buffered payload's claim() — computed
an empty deferral set and overtook the branch it was meant to follow. The live
registry still drops the entry immediately; a second short-lived map carries
ordering visibility across the gap.
A step result's buffered-hook skip is narrowed to the payload itself. A step
still never gates on an unclaimed payload, where the claim commonly sits
downstream of the step result, but it now gates on an earlier armed wait or
hook whatever that delivery is itself waiting for. Skipping those as well, via
the transitive resolvesOnItsOwn walk, was the buffered-claim corruption: in
Promise.all([step, sleep-then-read-hook]) the wait completion is precisely what
wakes the branch that goes on to claim the payload. awaitEarlierDeliveries no
longer consults that walk, leaving the per-delivery path a single linear pass;
the walk itself survives only in hasParkedCommittedDelivery, which gates
suspensions rather than deliveries.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
* feat(core): emit faas.instance span attribute for compute instance identity
Synthesize a per-warm-instance id (cinst_<ulid>) once at module load and emit it as the OTEL faas.instance attribute on the flow and step route spans. Vercel exposes no native per-instance id under Fluid compute, so this lets traces distinguish which compute instance handled each request. Purely additive telemetry.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
* test(core): assert faas.instance is stable across invocations
Covers the id format and, on the two-invocation warm-handler case, that both invocations report the same id — the module-scope minting contract that makes the attribute identify the instance rather than the invocation.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
---------
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
deleteJSON was the one mutation path in the fs layer that neither used
withWindowsRetry nor swallowed unlink errors. On Windows, unlink fails
with a share-violation EPERM while a concurrent reader briefly holds
the file open — hook polling races deleteAllHooksForRun by design — so
a transient EPERM surfaced as a failed operation, e.g. a failed
run.cancel().
Wrap the unlink in withWindowsRetry, matching the rename/link/unlink
guards the write pipeline already has. ENOENT is not in the retryable
set, so the already-deleted tolerance still short-circuits.
Signed-off-by: Andrew Barba <barba@hey.com>
Storm forensics (fence_count audit column) attributed every corrupted run
to a credit hole, not engine nondeterminism: two invocations that loaded
the identical prefix present byte-identical stateCursor+stateEventCount,
so the snapshot-keyed credit admitted both writers' decision batches and
their interleaved sets baked an order no replay derives (correlation
ordinals inverted against commit order at seqs the writers' fence_counts
prove they never saw).
The runtime now mints a random writerId per invocation delivery and sends
it with every precondition snapshot; the world's credit compares writerId
(falling back to stateCursor only for runtimes that don't send one). A
second writer with the same snapshot 412s and restarts against the
corrected log. Also adds the fence_count forensics column and a
WORKFLOW_POSTGRES_EVENT_FENCE=tail|decision experiment switch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The tail-based currency fence livelocked runs under steady inbound load:
every hook_received landing between a replay's load and its next write
412'd the write and forced a full replay restart, and the facts kept
coming (measured ~41 restarts/run in the step-storm repro; every attempt
timed out as stuck while zero corrupted).
Mirror Temporal's buffered-events model instead: facts (creates without a
precondition snapshot) get a commit-ordered position but never invalidate
a decision; only a foreign *decision* — a snapshot-carrying create from a
different snapshot — fences. The run row tracks lastFencedSeq (the last
decision's position); a fenced create is rejected iff a foreign decision
sits past its snapshot and the sibling credit doesn't match.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Events gain an optional dense per-run seq (EventSchema.seq) assigned at the
commit point. world-postgres serializes every append behind a run-scoped
advisory lock taken as the first lock of a single transaction that spans the
entity mutation and the event insert; event ids are re-minted inside that
critical section to dominate the run's tail, so seq order == id order ==
commit order == visibility order and a cursor reader can never skip a
late-committing event (the CORRUPTED_EVENT_LOG storage class).
Creates that carry the precondition snapshot (stateEventCount/stateCursor,
from the merged event-count-guard SDK machinery) are checked against a
transactional currency fence: a snapshot the log has moved past is rejected
with 412 and the whole transaction — entity mutation included — rolls back,
so a replay derived from a superseded view can never commit a decision.
Sibling creates of one suspension share a snapshot and are credited through
(writerSnapshot/writerBaseCount on the run row) instead of fencing each
other. The runtime asserts seq contiguity on every event load so any
remaining hole fails loudly instead of surfacing as a replay divergence
hundreds of events later.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The 12 tests that fail on main are right-reason failures documenting the
delivery-barrier idle collapse/starvation windows and the buffered-hook
claim-ordering race. Mark them it.fails so the suites merge as executable
documentation (the #3137 -> #3139 convention); the fix PR flips the
markers back to it. The 4 controls that pass on main keep plain it.
Also adds an empty changeset (test-only change).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
End-to-end reproduction on current main, from the event log alone — no
artificial hydration latency, no shared ReplayPayloadCache, failing at every
consumer hop count from 0 to 16 with the production error shape:
Replay divergence: step event step_created for step_<ULID> belongs to
"afterHook", but the current step consumer is "afterStep"
An unclaimed buffered hook payload makes resolvesOnItsOwn (private.ts:296-319)
report false for itself, and transitively for the armed wait barrier that
defers behind it. A later step result therefore skips BOTH via the
kind === 'step' escape at private.ts:368-371, computes an empty deferral set,
and resolves on microtasks while the wait is still parked — so the step branch
draws the ULID the log assigns to the hook branch.
Verified mechanism: neutralising that skip makes all six cases pass (and keeps
the 72 existing delivery-ordering assertions green), so the skip is the cause.
No existing suite covers this because every hook case in
step-delivery-ordering / step-delivery-hop-count / delivery-barrier-coverage
registers its awaiter before the drain, taking the armed path in hook.ts
rather than claim().
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
Four failing tests against the barrier primitives in private.ts:
- the idle safety net (private.ts:461) retires a barrier whose delivery is
still parked inside awaitEarlierDeliveries, because pendingDeliveries only
counts the hydration window (released at step.ts:322, before the detached
continuation at step.ts:324);
- one idle tick retires EVERY live barrier, not just abandoned ones;
- a later-in-log delivery consequently computes an empty deferral set, skips
the macrotask yield at private.ts:392, and is handed to workflow code before
an earlier-in-log one;
- the same inversion without the idle net at all: markDelivered() runs one
statement before resolve(), so the registry stops mentioning a delivery
while the branch it woke is still hops away from its next useStep.
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
* origin/peter/event-count-guard:
Derive the precondition watermark from the log's maximum, not its tail
Remove the temporary backend URL override
Temporarily point CI at the backend branch preview
Gate event creation on the loaded event count and restart replays in-process
* feat(nitro): embed observability dashboard in-process at /_workflow
Serve the @workflow/web observability UI inside the Nitro process at a
configurable route (default /_workflow) instead of spawning a separate
web server and 302-redirecting to it. Enabled in dev, omitted from
production builds by default (so prod bundles carry no @workflow/web
import). Never mounted on Vercel deploys (use the hosted dashboard).
- @workflow/web: add a framework-neutral `@workflow/web/handler`
(createWorkflowWebHandler) that serves SSR + static client assets +
RPC as one Web Request->Response handler under a runtime basename
(asset manifest URLs + publicPath are reprefixed so the dashboard is
self-contained under its mount). Add `@workflow/web/registry` for
embedded-dashboard discovery; make the RPC/stream client basename-aware.
- @workflow/nitro: mount the handler in-process (Nitro v2 h3 + v3 native
paths), gated by a new `dashboard` option (default = dev).
- @workflow/cli: `workflow web` / `inspect --web` defer to a running
embedded dashboard instead of starting a redundant server; pass
`--standalone` to force the standalone UI.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(nitro): normalize dashboard path once, use isNitroV2() helper
Address review feedback on the embedded dashboard:
- Normalize the dashboard mount path in one place before it feeds both
the Nitro route registration (`[path, path + '/**']`) and the handler
`basename`. Force a single leading slash, strip trailing slashes, and
reject the root mount, so a custom `path` can't make the route and the
handler's internal `normalizeBasename` disagree.
- Replace the handler-level `!nitro.routing` v2 checks with the existing
`isNitroV2()` helper for consistent v2/v3 detection.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
* fix(core): order step-result deliveries against wait/hook deliveries by event-log position
Two production runs on `@workflow/core@5.0.0-beta.36` burned all three
divergence-recovery replays at the same event and terminated with
CORRUPTED_EVENT_LOG:
wrun_41KYJENABV0GSF5YTE9EETV5DD (step vs wait)
wrun_41KYJEE01S0GPC9RWT5MEKVCX8 (step vs hook)
Replay divergence: step event step_created for step_X belongs to "A",
but the current step consumer is "B"
`useStep` proxies draw deterministic ULIDs in invocation order, so the
ULID -> stepName allocation is a function of the order in which promise
resolutions are delivered to workflow code. The delivery-barrier registry
pinned that order to event-log position for hook payloads and wait
completions, but step results were delivered straight off the serial
`promiseQueue` — and their latency varies between replays of the SAME
invocation, because the first replay pays full hydration while later
replays memo-hit primitive results in the shared `ReplayPayloadCache`.
A step completion adjacent in the log to a `wait_completed` was therefore
delivered wait-first on a cold replay and step-first on a warm one;
whichever order the invocation that wrote the follow-up `step_created`
events happened to see became law, and every replay computing the other
order diverged permanently.
Step results and step failures now register a 'step' delivery barrier at
their event-log index and resolve from a detached continuation after every
relevant earlier-in-log delivery, mirroring the hook payload path:
hydration stays inside the serial queue slot (which also releases
`pendingDeliveries`), while the barrier wait and the resolve run off the
queue so a queue slot never blocks on a resolution the queue itself drives.
Waits and hook payloads likewise defer behind earlier step results.
Two details are what actually make the ordering hold, and both were found
by testing rather than by reading the code:
The deferral set is captured while CONSUMING the event, not at the start of
the hydration slot. Captured at slot start it is not merely less
deterministic, it is usually empty: an earlier delivery whose own slot runs
first on the serial queue has typically already resolved and deregistered
its barrier before the later slot begins, so the later delivery does not
defer at all. Every event in one drain window is consumed before any slot
runs, so consumption time sees all of them.
A delivery that had to wait then yields a macrotask before resolving. An
earlier delivery being "delivered" only means its `resolve()` ran; the
branch it woke may need arbitrarily many further microtask hops before it
reaches its next `useStep` call (a `for await` over a hook resumes the
generator, settles the promise from `next()`, and only then runs the loop
body). Ordering the `resolve()` calls alone therefore buys a fixed hop or
two of margin and leaves a hop-count race that holds only for the shortest
consumers; yielding a macrotask lets the earlier branch drain completely,
whatever its shape.
One asymmetry is load-bearing: a step result skips any earlier delivery
that will not resolve on its own, i.e. one blocked directly or
transitively on a buffered hook payload no consumer has claimed. Such a
payload is delivered only when the workflow next reads the hook, and
reaching that read commonly requires the step result itself, so gating the
step on it stalls the run until the barrier's idle safety net fires — which
then releases every delivery queued behind that payload at once and loses
the very race the ordering exists to protect. Waits and hooks keep gating
on unclaimed payloads, where waiting for the claim IS the guarantee.
Tests come in two files. `step-delivery-ordering.test.ts` is byte-identical
to the file in the repro-only companion PR vercel/workflow#3137 apart from
two `it.fails` markers there (which let a repro-only branch have green CI);
`sed 's/it\.fails(/it(/g' | cmp` verifies it. Each of its five cases
replays one committed log twice through a shared `ReplayPayloadCache`, and
the two warm-replay cases fail on main with the production error text.
`step-delivery-hop-count.test.ts` exists because those five cases cannot
tell "delivered in log order" apart from "resolves a hop or two later than
before". It replays logs a live run legitimately produced — the live
invocation received the two events in separate deliveries, so the first
branch finished long before the second event existed — while the replay
receives both in one drain window, and pads the consumer with a varying
number of extra awaits so hop count is the only variable. It covers step
results against both wait completions and hook payloads, plus step
FAILURES against wait completions, since a rejection decides whether a
`catch` continuation runs and so which ULID the `useStep` there draws. All
18 cases fail on main; of the 12 that predate the macrotask, 9 still fail
with the resolve-ordering-only version of this fix; all 18 pass here.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
* fix(core): close remaining delivery-barrier ordering gaps
Follow-up on the step-delivery barrier work, addressing three cases the
registry did not yet cover. Each has a regression test in the new
`delivery-barrier-coverage.test.ts` that reproduces the production
`ReplayDivergenceError` when its fix is reverted.
- Step results now defer behind earlier STEP results. The old exclusion
assumed the serial `promiseQueue` fixes step-vs-step order, which stopped
holding once a step began resolving from a detached continuation instead
of its queue slot: two steps consumed in different drain windows can
disagree on their deferral set, and the earlier one — parked on the
macrotask yield — gets overtaken.
- `sleep.ts` and `hook.ts` (waiting-consumer path) now capture their
deferral at event-consumption time, as `step.ts` already does. Reading
the registry after their queue work misses an earlier step or hook that
delivered and retired its barrier in the meantime, skipping both the gate
and the macrotask yield. The buffered hook payload path deliberately
keeps evaluating at claim time; a consumption-time snapshot there stalls
the e2e `hookWithSleepWorkflow`.
- Abort deliveries participate in the registry. `_setAborted` fires the
signal's listeners, which may invoke a step and draw a ULID, so an abort
is as branch-deciding as any other delivery.
Also memoizes `resolvesOnItsOwn`. The walk is exponential in the number of
live hook/wait barriers, and the registry is not bounded — a fan-out of
`Promise.race([hook, sleep])` branches accumulates one barrier per branch
per kind (49 measured for 24 branches). At 40 barriers a single scan took
92s before, and is instant after.
---------
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Nathan Rajlich <n@n8.io>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
The Astro and SvelteKit webhook wrappers copied the incoming request via
`normalizeRequest()`, which buffers the whole body with `arrayBuffer()`,
before calling the handler that validates the webhook token. Requests
carrying an unknown token therefore did unnecessary work before being
rejected.
The copy turns out to be unnecessary: both frameworks already hand the
route a standard `Request`, so the webhook wrappers now pass it straight
to the handler. The body is left untouched until `resumeWebhook()` has
accepted the token.
The flow route keeps `normalizeRequest()` for now — it authenticates via
the queue trigger rather than a URL token, so the ordering does not
matter there, and whether the shim is needed at all is a separate
question.
Note that the Astro dev server buffers request bodies upstream of the
route handler, so the new behavior is only observable in built output;
the node adapter and Vercel builds both benefit.
The event-log merge no longer re-sorts by event id. A World's canonical
order is its own: world-vercel orders by event id, but world-local orders
by (createdAt, eventId) and deliberately re-mints keys (dominant-event and
claim canonicalization) so the two diverge. Re-sorting by event id there
produced an order no ordered load would ever return, reordering a terminal
event ahead of an accepted hook and breaking concurrent hook-token
arbitration.
The merge was only sorting so the snapshot could read its watermark off the
tail, so read the maximum ULID time across the log instead. That removes
the ordering dependency entirely and is exact rather than merely safe:
every loaded event is at or below the maximum, so stateEventCount is still
events.length whatever order the World returned.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Resolves the overlap with #3110, which introduced the same event-log merge
consolidation this branch had added as `mergeEvents`: `appendUniqueEvents`
now carries the optional id set from main plus the out-of-order re-sort and
warning, and `mergeEvents` is gone. Main's `withPreconditionRetry` edit drops
out with the function itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs: correct the workflow ID claim in publishing libraries
The consumer re-export file does not relocate a library's workflow and
step IDs into the consumer's source tree. An ID is derived from where the
file lives, so any export-reachable package file keeps a name@version ID
whether or not it is re-exported.
Rename the section to describe what the re-export actually does — put the
package's directive files on the compiler's discovery graph and give the
entry point a resolvable address — and add the upgrade guidance that
follows from the real behavior: a package version bump renames every
workflow and step it ships, so in-flight runs must drain first.
The wrong claim also appeared in the page summary and the CopyPrompt, so
it is corrected in all three places, in both the v4 and v5 copies.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
* docs: describe deployment pinning instead of drain guidance
Runs are pinned to the deployment that recorded their step IDs, so a
library version bump does not strand in-flight runs: new runs execute
the new version, in-flight runs keep replaying on their original
deployment. Replaces the incorrect drain-before-upgrade advice.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
* docs: qualify deployment pinning as world-dependent
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
---------
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The matching world-vercel guard shipped and is live in production, so the e2e
lanes exercise both halves against the default endpoint.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Hooks can carry an optional `resumeContext` mirrored from the run at
creation time. When present, `resumeHook`/`resumeWebhook` resume directly
from it instead of fetching the full run, saving a round trip per resume.
When the context also carries the run's `encryptionPublicKey`, the resume
seals its payload (`encp`) directly to that key. Combined with the sealed
envelope work (#3093-#3096), a default webhook resume then needs neither a
run read nor a cross-deployment run-key lookup: the key is resolved only
when the hook actually stores metadata that must be hydrated symmetrically.
Everything falls back transparently to the full run fetch and symmetric
key when the context (or the public key within it) is absent, so new
clients interoperate with old servers and vice versa.
- world: optional `encryptionPublicKey` on `HookResumeContext`
- world-postgres: `resume_context` column migration
- core: combined fast-path + seal in resume-hook; fast-path control-flow
suite split from the real-serialization crypto suite
- world-vercel: cover the `getEncryptionKeyForRun(runId, { deploymentId })`
overload the fast path relies on
- web-shared: render `resumeContext` in the attribute panel
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>