Commit Graph

1585 Commits

Author SHA1 Message Date
Pranay Prakash 94943d2f70 chore: harden the local storm harness and untrack its results checkpoint
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 13:47:55 -07:00
Pranay Prakash ef60c2a511 chore: remove temporary storm diagnostics and add fence changeset
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 12:44:17 -07:00
Pranay Prakash 02eed5bc5f feat(world): fence only decision writes
Add isDecisionEvent and gate the stateEventCount fence on it in
world-local and world-postgres: facts (completions, receipts, non-lazy
step_started claims) pass unfenced and never bump the fence even when
the runtime attaches a snapshot. A stale fact is byte-identical to a
fresh one and takes its meaning from its commit-assigned log position,
so fencing it prevents nothing and converts steady traffic into
replay-restart churn (measured: 776 wait_completed + 1107 step_started
rejections in one 48-run storm).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 12:39:13 -07:00
Pranay Prakash 50e607bb9a fix(world-local): keep directory-scan reads outside the append lock
The run_started preload and the step-terminal inline delta are
events-directory scans whose latency grows with the directory; holding
the per-run append serializer across them starved the run's other
writers and, under load, tripped the stale-lock breaker on a healthy
holder — minting duplicate positions. Attach both read pages after the
lock releases (their point-in-time semantics are unchanged), and err
the stale threshold far to the late side: breaking early corrupts,
breaking late merely stalls a crashed run's writers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 10:03:43 -07:00
Pranay Prakash a49b53ec1f fix(world-local): enforce the decision fence before any append side effect
A fenced create whose snapshot is stale was rejected at the publish
point — after entity writes, claim files, and the terminal marker had
already landed. Unlike postgres, those writes cannot roll back: a
rejected step_created stranded a step entity whose step executes with
no step_created in the log (an unreplayable completion), and a rejected
run_completed left the terminal marker behind, rejecting every later
hook resume. Admit or reject the session while the append is still a
pure no-op.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 17:32:22 -07:00
Pranay Prakash b39563f3ba feat(world-local): commit-ordered event positions with a cross-process append lock
Assign every event a dense per-run seq and a tail-dominant event key at
its publish point, under an on-disk per-run append serializer (the
filesystem analogue of world-postgres's advisory transaction lock), and
enforce the stateEventCount decision fence under the same lock. Extend
the SDK's contiguity assertion to the pre-replay convergence point so
the inline write-response delta and run_started preload are covered too.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 16:55:48 -07:00
Pranay Prakash b0e201caf3 Merge remote-tracking branch 'origin/pgp/replay-engine-determinism-tests' into pgp/event-seq
* origin/pgp/replay-engine-determinism-tests:
  fix(core): reset the idle poll's budget on delivery progress
  fix(core): do not start the abandon clock before a payload can be claimed
  fix(core): never retire a committed delivery on a timer, and count parked deliveries as in-flight
  fix(core): make delivery-barrier retirement deterministic
  test(core): mark known-failing determinism repros as expected failures
  test(core): reproduce delivery-barrier starvation under unrelated delivery traffic
  test(core): characterize divergence from a below-watermark event
  test(core): reproduce ReplayDivergenceError from a buffered hook payload
  test(core): reproduce delivery-barrier idle collapse and premature retirement
2026-07-30 16:08:52 -07:00
Pranay Prakash e1b5778974 Merge branch 'pgp/autoincrement-event-sort-ids' into pgp/event-seq
* pgp/autoincrement-event-sort-ids:
  Key the fence's sibling credit on a per-invocation writerId
  Fence decisions on foreign decisions, not on facts
  Assign commit-ordered event positions in world-postgres and fence stale writers
  Derive the precondition watermark from the log's maximum, not its tail
  Remove the temporary backend URL override
  Temporarily point CI at the backend branch preview
  Gate event creation on the loaded event count and restart replays in-process

# Conflicts:
#	packages/core/src/runtime/helpers.test.ts
#	packages/core/src/runtime/step-executor.ts
#	packages/world-vercel/src/events.ts
2026-07-30 16:08:45 -07:00
Nathan Rajlich 32ac8e73fd Fix Biome lint violations and add Biome CI check (#3222)
* Fix Biome lint violations and add Biome CI check

Biome was not configured to respect .gitignore, so ~92% of the 13,355
reported diagnostics came from gitignored build artifacts. Enable VCS
integration (useIgnoreFile), apply safe auto-fixes across the repo, fix
the remaining mechanical errors by hand, downgrade judgment-call a11y /
dangerouslySetInnerHTML rules to warnings, and add a 'biome ci' job to
the Lint workflow so violations block PRs going forward.

* Use an empty changeset (no behavior change, no release needed)
2026-07-30 22:32:12 +00:00
Alex Langenfeld 4a9d26b1cb feat(world): persist the compute instance that ran each step attempt (#3186)
* feat(world): persist the compute instance that ran each step attempt

Add CreateEventParams.computeInstanceId (ambient per-event identity, mirroring requestId) and a readable Event.computeInstanceId. Core stamps it on every step_started write; world-vercel forwards it in the v4 frame meta next to vercelId. Lets observability distinguish steps sharing a compute instance from those on different instances or invocations.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>

* test: cover computeInstanceId threading from params to v4 frame meta

world-vercel: computeInstanceId reaches the v4 frame meta, rides alongside vercelId rather than replacing it, and is omitted when unset. core: step_started carries it without displacing the stateUpdatedAt precondition guard (both share one params object).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>

* fix(web-shared): render computeInstanceId in the attribute panel

AttributeKey derives from keyof Event, so adding computeInstanceId to the event schema widened it and left the exhaustive attributeToDisplayFn map incomplete (TS2741). Renders it beside deploymentId as 'Compute Instance ID', copyable like the other opaque ids.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>

* fix(world-postgres): exclude computeInstanceId from the events column contract

The events table asserts satisfies DrizzlishOfType<...Omit<Event, 'occurredAt'>...>, so adding computeInstanceId to the event schema broke the build (TS1360). This world does not persist it, matching how occurredAt is already handled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>

* refactor: move computeInstanceId to the analytics read contract

The server routes computeInstanceId into ClickHouse and returns it on AnalyticsEvent/AnalyticsStep, never on the event record — so Event.computeInstanceId was dead on read and zod would strip the field off the analytics wire. Move it to AnalyticsEventSchema/AnalyticsStepSchema (beside vercelId/requestId, the same class of ambient provenance), which also drops the world-postgres column-contract exclusion entirely.

Also: hoist the duplicated step_started params into one local, extract the repeated mock-agent harness in events.test.ts, and use vi.spyOn plus an identity assertion against COMPUTE_INSTANCE_ID.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>

---------

Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 17:20:18 -05:00
Pranay Prakash b397eeafcf fix(core): reset the idle poll's budget on delivery progress
Counted flat, `IDLE_POLL_DEADLINE_ROUNDS` is not a deadline but a cap on
how many events one drain window may deliver.

Delivering a window in log order costs one macrotask per deferring
delivery, so a window of K events needs K poll rounds to clear. Past the
budget `scheduleWhenIdle` fires anyway, into the middle of the batch — and
for a hook consumer that callback raises a `WorkflowSuspension`, so the run
suspends carrying none of the work the remaining deliveries were about to
create. A new test in `delivery-barrier-idle-starvation.test.ts` puts 40
deliveries in one window and shows the callback landing after exactly 16 of
them.

The budget now restarts whenever a delivery reaches workflow code, tracked
as `deliveryProgress` and incremented by `markDelivered()` only. Excluding
the abandonment safety net's retirements is what keeps the escape hatch
working: a stream of pokes to a hook nobody reads retires barrier after
barrier and delivers nothing, so the budget still runs out and the
suspension still fires. That is the 3m19s stall the deadline was added for,
and its test still passes.

Ordering machinery that slows delivery down is only safe if the liveness
timers around it measure progress rather than elapsed rounds; this is the
one that did not.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
2026-07-30 15:12:56 -07:00
Pranay Prakash c8235f4958 fix(core): do not start the abandon clock before a payload can be claimed
An unarmed delivery barrier is retired on the inference that no consumer
has claimed the payload. That inference is not available until a consumer
could have: `claim()` cannot hand anything over until the payload's own
hydration resolves, so while it is still hydrating, "nobody has claimed
it" carries no information.

Left unguarded, the two retirement conditions conspire during that window.
The payload's own hydration holds `pendingDeliveries` above zero, so the
idle route cannot fire, and the abandon deadline retires the barrier
because its hydration was slow. Against a production hydration (object
storage fetch plus decrypt, 10-500ms) and a deadline of a few dozen
`setTimeout(0)` ticks, that is the common case for a buffered payload
rather than an edge one.

`registerDeliveryBarrier` now takes `abandonableAfter` and does not start
the retirement poll until it settles; `workflow/hook.ts` passes the
buffered payload's own hydration. The waiting-consumer path is unaffected:
it registers armed, and an armed barrier leaves the poll on its first
check either way.

This is the same slow-hydration trap b58fdcf0d had on the armed side,
found by re-auditing the deadline against production latency after that
commit's storm run. The regression test holds hydration open well past the
deadline and asserts both that the barrier survives and that a later step
result still orders behind the claim.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
2026-07-30 15:00:51 -07:00
Pranay Prakash 977cd332bd fix(core): never retire a committed delivery on a timer, and count parked deliveries as in-flight
Follow-up to the previous commit, from review of #3196 and from an e2e
regression that review work exposed.

Retirement by the safety net is now UNARMED-only. The abandon deadline used to
fire for any barrier, which reintroduced the collapse at a longer timescale:
production hydration (object-storage fetch plus decrypt) runs 10-500ms, far
past any tick budget worth setting, so the slowest deliveries would have been
exactly the ones force-retired mid-flight. Only an unclaimed buffered hook
payload — the one delivery that can be abandoned at the root — is retirable by
the net; every other barrier leaves the registry from its own delivery chain.
All five call sites were audited against that invariant: step_completed,
step_failed, wait_completed, the waiting-consumer and claim() hook paths, and
the abort hook all attach their chain unconditionally.

Deliveries parked in awaitEarlierDeliveries are now counted, and the count
gates scheduleWhenIdle. A delivery whose only remaining gate is a
recentlyDeliveredBarriers entry is invisible to both pendingDeliveries (already
released in its hydration slot) and to a scan of the live registry (already
empty), so a suspension armed in that window preempted it and the run suspended
carrying none of the work the delivery was about to create. That is how the
second payload of the hookWithSleepWorkflow e2e went missing. Guarded by a unit
test that fails without the counter.

Review fixes:

- Starvation suite: the poke storm held pendingDeliveries above zero with a
  single unmatched increment, so it modelled a stuck hydration rather than
  overlapping traffic. It now overlaps genuinely, leaks nothing, and asserts
  both properties. The header no longer attributes the observed 3m19s stall to
  this mechanism, which the suite does not establish.
- Added the liveness guard asked for in review: 228 unclaimed payloads gating a
  later delivery still drain in a handful of ticks.
- Added the mirrored-log control asked for in review: the interleaving where
  the step result legitimately precedes the wait completion must keep replaying
  cleanly, so the ordering assertions cannot be satisfied by flipping the bias.
- Below-watermark suite: the read shape is not a world-postgres class.
  world-vercel paginates DynamoDB by event ULID with an `eid:` cursor over a
  strictly-after range, and workflow-server documents its IDs as monotonic only
  intra-instance. Recorded in the header; the second test is renamed to say it
  characterizes current behavior rather than pinning a contract.
- Call-site comments that documented idle retirement of parked barriers as a
  deliberate tradeoff now describe what the net actually does.
- Stale private.ts / step.ts line references replaced with symbol names.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
2026-07-30 14:29:36 -07:00
Pranay Prakash 7c5c1d37c0 Merge remote-tracking branch 'origin/main' into pgp/autoincrement-event-sort-ids
* origin/main:
  [benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main (#3213)
  feat(core): emit faas.instance span attribute for compute instance identity (#2989)
  [world-local] Retry transient EPERM unlink failures on Windows (#3215)
  [world-vercel] Raise H2 receive windows on the events agent (#3212)
  Version Packages (beta) (#3185)
  Align Streams UI with trace viewer (#3197)
  fix(core): don't observe idle while a committed delivery is parked behind its deferral (#3198)
  [world-vercel] Make HTTP/2 actually multiplex on the events path (#3190)
  [RFC] feat(nitro): embed observability dashboard in-process at /_workflow (#2548)
2026-07-30 14:03:50 -07:00
Shalabh Chaturvedi 8bda7cef79 [benchmarks] Split STSO by inline vs queue-hop steps, add distribution diffs vs main (#3213)
* Split STSO by inline vs queue-hop steps, add distribution diffs vs main

The sequential-steps benchmark's STSO metric mixed two unrelated
phenomena: gaps between steps running back-to-back in the same warm
process, and gaps across an invocation boundary (queue dispatch, client
reinit, event-log replay), which cost ~10x more. The old step-index
windows (1-20 / 101-120 / 1001-1020) sampled 19 gaps each and captured
neither cleanly: whether a boundary happened to land inside a window
moved that window's P99 by hundreds of percent, which is most of the
run-to-run variance the benchmark comment was reporting.

The workflow now tags each step with whether it was the first step body
executed in its process ('queue-hop') or a later one in the same warm
process ('inline') via a process-global, so the split is ground truth
rather than inferred from step index or trace timestamps. STSO is
reported as two rows over *every* gap in the run instead of three
sampled windows. No targets on the new rows — the old ones described the
index-bucketed grouping.

computeStats now keeps the full sorted sample array alongside the
percentiles, and the comment renders a histogram + cumulative-time diff
against `main` under the table, one per STSO kind. Percentiles alone
hide how many samples moved and by how much, which is exactly where the
variance lives. Inline rows use a fixed 50ms bin width (the adaptive
width is coarse enough to hide structure inside that cluster); queue-hop
rows keep the adaptive width. Negative gaps (clock skew between two step
bodies' clocks) get their own bucket rather than being counted with the
slow tail.

Raw samples are stripped from the comment's embedded data block — ~1000
per run would exceed GitHub's comment size limit within a couple of
history entries — so the histogram renders for the current run only,
while collapsed history keeps its tables. Until this lands on `main` no
baseline has raw samples, so the section renders this run's distribution
as a single series.

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Clarify what stripping raw samples from the data block does not affect

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Drop the bucket tables; fold counts and deltas into the histogram bars

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Collapse the STSO distribution section into a dropdown

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Fix footer assertion after the dropdown wording change

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

---------

Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
2026-07-30 13:38:25 -07:00
Pranay Prakash b58fdcf0d7 fix(core): make delivery-barrier retirement deterministic
Three defects in the delivery-barrier registry could hand branch-deciding
deliveries to workflow code out of event-log order, surfacing as
CORRUPTED_EVENT_LOG on replay. All three are fixed here, and the twelve
expected-fail repros that documented them are flipped back to plain `it`.

Barriers no longer retire on a global idle tick. #3198 stopped the idle CHECK
from observing idle while a committed delivery is parked; this stops the same
predicate from retiring the barriers themselves. `pendingDeliveries` tracks
only the host-side hydration window — released inside the promiseQueue slot,
before the detached continuation that hands the value over, and never touched
at all by a wait_completed — so one idle tick used to retire every live
barrier while their deliveries were still parked in awaitEarlierDeliveries.
Retirement moves to its own poll, which retires only an UNARMED barrier: an
unclaimed buffered hook payload, the one delivery that can be abandoned at the
root. Every other kind is committed and retires from its own chain.

Both polls get a deadline. Without one neither has any escape from continuous
unrelated traffic: an abandoned barrier under a stream of deliveries to a
never-read hook starved indefinitely, observed live as a run making no
progress for 3m19s across 228 pokes. The barrier deadline is counted in raw
ticks, so the traffic keeping the system busy cannot stretch it;
scheduleWhenIdle's is counted in poll rounds, each of which waits out a full
promiseQueue drain, because that function also schedules suspensions, where
firing early preempts data delivery.

A retired delivery stays visible to the registry for one more macrotask.
markDelivered() resolved a barrier one statement before the resolve() that
wakes the branch, so anything reading the registry in between — a delivery
consumed in a later drain window, or a buffered payload's claim() — computed
an empty deferral set and overtook the branch it was meant to follow. The live
registry still drops the entry immediately; a second short-lived map carries
ordering visibility across the gap.

A step result's buffered-hook skip is narrowed to the payload itself. A step
still never gates on an unclaimed payload, where the claim commonly sits
downstream of the step result, but it now gates on an earlier armed wait or
hook whatever that delivery is itself waiting for. Skipping those as well, via
the transitive resolvesOnItsOwn walk, was the buffered-claim corruption: in
Promise.all([step, sleep-then-read-hook]) the wait completion is precisely what
wakes the branch that goes on to claim the payload. awaitEarlierDeliveries no
longer consults that walk, leaving the per-delivery path a single linear pass;
the walk itself survives only in hasParkedCommittedDelivery, which gates
suspensions rather than deliveries.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
2026-07-30 13:36:48 -07:00
Alex Langenfeld 9cc11f5329 feat(core): emit faas.instance span attribute for compute instance identity (#2989)
* feat(core): emit faas.instance span attribute for compute instance identity

Synthesize a per-warm-instance id (cinst_<ulid>) once at module load and emit it as the OTEL faas.instance attribute on the flow and step route spans. Vercel exposes no native per-instance id under Fluid compute, so this lets traces distinguish which compute instance handled each request. Purely additive telemetry.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>

* test(core): assert faas.instance is stable across invocations

Covers the id format and, on the two-invocation warm-handler case, that both invocations report the same id — the module-scope minting contract that makes the attribute identify the instance rather than the invocation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>

---------

Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-30 15:19:04 -05:00
Andrew Barba f05f642e89 [world-local] Retry transient EPERM unlink failures on Windows (#3215)
deleteJSON was the one mutation path in the fs layer that neither used
withWindowsRetry nor swallowed unlink errors. On Windows, unlink fails
with a share-violation EPERM while a concurrent reader briefly holds
the file open — hook polling races deleteAllHooksForRun by design — so
a transient EPERM surfaced as a failed operation, e.g. a failed
run.cancel().

Wrap the unlink in withWindowsRetry, matching the rename/link/unlink
guards the write pipeline already has. ENOENT is not in the retryable
set, so the already-deleted tolerance still short-circuits.

Signed-off-by: Andrew Barba <barba@hey.com>
2026-07-30 12:23:20 -07:00
Peter Wielander c93f6f7bd0 [world-vercel] Raise H2 receive windows on the events agent (#3212) 2026-07-30 11:54:18 -07:00
Peter Wielander dc0d25eaf2 Merge branch 'main' into pgp/replay-engine-determinism-tests 2026-07-30 09:37:45 -07:00
github-actions[bot] b12f248b66 Version Packages (beta) (#3185)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
workflow@5.0.0-beta.38
2026-07-30 08:40:06 -07:00
Mitul Shah e181f64b72 Align Streams UI with trace viewer (#3197)
* Align streams UI with trace viewer

Co-authored-by: Cursor <cursoragent@cursor.com>

* cleanupp

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-30 08:00:45 -07:00
Pranay Prakash 3064b8f13d Key the fence's sibling credit on a per-invocation writerId
Storm forensics (fence_count audit column) attributed every corrupted run
to a credit hole, not engine nondeterminism: two invocations that loaded
the identical prefix present byte-identical stateCursor+stateEventCount,
so the snapshot-keyed credit admitted both writers' decision batches and
their interleaved sets baked an order no replay derives (correlation
ordinals inverted against commit order at seqs the writers' fence_counts
prove they never saw).

The runtime now mints a random writerId per invocation delivery and sends
it with every precondition snapshot; the world's credit compares writerId
(falling back to stateCursor only for runtimes that don't send one). A
second writer with the same snapshot 412s and restarts against the
corrected log. Also adds the fence_count forensics column and a
WORKFLOW_POSTGRES_EVENT_FENCE=tail|decision experiment switch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 17:10:46 -07:00
Pranay Prakash cd6b2df5a8 Fence decisions on foreign decisions, not on facts
The tail-based currency fence livelocked runs under steady inbound load:
every hook_received landing between a replay's load and its next write
412'd the write and forced a full replay restart, and the facts kept
coming (measured ~41 restarts/run in the step-storm repro; every attempt
timed out as stuck while zero corrupted).

Mirror Temporal's buffered-events model instead: facts (creates without a
precondition snapshot) get a commit-ordered position but never invalidate
a decision; only a foreign *decision* — a snapshot-carrying create from a
different snapshot — fences. The run row tracks lastFencedSeq (the last
decision's position); a fenced create is rejected iff a foreign decision
sits past its snapshot and the sibling credit doesn't match.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 16:27:08 -07:00
Nathan Rajlich b92c23ccb4 fix(core): don't observe idle while a committed delivery is parked behind its deferral (#3198) 2026-07-29 16:12:45 -07:00
Pranay Prakash 5849388161 Assign commit-ordered event positions in world-postgres and fence stale writers
Events gain an optional dense per-run seq (EventSchema.seq) assigned at the
commit point. world-postgres serializes every append behind a run-scoped
advisory lock taken as the first lock of a single transaction that spans the
entity mutation and the event insert; event ids are re-minted inside that
critical section to dominate the run's tail, so seq order == id order ==
commit order == visibility order and a cursor reader can never skip a
late-committing event (the CORRUPTED_EVENT_LOG storage class).

Creates that carry the precondition snapshot (stateEventCount/stateCursor,
from the merged event-count-guard SDK machinery) are checked against a
transactional currency fence: a snapshot the log has moved past is rejected
with 412 and the whole transaction — entity mutation included — rolls back,
so a replay derived from a superseded view can never commit a decision.
Sibling creates of one suspension share a snapshot and are credited through
(writerSnapshot/writerBaseCount on the run row) instead of fencing each
other. The runtime asserts seq contiguity on every event load so any
remaining hole fails loudly instead of surfacing as a replay divergence
hundreds of events later.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 15:11:18 -07:00
Pranay Prakash 51de58490c test(core): mark known-failing determinism repros as expected failures
The 12 tests that fail on main are right-reason failures documenting the
delivery-barrier idle collapse/starvation windows and the buffered-hook
claim-ordering race. Mark them it.fails so the suites merge as executable
documentation (the #3137 -> #3139 convention); the fix PR flips the
markers back to it. The 4 controls that pass on main keep plain it.

Also adds an empty changeset (test-only change).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
2026-07-29 14:39:55 -07:00
Pranay Prakash fe58cf0f0a test(core): reproduce delivery-barrier starvation under unrelated delivery traffic
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
2026-07-29 14:39:54 -07:00
Pranay Prakash a7e0cd8084 test(core): characterize divergence from a below-watermark event
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
2026-07-29 14:39:53 -07:00
Pranay Prakash 8274a57a55 test(core): reproduce ReplayDivergenceError from a buffered hook payload
End-to-end reproduction on current main, from the event log alone — no
artificial hydration latency, no shared ReplayPayloadCache, failing at every
consumer hop count from 0 to 16 with the production error shape:

  Replay divergence: step event step_created for step_<ULID> belongs to
  "afterHook", but the current step consumer is "afterStep"

An unclaimed buffered hook payload makes resolvesOnItsOwn (private.ts:296-319)
report false for itself, and transitively for the armed wait barrier that
defers behind it. A later step result therefore skips BOTH via the
kind === 'step' escape at private.ts:368-371, computes an empty deferral set,
and resolves on microtasks while the wait is still parked — so the step branch
draws the ULID the log assigns to the hook branch.

Verified mechanism: neutralising that skip makes all six cases pass (and keeps
the 72 existing delivery-ordering assertions green), so the skip is the cause.
No existing suite covers this because every hook case in
step-delivery-ordering / step-delivery-hop-count / delivery-barrier-coverage
registers its awaiter before the drain, taking the armed path in hook.ts
rather than claim().

Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
2026-07-29 14:39:52 -07:00
Pranay Prakash 0de98e1816 test(core): reproduce delivery-barrier idle collapse and premature retirement
Four failing tests against the barrier primitives in private.ts:

- the idle safety net (private.ts:461) retires a barrier whose delivery is
  still parked inside awaitEarlierDeliveries, because pendingDeliveries only
  counts the hydration window (released at step.ts:322, before the detached
  continuation at step.ts:324);
- one idle tick retires EVERY live barrier, not just abandoned ones;
- a later-in-log delivery consequently computes an empty deferral set, skips
  the macrotask yield at private.ts:392, and is handed to workflow code before
  an earlier-in-log one;
- the same inversion without the idle net at all: markDelivered() runs one
  statement before resolve(), so the registry stops mentioning a delivery
  while the branch it woke is still hops away from its next useStep.

Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
2026-07-29 14:39:52 -07:00
Peter Wielander 34975f6b7d [world-vercel] Make HTTP/2 actually multiplex on the events path (#3190) 2026-07-29 14:38:17 -07:00
Pranay Prakash b796d42f4e Merge remote-tracking branch 'origin/peter/event-count-guard' into pgp/autoincrement-event-sort-ids
* origin/peter/event-count-guard:
  Derive the precondition watermark from the log's maximum, not its tail
  Remove the temporary backend URL override
  Temporarily point CI at the backend branch preview
  Gate event creation on the loaded event count and restart replays in-process
2026-07-29 14:14:01 -07:00
Pranay Prakash 25715d4521 [RFC] feat(nitro): embed observability dashboard in-process at /_workflow (#2548)
* feat(nitro): embed observability dashboard in-process at /_workflow

Serve the @workflow/web observability UI inside the Nitro process at a
configurable route (default /_workflow) instead of spawning a separate
web server and 302-redirecting to it. Enabled in dev, omitted from
production builds by default (so prod bundles carry no @workflow/web
import). Never mounted on Vercel deploys (use the hosted dashboard).

- @workflow/web: add a framework-neutral `@workflow/web/handler`
  (createWorkflowWebHandler) that serves SSR + static client assets +
  RPC as one Web Request->Response handler under a runtime basename
  (asset manifest URLs + publicPath are reprefixed so the dashboard is
  self-contained under its mount). Add `@workflow/web/registry` for
  embedded-dashboard discovery; make the RPC/stream client basename-aware.
- @workflow/nitro: mount the handler in-process (Nitro v2 h3 + v3 native
  paths), gated by a new `dashboard` option (default = dev).
- @workflow/cli: `workflow web` / `inspect --web` defer to a running
  embedded dashboard instead of starting a redundant server; pass
  `--standalone` to force the standalone UI.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(nitro): normalize dashboard path once, use isNitroV2() helper

Address review feedback on the embedded dashboard:

- Normalize the dashboard mount path in one place before it feeds both
  the Nitro route registration (`[path, path + '/**']`) and the handler
  `basename`. Force a single leading slash, strip trailing slashes, and
  reject the root mount, so a custom `path` can't make the route and the
  handler's internal `normalizeBasename` disagree.
- Replace the handler-level `!nitro.routing` v2 checks with the existing
  `isNitroV2()` helper for consistent v2/v3 detection.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
2026-07-29 13:58:59 -07:00
Douglas Harcourt Parsons e8934ade9c Remove self-attribution from README (#3182)
Signed-off-by: Douglas Harcourt Parsons <dglsparsons@users.noreply.github.com>
2026-07-29 10:28:11 -07:00
Peter Wielander a09d00135b Revert "Statically inject workflow world target" (#2752) (#3142) 2026-07-29 08:55:29 -07:00
github-actions[bot] 741a0d9eaf Version Packages (beta) (#3087)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
workflow@5.0.0-beta.37
2026-07-28 17:20:21 -07:00
Pranay Prakash 2941b1c360 fix(core): order step-result deliveries against wait/hook deliveries by event-log position (#3139)
* fix(core): order step-result deliveries against wait/hook deliveries by event-log position

Two production runs on `@workflow/core@5.0.0-beta.36` burned all three
divergence-recovery replays at the same event and terminated with
CORRUPTED_EVENT_LOG:

  wrun_41KYJENABV0GSF5YTE9EETV5DD  (step vs wait)
  wrun_41KYJEE01S0GPC9RWT5MEKVCX8  (step vs hook)

  Replay divergence: step event step_created for step_X belongs to "A",
  but the current step consumer is "B"

`useStep` proxies draw deterministic ULIDs in invocation order, so the
ULID -> stepName allocation is a function of the order in which promise
resolutions are delivered to workflow code. The delivery-barrier registry
pinned that order to event-log position for hook payloads and wait
completions, but step results were delivered straight off the serial
`promiseQueue` — and their latency varies between replays of the SAME
invocation, because the first replay pays full hydration while later
replays memo-hit primitive results in the shared `ReplayPayloadCache`.
A step completion adjacent in the log to a `wait_completed` was therefore
delivered wait-first on a cold replay and step-first on a warm one;
whichever order the invocation that wrote the follow-up `step_created`
events happened to see became law, and every replay computing the other
order diverged permanently.

Step results and step failures now register a 'step' delivery barrier at
their event-log index and resolve from a detached continuation after every
relevant earlier-in-log delivery, mirroring the hook payload path:
hydration stays inside the serial queue slot (which also releases
`pendingDeliveries`), while the barrier wait and the resolve run off the
queue so a queue slot never blocks on a resolution the queue itself drives.
Waits and hook payloads likewise defer behind earlier step results.

Two details are what actually make the ordering hold, and both were found
by testing rather than by reading the code:

The deferral set is captured while CONSUMING the event, not at the start of
the hydration slot. Captured at slot start it is not merely less
deterministic, it is usually empty: an earlier delivery whose own slot runs
first on the serial queue has typically already resolved and deregistered
its barrier before the later slot begins, so the later delivery does not
defer at all. Every event in one drain window is consumed before any slot
runs, so consumption time sees all of them.

A delivery that had to wait then yields a macrotask before resolving. An
earlier delivery being "delivered" only means its `resolve()` ran; the
branch it woke may need arbitrarily many further microtask hops before it
reaches its next `useStep` call (a `for await` over a hook resumes the
generator, settles the promise from `next()`, and only then runs the loop
body). Ordering the `resolve()` calls alone therefore buys a fixed hop or
two of margin and leaves a hop-count race that holds only for the shortest
consumers; yielding a macrotask lets the earlier branch drain completely,
whatever its shape.

One asymmetry is load-bearing: a step result skips any earlier delivery
that will not resolve on its own, i.e. one blocked directly or
transitively on a buffered hook payload no consumer has claimed. Such a
payload is delivered only when the workflow next reads the hook, and
reaching that read commonly requires the step result itself, so gating the
step on it stalls the run until the barrier's idle safety net fires — which
then releases every delivery queued behind that payload at once and loses
the very race the ordering exists to protect. Waits and hooks keep gating
on unclaimed payloads, where waiting for the claim IS the guarantee.

Tests come in two files. `step-delivery-ordering.test.ts` is byte-identical
to the file in the repro-only companion PR vercel/workflow#3137 apart from
two `it.fails` markers there (which let a repro-only branch have green CI);
`sed 's/it\.fails(/it(/g' | cmp` verifies it. Each of its five cases
replays one committed log twice through a shared `ReplayPayloadCache`, and
the two warm-replay cases fail on main with the production error text.

`step-delivery-hop-count.test.ts` exists because those five cases cannot
tell "delivered in log order" apart from "resolves a hop or two later than
before". It replays logs a live run legitimately produced — the live
invocation received the two events in separate deliveries, so the first
branch finished long before the second event existed — while the replay
receives both in one drain window, and pads the consumer with a varying
number of extra awaits so hop count is the only variable. It covers step
results against both wait completions and hook payloads, plus step
FAILURES against wait completions, since a rejection decides whether a
`catch` continuation runs and so which ULID the `useStep` there draws. All
18 cases fail on main; of the 12 that predate the macrotask, 9 still fail
with the resolve-ordering-only version of this fix; all 18 pass here.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>

* fix(core): close remaining delivery-barrier ordering gaps

Follow-up on the step-delivery barrier work, addressing three cases the
registry did not yet cover. Each has a regression test in the new
`delivery-barrier-coverage.test.ts` that reproduces the production
`ReplayDivergenceError` when its fix is reverted.

- Step results now defer behind earlier STEP results. The old exclusion
  assumed the serial `promiseQueue` fixes step-vs-step order, which stopped
  holding once a step began resolving from a detached continuation instead
  of its queue slot: two steps consumed in different drain windows can
  disagree on their deferral set, and the earlier one — parked on the
  macrotask yield — gets overtaken.

- `sleep.ts` and `hook.ts` (waiting-consumer path) now capture their
  deferral at event-consumption time, as `step.ts` already does. Reading
  the registry after their queue work misses an earlier step or hook that
  delivered and retired its barrier in the meantime, skipping both the gate
  and the macrotask yield. The buffered hook payload path deliberately
  keeps evaluating at claim time; a consumption-time snapshot there stalls
  the e2e `hookWithSleepWorkflow`.

- Abort deliveries participate in the registry. `_setAborted` fires the
  signal's listeners, which may invoke a step and draw a ULID, so an abort
  is as branch-deciding as any other delivery.

Also memoizes `resolvesOnItsOwn`. The walk is exponential in the number of
live hook/wait barriers, and the registry is not bounded — a fan-out of
`Promise.race([hook, sleep])` branches accumulates one barrier per branch
per kind (49 measured for 24 branches). At 40 barriers a single scan took
92s before, and is instant after.

---------

Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Nathan Rajlich <n@n8.io>
Co-authored-by: Peter Wielander <mittgfu@gmail.com>
2026-07-28 15:54:50 -07:00
Nathan Rajlich 0d6ec43877 Pass the webhook request through without buffering its body (#3166)
The Astro and SvelteKit webhook wrappers copied the incoming request via
`normalizeRequest()`, which buffers the whole body with `arrayBuffer()`,
before calling the handler that validates the webhook token. Requests
carrying an unknown token therefore did unnecessary work before being
rejected.

The copy turns out to be unnecessary: both frameworks already hand the
route a standard `Request`, so the webhook wrappers now pass it straight
to the handler. The body is left untouched until `resumeWebhook()` has
accepted the token.

The flow route keeps `normalizeRequest()` for now — it authenticates via
the queue trigger rather than a URL token, so the ordering does not
matter there, and whether the shim is needed at all is a separate
question.

Note that the Astro dev server buffers request bodies upstream of the
route handler, so the new behavior is only observable in built output;
the node adapter and Vercel builds both benefit.
2026-07-28 15:08:33 -07:00
Peter Wielander e81e5305d3 Derive the precondition watermark from the log's maximum, not its tail
The event-log merge no longer re-sorts by event id. A World's canonical
order is its own: world-vercel orders by event id, but world-local orders
by (createdAt, eventId) and deliberately re-mints keys (dominant-event and
claim canonicalization) so the two diverge. Re-sorting by event id there
produced an order no ordered load would ever return, reordering a terminal
event ahead of an accepted hook and breaking concurrent hook-token
arbitration.

The merge was only sorting so the snapshot could read its watermark off the
tail, so read the maximum ULID time across the log instead. That removes
the ordering dependency entirely and is exact rather than merely safe:
every loaded event is at or below the maximum, so stateEventCount is still
events.length whatever order the World returned.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-28 13:35:19 -07:00
Peter Wielander 55c3746d28 Merge origin/main into peter/event-count-guard
Resolves the overlap with #3110, which introduced the same event-log merge
consolidation this branch had added as `mergeEvents`: `appendUniqueEvents`
now carries the optional id set from main plus the out-of-order re-sort and
warning, and `mergeEvents` is gone. Main's `withPreconditionRetry` edit drops
out with the function itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-28 12:55:29 -07:00
Nathan Colosimo 5d17c609b1 Reuse one event-log deduplication helper (#3110)
* refactor(core): reuse event log merge helper

* refactor(core): simplify event merge helper
2026-07-28 12:42:17 -07:00
Nathan Colosimo fba26fd9bf Correct step registration documentation (#3129) 2026-07-28 18:57:53 +00:00
Pranay Prakash 7d7effd49f docs: correct the workflow ID claim in publishing libraries (#3153)
* docs: correct the workflow ID claim in publishing libraries

The consumer re-export file does not relocate a library's workflow and
step IDs into the consumer's source tree. An ID is derived from where the
file lives, so any export-reachable package file keeps a name@version ID
whether or not it is re-exported.

Rename the section to describe what the re-export actually does — put the
package's directive files on the compiler's discovery graph and give the
entry point a resolvable address — and add the upgrade guidance that
follows from the real behavior: a package version bump renames every
workflow and step it ships, so in-flight runs must drain first.

The wrong claim also appeared in the page summary and the CopyPrompt, so
it is corrected in all three places, in both the v4 and v5 copies.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>

* docs: describe deployment pinning instead of drain guidance

Runs are pinned to the deployment that recorded their step IDs, so a
library version bump does not strand in-flight runs: new runs execute
the new version, in-flight runs keep replaying on their original
deployment. Replaces the incorrect drain-before-upgrade advice.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>

* docs: qualify deployment pinning as world-dependent

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>

---------

Signed-off-by: Pranay Prakash <pranay.gp@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 11:42:15 -07:00
Nathan Colosimo 7959acc8cf Remove deprecated setAttributes aliases (#3128) 2026-07-28 18:36:22 +00:00
Peter Wielander f00518b086 Remove the temporary backend URL override
The matching world-vercel guard shipped and is live in production, so the e2e
lanes exercise both halves against the default endpoint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-28 11:31:19 -07:00
Peter Wielander 49276f2d0b [utils] Fix vercel world not being selected when running build on external CI (#3144) 2026-07-28 11:26:30 -07:00
Karthik Kalyan d24c91cfde feat(core): resume hooks from stored resumeContext and seal to the run key (#3125)
Hooks can carry an optional `resumeContext` mirrored from the run at
creation time. When present, `resumeHook`/`resumeWebhook` resume directly
from it instead of fetching the full run, saving a round trip per resume.

When the context also carries the run's `encryptionPublicKey`, the resume
seals its payload (`encp`) directly to that key. Combined with the sealed
envelope work (#3093-#3096), a default webhook resume then needs neither a
run read nor a cross-deployment run-key lookup: the key is resolved only
when the hook actually stores metadata that must be hydrated symmetrically.

Everything falls back transparently to the full run fetch and symmetric
key when the context (or the public key within it) is absent, so new
clients interoperate with old servers and vice versa.

- world: optional `encryptionPublicKey` on `HookResumeContext`
- world-postgres: `resume_context` column migration
- core: combined fast-path + seal in resume-hook; fast-path control-flow
  suite split from the real-serialization crypto suite
- world-vercel: cover the `getEncryptionKeyForRun(runId, { deploymentId })`
  overload the fast path relies on
- web-shared: render `resumeContext` in the attribute panel

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 09:55:26 -07:00
Peter Wielander 0dabdf893e Merge branch 'main' into peter/event-count-guard 2026-07-28 08:33:42 -07:00
Peter Wielander 62c01d94b0 [e2e] Report partial results when the event-log race repro is cut short (#3148) 2026-07-28 08:32:59 -07:00