Files
vercel__workflow/packages/world-sim/DESIGN.md
Shalabh Chaturvedi b7591fbf21 sim-world - deterministic scenario testing for race conditions (#3328)
* Add world-sim: a deterministic simulation World for concurrency scenarios

`@workflow/world-sim` is an in-process World implementation whose point is
that nothing in it races. Every world call is a stoppable point with
`before`/`positioned`/`after` phases, every call is attributed to a writer
(orchestrator, a named step body, or an out-of-band external client), and
the clock is virtual — a thirty-day sleep costs no wall time. A scenario
script says "stop this writer here, do that, let it go", so an interleaving
that a real deployment leaves to chance becomes something you can name.

The model follows workflow-server where it matters:

- Event ids are minted at the handler boundary, not at storage append.
  DynamoDB does not generate ids, so the server does, in the request
  handler — and that id is the log's sort key. This is what makes "earlier
  log position, later commit" expressible, and it is the hazard the fence
  scenarios are about.
- The out-of-band write marker is keyed on the event's own ULID time and is
  forward-only, over `hook_received`, `step_completed` and `step_failed`.
- Both halves of the staleness fence are modelled behind flags: the
  watermark (`preconditionGuard`) that clients send today, and the count
  (`countGuard`) that they do not. The count is synthesized on the caller's
  behalf and keyed by run rather than by writer, since an orchestrator and
  its inline step bodies are one process sharing one loaded log.

`workbench/sim-world` is the scenario book — 39 of them, each a workflow
plus a script. Several are pinned corruptions rather than passing
assertions: they record what the runtime does today, so that a change in
behaviour shows up as a diff. The doc-29/30/31 trio is the argument for the
count guard, one flag apart: (B, A, C) corrupts under the watermark alone,
is fenced once the count is on, and (B, C, A) corrupts with both on. That
last one suspends mid-run on purpose. The fence is a conditional append
evaluated inside the storage write, so there is no checked-but-uncommitted
moment to slip past; the window nothing can close is the quiescent gap
between deliveries, where the run makes no writes and so meets no checks.

Both packages are private and unpublished, so no changeset.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Add the new sim-world workspaces to the lockfile

Kept separate from the source commit because it is not a clean diff. The
two new importers are the part that belongs to this branch; the rest is
re-resolution churn from running `pnpm install` at a later date than
whoever last touched the file — `latest` specifiers like docs' `radix-ui`
move on their own. Drop or regenerate this commit if the churn is
unwelcome; the source commit stands on its own.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Make a consistency violation fail the scenario that tripped it

`expect.violations` let a scenario declare the corruption it reproduced and
pass for reproducing it. That is a green suite describing a broken system,
and it goes red on the day someone *fixes* the bug — backwards, and the
opposite of what a test is for.

The field is gone. A scenario now states the outcome the run should have
reached, which for a corruption means the branch its own durable log
implies, and stays red until the runtime gets there. Any violation fails the
scenario. The six reproductions read as ordinary failures now:

  expected output "afterSlow:doc-26", got "afterFast:doc-26"

Six scenarios are therefore red, and `run.ts` exits non-zero. The count is
the signal: seven is a regression, five means something was fixed and a
scenario is ready to retire. Five of the six have known fixes — four predate
the count guard, and doc-29 goes green the moment a client sends
`stateEventCount`. doc-31 has none, which is the point of it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Document the world-sim implementation in DESIGN.md

Folds the working notes into a checked-in design doc: module map, the
runtime model the simulator has to match, the interception model and its
three call phases, the determinism machinery, the store's guards and
fault injection, the writer vocabulary, the termination budgets, both
consistency checkers, and current test status.

Links it from both READMEs, and updates the workbench's doc-31 note now
that the append-tail fence it needs exists as a proposal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Say which fix closes each red scenario, and which are demonstrated

The fix column read "predates the count guard" on three rows, which scans as
"the count guard is involved" when it meant the opposite. It was also wrong on
one: `two racing STEPS, no hook anywhere` cannot be fixed by anything predating
the count guard, since the row below it is that same fault with
`preconditionGuard` on and failing identically.

The column now names the specific change, and a second column separates a fix
that is argued for from one that is shown — a passing scenario that is the red
one with the fix armed, same tempo, one flag apart. Two of the six have that;
three name a fix with no paired scenario yet; doc-31 has none.

Also states the thing the column could imply but does not mean: `countGuard`
requires `stateEventCount`, which no client sends, so three of the five
identified fixes are dark in production.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Give scenarios stable ids, refer to events by log position, add colour

Four changes to make a rendered scenario something you can cite and read.

Scenario ids. Every scenario gets a hyphenated `id` next to its prose name —
`in-flight-after-decision`, `stale-read-step-count-fork`. The name is a
sentence and will be reworded; the id is what a commit message, a bug report
or `pnpm sim <id>` refers to. The runner matches the id first and falls back
to the name, so existing invocations keep working, and the table of the six
red scenarios in DESIGN.md now cites ids.

One reference scheme. Events were referred to three ways at once: the trace
printed raw correlation ULIDs, violation messages carried `evnt_<ulid>` event
ids, and the docs spoke of "position 7". Now there is only log position. `#12`
is the twelfth event in the durable log sorted the way `events.list` sorts it,
`@7` is the resource created at position 7, and ids inside violation messages
are rewritten on the way out.

That numbering is deliberately in *log* order while the trace prints in
*commit* order, which makes the subject of the red scenarios visible on the
page: a run whose log disagrees with the order its writers committed in shows
positions counting backwards.

Colour, by event family, and only when the destination is a terminal. Off
under `NO_COLOR` or `--no-color`, forced on with `--color`. With colour off
the output is the same plain ASCII, so it stays usable as a golden file.

The workflows under test are now one file. Nothing imported across them and
scenarios read them alongside the tempo that steers them, so five modules were
a two-file hop for no benefit.

Unit tests 61/61. Scenario book unchanged at 33 passed, 6 failed.

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Add an append-only log world, and split the book into one file per scenario

Three things the book could not do before.

`appendOnlyLog` gives an event its log position at *commit* instead of when
its handler minted one. That is the single change that makes a stale read
impossible — the log can be behind, never wrong — so playing the same 39
scenarios with and without it is how you tell which of the six reds the change
would actually close. All six: 33 pass / 6 violations mint-ordered, 39 pass /
0 violations append-only. It bundles two effects that have to move together:
an overtaken write re-mints at the tail, and a withheld read returns a prefix
rather than a hole, which is why `onStaleRead` now carries `{eventId, hidden,
truncated}` and the trace distinguishes a lagging read from a stale one.

`preconditionGuard` on `RunScenarioOptions` forces the fence off across the
book, asking whether anything relies on it. Violations 6 -> 8 mint-ordered, so
it is load-bearing there; 0 -> 0 append-only, so it is dead weight once
positions are assigned at commit. Both flags are tri-state: `undefined` leaves
it to the spec, which is not the same as `false`.

The scenario book was a 1020-line file; it is now 39 files and an index that
only decides reading order. Same 39 ids. This is the whole answer to "how do I
add a scenario" — copy the file next door — and it is why the README can be
short. Three `in-flight-*` scenarios are reworded so that no expectation is
restated per world: a scenario is one sequence of movements, the only thing a
world changes is what a read returns, and what catches the fault in both is
the invariant that a run's log must replay back into that run.

Also: `loadFlowHandler` moves out of `build.ts` into `load.ts`, so playing
scenarios no longer drags SWC and esbuild into the module graph; the CLI gains
`--report-only`, `--summary-file` and `--detail-file` for CI, with the default
still exiting non-zero; and the workbench's `test` script points at
`--report-only` so a recursive `pnpm -r test` does not go red for the six.

Docs are split by task: the workbench README is how to add a scenario, and the
package README plus DESIGN.md are how to change the simulator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Publish the scenario book from CI, without blocking a merge

Plays the book on every pull request, once per log world, and posts both
summaries as one sticky comment.

It never gates. Six scenarios are red on purpose — each is a reproduction of a
corruption the runtime can still produce, stating the outcome its own durable
log implies and staying red until the runtime gets there — so a lane that
failed on them would be red on every PR and read as broken rather than as
informative. What it publishes is the pair of counts, and the thing to look at
is whether they still say 33/6 and 39/0. A seventh red is a regression; five
means something got fixed and a scenario is ready to retire. The append-only
column is the measurement the pair exists for: it says which of the six would
close if positions were assigned at commit.

Non-blocking at the job level rather than only on the two sim steps, so that a
failed install or a PR comment the token cannot write does not turn this into
a red X either. The steps themselves keep their real exit codes, so the run
still says which world was clean.

`--title` is new on the CLI: two summaries land in one comment, and two
headings reading "world-sim scenario book" would leave the chips line as the
only way to tell them apart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* Make the Sim World comment four lines until you open it

Two collapsed folds, one per world, each with its count and a green or orange
dot on the visible line; the table is behind them. Nothing else above the fold.

The failures list is gone. Six scenarios are red on purpose, so a comment that
led with them led with the part that was not news, and grew a wall of text on
exactly the PRs that changed nothing. The count is the signal — and anyone who
wants the names can open the table, which has always had them.

`renderMarkdownSummary` now renders no heading of its own, since it is built
to be stacked under one; the workflow supplies the heading and the one-line
description. `mdCell`, `clip` and `dedupe` existed only for the failures table
and go with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* docs(sim-world): drop the six-red list, add an API reference

The "six red scenarios" section was a second copy of something `pnpm sim`
already prints exactly, and the copy that goes stale. What replaces it says how
to read a red — an open bug stating the outcome its own durable log implies,
green when fixed rather than when seen again — and points at DESIGN.md for the
part that is not re-derivable from a run.

The reference has three tables. Writers: what a concurrent thread of execution
is here, and the four ids one can have. Movements: the instruction that moves a
writer to a place and holds it, described against the world boundary, the
position in the event log, and the commit to storage — three moments, which is
why there are three stops rather than one. Withholdings: the two ways to change
what a reader sees without holding anybody.

Terminology unified on hold/held (the word `Held` and the trace already use)
and on "assigned a position in the event log" / "committed to storage", which
also fixes the `runToEventProduced` docstring: the event has crossed the world
boundary there, so "submitted" was the misleading half.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* docs(sim-world): move the API reference into the package, uniform vocabulary

The reference documents `@workflow/world-sim`, so it belongs beside the API and
not in the workbench that happens to be its first caller. It replaces the old
`## Writers` section, which said most of the same things in a different order
and a different vocabulary.

Three nouns, one meaning each. A *writer* is a thread of execution — the table
lists all three kinds with the handle that names each and the events each one
writes. An *advance* moves one writer to a named place and holds it; the three
`runTo*` stops are the three moments a write has (crossed the world boundary,
assigned a position in the event log, committed to storage), which is why there
are three rather than one, and the table says which way a concurrent commit
sorts at each. A *withholding* hides something from readers without holding
anybody, which is what `withholdNextEvent` and `beginHookDelivery` have in
common and why the latter is not an advance.

"Movements" is gone in favour of "advances", which the code and the older prose
already used. "Held" replaces "stopped"/"paused" throughout, matching `Held` and
`isHeld()`.

The workbench README loses the duplicate tables and keeps the four rules that
bite on a first scenario, so it is a guide to adding one and links out for the
rest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* docs(world-sim): rename "Two logs" to "World behaviors", one-line motivations

The section described two logs, but there are four world behaviors a scenario
can pick between: the two log orderings and the two guards. Renamed and
regrouped so each one is described by what it does rather than by which of the
current reds it closes.

Measured counts and "production" / "what happens today" framing are out of this
README throughout — they belong to a run of the book, and a doc that carries
them is stale the first time a scenario changes colour. The workbench README
and the CI lane still print them, where they read as measurements.

Every section's motivation is one line; the file-top motivation is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* docs(world-sim): Usage shows both halves — the workflow and the script

The example was a script with no workflow beside it, which hid the fact that
you write both. Replaced with the smallest pair that has something to control:
two steps in flight at once, and a script that decides which of them reaches
the log first.

The scenario in it was run before it was written down — `prepare` held before
it takes a position, `finalize` committed into the earlier slot, #6 then #7 —
so the positions the prose cites are the ones the trace prints.

Link to the workbench moved to the end of the section, where it reads as
"where to go next" rather than as an aside mid-example.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* docs(world-sim): say what "arm" meant instead of using the word

The READMEs told a reader to arm a wait without ever saying that calling an
advance and awaiting it are two separate things. That is the whole mechanism —
`runTo` registers its watch synchronously when called, and the returned promise
only reports arrival — so it is stated plainly once in Advances and the jargon
is dropped everywhere it stood in for the explanation.

"Armed" survives only where it means a guard is switched on, which is a
different word doing a different job.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* docs(world-sim): two symmetric watches in the Usage example

The example started one watch and awaited the other inline, which is exactly
the shape a reader cannot generalise from — it looks like the two calls do
different things. Both are now started, then both awaited, so the sentence
above the block and the code below it say the same thing.

Re-run before committing: same log, finalize at #6 and prepare at #7.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* docs(world-sim): neutral names in the Usage example

`prepare` / `finalize` / `held` / `committed` made a reader decode four
domain-ish names to see a mechanism that has nothing to do with any of them.
stepA and stepB, watchA and watchB: the only thing left to notice is that the
two watches stop at different points, which is the whole lesson.

This detaches the example from the workbench's `parallelStepsWorkflow`, so it
was verified against a throwaway copy of the workflow rather than assumed —
same shape, stepB at #6 and stepA at #7, replay ok.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* refactor(world-sim): one event fold, two call phases, a smaller entry

Three reductions to the same end — less surface to keep consistent — with
the scenario book as the check that none of them changed behaviour.

**One fold.** `create` validated *and* applied an event in a 17-case switch;
`foldSeededEvent` applied it again in 15 cases without validating. Two copies
of the event -> entity state machine that only a replay compares, and any
divergence between them makes a replay disagree with the run it is checking
for a reason that is not the runtime's fault. Both paths now end in one total
`applyEvent`: the write path validates first and calls it, `seedFromLog`
calls it with no validation at all, because those events were accepted once
already and re-litigating them would reject legitimate history. store.ts
loses 175 lines. Characterization tests for the seeded path went in first and
were confirmed to bite before the refactor started.

**Two phases.** `CallPhase` had a `positioned` phase between `before` and
`after`, and `runToPositionMinted` / `runToCall` to park on it. Nothing in the
book used any of the three: the mint/commit gap they existed for is reached
through `beginHookDelivery`, which owns both halves explicitly and does not
block the writer that made the write. Removing the phase also removes an
`await` from the interception path, so this was measured rather than reasoned
about — all four book runs are unchanged.

**A smaller entry.** `index.ts` re-exported the construction kit —
`createSimWorld`, `createSimStore`, `driveQueue`, `verifyReplay`,
`checkInvariants`, the clock — which nothing outside the package imports and
which made every one of their signatures a compatibility promise. The entry
is now the scenario surface; extenders import from the module. `InFlightWrite`
joins it, having been missing though `beginHookDelivery` returns it.

Unchanged, and the point of saying so: 33/6/6 mint-ordered, 39/0/0
append-only, 8 violations with `--no-fence`, 0 with both. 70 unit tests green
(up from 67), tsc and biome clean in both workspaces.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* fix(world-sim): follow the new event-result contract

Rebase fixups for four core replay commits, no behaviour change of our own.

`EventResult` is now a union — a populated page (`events` + `cursor` +
`hasMore`) or all three absent — so the three fields cannot be assembled from
three independent variables the type has no way to see agree. They travel as
one `deltaPage` object, spread as a whole or not at all. The spread has to be
conditional rather than optional: `...maybe` widens each field to `T |
undefined`, which is neither arm.

`processExitTriggersQueueRedelivery` is gone from the `World` interface —
#3385 stopped exiting on an exhausted replay budget, so there is nothing left
to tell it not to.

The book is unmoved across the rebase, which is the thing worth reporting:
33/6/6 mint-ordered, 39/0/0 append-only, 8 violations with `--no-fence`, 0
with both — and the same six ids red, not a swap that nets to the same count.
70 unit tests green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* feat(world-sim): a scenario for the unclaimed-payload delivery race

Reproduces vercel/workflow#3406 as a pair of scenarios. A hook payload
nobody reads registers an unarmed delivery barrier, and a step result is
allowed to skip it — otherwise the workflow stalls until the barrier
registry idles. The skip is transitive, so the step also skips an armed
`wait_completed` merely parked behind that payload. Both branches then
draw their next `step_created` id in the order the log does not record.

The replay invariant cannot see this: live and replay run the same code,
make the same mistake, and agree. What is checkable is the log
disagreeing with itself — `wait_completed` is committed first, yet the
step branch draws the earlier id. Verified red on the current runtime
and green with #3406's diff applied.

`claimed-payload-under-fork` is the control: same steps, same tempo,
same log order through the step result, one branch awaiting the hook so
the payload lands claimed. It passes in both cases.

Getting there needed one new primitive. The delivery loop is serial, so
a held inline step body stops the loop inside its own delivery and no
timer can fire — every interleaving where a `wait_completed` lands while
a step result is outstanding was unreachable. `sim.deliverQueued(select?)`
takes a message out of the pending set and delivers it from the script,
concurrently with the hold; `takeById` removes it first, so the loop can
never pick up the same message.

The book is now 41 scenarios: 34/7 with 6 violations mint-ordered,
40/1 with 0 violations append-only. The six violation reds and the
--no-fence 6 -> 8 / 0 -> 0 measurements are unchanged. The new red is
the first that stays red in both worlds, because no log position is
wrong in it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* fix(world-sim): address review on the simulation world

Fifteen findings from review, and one measurement that moved because of
them.

Correctness:

- tempo: `abort()` stored the raw `reject` rather than the wrapper that
  also disposes the watch, so a point reached after teardown blocked a
  world call on a promise nobody could release. The deadline is one-shot
  and already spent by then, so that is a hang. Route abort through the
  disposing path and make the action a no-op once aborted.
- world: `externalDepth` was a plain counter, so a step body writing
  concurrently with a scripted `deliverHook` was attributed to the
  script — invisible to any armed `runTo`, and `ext` in the trace.
  Measured: one misattribution book-wide. Scope it with
  `AsyncLocalStorage` instead.
- writers: `runToEventCommitted` matched a *rejected* create, so in
  fence scenarios a script could wake believing a write was durable when
  it had 412'd. Record `failed` on the observed point and require
  `failed: false`.
- clock/ids: reject fractional `advanceBy`, and floor in `ulid()` — a
  fractional millisecond silently minted `undefinedundefined…` ids that
  failed much later at `z.ulid()`.
- streams: literal NUL bytes as key separators, `limit: 0` paging
  forever, and a garbled cursor slicing the array away as `NaN`.
- scenario: register the handler under the exported
  `WORKFLOW_QUEUE_PREFIX`.

The count guard, which is the substantive one. Since #3145 `@workflow/core`
sends `stateEventCount` on every replay-context create and the server's
count guard defaults on, so the sim's "no client sends the count" claim
was stale in six places. `countGuard` now follows the fence, the sim
prefers the runtime's own count over its reconstruction, and the two
scenarios whose subject is isolating the watermark half say
`countGuard: false` explicitly. The scoreboard does not move under the
production-shaped default, which is itself the answer to the review's
question.

`log.monotonic-order` could never fire: its only caller fed it the sorted
array. It now takes commit order, supplied only by a world that promises
the two agree — under a mint-ordered log an out-of-order commit is the
premise the scenario injected, not a defect.

Docs: DESIGN §5 gains the two ways the sim's guards are stronger than
production's (FIFO-vs-mint-order pruning, and a fence that is exact
in-process where production's is region-local and fails open); §9's
scoreboard is regenerated and its "no fix armed anywhere real" claim
corrected; §10 gains the parallel hook-resume path, which no sim delivery
takes. `step-vs-step-fork` and its twin now say who the withheld reader
is in production.

`unclaimed-payload-under-fork` is green after the rebase — #3406 fixed
the delivery-barrier ordering — so the book is back to six reds and
stays there. Counts refreshed everywhere they appear: 35/6/6
mint-ordered, 41/0/0 append-only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

* fix(world-sim): make the README's code samples type-check

`packages/docs-typecheck` globs `packages/*/README.md`, so the new README
went into the Docs Code Samples job and five of its `ts` blocks failed.

Four were genuine excerpts with an unbound `sim` or `spec`; one was a match-
object *sketch* containing a `…`, so not TypeScript at all — that one loses
its `ts` tag, joining the eleven other untagged blocks in the file.

The rest name what they use. That needed two `paths` entries in the checker:
`@workflow/world-sim` was unmapped, and an unresolved import is deliberately
tolerated there (`isExpectedMissingModule`), so every identifier in those
samples was `any` — they would have gone green while checking nothing.
Mapped, they are checked against the real declarations: seeding `id: 12345`,
`readyAtMsTYPO`, `dirsTYPO` and `runToEventCommittedTYPO` is caught against
`ScenarioSpec`, `PendingMessageView`, `SimBuildOptions` and `Writer`.

The lead teaser keeps its four unadorned lines and takes a `@skip-typecheck`
marker instead; the same calls appear in full, checked form further down.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com>

---------

Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
2026-08-11 08:37:16 -07:00

42 KiB

world-sim: design

How @workflow/world-sim is built and why it is built that way. The package README is the introduction; this is the implementation.

Two workspaces:

path what it is
packages/world-sim the World implementation, the scenario runner, the checkers
workbench/sim-world the scenario book and the workflows it runs (pnpm sim)

1. Module map

module responsibility
world.ts the World implementation; wraps every method as a call point and attributes it to a writer
store.ts in-memory event store — the event → entity state machine, plus the write-time guards
queue.ts deterministic queue: records messages, never delivers on its own
clock.ts virtual clock; patches Date.now() readings, not timers
ids.ts deterministic ULID minting from (virtual time, counter)
drive.ts the scheduler loop and the scenario budgets
tempo.ts the scripting layer: park / permit, and the wait bookkeeping under runTo
writers.ts named writers and their level-triggered runTo* vocabulary
scenario.ts runs one ScenarioSpec end to end and produces a ScenarioResult
replay.ts cold-start replay verification of a finished log
invariants.ts consistency checks re-derived from the event log alone
report.ts renders a scenario; log positions for every event reference, colour only when the destination is a terminal
streams.ts in-memory streamer
build.ts bundles a project's workflows so the runtime can be handed real compiled code. Its own entry (@workflow/world-sim/build), because it reaches a compiler and playing a scenario should not
load.ts loads a built bundle's flow handler — the half of the old build.ts that needs no compiler
types.ts the public vocabulary

2. What the simulator has to model

The design follows from four properties of the runtime. None of them are choices this package made; they are the constraints it works inside.

The orchestrator is re-run from the top, every time

A workflow function is not a coroutine that parks and resumes. It is re-executed from its first line on every replay pass, inside a fresh node:vm context (createContext in packages/core/src/vm/index.ts). The VM seals the two obvious sources of nondeterminism: Math.random is seeded from ${runId}:${workflowName}:${deploymentId}, and Date.now() / new Date() return a fixed timestamp advanced only from each consumed event's createdAt.

A pass is therefore a pure function of (workflow code, run identity, event log prefix). Stated precisely: same log prefix → same decisions. That is the whole basis of durability, and it is also the property the simulator exists to attack — the interesting bug class is not "the workflow behaved randomly" but "the decisions and the persisted log disagree", which requires the decisions to have been made against a different log than the one that ended up durable.

Entity identity is positional

Correlation ids come from ctx.generateUlid(), driven by the VM's seeded Math.random. They are positional ordinals of one seeded sequence: the Nth entity the workflow asks for gets the same id in every pass. Steps, hooks and waits all draw from that one sequence.

This is why a flipped branch is dangerous. It does not produce a different step; it produces a different step wearing the same name badge. The runtime's divergence check is a step-name comparison at the same ordinal:

Replay divergence: step event step_created for step_…445J belongs to
"…//settle", but the current step consumer is "…//recoverFirst"

Suspension is the unit of progress

useStep does the same thing on every pass: mint the correlation id, register a StepInvocationQueueItem, subscribe a consumer, return a promise. What differs is what the consumer finds. On replay the log holds step_created / step_started / step_completed for that correlation id, so the consumer hydrates the recorded result and the step body is never called. First time through, the consumer reaches the end of the log, returns NotConsumed, and the promise never resolves — so the workflow cannot proceed.

When nothing can make further progress a WorkflowSuspension is raised carrying the whole invocationsQueue. The runtime commits the pending *_created events, executes what it can, and runs the workflow again from the top against a longer log:

load log → run workflow from top → suspend → commit + execute → run from top → …

Hooks follow the same shape with a different event family. hook_created is committed at the next suspension rather than at the call. An out-of-band resumeHook(token, payload) writes hook_received and enqueues a flow message, so the run wakes up. Payloads landing before the workflow awaits are buffered in a payloadsQueue, which is why a duplicate delivery is absorbed rather than lost.

There is no hook state a workflow can read — the surface is token, getConflict(), dispose(), then, [Symbol.asyncIterator]. The only way to observe a hook is to attach a continuation and see whether it resolves, which is a timing observation, not a state read. That is why the scenario API steers time rather than poking state.

Nothing sleeps

sleep() registers a WaitInvocationQueueItem; wait_created records a resumeAt; the runtime enqueues a delayed queue message. When that message is delivered, the flow handler's "complete elapsed waits" pass compares Date.now() >= resumeAt and writes wait_completed.

A timeout is a delayed message plus a clock comparison. That is exactly why virtual time works here, and why a thirty-day sleep costs microseconds.

Where the concurrency actually is

The workflow body is single-threaded JS and stays that way. The interleaving that matters lives in three other places:

  1. Between passes. The log grows. A branch decided in pass N was decided against pass N's prefix.
  2. Between event deliveries inside one pass. Several awaits can be pending at once. The runtime forces resolution order to match log position via the delivery-barrier registry (registerDeliveryBarrier / awaitEarlierDeliveries in private.ts), so this one is reproducible.
  3. Between invocations. Two flow deliveries for one run can execute simultaneously in different processes. This is the one the SDK cannot make deterministic by itself.

Two step bodies that suspend together are already concurrent writers to one log, inside a single delivery. That is cheaper to reach than "concurrent writers" suggests, and it is the case the writer vocabulary is built around.


3. Interception

Every World method is a call point

Each method on the World is wrapped so a scenario can stop it. The wrapper records the call, fires any matching watches, runs the underlying implementation, fires the watches again on the way out, and only then resumes the caller.

A watch action returns a promise, and the intercepted call awaits it. That is the entire hold mechanism: a "held" writer is a caller blocked inside a World method. release() resolves that promise.

Two phases, and a third hold that is not one

type CallPhase = 'before' | 'after';

before and after bracket the call:

held at what a competing write does
before no position taken yet, so a write landing during the hold sorts ahead
after durable, and the writer has not been resumed yet

after is the window the package was originally built for: "the hook arrives after step_started is durable and before the workflow resumes".

Neither phase produces the opposite order, and a write is not atomic, so the opposite order is reachable: a real backend mints the event id first — DynamoDB does not generate ids, and that id is the log's sort key — and only then attempts the storage write. Between the two the event has a position but no visibility, and a write that commits in that window sorts behind it.

That gap is the point, not a detail. It is the only way to produce an event behind a position a reader has already read past — a complete, consistent log prefix that is simply missing an event still in flight. No high-water-mark fence can represent that shape.

It is not a phase, though, because the writer holding it is not blocked inside a World method: the script owns the two halves explicitly, via reservePosition / withReservedPosition under sim.beginHookDelivery (§Withholdings below). A phase would have been a second way to say the same thing, and it went unused.

Watches do not fire inside watches

Calls made from inside another watch's action are not call points. Without that rule a watch on events.create would re-trigger on the hook_received it just wrote, and every scenario using deliverHook would recurse forever. The depth is tracked and surfaced in the trace, so a line committed from inside a held call is visibly at depth > 0.

A related rule is easy to get wrong: the depth counter must be raised only by asExternal, which brackets exactly one call. Raising it for the whole duration of a watch action is correct for something that returns immediately and wrong for a hold, which does not return until the scenario releases it — under that rule, holding one writer makes every other writer's call stop being a call point, so a held step body's sibling becomes invisible and unsteerable.

Writer attribution is derived, not instrumented

Writer identity comes from the intercepted call plus its request. No runtime hook is needed:

write writer
step_created / hook_created / wait_created / run_* / attr_* orchestrator
step_started orchestrator (the executor, which precedes the body)
step_completed / step_failed @ correlationId C step:<name of C>
hook_received external
wait_completed the wait-continuation delivery
events.list / runs.get orchestrator (the read half)

The writer is printed as a column in every event stream, so a rendered log says who wrote each line.


4. Determinism machinery

Clock

install() patches Date.now() and the zero-argument Date constructor to read the virtual clock. Timers are deliberately not patched: @workflow/core uses setTimeout(fn, 0) as a macrotask barrier in several ordering-sensitive places (events-consumer.ts, private.ts), and swapping those for fake timers would change the very interleavings the simulation exists to observe. Real zero-delay timers stay real; only the readings of wall time move.

The clock never moves on its own. Only the scheduler calls advanceTo / advanceBy, so two runs of a scenario see the same sequence of timestamps.

Ids

Every id is a function of (virtual time, per-scenario counter) — never Math.random() or the host clock. They still have to be real ULIDs, because @workflow/world validates run ids with z.string().ulid() and decodes the embedded timestamp, so the encoding is standard Crockford base32 with the 16 "random" characters filled from the counter.

Byte-identical ids run to run are what make an event-stream dump usable as a golden file.

Queue

@workflow/world-local's queue fires a detached delivery loop from inside queue(), so a message races whatever the caller does next. Faithful to production, useless for a simulation.

Here queue() only records. Delivery happens when the scheduler asks, and it always takes the same message: the minimum by (readyAtMs, enqueueSeq). Delays are virtual — a message 23 hours out is delivered by jumping the clock. ScenarioSpec.selectNext can override the choice to pin an order the default would not produce.

Scheduler

take the next message → jump the clock to its delivery time → hand it to the
flow handler → wait → repeat until the queue is empty or a budget stops it

Between deliveries the loop drains the event loop for several rounds (settle()), because the runtime uses zero-delay macrotasks as ordering barriers and waitUntil-style background work is not awaited by anyone. Without that drain, a message enqueued from a trailing microtask would be missed and the scenario would report a spurious stall.

The scheduler lives apart from scenario.ts because two things drive it: a scenario, and the replay verification that cold-starts a second world.

One delivery at a time. This is the deliberate limit of the model — see §9.


5. The store

A reference implementation of the World storage contract: the same event → entity state machine @workflow/world-local implements on the filesystem, minus every mechanism that exists purely to make that state machine safe against concurrent processes (exclusive-create claim files, per-entity locks, staged/promoted hook events, canonical event-id pinning after a crash). One delivery at a time in one process means those races cannot occur, and their absence keeps the file small enough to audit.

What is deliberately kept is every validation that rejects an event — terminal-run guards, step lifecycle ordering, hook token uniqueness, wait duplication. Those rejections are the observable contract the runtime is written against; a simulation that relaxed them would agree with the runtime about nothing interesting.

The two guards

The fence is off by default and set per scenario (ScenarioSpec.preconditionGuard, which flows into SimWorldOptions), which is how a scenario can be run one flag apart from its neighbour. countGuard follows the fence unless a spec says otherwise, because that is what production does — see below.

preconditionGuard models WorldCapabilities.preconditionGuard: reject a replay-context write whose stateUpdatedAt snapshot predates the newest externally-originated event. In the SDK this is declared by world-vercel only (packages/world-vercel/src/index.ts:40); world-local and world-postgres declare neither it nor maxConcurrency.

Its predicate is narrower than the bug class, and the reason is its shape, not the event type it watches. The marker advances on hook_received or step_completed, but it is a high-water mark — the newest such write — and the test is stateUpdatedAt < marker, strictly. So it detects a log truncated at the end and is blind to a hole in the middle: when the withheld event is older than one the reader can see, the reader's snapshot is never strictly older than the mark. The hook direction is caught for the mirror-image reason — the withheld hook_received is the newest out-of-band write and the orchestrator's snapshot predates the sleep, so the fence fires and the run reconciles.

countGuard adds the count half: how many events the log holds at or below stateUpdatedAt, compared against how many the caller loaded. It closes the hole the watermark cannot see. It requires the caller to send stateEventCount — and since #3145 (1471f252f) @workflow/core sends it on every replay-context create, gated only by the WORKFLOW_PRECONDITION_GUARD kill-switch, with workflow-server's own count guard defaulting on. So both halves are armed together in production, and countGuard defaults here to whatever the fence is set to. A run with the fence on and the count off is a world that exists nowhere; the two scenarios that ask for it (step-vs-step-fork-fenced, in-flight-before-decision) do so explicitly, because isolating the watermark half is their whole subject.

Where the runtime sends a count, the sim uses that value rather than its own reconstruction (loadedCount()), so the guard is tested against the number the real client computes; the reconstruction is the fallback for writes core does not count.

Two ways the sim's guards are stronger than production's. Both are deliberate, and both mean a fenced green here is a claim about the predicate, not about production's deployment of it:

  • The server's retained-id window is a FIFO in insertion (commit) order — its Lua script prunes with table.remove(ids, 1), oldest-inserted — while pruneRunEventIndex here sorts by id and drops the smallest, i.e. mint order. The two differ exactly when commits happen out of mint order, which is these scenarios' whole subject, and they differ in when countRecordedAtOrBelow goes indeterminate once a run passes 16 events.
  • Production's watermark is best-effort: region-local Redis, failing open on Redis errors, and blind to a webhook served in another region entirely (see outside-event-tracker.ts's own docs). The sim's is exact and in-process.

Overriding them for a whole run. RunScenarioOptions.preconditionGuard (pnpm sim --fence / --no-fence) forces the fence on or off for every scenario, undefined leaving each spec to decide. Both halves move together: the count guard is evaluated inside the same predicate, so disarming the fence disarms it too.

Forcing it off across the book asks whether anything relies on it. Violations go 6 → 8 mint-ordered, so it is load-bearing there; 0 → 0 against an append-only log, so it is dead weight once positions are assigned at commit. This is a diagnostic rather than a world, and the number to read is the violation count: a scenario whose subject is that the guard fired asserts that with sim.check and fails by design when it does not (in-flight-before-decision-counted today).

Fault injection

withholdNextEvent(reads = 1) hides the next committed event from the following N event-log reads. This is the only way a serial simulation can produce "a write derived from an incomplete event load", the precondition a real deployment reaches through concurrency.

It is a faithful model rather than an approximation, because production reaches the same ordering natively: world-local mints evnt_${monotonicUlid()} near the top of createImpl (packages/world-local/src/storage/events-storage.ts) and writes the file much later, so two concurrent creates take positions N and N+1 and can land in the opposite order. Postgres does the same via nextval before COMMIT. world-local defends this with mintRunDominantEventKey (src/storage/helpers.ts) — but only for terminal run events; wait_completed gets no re-derivation.

One withheld read poisons a whole invocation, which is worth knowing when reading a trace: after the next step_completed the runtime continues from its cursor, fetching only events written strictly after that position. A withheld event sitting before the cursor can never re-enter that invocation's view. Incremental reads make the hole permanent.

beginHookDelivery(token, payload) returns an InFlightWrite — a write held between mint and commit, with eventId already fixed and commit() still pending. Unlike a held writer, nothing is blocked meanwhile, because the receiver is a separate process from the run's invocation. Holding an inline write instead would stall the delivery that made it, and thus the reader too, which is why the out-of-band writer is the one that can express this shape.

Changing the world instead of the runtime

appendOnlyLog is the one option that alters the store's contract rather than its strictness. With it on, an event takes its position in append instead of at the handler boundary: a write that is still the newest when it commits keeps the id it was already handed out under, and one that was overtaken while it was held re-mints and takes the tail.

That single move collapses both faults above into the same, weaker one. A hold between mint and commit can no longer open a hole, because the held write is not claiming a position while it waits — it has none until it lands. And withholdNextEvent degrades from serving a read around the withheld event to stopping it at the event, because a hole is not expressible in a log whose order is its commit order. Both leave the reader short rather than wrong, and short is precisely what the fence's watermark was designed to catch.

Off by default: the sim exists to model the world that exists, and production mints at the boundary because DynamoDB does not generate ids. The value of the switch is differential — play the book both ways and the diff separates "fails because of the mint-before-commit window" from "fails for some other reason". No scenario in the book sets it; it is meant to be driven from RunScenarioOptions or pnpm sim --append-only.


6. The scenario surface

Spec

interface ScenarioSpec {
  id: string;                   // stable hyphenated handle; what `pnpm sim <id>` selects
  name: string;                 // prose, expected to be reworded; the id is not
  description?: string;
  workflow: string | { workflowId: string };  // plain fn name, resolved via the build manifest
  input?: unknown[];
  script?: ScenarioScript;      // omitted = a control: run on the default schedule
  selectNext?: SelectNext;      // override queue delivery order
  verifyReplay?: boolean;       // default on for runs reaching completed/failed
  expect?: { status?: ScenarioOutcome; output?: unknown };  // output: deep equality
  limits?: ScenarioLimits;
  preconditionGuard?: boolean;  // advertise + enforce the optimistic-concurrency fence
  countGuard?: boolean;         // also enforce its count half
  appendOnlyLog?: boolean;      // position at commit, not at mint; see §5
}

RunScenarioOptions.appendOnlyLog overrides the last of those for every scenario in a run, which is how the whole book gets played both ways; undefined there leaves each spec to decide. The mode a result was produced under is recorded on ScenarioResult.appendOnlyLog rather than left to the reader's memory.

expect.status accepts the non-run outcomes (stalled, budget-exceeded) because "this workflow deadlocks when the hook never arrives" is a property worth pinning down rather than an accident to tolerate.

There is deliberately no way to expect a consistency violation. A scenario reproducing a corruption states the outcome the run should have had and fails until the runtime delivers it. A red is an open bug, not a recorded observation, and it goes green when the bug is fixed rather than when the bug is seen once more.

Scripting

ScenarioApi is the complete set of sanctioned external inputs — anything a real deployment could do out-of-band has an entry, so the script is a complete description of what happened:

deliverHook · beginHookDelivery · cancelRun · advanceTime · withholdNextEvent · note · check · world (read-only snapshot) · runId

Tempo adds the steering: writer handles, plus the raw park / until / during primitives. The vocabulary is borrowed from Python's blanket, which does the same for threading primitives — the call parks, the script issues the permit, and the resulting order of permits is the tempo.

Writers

sim.writer.orchestrator() / .step(shortName) / .anyStep() / .any() return a Writer. A handle is a name, not a live object: it can be taken before the step exists and binds to whichever writer shows up under it.

method phase meaning
runToEventProduced before decided and submitted, nothing in the log
runToEventCommitted after durable, writer not yet resumed
release — let it go; idempotent
isHeld / history — inspection

Two implementation details of release() matter to scenario authors. It is guarded by a done flag so double release is a no-op. And it awaits a full macrotask turn before resolving — without that, await release() returns while the resumed call is still queued as a microtask, and a scenario reading the log on the next line sees the state it was trying to leave.

runTo is level-triggered

It consults the history of points the writer has already reached before arming anything, and throws AlreadyPassedError naming the call it happened at if the point has gone by.

The alternative — arm a watch and wait — is a hang. A held call blocks its writer, and when that writer is the one the scheduler is inside, it blocks the loop; so there is no quiescence to fall back on and no timer to eventually fire. An edge-triggered wait on an edge that has passed is the one way to lose this package's termination guarantee, so it is made impossible rather than documented.

Three consequences:

  • Holds must be armed before they are needed. To catch two writers at the same point, start both waits and then await them. Awaiting the first before starting the second yields the event loop, and the other writer may sail past.
  • runTo on an already-held writer releases it first, and arms the new watch before releasing. That order is load-bearing: the released writer can reach the next point within the same turn — the after phase of the very call it was held in is the common case — and a watch armed afterwards would miss it. The same rule applies to authors sequencing two writers: arm B before releasing A.
  • A call is two records, so seq cannot order them. The before and after phases of one call share a seq, so each recorded point carries its own ordinal and the level check compares against that.

A watermark tracks how far each writer has been advanced. Points at or before it are "already consumed" and do not count as already-passed — asking twice for step_completed means the next one, which is what the duplicate-delivery scenarios need.

What is not offered

Writers form a dependency graph — the orchestrator awaits its own step bodies — so not every interleaving exists to be asked for, and an unsatisfiable runTo can only be reported, not prevented. The runtime's await graph is not visible from here, so true deadlock detection is out of reach; the substitute is a per-runTo watchdog that reports where every writer was standing.


7. Termination

Every scenario terminates. Four budgets, layered so the most specific one reports first:

budget default catches
maxRunToWallMs 5 s one runTo that will never be satisfied
maxDeliveries 200 a run that keeps re-enqueueing itself
maxVirtualMs 365 d while (true) { await sleep('1d') }
maxWallMs 60 s a genuinely non-terminating step body

maxRunToWallMs sits far below maxWallMs on purpose: it can name which writer failed to reach which point and where the others were standing, and that diagnosis is worth more than the generic "ran out of wall clock" the global deadline can offer. It is clamped to maxWallMs so lowering the global budget does not require remembering to lower this one.

The scenario's global deadline must not be unref'd. An unref'd timer does not hold the event loop open, so a total deadlock — every writer held, scheduler blocked inside a held call, script awaiting the impossible — empties the loop and exits Node with a bare "unsettled top-level await" instead of firing the watchdog, which is precisely the case the watchdog exists for. The finally already clears it, so it cannot outlive a scenario.

Stream readers get the same treatment: a reader that parked on an unfinished stream would deadlock the scenario, so readers park on a promise the writer resolves and abortOpenReaders() releases any still parked at teardown — turning a hang into a reported diagnostic.

Outcomes are WorkflowRunStatus | 'stalled' | 'budget-exceeded' | 'error'. A hook that never arrives is reported as a stall naming the undelivered token, not a hang.


8. Consistency checking

Two independent checkers run over every scenario.

Invariants

The store enforces most rules at write time by rejecting bad events — but "the store rejected it" and "the log is actually consistent" are different claims, and only the second is worth trusting. So invariants.ts re-derives everything from the event log alone and compares against the entity rows.

25 rules, grouped:

log.monotonic-order          log.unique-event-id
run.created-first            run.created-once           run.terminal-is-last
run.entity-matches-log       run.attributes-match-log   run.output-materialized
run.resources-released
step.created-once            step.started-after-created step.terminal-after-created
step.terminal-once           step.no-restart-after-terminal
step.entity-has-log          step.entity-matches-log    step.attempt-matches-log
hook.token-unique            hook.received-after-created
hook.dispose-once            hook.no-receive-after-dispose
wait.created-once            wait.completed-after-created
wait.completed-once          wait.resume-at-stable

A violation is a bug somewhere — in the runtime that produced the sequence, in the store that accepted it, or in the scenario that injected something impossible. Which one is a question for the reader; the checker's job is only to notice.

Replay verification

The invariants check the log's shape. None of that answers the question durability actually rests on: if a fresh process picked up this log tomorrow, would it reconstruct the same run?

The check is a cold start with the answer withheld. Take the finished log, drop its terminal run_* event, load the rest into an empty world as durable history, and deliver one queue message. The real runtime — the same workflowEntrypoint a deployment serves — replays from the log and must re-derive the event that was removed, with the same output. No step body re-executes, since every step_completed is in the log and the step consumer resolves from it, so anything the replay produces came from the log alone.

In this frame, replay is the serializability check. A pass is pure, so re-running it over the committed log asks whether the schedule had a serial equivalent. Six failure ids:

replay.diverged · replay.suspended · replay.output-differs · replay.log-differs · replay.status-differs · replay.budget

replay.diverged is the runtime raising ReplayDivergenceError, exhausting its recovery replays, and failing the run with CorruptedEventLogError. replay.suspended means the replay ran out of log before the workflow finished — the log did not contain enough to rebuild the run.


9. Current status

Measured on branch sim-world.

Unit tests — 72 passing across 8 files (pnpm --filter @workflow/world-sim test).

Scenarios — pnpm sim in workbench/sim-world:

41 scenario(s): 35 passed, 6 failed, 6 consistency violation(s)

And the same book against an append-only log (pnpm sim --append-only):

41 scenario(s): 41 passed, 0 failed, 0 consistency violation(s)

Both numbers are the intended steady state; see "The six" below for which of the six violations that second line closes on the merits and which close because the correct answer itself changes.

There was a seventh red until recently, unclaimed-payload-under-fork, and it was a different animal: it tripped a sim.check rather than the replay invariant, and it was red in both worlds, because nothing was wrong with its log's positions — the runtime handed the workflow two resolutions in an order the log did not record, so live and replay ran the same code, made the same mistake, and agreed. Only the log disagreeing with itself caught it. #3406 fixed the delivery-barrier ordering and it is now green in both worlds; the scenario stays as that fix's regression test.

With the fence forced off (pnpm sim --no-fence), violations go to 8 mint-ordered and stay at 0 append-only — see §5.

Replay verification across the book: 33 ok, 6 MISMATCH, 2 skipped (skipped where the run did not reach a terminal status).

run.ts exits non-zero, and that is the intended steady state. The six failures are reproductions of corruptions the runtime can still produce; each states the outcome its own durable log implies and fails until the runtime gets there, so the failure line names both sides (expected "afterSlow:doc-26", got "afterFast:doc-26").

The number is the thing to watch: six today. A seventh is a regression; five means something got fixed and a scenario is ready to retire.

That makes the book a poor plain CI gate, which is what --report-only is for: it prints every failure and exits 0, so a job can publish the book's current state rather than block on it. --summary-file writes one collapsed <details> — a visible line carrying the count and a green or orange dot, the whole table behind it — for a PR comment or $GITHUB_STEP_SUMMARY, and --detail-file writes the full colour-free trace as an artifact to read when a number moves. Deliberately nothing above the fold but the count: six are red on purpose, so a comment that leads with the failures leads with the part that is not news, and grows a wall of text on exactly the PRs that changed nothing. The workbench's pnpm test is --report-only --summary-file, so a recursive pnpm -r test stays green and still says what happened; pnpm sim stays strict, so running it by hand fails loudly.

The six

The fix column names the specific change that closes the scenario. shown green by is the stronger claim: a passing scenario that is this one with that fix armed, same workflow and same tempo, one flag apart. Where it says "none yet", the fix is identified by argument but nothing in the book proves it.

scenario mechanism fix shown green by
stale-read-step-count-fork (doc-23) withholdNextEvent(1) + deliverHook; hook at #7, wait_completed at #8, no-hook branch at #9 preconditionGuard — the withheld hook_received is the newest out-of-band write and the orchestrator's snapshot predates the sleep, so the watermark fires stale-read-step-count-fork-fenced (doc-24)
stale-read-equal-step-counts (doc-25) same fault on a fork whose branches emit one step each preconditionGuard, for the same reason none yet
step-vs-step-fork (doc-26) two of the run's own step_completed events, one delivery countGuard. Not preconditionGuard: the withheld completion is a hole in the middle of the log, which moves no high-water mark (§5) none yet
step-vs-step-fork-fenced (doc-27) same, preconditionGuard: true, zero rejections countGuard. This row is the proof that the watermark half does not fix doc-26 none yet
in-flight-before-decision (doc-29) beginHookDelivery, committed before the decision is written countGuard in-flight-before-decision-counted (doc-30)
in-flight-after-decision (doc-31) beginHookDelivery, committed after the decision none in the SDK. Needs an append-tail fence — assertSlotAboveTail, vercel/workflow-server#692 —

Those handles are ScenarioSpec.id, and they select: pnpm sim in-flight-after-decision plays one row of this table.

So: two of the six have their fix demonstrated by a paired green scenario, three have a fix identified but unproven here, and one has no fix at all. Writing the three missing pairs is the obvious next increment.

Note what the fix column does not mean. countGuard closing doc-29 is a statement about the World implementation and about production's predicate: core has sent stateEventCount on every replay-context create since #3145 and the server's count guard defaults on (§5). What it is not is a statement about production's deployment of that predicate, which is region-local, fails open, and prunes its window in a different order than this store does — all three noted in §5. So the honest reading of the column is: four of the six have a fix whose predicate is armed in production today, and whether it fires there depends on conditions the sim does not model.

The append-only log closes all six, in two different senses — and the split is four and two, not three and three. Four (doc-23, doc-25, doc-26, doc-27) close on the merits, with the book asking them exactly what it asked before: the reordering was the fault, and once positions are assigned at commit the withheld read degrades from a hole to a truncation, which the fence can see. The remaining two (doc-29, doc-31) close because the branch the run ends on changes. A hook that commits after the timeout genuinely is after it when the tail is the only place a write can land, so the log records the timeout first and the run that settled is the run the log describes.

No expectation is restated per world, and there is no mechanism to. The first cut of this had one — an expectAppendOnly field on three scenarios, naming a second correct output. It was the wrong instrument, for a reason worth keeping written down. A scenario is one sequence of advances. The only thing a world changes is what a read returns. The branch a run ends on is decided by what it read, so pinning the branch pins a consequence of the world rather than a property of the run, and any expectation that then has to be restated per world is evidence the pin was wrong — not evidence that a second answer is needed. The three now assert what holds in both worlds (the run completes) and report the branch in the trace.

That costs nothing, because the expectations were never what caught the fault. The load-bearing assertion is the invariant: the log a run wrote must be a log the runtime can replay back into that same run. It is world-independent, on by default (verifyReplay), and it is what all six reds trip. Measured, not assumed: strip every expect in the book and the violation counts do not move — 6 mint-ordered, 0 append-only, the same six by name. (Pass/fail does move by one, and only for a bookkeeping reason: hook-never-arrives expects stalled, and a stall's reason is reported as a problem unless the scenario said it was expecting one.) That also removes the one place where the flag's scoreboard rested on a judgement about what the right answer is rather than on something the harness checks on its own.

doc-30 is worth a line because it was the third expectAppendOnly and is not one of the six — mint-ordered it already passes, since countGuard catches there what the watermark half misses. Its branch moves under the flag for the same reason its uncounted twin's does, so the old pinned output would have turned a green scenario red. What made it distinct from doc-29 was never the branch anyway; it is that the count half of the fence fires at all. That is now asserted directly, matched on the guard's own message and true in both worlds — and it fails if countGuard is turned off, which is the check that a bare rejections().length > 0 would have missed, since doc-29 rejects too.

Two details worth keeping:

  • Hook delivery participates. beginHookDelivery still reserves a position at the handler boundary; under the flag the reservation stops being binding and the write re-mints at the tail if anything overtook it (positionAtCommit).
  • doc-30's 412 still fires, and now saves nothing. The count guard counts events at or below the caller's watermark, the watermark is a millisecond, and a hook committing after the timeout within the same virtual millisecond is still "at or below" it. The reload finds nothing to correct and the run settles anyway. That false positive is the standing cost of the count half of the fence once the log is append-only, and doc-30's trace is where to see it.

Four of the six are hook-driven and two deliberately are not — the pair proves the corruption needs no out-of-band event type. All the pure hook-timing scenarios pass: placing a hook precisely is what works. What fails is a hook that is durable in the log but absent from the read the live pass decided on.

The last row is the only one with no fix in the SDK: the hole opens after the write that should have fenced it, in the quiescent gap between deliveries where the run makes no writes and so meets no checks. assertSlotAboveTail in vercel/workflow-server#692 is the append-tail fence for it.


10. Limits

Concurrent invocations are out of reach. The scheduler does await deliver(...), so two flow deliveries for one run cannot overlap. Reaching that would need concurrent delivery with hold points to pin the interleaving. The gap matters because it is a real production route: resumeHook writes hook_received and enqueues a flow message, so two deliveries end up in flight — one writing wait_completed and deciding no-hook, one seeing the hook and deciding hook-branch — racing to create the same ordinal, with every reader holding a perfectly consistent view. Just different ones.

Two step bodies inside one delivery are genuinely concurrent and separately steerable, which is enough to reach the interesting corruption without a second invocation. That is why the limit has been acceptable so far.

The parallel hook-resume path is never exercised. resumeHook picks parallel ("lazy") vs sequential from world.capabilities.hookResumeDedup (or a fresh server attestation). world-local declares it and world-vercel attests it per lookup, so every real world takes the parallel path, where the queue publish races the hook_received write and the consumer re-ensures the event through the durable (runId, resumeId) claim. The sim advertises neither the capability nor a resumeId dedupe, so every sim hook delivery takes the sequential path — meaning the hook-timing shapes in this book are the legacy shape, not the one production runs. Closing this needs (runId, resumeId) dedupe in the store plus the capability; it is the largest single gap for a package about hook races.

Also untested: turbo / optimistic-inline-start, which skip replays and so give a stale branch somewhere to hide; and the fence's same-millisecond behaviour, where an equal stateUpdatedAt passes by design as anti-livelock.

Not modelled at all: the concurrency machinery world-local needs and this store omits — claim files, per-entity locks, staged/promoted hook events, canonical event-id pinning after a crash. Bugs in those are invisible here.


11. A caveat worth stating

A simulated world only produces trustworthy results while its model matches reality. Every simplification in §5 and every limit in §10 is a place where a green scenario could be green for the wrong reason. The mitigations are that the store keeps every rejection the real one performs, that the runtime under test is the real workflowEntrypoint running real compiled workflow code, and that every scenario ends by replaying its own log through that same runtime — but none of those is a proof, and a red here is worth more than a green.