* Add world-sim: a deterministic simulation World for concurrency scenarios `@workflow/world-sim` is an in-process World implementation whose point is that nothing in it races. Every world call is a stoppable point with `before`/`positioned`/`after` phases, every call is attributed to a writer (orchestrator, a named step body, or an out-of-band external client), and the clock is virtual — a thirty-day sleep costs no wall time. A scenario script says "stop this writer here, do that, let it go", so an interleaving that a real deployment leaves to chance becomes something you can name. The model follows workflow-server where it matters: - Event ids are minted at the handler boundary, not at storage append. DynamoDB does not generate ids, so the server does, in the request handler — and that id is the log's sort key. This is what makes "earlier log position, later commit" expressible, and it is the hazard the fence scenarios are about. - The out-of-band write marker is keyed on the event's own ULID time and is forward-only, over `hook_received`, `step_completed` and `step_failed`. - Both halves of the staleness fence are modelled behind flags: the watermark (`preconditionGuard`) that clients send today, and the count (`countGuard`) that they do not. The count is synthesized on the caller's behalf and keyed by run rather than by writer, since an orchestrator and its inline step bodies are one process sharing one loaded log. `workbench/sim-world` is the scenario book — 39 of them, each a workflow plus a script. Several are pinned corruptions rather than passing assertions: they record what the runtime does today, so that a change in behaviour shows up as a diff. The doc-29/30/31 trio is the argument for the count guard, one flag apart: (B, A, C) corrupts under the watermark alone, is fenced once the count is on, and (B, C, A) corrupts with both on. That last one suspends mid-run on purpose. The fence is a conditional append evaluated inside the storage write, so there is no checked-but-uncommitted moment to slip past; the window nothing can close is the quiescent gap between deliveries, where the run makes no writes and so meets no checks. Both packages are private and unpublished, so no changeset. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Add the new sim-world workspaces to the lockfile Kept separate from the source commit because it is not a clean diff. The two new importers are the part that belongs to this branch; the rest is re-resolution churn from running `pnpm install` at a later date than whoever last touched the file — `latest` specifiers like docs' `radix-ui` move on their own. Drop or regenerate this commit if the churn is unwelcome; the source commit stands on its own. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Make a consistency violation fail the scenario that tripped it `expect.violations` let a scenario declare the corruption it reproduced and pass for reproducing it. That is a green suite describing a broken system, and it goes red on the day someone *fixes* the bug — backwards, and the opposite of what a test is for. The field is gone. A scenario now states the outcome the run should have reached, which for a corruption means the branch its own durable log implies, and stays red until the runtime gets there. Any violation fails the scenario. The six reproductions read as ordinary failures now: expected output "afterSlow:doc-26", got "afterFast:doc-26" Six scenarios are therefore red, and `run.ts` exits non-zero. The count is the signal: seven is a regression, five means something was fixed and a scenario is ready to retire. Five of the six have known fixes — four predate the count guard, and doc-29 goes green the moment a client sends `stateEventCount`. doc-31 has none, which is the point of it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Document the world-sim implementation in DESIGN.md Folds the working notes into a checked-in design doc: module map, the runtime model the simulator has to match, the interception model and its three call phases, the determinism machinery, the store's guards and fault injection, the writer vocabulary, the termination budgets, both consistency checkers, and current test status. Links it from both READMEs, and updates the workbench's doc-31 note now that the append-tail fence it needs exists as a proposal. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Say which fix closes each red scenario, and which are demonstrated The fix column read "predates the count guard" on three rows, which scans as "the count guard is involved" when it meant the opposite. It was also wrong on one: `two racing STEPS, no hook anywhere` cannot be fixed by anything predating the count guard, since the row below it is that same fault with `preconditionGuard` on and failing identically. The column now names the specific change, and a second column separates a fix that is argued for from one that is shown — a passing scenario that is the red one with the fix armed, same tempo, one flag apart. Two of the six have that; three name a fix with no paired scenario yet; doc-31 has none. Also states the thing the column could imply but does not mean: `countGuard` requires `stateEventCount`, which no client sends, so three of the five identified fixes are dark in production. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Give scenarios stable ids, refer to events by log position, add colour Four changes to make a rendered scenario something you can cite and read. Scenario ids. Every scenario gets a hyphenated `id` next to its prose name — `in-flight-after-decision`, `stale-read-step-count-fork`. The name is a sentence and will be reworded; the id is what a commit message, a bug report or `pnpm sim <id>` refers to. The runner matches the id first and falls back to the name, so existing invocations keep working, and the table of the six red scenarios in DESIGN.md now cites ids. One reference scheme. Events were referred to three ways at once: the trace printed raw correlation ULIDs, violation messages carried `evnt_<ulid>` event ids, and the docs spoke of "position 7". Now there is only log position. `#12` is the twelfth event in the durable log sorted the way `events.list` sorts it, `@7` is the resource created at position 7, and ids inside violation messages are rewritten on the way out. That numbering is deliberately in *log* order while the trace prints in *commit* order, which makes the subject of the red scenarios visible on the page: a run whose log disagrees with the order its writers committed in shows positions counting backwards. Colour, by event family, and only when the destination is a terminal. Off under `NO_COLOR` or `--no-color`, forced on with `--color`. With colour off the output is the same plain ASCII, so it stays usable as a golden file. The workflows under test are now one file. Nothing imported across them and scenarios read them alongside the tempo that steers them, so five modules were a two-file hop for no benefit. Unit tests 61/61. Scenario book unchanged at 33 passed, 6 failed. Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Add an append-only log world, and split the book into one file per scenario Three things the book could not do before. `appendOnlyLog` gives an event its log position at *commit* instead of when its handler minted one. That is the single change that makes a stale read impossible — the log can be behind, never wrong — so playing the same 39 scenarios with and without it is how you tell which of the six reds the change would actually close. All six: 33 pass / 6 violations mint-ordered, 39 pass / 0 violations append-only. It bundles two effects that have to move together: an overtaken write re-mints at the tail, and a withheld read returns a prefix rather than a hole, which is why `onStaleRead` now carries `{eventId, hidden, truncated}` and the trace distinguishes a lagging read from a stale one. `preconditionGuard` on `RunScenarioOptions` forces the fence off across the book, asking whether anything relies on it. Violations 6 -> 8 mint-ordered, so it is load-bearing there; 0 -> 0 append-only, so it is dead weight once positions are assigned at commit. Both flags are tri-state: `undefined` leaves it to the spec, which is not the same as `false`. The scenario book was a 1020-line file; it is now 39 files and an index that only decides reading order. Same 39 ids. This is the whole answer to "how do I add a scenario" — copy the file next door — and it is why the README can be short. Three `in-flight-*` scenarios are reworded so that no expectation is restated per world: a scenario is one sequence of movements, the only thing a world changes is what a read returns, and what catches the fault in both is the invariant that a run's log must replay back into that run. Also: `loadFlowHandler` moves out of `build.ts` into `load.ts`, so playing scenarios no longer drags SWC and esbuild into the module graph; the CLI gains `--report-only`, `--summary-file` and `--detail-file` for CI, with the default still exiting non-zero; and the workbench's `test` script points at `--report-only` so a recursive `pnpm -r test` does not go red for the six. Docs are split by task: the workbench README is how to add a scenario, and the package README plus DESIGN.md are how to change the simulator. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Publish the scenario book from CI, without blocking a merge Plays the book on every pull request, once per log world, and posts both summaries as one sticky comment. It never gates. Six scenarios are red on purpose — each is a reproduction of a corruption the runtime can still produce, stating the outcome its own durable log implies and staying red until the runtime gets there — so a lane that failed on them would be red on every PR and read as broken rather than as informative. What it publishes is the pair of counts, and the thing to look at is whether they still say 33/6 and 39/0. A seventh red is a regression; five means something got fixed and a scenario is ready to retire. The append-only column is the measurement the pair exists for: it says which of the six would close if positions were assigned at commit. Non-blocking at the job level rather than only on the two sim steps, so that a failed install or a PR comment the token cannot write does not turn this into a red X either. The steps themselves keep their real exit codes, so the run still says which world was clean. `--title` is new on the CLI: two summaries land in one comment, and two headings reading "world-sim scenario book" would leave the chips line as the only way to tell them apart. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * Make the Sim World comment four lines until you open it Two collapsed folds, one per world, each with its count and a green or orange dot on the visible line; the table is behind them. Nothing else above the fold. The failures list is gone. Six scenarios are red on purpose, so a comment that led with them led with the part that was not news, and grew a wall of text on exactly the PRs that changed nothing. The count is the signal — and anyone who wants the names can open the table, which has always had them. `renderMarkdownSummary` now renders no heading of its own, since it is built to be stacked under one; the workflow supplies the heading and the one-line description. `mdCell`, `clip` and `dedupe` existed only for the failures table and go with it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * docs(sim-world): drop the six-red list, add an API reference The "six red scenarios" section was a second copy of something `pnpm sim` already prints exactly, and the copy that goes stale. What replaces it says how to read a red — an open bug stating the outcome its own durable log implies, green when fixed rather than when seen again — and points at DESIGN.md for the part that is not re-derivable from a run. The reference has three tables. Writers: what a concurrent thread of execution is here, and the four ids one can have. Movements: the instruction that moves a writer to a place and holds it, described against the world boundary, the position in the event log, and the commit to storage — three moments, which is why there are three stops rather than one. Withholdings: the two ways to change what a reader sees without holding anybody. Terminology unified on hold/held (the word `Held` and the trace already use) and on "assigned a position in the event log" / "committed to storage", which also fixes the `runToEventProduced` docstring: the event has crossed the world boundary there, so "submitted" was the misleading half. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * docs(sim-world): move the API reference into the package, uniform vocabulary The reference documents `@workflow/world-sim`, so it belongs beside the API and not in the workbench that happens to be its first caller. It replaces the old `## Writers` section, which said most of the same things in a different order and a different vocabulary. Three nouns, one meaning each. A *writer* is a thread of execution — the table lists all three kinds with the handle that names each and the events each one writes. An *advance* moves one writer to a named place and holds it; the three `runTo*` stops are the three moments a write has (crossed the world boundary, assigned a position in the event log, committed to storage), which is why there are three rather than one, and the table says which way a concurrent commit sorts at each. A *withholding* hides something from readers without holding anybody, which is what `withholdNextEvent` and `beginHookDelivery` have in common and why the latter is not an advance. "Movements" is gone in favour of "advances", which the code and the older prose already used. "Held" replaces "stopped"/"paused" throughout, matching `Held` and `isHeld()`. The workbench README loses the duplicate tables and keeps the four rules that bite on a first scenario, so it is a guide to adding one and links out for the rest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * docs(world-sim): rename "Two logs" to "World behaviors", one-line motivations The section described two logs, but there are four world behaviors a scenario can pick between: the two log orderings and the two guards. Renamed and regrouped so each one is described by what it does rather than by which of the current reds it closes. Measured counts and "production" / "what happens today" framing are out of this README throughout — they belong to a run of the book, and a doc that carries them is stale the first time a scenario changes colour. The workbench README and the CI lane still print them, where they read as measurements. Every section's motivation is one line; the file-top motivation is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * docs(world-sim): Usage shows both halves — the workflow and the script The example was a script with no workflow beside it, which hid the fact that you write both. Replaced with the smallest pair that has something to control: two steps in flight at once, and a script that decides which of them reaches the log first. The scenario in it was run before it was written down — `prepare` held before it takes a position, `finalize` committed into the earlier slot, #6 then #7 — so the positions the prose cites are the ones the trace prints. Link to the workbench moved to the end of the section, where it reads as "where to go next" rather than as an aside mid-example. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * docs(world-sim): say what "arm" meant instead of using the word The READMEs told a reader to arm a wait without ever saying that calling an advance and awaiting it are two separate things. That is the whole mechanism — `runTo` registers its watch synchronously when called, and the returned promise only reports arrival — so it is stated plainly once in Advances and the jargon is dropped everywhere it stood in for the explanation. "Armed" survives only where it means a guard is switched on, which is a different word doing a different job. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * docs(world-sim): two symmetric watches in the Usage example The example started one watch and awaited the other inline, which is exactly the shape a reader cannot generalise from — it looks like the two calls do different things. Both are now started, then both awaited, so the sentence above the block and the code below it say the same thing. Re-run before committing: same log, finalize at #6 and prepare at #7. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * docs(world-sim): neutral names in the Usage example `prepare` / `finalize` / `held` / `committed` made a reader decode four domain-ish names to see a mechanism that has nothing to do with any of them. stepA and stepB, watchA and watchB: the only thing left to notice is that the two watches stop at different points, which is the whole lesson. This detaches the example from the workbench's `parallelStepsWorkflow`, so it was verified against a throwaway copy of the workflow rather than assumed — same shape, stepB at #6 and stepA at #7, replay ok. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * refactor(world-sim): one event fold, two call phases, a smaller entry Three reductions to the same end — less surface to keep consistent — with the scenario book as the check that none of them changed behaviour. **One fold.** `create` validated *and* applied an event in a 17-case switch; `foldSeededEvent` applied it again in 15 cases without validating. Two copies of the event -> entity state machine that only a replay compares, and any divergence between them makes a replay disagree with the run it is checking for a reason that is not the runtime's fault. Both paths now end in one total `applyEvent`: the write path validates first and calls it, `seedFromLog` calls it with no validation at all, because those events were accepted once already and re-litigating them would reject legitimate history. store.ts loses 175 lines. Characterization tests for the seeded path went in first and were confirmed to bite before the refactor started. **Two phases.** `CallPhase` had a `positioned` phase between `before` and `after`, and `runToPositionMinted` / `runToCall` to park on it. Nothing in the book used any of the three: the mint/commit gap they existed for is reached through `beginHookDelivery`, which owns both halves explicitly and does not block the writer that made the write. Removing the phase also removes an `await` from the interception path, so this was measured rather than reasoned about — all four book runs are unchanged. **A smaller entry.** `index.ts` re-exported the construction kit — `createSimWorld`, `createSimStore`, `driveQueue`, `verifyReplay`, `checkInvariants`, the clock — which nothing outside the package imports and which made every one of their signatures a compatibility promise. The entry is now the scenario surface; extenders import from the module. `InFlightWrite` joins it, having been missing though `beginHookDelivery` returns it. Unchanged, and the point of saying so: 33/6/6 mint-ordered, 39/0/0 append-only, 8 violations with `--no-fence`, 0 with both. 70 unit tests green (up from 67), tsc and biome clean in both workspaces. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * fix(world-sim): follow the new event-result contract Rebase fixups for four core replay commits, no behaviour change of our own. `EventResult` is now a union — a populated page (`events` + `cursor` + `hasMore`) or all three absent — so the three fields cannot be assembled from three independent variables the type has no way to see agree. They travel as one `deltaPage` object, spread as a whole or not at all. The spread has to be conditional rather than optional: `...maybe` widens each field to `T | undefined`, which is neither arm. `processExitTriggersQueueRedelivery` is gone from the `World` interface — #3385 stopped exiting on an exhausted replay budget, so there is nothing left to tell it not to. The book is unmoved across the rebase, which is the thing worth reporting: 33/6/6 mint-ordered, 39/0/0 append-only, 8 violations with `--no-fence`, 0 with both — and the same six ids red, not a swap that nets to the same count. 70 unit tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * feat(world-sim): a scenario for the unclaimed-payload delivery race Reproduces vercel/workflow#3406 as a pair of scenarios. A hook payload nobody reads registers an unarmed delivery barrier, and a step result is allowed to skip it — otherwise the workflow stalls until the barrier registry idles. The skip is transitive, so the step also skips an armed `wait_completed` merely parked behind that payload. Both branches then draw their next `step_created` id in the order the log does not record. The replay invariant cannot see this: live and replay run the same code, make the same mistake, and agree. What is checkable is the log disagreeing with itself — `wait_completed` is committed first, yet the step branch draws the earlier id. Verified red on the current runtime and green with #3406's diff applied. `claimed-payload-under-fork` is the control: same steps, same tempo, same log order through the step result, one branch awaiting the hook so the payload lands claimed. It passes in both cases. Getting there needed one new primitive. The delivery loop is serial, so a held inline step body stops the loop inside its own delivery and no timer can fire — every interleaving where a `wait_completed` lands while a step result is outstanding was unreachable. `sim.deliverQueued(select?)` takes a message out of the pending set and delivers it from the script, concurrently with the hold; `takeById` removes it first, so the loop can never pick up the same message. The book is now 41 scenarios: 34/7 with 6 violations mint-ordered, 40/1 with 0 violations append-only. The six violation reds and the --no-fence 6 -> 8 / 0 -> 0 measurements are unchanged. The new red is the first that stays red in both worlds, because no log position is wrong in it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * fix(world-sim): address review on the simulation world Fifteen findings from review, and one measurement that moved because of them. Correctness: - tempo: `abort()` stored the raw `reject` rather than the wrapper that also disposes the watch, so a point reached after teardown blocked a world call on a promise nobody could release. The deadline is one-shot and already spent by then, so that is a hang. Route abort through the disposing path and make the action a no-op once aborted. - world: `externalDepth` was a plain counter, so a step body writing concurrently with a scripted `deliverHook` was attributed to the script — invisible to any armed `runTo`, and `ext` in the trace. Measured: one misattribution book-wide. Scope it with `AsyncLocalStorage` instead. - writers: `runToEventCommitted` matched a *rejected* create, so in fence scenarios a script could wake believing a write was durable when it had 412'd. Record `failed` on the observed point and require `failed: false`. - clock/ids: reject fractional `advanceBy`, and floor in `ulid()` — a fractional millisecond silently minted `undefinedundefined…` ids that failed much later at `z.ulid()`. - streams: literal NUL bytes as key separators, `limit: 0` paging forever, and a garbled cursor slicing the array away as `NaN`. - scenario: register the handler under the exported `WORKFLOW_QUEUE_PREFIX`. The count guard, which is the substantive one. Since #3145 `@workflow/core` sends `stateEventCount` on every replay-context create and the server's count guard defaults on, so the sim's "no client sends the count" claim was stale in six places. `countGuard` now follows the fence, the sim prefers the runtime's own count over its reconstruction, and the two scenarios whose subject is isolating the watermark half say `countGuard: false` explicitly. The scoreboard does not move under the production-shaped default, which is itself the answer to the review's question. `log.monotonic-order` could never fire: its only caller fed it the sorted array. It now takes commit order, supplied only by a world that promises the two agree — under a mint-ordered log an out-of-order commit is the premise the scenario injected, not a defect. Docs: DESIGN §5 gains the two ways the sim's guards are stronger than production's (FIFO-vs-mint-order pruning, and a fence that is exact in-process where production's is region-local and fails open); §9's scoreboard is regenerated and its "no fix armed anywhere real" claim corrected; §10 gains the parallel hook-resume path, which no sim delivery takes. `step-vs-step-fork` and its twin now say who the withheld reader is in production. `unclaimed-payload-under-fork` is green after the rebase — #3406 fixed the delivery-barrier ordering — so the book is back to six reds and stays there. Counts refreshed everywhere they appear: 35/6/6 mint-ordered, 41/0/0 append-only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> * fix(world-sim): make the README's code samples type-check `packages/docs-typecheck` globs `packages/*/README.md`, so the new README went into the Docs Code Samples job and five of its `ts` blocks failed. Four were genuine excerpts with an unbound `sim` or `spec`; one was a match- object *sketch* containing a `…`, so not TypeScript at all — that one loses its `ts` tag, joining the eleven other untagged blocks in the file. The rest name what they use. That needed two `paths` entries in the checker: `@workflow/world-sim` was unmapped, and an unresolved import is deliberately tolerated there (`isExpectedMissingModule`), so every identifier in those samples was `any` — they would have gone green while checking nothing. Mapped, they are checked against the real declarations: seeding `id: 12345`, `readyAtMsTYPO`, `dirsTYPO` and `runToEventCommittedTYPO` is caught against `ScenarioSpec`, `PendingMessageView`, `SimBuildOptions` and `Writer`. The lead teaser keeps its four unadorned lines and takes a `@skip-typecheck` marker instead; the same calls appear in full, checked form further down. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: shalabhc <shalabh.chaturvedi@vercel.com> --------- Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
42 KiB
world-sim: design
How @workflow/world-sim is built and why it is built that way. The package
README is the introduction; this is the implementation.
Two workspaces:
| path | what it is |
|---|---|
packages/world-sim |
the World implementation, the scenario runner, the checkers |
workbench/sim-world |
the scenario book and the workflows it runs (pnpm sim) |
1. Module map
| module | responsibility |
|---|---|
world.ts |
the World implementation; wraps every method as a call point and attributes it to a writer |
store.ts |
in-memory event store — the event → entity state machine, plus the write-time guards |
queue.ts |
deterministic queue: records messages, never delivers on its own |
clock.ts |
virtual clock; patches Date.now() readings, not timers |
ids.ts |
deterministic ULID minting from (virtual time, counter) |
drive.ts |
the scheduler loop and the scenario budgets |
tempo.ts |
the scripting layer: park / permit, and the wait bookkeeping under runTo |
writers.ts |
named writers and their level-triggered runTo* vocabulary |
scenario.ts |
runs one ScenarioSpec end to end and produces a ScenarioResult |
replay.ts |
cold-start replay verification of a finished log |
invariants.ts |
consistency checks re-derived from the event log alone |
report.ts |
renders a scenario; log positions for every event reference, colour only when the destination is a terminal |
streams.ts |
in-memory streamer |
build.ts |
bundles a project's workflows so the runtime can be handed real compiled code. Its own entry (@workflow/world-sim/build), because it reaches a compiler and playing a scenario should not |
load.ts |
loads a built bundle's flow handler — the half of the old build.ts that needs no compiler |
types.ts |
the public vocabulary |
2. What the simulator has to model
The design follows from four properties of the runtime. None of them are choices this package made; they are the constraints it works inside.
The orchestrator is re-run from the top, every time
A workflow function is not a coroutine that parks and resumes. It is
re-executed from its first line on every replay pass, inside a fresh node:vm
context (createContext in packages/core/src/vm/index.ts). The VM seals the
two obvious sources of nondeterminism: Math.random is seeded from
${runId}:${workflowName}:${deploymentId}, and Date.now() / new Date()
return a fixed timestamp advanced only from each consumed event's createdAt.
A pass is therefore a pure function of (workflow code, run identity, event log prefix). Stated precisely: same log prefix → same decisions. That is the whole basis of durability, and it is also the property the simulator exists to attack — the interesting bug class is not "the workflow behaved randomly" but "the decisions and the persisted log disagree", which requires the decisions to have been made against a different log than the one that ended up durable.
Entity identity is positional
Correlation ids come from ctx.generateUlid(), driven by the VM's seeded
Math.random. They are positional ordinals of one seeded sequence: the Nth
entity the workflow asks for gets the same id in every pass. Steps, hooks and
waits all draw from that one sequence.
This is why a flipped branch is dangerous. It does not produce a different step; it produces a different step wearing the same name badge. The runtime's divergence check is a step-name comparison at the same ordinal:
Replay divergence: step event step_created for step_…445J belongs to
"…//settle", but the current step consumer is "…//recoverFirst"
Suspension is the unit of progress
useStep does the same thing on every pass: mint the correlation id, register
a StepInvocationQueueItem, subscribe a consumer, return a promise. What
differs is what the consumer finds. On replay the log holds
step_created / step_started / step_completed for that correlation id, so
the consumer hydrates the recorded result and the step body is never called.
First time through, the consumer reaches the end of the log, returns
NotConsumed, and the promise never resolves — so the workflow cannot proceed.
When nothing can make further progress a WorkflowSuspension is raised
carrying the whole invocationsQueue. The runtime commits the pending
*_created events, executes what it can, and runs the workflow again from the
top against a longer log:
load log → run workflow from top → suspend → commit + execute → run from top → …
Hooks follow the same shape with a different event family. hook_created is
committed at the next suspension rather than at the call. An out-of-band
resumeHook(token, payload) writes hook_received and enqueues a flow
message, so the run wakes up. Payloads landing before the workflow awaits are
buffered in a payloadsQueue, which is why a duplicate delivery is absorbed
rather than lost.
There is no hook state a workflow can read — the surface is token,
getConflict(), dispose(), then, [Symbol.asyncIterator]. The only way to
observe a hook is to attach a continuation and see whether it resolves, which
is a timing observation, not a state read. That is why the scenario API
steers time rather than poking state.
Nothing sleeps
sleep() registers a WaitInvocationQueueItem; wait_created records a
resumeAt; the runtime enqueues a delayed queue message. When that message is
delivered, the flow handler's "complete elapsed waits" pass compares
Date.now() >= resumeAt and writes wait_completed.
A timeout is a delayed message plus a clock comparison. That is exactly why virtual time works here, and why a thirty-day sleep costs microseconds.
Where the concurrency actually is
The workflow body is single-threaded JS and stays that way. The interleaving that matters lives in three other places:
- Between passes. The log grows. A branch decided in pass N was decided against pass N's prefix.
- Between event deliveries inside one pass. Several awaits can be pending
at once. The runtime forces resolution order to match log position via the
delivery-barrier registry (
registerDeliveryBarrier/awaitEarlierDeliveriesinprivate.ts), so this one is reproducible. - Between invocations. Two flow deliveries for one run can execute simultaneously in different processes. This is the one the SDK cannot make deterministic by itself.
Two step bodies that suspend together are already concurrent writers to one log, inside a single delivery. That is cheaper to reach than "concurrent writers" suggests, and it is the case the writer vocabulary is built around.
3. Interception
Every World method is a call point
Each method on the World is wrapped so a scenario can stop it. The wrapper records the call, fires any matching watches, runs the underlying implementation, fires the watches again on the way out, and only then resumes the caller.
A watch action returns a promise, and the intercepted call awaits it. That is
the entire hold mechanism: a "held" writer is a caller blocked inside a World
method. release() resolves that promise.
Two phases, and a third hold that is not one
type CallPhase = 'before' | 'after';
before and after bracket the call:
| held at | what a competing write does |
|---|---|
before |
no position taken yet, so a write landing during the hold sorts ahead |
after |
durable, and the writer has not been resumed yet |
after is the window the package was originally built for: "the hook arrives
after step_started is durable and before the workflow resumes".
Neither phase produces the opposite order, and a write is not atomic, so the opposite order is reachable: a real backend mints the event id first — DynamoDB does not generate ids, and that id is the log's sort key — and only then attempts the storage write. Between the two the event has a position but no visibility, and a write that commits in that window sorts behind it.
That gap is the point, not a detail. It is the only way to produce an event behind a position a reader has already read past — a complete, consistent log prefix that is simply missing an event still in flight. No high-water-mark fence can represent that shape.
It is not a phase, though, because the writer holding it is not blocked inside a
World method: the script owns the two halves explicitly, via reservePosition /
withReservedPosition under sim.beginHookDelivery (§Withholdings below). A
phase would have been a second way to say the same thing, and it went unused.
Watches do not fire inside watches
Calls made from inside another watch's action are not call points. Without
that rule a watch on events.create would re-trigger on the hook_received
it just wrote, and every scenario using deliverHook would recurse forever.
The depth is tracked and surfaced in the trace, so a line committed from inside
a held call is visibly at depth > 0.
A related rule is easy to get wrong: the depth counter must be raised only by
asExternal, which brackets exactly one call. Raising it for the whole
duration of a watch action is correct for something that returns immediately
and wrong for a hold, which does not return until the scenario releases it —
under that rule, holding one writer makes every other writer's call stop being
a call point, so a held step body's sibling becomes invisible and unsteerable.
Writer attribution is derived, not instrumented
Writer identity comes from the intercepted call plus its request. No runtime hook is needed:
| write | writer |
|---|---|
step_created / hook_created / wait_created / run_* / attr_* |
orchestrator |
step_started |
orchestrator (the executor, which precedes the body) |
step_completed / step_failed @ correlationId C |
step:<name of C> |
hook_received |
external |
wait_completed |
the wait-continuation delivery |
events.list / runs.get |
orchestrator (the read half) |
The writer is printed as a column in every event stream, so a rendered log says who wrote each line.
4. Determinism machinery
Clock
install() patches Date.now() and the zero-argument Date constructor to
read the virtual clock. Timers are deliberately not patched:
@workflow/core uses setTimeout(fn, 0) as a macrotask barrier in several
ordering-sensitive places (events-consumer.ts, private.ts), and swapping
those for fake timers would change the very interleavings the simulation exists
to observe. Real zero-delay timers stay real; only the readings of wall time
move.
The clock never moves on its own. Only the scheduler calls advanceTo /
advanceBy, so two runs of a scenario see the same sequence of timestamps.
Ids
Every id is a function of (virtual time, per-scenario counter) — never
Math.random() or the host clock. They still have to be real ULIDs, because
@workflow/world validates run ids with z.string().ulid() and decodes the
embedded timestamp, so the encoding is standard Crockford base32 with the
16 "random" characters filled from the counter.
Byte-identical ids run to run are what make an event-stream dump usable as a golden file.
Queue
@workflow/world-local's queue fires a detached delivery loop from inside
queue(), so a message races whatever the caller does next. Faithful to
production, useless for a simulation.
Here queue() only records. Delivery happens when the scheduler asks, and it
always takes the same message: the minimum by (readyAtMs, enqueueSeq). Delays
are virtual — a message 23 hours out is delivered by jumping the clock.
ScenarioSpec.selectNext can override the choice to pin an order the default
would not produce.
Scheduler
take the next message → jump the clock to its delivery time → hand it to the
flow handler → wait → repeat until the queue is empty or a budget stops it
Between deliveries the loop drains the event loop for several rounds
(settle()), because the runtime uses zero-delay macrotasks as ordering
barriers and waitUntil-style background work is not awaited by anyone. Without
that drain, a message enqueued from a trailing microtask would be missed and the
scenario would report a spurious stall.
The scheduler lives apart from scenario.ts because two things drive it: a
scenario, and the replay verification that cold-starts a second world.
One delivery at a time. This is the deliberate limit of the model — see §9.
5. The store
A reference implementation of the World storage contract: the same event →
entity state machine @workflow/world-local implements on the filesystem,
minus every mechanism that exists purely to make that state machine safe
against concurrent processes (exclusive-create claim files, per-entity locks,
staged/promoted hook events, canonical event-id pinning after a crash). One
delivery at a time in one process means those races cannot occur, and their
absence keeps the file small enough to audit.
What is deliberately kept is every validation that rejects an event — terminal-run guards, step lifecycle ordering, hook token uniqueness, wait duplication. Those rejections are the observable contract the runtime is written against; a simulation that relaxed them would agree with the runtime about nothing interesting.
The two guards
The fence is off by default and set per scenario
(ScenarioSpec.preconditionGuard, which flows into SimWorldOptions), which is
how a scenario can be run one flag apart from its neighbour. countGuard
follows the fence unless a spec says otherwise, because that is what
production does — see below.
preconditionGuard models WorldCapabilities.preconditionGuard: reject a
replay-context write whose stateUpdatedAt snapshot predates the newest
externally-originated event. In the SDK this is declared by world-vercel
only (packages/world-vercel/src/index.ts:40); world-local and
world-postgres declare neither it nor maxConcurrency.
Its predicate is narrower than the bug class, and the reason is its shape,
not the event type it watches. The marker advances on hook_received or
step_completed, but it is a high-water mark — the newest such write — and
the test is stateUpdatedAt < marker, strictly. So it detects a log truncated
at the end and is blind to a hole in the middle: when the withheld event is
older than one the reader can see, the reader's snapshot is never strictly
older than the mark. The hook direction is caught for the mirror-image reason —
the withheld hook_received is the newest out-of-band write and the
orchestrator's snapshot predates the sleep, so the fence fires and the run
reconciles.
countGuard adds the count half: how many events the log holds at or below
stateUpdatedAt, compared against how many the caller loaded. It closes the
hole the watermark cannot see. It requires the caller to send
stateEventCount — and since #3145 (1471f252f) @workflow/core sends it on
every replay-context create, gated only by the WORKFLOW_PRECONDITION_GUARD
kill-switch, with workflow-server's own count guard defaulting on. So both
halves are armed together in production, and countGuard defaults here to
whatever the fence is set to. A run with the fence on and the count off is a
world that exists nowhere; the two scenarios that ask for it
(step-vs-step-fork-fenced, in-flight-before-decision) do so explicitly,
because isolating the watermark half is their whole subject.
Where the runtime sends a count, the sim uses that value rather than its own
reconstruction (loadedCount()), so the guard is tested against the number the
real client computes; the reconstruction is the fallback for writes core does
not count.
Two ways the sim's guards are stronger than production's. Both are deliberate, and both mean a fenced green here is a claim about the predicate, not about production's deployment of it:
- The server's retained-id window is a FIFO in insertion (commit) order —
its Lua script prunes with
table.remove(ids, 1), oldest-inserted — whilepruneRunEventIndexhere sorts by id and drops the smallest, i.e. mint order. The two differ exactly when commits happen out of mint order, which is these scenarios' whole subject, and they differ in whencountRecordedAtOrBelowgoes indeterminate once a run passes 16 events. - Production's watermark is best-effort: region-local Redis, failing open on
Redis errors, and blind to a webhook served in another region entirely (see
outside-event-tracker.ts's own docs). The sim's is exact and in-process.
Overriding them for a whole run. RunScenarioOptions.preconditionGuard
(pnpm sim --fence / --no-fence) forces the fence on or off for every
scenario, undefined leaving each spec to decide. Both halves move together:
the count guard is evaluated inside the same predicate, so disarming the fence
disarms it too.
Forcing it off across the book asks whether anything relies on it. Violations
go 6 → 8 mint-ordered, so it is load-bearing there; 0 → 0 against an
append-only log, so it is dead weight once positions are assigned at commit.
This is a diagnostic rather than a world, and the number to read is the
violation count: a scenario whose subject is that the guard fired asserts that
with sim.check and fails by design when it does not
(in-flight-before-decision-counted today).
Fault injection
withholdNextEvent(reads = 1) hides the next committed event from the
following N event-log reads. This is the only way a serial simulation can
produce "a write derived from an incomplete event load", the precondition a
real deployment reaches through concurrency.
It is a faithful model rather than an approximation, because production reaches
the same ordering natively: world-local mints evnt_${monotonicUlid()} near
the top of createImpl (packages/world-local/src/storage/events-storage.ts)
and writes the file much later, so two concurrent creates take positions N and
N+1 and can land in the opposite order. Postgres does the same via nextval
before COMMIT. world-local defends this with mintRunDominantEventKey
(src/storage/helpers.ts) — but only for terminal run events;
wait_completed gets no re-derivation.
One withheld read poisons a whole invocation, which is worth knowing when
reading a trace: after the next step_completed the runtime continues from its
cursor, fetching only events written strictly after that position. A withheld
event sitting before the cursor can never re-enter that invocation's view.
Incremental reads make the hole permanent.
beginHookDelivery(token, payload) returns an InFlightWrite — a write
held between mint and commit, with eventId already fixed and commit()
still pending. Unlike a held writer, nothing is blocked meanwhile, because the
receiver is a separate process from the run's invocation. Holding an inline
write instead would stall the delivery that made it, and thus the reader too,
which is why the out-of-band writer is the one that can express this shape.
Changing the world instead of the runtime
appendOnlyLog is the one option that alters the store's contract rather
than its strictness. With it on, an event takes its position in append instead
of at the handler boundary: a write that is still the newest when it commits
keeps the id it was already handed out under, and one that was overtaken while
it was held re-mints and takes the tail.
That single move collapses both faults above into the same, weaker one. A hold
between mint and commit can no longer open a hole, because the held write is not
claiming a position while it waits — it has none until it lands. And
withholdNextEvent degrades from serving a read around the withheld event to
stopping it at the event, because a hole is not expressible in a log whose
order is its commit order. Both leave the reader short rather than wrong, and
short is precisely what the fence's watermark was designed to catch.
Off by default: the sim exists to model the world that exists, and production
mints at the boundary because DynamoDB does not generate ids. The value of the
switch is differential — play the book both ways and the diff separates "fails
because of the mint-before-commit window" from "fails for some other reason".
No scenario in the book sets it; it is meant to be driven from
RunScenarioOptions or pnpm sim --append-only.
6. The scenario surface
Spec
interface ScenarioSpec {
id: string; // stable hyphenated handle; what `pnpm sim <id>` selects
name: string; // prose, expected to be reworded; the id is not
description?: string;
workflow: string | { workflowId: string }; // plain fn name, resolved via the build manifest
input?: unknown[];
script?: ScenarioScript; // omitted = a control: run on the default schedule
selectNext?: SelectNext; // override queue delivery order
verifyReplay?: boolean; // default on for runs reaching completed/failed
expect?: { status?: ScenarioOutcome; output?: unknown }; // output: deep equality
limits?: ScenarioLimits;
preconditionGuard?: boolean; // advertise + enforce the optimistic-concurrency fence
countGuard?: boolean; // also enforce its count half
appendOnlyLog?: boolean; // position at commit, not at mint; see §5
}
RunScenarioOptions.appendOnlyLog overrides the last of those for every
scenario in a run, which is how the whole book gets played both ways;
undefined there leaves each spec to decide. The mode a result was produced
under is recorded on ScenarioResult.appendOnlyLog rather than left to the
reader's memory.
expect.status accepts the non-run outcomes (stalled, budget-exceeded)
because "this workflow deadlocks when the hook never arrives" is a property
worth pinning down rather than an accident to tolerate.
There is deliberately no way to expect a consistency violation. A scenario reproducing a corruption states the outcome the run should have had and fails until the runtime delivers it. A red is an open bug, not a recorded observation, and it goes green when the bug is fixed rather than when the bug is seen once more.
Scripting
ScenarioApi is the complete set of sanctioned external inputs — anything a
real deployment could do out-of-band has an entry, so the script is a complete
description of what happened:
deliverHook · beginHookDelivery · cancelRun · advanceTime ·
withholdNextEvent · note · check · world (read-only snapshot) · runId
Tempo adds the steering: writer handles, plus the raw park / until /
during primitives. The vocabulary is borrowed from Python's blanket, which
does the same for threading primitives — the call parks, the script issues
the permit, and the resulting order of permits is the tempo.
Writers
sim.writer.orchestrator() / .step(shortName) / .anyStep() / .any()
return a Writer. A handle is a name, not a live object: it can be taken
before the step exists and binds to whichever writer shows up under it.
| method | phase | meaning |
|---|---|---|
runToEventProduced |
before |
decided and submitted, nothing in the log |
runToEventCommitted |
after |
durable, writer not yet resumed |
release |
— | let it go; idempotent |
isHeld / history |
— | inspection |
Two implementation details of release() matter to scenario authors. It is
guarded by a done flag so double release is a no-op. And it awaits a full
macrotask turn before resolving — without that, await release() returns while
the resumed call is still queued as a microtask, and a scenario reading the log
on the next line sees the state it was trying to leave.
runTo is level-triggered
It consults the history of points the writer has already reached before
arming anything, and throws AlreadyPassedError naming the call it happened at
if the point has gone by.
The alternative — arm a watch and wait — is a hang. A held call blocks its writer, and when that writer is the one the scheduler is inside, it blocks the loop; so there is no quiescence to fall back on and no timer to eventually fire. An edge-triggered wait on an edge that has passed is the one way to lose this package's termination guarantee, so it is made impossible rather than documented.
Three consequences:
- Holds must be armed before they are needed. To catch two writers at the same point, start both waits and then await them. Awaiting the first before starting the second yields the event loop, and the other writer may sail past.
runToon an already-held writer releases it first, and arms the new watch before releasing. That order is load-bearing: the released writer can reach the next point within the same turn — theafterphase of the very call it was held in is the common case — and a watch armed afterwards would miss it. The same rule applies to authors sequencing two writers: arm B before releasing A.- A call is two records, so
seqcannot order them. Thebeforeandafterphases of one call share aseq, so each recorded point carries its ownordinaland the level check compares against that.
A watermark tracks how far each writer has been advanced. Points at or before
it are "already consumed" and do not count as already-passed — asking twice for
step_completed means the next one, which is what the duplicate-delivery
scenarios need.
What is not offered
Writers form a dependency graph — the orchestrator awaits its own step bodies —
so not every interleaving exists to be asked for, and an unsatisfiable runTo
can only be reported, not prevented. The runtime's await graph is not visible
from here, so true deadlock detection is out of reach; the substitute is a
per-runTo watchdog that reports where every writer was standing.
7. Termination
Every scenario terminates. Four budgets, layered so the most specific one reports first:
| budget | default | catches |
|---|---|---|
maxRunToWallMs |
5 s | one runTo that will never be satisfied |
maxDeliveries |
200 | a run that keeps re-enqueueing itself |
maxVirtualMs |
365 d | while (true) { await sleep('1d') } |
maxWallMs |
60 s | a genuinely non-terminating step body |
maxRunToWallMs sits far below maxWallMs on purpose: it can name which
writer failed to reach which point and where the others were standing, and that
diagnosis is worth more than the generic "ran out of wall clock" the global
deadline can offer. It is clamped to maxWallMs so lowering the global budget
does not require remembering to lower this one.
The scenario's global deadline must not be unref'd. An unref'd timer does
not hold the event loop open, so a total deadlock — every writer held, scheduler
blocked inside a held call, script awaiting the impossible — empties the loop
and exits Node with a bare "unsettled top-level await" instead of firing the
watchdog, which is precisely the case the watchdog exists for. The finally
already clears it, so it cannot outlive a scenario.
Stream readers get the same treatment: a reader that parked on an unfinished
stream would deadlock the scenario, so readers park on a promise the writer
resolves and abortOpenReaders() releases any still parked at teardown —
turning a hang into a reported diagnostic.
Outcomes are WorkflowRunStatus | 'stalled' | 'budget-exceeded' | 'error'. A
hook that never arrives is reported as a stall naming the undelivered token,
not a hang.
8. Consistency checking
Two independent checkers run over every scenario.
Invariants
The store enforces most rules at write time by rejecting bad events — but "the
store rejected it" and "the log is actually consistent" are different claims,
and only the second is worth trusting. So invariants.ts re-derives everything
from the event log alone and compares against the entity rows.
25 rules, grouped:
log.monotonic-order log.unique-event-id
run.created-first run.created-once run.terminal-is-last
run.entity-matches-log run.attributes-match-log run.output-materialized
run.resources-released
step.created-once step.started-after-created step.terminal-after-created
step.terminal-once step.no-restart-after-terminal
step.entity-has-log step.entity-matches-log step.attempt-matches-log
hook.token-unique hook.received-after-created
hook.dispose-once hook.no-receive-after-dispose
wait.created-once wait.completed-after-created
wait.completed-once wait.resume-at-stable
A violation is a bug somewhere — in the runtime that produced the sequence, in the store that accepted it, or in the scenario that injected something impossible. Which one is a question for the reader; the checker's job is only to notice.
Replay verification
The invariants check the log's shape. None of that answers the question durability actually rests on: if a fresh process picked up this log tomorrow, would it reconstruct the same run?
The check is a cold start with the answer withheld. Take the finished log,
drop its terminal run_* event, load the rest into an empty world as durable
history, and deliver one queue message. The real runtime — the same
workflowEntrypoint a deployment serves — replays from the log and must
re-derive the event that was removed, with the same output. No step body
re-executes, since every step_completed is in the log and the step consumer
resolves from it, so anything the replay produces came from the log alone.
In this frame, replay is the serializability check. A pass is pure, so re-running it over the committed log asks whether the schedule had a serial equivalent. Six failure ids:
replay.diverged · replay.suspended · replay.output-differs ·
replay.log-differs · replay.status-differs · replay.budget
replay.diverged is the runtime raising ReplayDivergenceError, exhausting
its recovery replays, and failing the run with CorruptedEventLogError.
replay.suspended means the replay ran out of log before the workflow
finished — the log did not contain enough to rebuild the run.
9. Current status
Measured on branch sim-world.
Unit tests — 72 passing across 8 files (pnpm --filter @workflow/world-sim test).
Scenarios — pnpm sim in workbench/sim-world:
41 scenario(s): 35 passed, 6 failed, 6 consistency violation(s)
And the same book against an append-only log (pnpm sim --append-only):
41 scenario(s): 41 passed, 0 failed, 0 consistency violation(s)
Both numbers are the intended steady state; see "The six" below for which of the six violations that second line closes on the merits and which close because the correct answer itself changes.
There was a seventh red until recently, unclaimed-payload-under-fork, and it
was a different animal: it tripped a sim.check rather than the replay
invariant, and it was red in both worlds, because nothing was wrong with its
log's positions — the runtime handed the workflow two resolutions in an order
the log did not record, so live and replay ran the same code, made the same
mistake, and agreed. Only the log disagreeing with itself caught it. #3406
fixed the delivery-barrier ordering and it is now green in both worlds; the
scenario stays as that fix's regression test.
With the fence forced off (pnpm sim --no-fence), violations go to 8
mint-ordered and stay at 0 append-only — see §5.
Replay verification across the book: 33 ok, 6 MISMATCH, 2 skipped
(skipped where the run did not reach a terminal status).
run.ts exits non-zero, and that is the intended steady state. The six
failures are reproductions of corruptions the runtime can still produce; each
states the outcome its own durable log implies and fails until the runtime gets
there, so the failure line names both sides (expected "afterSlow:doc-26", got "afterFast:doc-26").
The number is the thing to watch: six today. A seventh is a regression; five means something got fixed and a scenario is ready to retire.
That makes the book a poor plain CI gate, which is what --report-only is for:
it prints every failure and exits 0, so a job can publish the book's current
state rather than block on it. --summary-file writes one collapsed
<details> — a visible line carrying the count and a green or orange dot, the
whole table behind it — for a PR comment or $GITHUB_STEP_SUMMARY, and
--detail-file writes the full colour-free trace as an artifact to read when a
number moves. Deliberately nothing above the fold but the count: six are red on
purpose, so a comment that leads with the failures leads with the part that is
not news, and grows a wall of text on exactly the PRs that changed nothing. The
workbench's pnpm test is --report-only --summary-file, so a recursive
pnpm -r test stays green and still says what happened; pnpm sim stays
strict, so running it by hand fails loudly.
The six
The fix column names the specific change that closes the scenario. shown green by is the stronger claim: a passing scenario that is this one with that
fix armed, same workflow and same tempo, one flag apart. Where it says "none
yet", the fix is identified by argument but nothing in the book proves it.
| scenario | mechanism | fix | shown green by |
|---|---|---|---|
stale-read-step-count-fork (doc-23) |
withholdNextEvent(1) + deliverHook; hook at #7, wait_completed at #8, no-hook branch at #9 |
preconditionGuard — the withheld hook_received is the newest out-of-band write and the orchestrator's snapshot predates the sleep, so the watermark fires |
stale-read-step-count-fork-fenced (doc-24) |
stale-read-equal-step-counts (doc-25) |
same fault on a fork whose branches emit one step each | preconditionGuard, for the same reason |
none yet |
step-vs-step-fork (doc-26) |
two of the run's own step_completed events, one delivery |
countGuard. Not preconditionGuard: the withheld completion is a hole in the middle of the log, which moves no high-water mark (§5) |
none yet |
step-vs-step-fork-fenced (doc-27) |
same, preconditionGuard: true, zero rejections |
countGuard. This row is the proof that the watermark half does not fix doc-26 |
none yet |
in-flight-before-decision (doc-29) |
beginHookDelivery, committed before the decision is written |
countGuard |
in-flight-before-decision-counted (doc-30) |
in-flight-after-decision (doc-31) |
beginHookDelivery, committed after the decision |
none in the SDK. Needs an append-tail fence — assertSlotAboveTail, vercel/workflow-server#692 |
— |
Those handles are ScenarioSpec.id, and they select: pnpm sim in-flight-after-decision plays one row of this table.
So: two of the six have their fix demonstrated by a paired green scenario, three have a fix identified but unproven here, and one has no fix at all. Writing the three missing pairs is the obvious next increment.
Note what the fix column does not mean. countGuard closing doc-29 is a
statement about the World implementation and about production's predicate:
core has sent stateEventCount on every replay-context create since #3145 and
the server's count guard defaults on (§5). What it is not is a statement about
production's deployment of that predicate, which is region-local, fails open,
and prunes its window in a different order than this store does — all three
noted in §5. So the honest reading of the column is: four of the six have a fix
whose predicate is armed in production today, and whether it fires there depends
on conditions the sim does not model.
The append-only log closes all six, in two different senses — and the split is four and two, not three and three. Four (doc-23, doc-25, doc-26, doc-27) close on the merits, with the book asking them exactly what it asked before: the reordering was the fault, and once positions are assigned at commit the withheld read degrades from a hole to a truncation, which the fence can see. The remaining two (doc-29, doc-31) close because the branch the run ends on changes. A hook that commits after the timeout genuinely is after it when the tail is the only place a write can land, so the log records the timeout first and the run that settled is the run the log describes.
No expectation is restated per world, and there is no mechanism to. The
first cut of this had one — an expectAppendOnly field on three scenarios,
naming a second correct output. It was the wrong instrument, for a reason worth
keeping written down. A scenario is one sequence of advances. The only thing a
world changes is what a read returns. The branch a run ends on is decided by
what it read, so pinning the branch pins a consequence of the world rather than
a property of the run, and any expectation that then has to be restated per
world is evidence the pin was wrong — not evidence that a second answer is
needed. The three now assert what holds in both worlds (the run completes) and
report the branch in the trace.
That costs nothing, because the expectations were never what caught the fault.
The load-bearing assertion is the invariant: the log a run wrote must be a log
the runtime can replay back into that same run. It is world-independent, on by
default (verifyReplay), and it is what all six reds trip. Measured, not
assumed: strip every expect in the book and the violation counts do not move
— 6 mint-ordered, 0 append-only, the same six by name. (Pass/fail does move by
one, and only for a bookkeeping reason: hook-never-arrives expects stalled,
and a stall's reason is reported as a problem unless the scenario said it was
expecting one.) That also removes
the one place where the flag's scoreboard rested on a judgement about what the
right answer is rather than on something the harness checks on its own.
doc-30 is worth a line because it was the third expectAppendOnly and is not
one of the six — mint-ordered it already passes, since countGuard catches
there what the watermark half misses. Its branch moves under the flag for the
same reason its uncounted twin's does, so the old pinned output would have
turned a green scenario red. What made it distinct from doc-29 was never the
branch anyway; it is that the count half of the fence fires at all. That is now
asserted directly, matched on the guard's own message and true in both worlds —
and it fails if countGuard is turned off, which is the check that a bare
rejections().length > 0 would have missed, since doc-29 rejects too.
Two details worth keeping:
- Hook delivery participates.
beginHookDeliverystill reserves a position at the handler boundary; under the flag the reservation stops being binding and the write re-mints at the tail if anything overtook it (positionAtCommit). - doc-30's 412 still fires, and now saves nothing. The count guard counts events at or below the caller's watermark, the watermark is a millisecond, and a hook committing after the timeout within the same virtual millisecond is still "at or below" it. The reload finds nothing to correct and the run settles anyway. That false positive is the standing cost of the count half of the fence once the log is append-only, and doc-30's trace is where to see it.
Four of the six are hook-driven and two deliberately are not — the pair proves the corruption needs no out-of-band event type. All the pure hook-timing scenarios pass: placing a hook precisely is what works. What fails is a hook that is durable in the log but absent from the read the live pass decided on.
The last row is the only one with no fix in the SDK: the hole opens after the
write that should have fenced it, in the quiescent gap between deliveries where
the run makes no writes and so meets no checks. assertSlotAboveTail in
vercel/workflow-server#692 is the append-tail fence for it.
10. Limits
Concurrent invocations are out of reach. The scheduler does
await deliver(...), so two flow deliveries for one run cannot overlap.
Reaching that would need concurrent delivery with hold points to pin the
interleaving. The gap matters because it is a real production route:
resumeHook writes hook_received and enqueues a flow message, so two
deliveries end up in flight — one writing wait_completed and deciding
no-hook, one seeing the hook and deciding hook-branch — racing to create the
same ordinal, with every reader holding a perfectly consistent view. Just
different ones.
Two step bodies inside one delivery are genuinely concurrent and separately steerable, which is enough to reach the interesting corruption without a second invocation. That is why the limit has been acceptable so far.
The parallel hook-resume path is never exercised. resumeHook picks
parallel ("lazy") vs sequential from world.capabilities.hookResumeDedup (or a
fresh server attestation). world-local declares it and world-vercel attests
it per lookup, so every real world takes the parallel path, where the queue
publish races the hook_received write and the consumer re-ensures the event
through the durable (runId, resumeId) claim. The sim advertises neither the
capability nor a resumeId dedupe, so every sim hook delivery takes the
sequential path — meaning the hook-timing shapes in this book are the legacy
shape, not the one production runs. Closing this needs (runId, resumeId)
dedupe in the store plus the capability; it is the largest single gap for a
package about hook races.
Also untested: turbo / optimistic-inline-start, which skip replays and so
give a stale branch somewhere to hide; and the fence's same-millisecond
behaviour, where an equal stateUpdatedAt passes by design as anti-livelock.
Not modelled at all: the concurrency machinery world-local needs and this
store omits — claim files, per-entity locks, staged/promoted hook events,
canonical event-id pinning after a crash. Bugs in those are invisible here.
11. A caveat worth stating
A simulated world only produces trustworthy results while its model matches
reality. Every simplification in §5 and every limit in §10 is a place where a
green scenario could be green for the wrong reason. The mitigations are that the
store keeps every rejection the real one performs, that the runtime under test
is the real workflowEntrypoint running real compiled workflow code, and that
every scenario ends by replaying its own log through that same runtime — but
none of those is a proof, and a red here is worth more than a green.