Skip to content

[workflow] Load per-source replay bundles - #3551

Draft
NathanColosimo wants to merge 9 commits into
codex/replay-prune-runtime-graphfrom
codex/replay-lazy-source-bundles
Draft

[workflow] Load per-source replay bundles#3551
NathanColosimo wants to merge 9 commits into
codex/replay-prune-runtime-graphfrom
codex/replay-lazy-source-bundles

Conversation

@NathanColosimo

@NathanColosimo NathanColosimo commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Summary

  • emit one VM bundle per workflow source; workflows in the same source share one bundle
  • keep one flow route with a static workflow-ID → dynamic-import loader map
  • overlap the selected bundle load with world initialization and event setup
  • skip bundle loading for queue-only step deliveries
  • retain compiled scripts for every immutable production source bundle
  • keep watch loaders cache-safe and derive loader IDs from the same transform that emitted each bundle
  • encode generated VM modules as opaque artifacts so framework plugins cannot rewrite inert source

Result

For the Next/Turbopack benchmark app's 97_bench source, measured against the monolithic bundle at the base of this PR:

  • generated flow route: 1,327,514 B → 18,264 B (-98.6%)
  • selected benchmark VM code: 1,307,595 B → 133,859 B (-89.8%)
  • cold compile p50: 10.405 ms → 1.089 ms (-89.5%)
  • fresh-context evaluation p50: 7.484 ms → 0.356 ms (-95.2%)
  • one-time base64 artifact decode p50/p90: 0.037/0.044 ms

The compile/evaluation microbenchmarks use 60 uncached vm.Script compilations and 80 evaluations in fresh VM contexts on Node. The deployment benchmark and APM traces remain the end-to-end signal.

The correctness-preserving split deliberately includes any source that defines both workflows and custom serializers in every other source bundle. The workbench's unusually large 99_e2e.ts is such a hybrid, so the selected benchmark bundle is ~134 KB rather than the ~28 KB it would be without that shared registration. The next structural improvement is to extract serializer registration into a dedicated small VM prelude; omitting it would make replay deserialization incorrect.

The 19 split bundles total 5.28 MB decoded versus the old 1.31 MB monolith because shared sandbox dependencies and the hybrid serializer source repeat. Opaque base64 transport is 7.04 MB on disk before compression (1.33× decoded; gzip is close to the original text). This is a build/storage tradeoff: a replay imports, decodes, compiles, and evaluates only its selected source bundle, and the route caches that decoded promise.

Deployment benchmark

The final benchmark run completed successfully on deployment dpl_CHJqdRSi52QvbKm1KzXhEs4Aa1B2; the sticky comparison reports:

  • inline STSO p75 191 → 169 ms (-12%), p90 229 → 191 ms (-17%), p99 580 → 268 ms (-54%)
  • whole-run overhead for 1,020 steps 195,405 → 167,186 ms (-14%)
  • an earlier run of the same runtime code measured STSO p75/p90/p99 at 155/172/264 ms and whole-run overhead at 153,275 ms, so the improvement survives expected cold-start and infrastructure variance
  • TTFS p75 remains 1.12–1.50 s across scenarios, showing that dispatch, route cold start, and workflow-server work still dominate end-to-end startup

A final cold compile-miss trace has a 22.52 ms workflow.run; the surrounding fresh-replay path shows 6.88 ms bundle load (overlapped with 2.37 ms world init), followed by 3.00 ms context creation, 4.00 ms compile, 2.63 ms evaluate, 1.53 ms input hydrate, and 8.82 ms replay execute. Its 146 ms payload-preparation wall span begins before the replay and overlaps event loading and execution rather than representing 146 ms of blocking CPU work.

After the bundle and script caches are warm, the final sequential benchmark's first fresh VM replay trace is 6.88 ms: context creation 2.00 ms, compile hit 0.14 ms, evaluation 1.94 ms, input hydration 0.12 ms, and replay execution 2.41 ms. The comparable main trace is 186.10 ms (-96.3% for this representative warm-cache replay). Subsequent retained invocations in the final trace are generally 1–2.4 ms.

Validation

  • 295 focused tests across builders, runtime, tracing, payload cache, and stream recovery
  • complete builders suite: 226/226
  • affected package typechecks and builds
  • Next/Turbopack, Vite, and Nuxt production builds
  • exact Next HMR regression plus same-size/same-mtime workflow-ID rename regression
  • all eight stable/canary × Webpack/Turbopack × Node/QuickJS Next dev E2E lanes green
  • Linux and Windows unit lanes plus Node/QuickJS Windows Next E2E lanes green
  • standalone/BOPA import smoke test
  • final benchmark workflow green
  • final Codex and Claude autoreviews clean

@changeset-bot

changeset-bot Bot commented Aug 14, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: 6ff2c93

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
Name Type
@workflow/builders Patch
@workflow/core Patch
@workflow/next Patch
@workflow/world-testing Patch
@workflow/astro Patch
@workflow/cli Patch
@workflow/nest Patch
@workflow/nitro Patch
@workflow/rollup Patch
@workflow/sveltekit Patch
@workflow/vite Patch
@workflow/vitest Patch
@workflow/web-shared Patch
@workflow/web Patch
workflow Patch
@workflow/nuxt Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
example-nextjs-workflow-turbopack Ready Ready Preview Aug 15, 2026 6:18am
example-nextjs-workflow-webpack Ready Ready Preview Aug 15, 2026 6:18am
example-workflow Ready Ready Preview Aug 15, 2026 6:18am
workbench-astro-workflow Ready Ready Preview Aug 15, 2026 6:18am
workbench-express-workflow Ready Ready Preview Aug 15, 2026 6:18am
workbench-fastify-workflow Ready Ready Preview Aug 15, 2026 6:18am
workbench-hono-workflow Ready Ready Preview Aug 15, 2026 6:18am
workbench-nestjs-workflow Ready Ready Preview Aug 15, 2026 6:18am
workbench-nitro-workflow Ready Ready Preview Aug 15, 2026 6:18am
workbench-nuxt-workflow Ready Ready Preview Aug 15, 2026 6:18am
workbench-python-workflow Ready Ready Preview Aug 15, 2026 6:18am
workbench-sveltekit-workflow Ready Ready Preview Aug 15, 2026 6:18am
workbench-tanstack-start-workflow Ready Ready Preview Aug 15, 2026 6:18am
workbench-vite-workflow Ready Ready Preview Aug 15, 2026 6:18am
workflow-docs Ready Ready Preview, v0 Aug 15, 2026 6:18am
workflow-swc-playground Ready Ready Preview Aug 15, 2026 6:18am
workflow-tarballs Ready Ready Preview Aug 15, 2026 6:18am
workflow-web Ready Ready Preview Aug 15, 2026 6:18am

@github-actions

github-actions Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • webhookWorkflow (nextjs-turbopack)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · addTenWorkflow (tanstack-start) · at 06:18:55Z · abandoned wrun_01M02153Y378GKBC9VYN5HC0GM
  • run-pickup-stall · addTenWorkflow (vite) · at 06:20:04Z · abandoned wrun_01M02177TA6FBCMD00RWP704ES
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (vite) · at 06:20:39Z · abandoned wrun_01M021897GFMGV00T0JHWZ23E6
  • run-pickup-stall · promiseAnyWorkflow (tanstack-start) · at 06:20:52Z · abandoned wrun_01M0218PDPF8BBFKD0W48QDRKY
  • run-pickup-stall · promiseAllWorkflow (vite) · at 06:20:58Z · abandoned wrun_01M0218W9E9FQ2P7S215WNW8P1
  • run-pickup-stall · promiseRaceWorkflow (vite) · at 06:21:33Z · abandoned wrun_01M0219YGQTSXB08EADKP8B90M
  • run-pickup-stall · retainedInterleavingWorkflow (vite) · at 06:22:45Z · abandoned wrun_01M021C4XM8E4NZ9T0F8G6PA9J
  • run-pickup-stall · hookWorkflow (vite) · at 06:23:36Z · abandoned wrun_01M021DP343TGC8S632B07YRD2

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3474 0 738 4212
✅ 💻 Local Development 3087 0 501 3588
✅ 📦 Local Production 3810 0 558 4368
✅ 🐘 Local Postgres 3810 0 558 4368
✅ 🪟 Windows 312 0 0 312
✅ 🌐 Cross-language Conformance 9 0 128 137
✅ vercel-multi-region 27 0 0 27
Total 14529 0 2483 17012
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 128 0 28
✅ astro-quickjs 128 0 28
✅ example-node 128 0 28
✅ example-quickjs 128 0 28
✅ express-node 128 0 28
✅ express-quickjs 128 0 28
✅ fastify-node 128 0 28
✅ fastify-quickjs 128 0 28
✅ hono-node 128 0 28
✅ hono-quickjs 128 0 28
✅ nest-node 128 0 28
✅ nest-quickjs 128 0 28
✅ nextjs-turbopack-node 153 0 3
✅ nextjs-turbopack-quickjs 153 0 3
✅ nextjs-webpack-node 153 0 3
✅ nextjs-webpack-quickjs 153 0 3
✅ nitro-node 128 0 28
✅ nitro-quickjs 128 0 28
✅ nuxt-node 128 0 28
✅ nuxt-quickjs 128 0 28
✅ python-node 8 0 148
✅ sveltekit-node 147 0 9
✅ sveltekit-quickjs 147 0 9
✅ tanstack-start-node 128 0 28
✅ tanstack-start-quickjs 128 0 28
✅ vite-node 128 0 28
✅ vite-quickjs 128 0 28

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 156 0 0
✅ nextjs-turbopack-quickjs 156 0 0

✅ 🌐 Cross-language Conformance

App Passed Failed Skipped
✅ python 9 0 128

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

📋 View full workflow run

@github-actions

github-actions Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 6ff2c93 · Sat, 15 Aug 2026 06:37:29 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 1147 (+405%) 🔻 1201 🔴 (+4.7%) 1208 🔴 (±0%) 1244 🔴 (-8.2%) 30
TTFS stream 1162 (+394%) 🔻 1206 🔴 (+0.9%) 1225 🔴 (-3.2%) 1234 🔴 (-8.7%) 30
TTFS hook + stream 1425 (+243%) 🔻 1528 🔴 (+7.2%) 1622 🔴 (+11%) 1715 🔴 (+11%) 30
Fan-out TTFS Promise.all(100 steps) 954 (-28%) 💚 1068 (-58%) 💚 1100 (-59%) 💚 2439 (-24%) 💚 10
Fan-out TTLS Promise.all(100 steps) 9758 (-10%) 10385 (-91%) 💚 12871 (-93%) 💚 18558 (-90%) 💚 10
STSO 1020 steps (inline) 121 (-21%) 💚 413 (-17%) 💚 442 (-21%) 💚 516 (-34%) 💚 1018
STSO 1020 steps (queue-hop) 2786 2786 2786 2786 1
WO 1020 steps 376196 (-9.3%) 376196 (-9.3%) 376196 (-9.3%) 376196 (-9.3%) 1
CRTT first chunk (pooled) 77 (-21%) 💚 105 (-32%) 💚 158 (-33%) 💚 295 (+2.1%) 28

Streams

Scenario wr c/s rd c/s wr KiB/s rd KiB/s CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 100 (±0%) 100 (-2%) 5 (±0%) 5 (-2%) 95.5 (-29%) 130 (-56%) 195 (-81%) 334 (-77%) 156 (-25%) 10
size sweep (100/s, 160B-12KB) 100 (±0%) 99.9 (±0%) 334 (±0%) 333 (±0%) 90.5 (-22%) 125 (-64%) 245 (-68%) 463 (-63%) 146 (-53%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 89.2 (±0%) 89.7 (±0%) 16.2 (±0%) 16.3 (±0%) 100 (±0%) 129 (-8%) 193 (±0%) 445 (-49%) 398 (+10%) 3
replay eve-gpt-5.6-sol-2000t (1x) 54.7 (±0%) 54.7 (±0%) 355 (±0%) 355 (±0%) 211 (+26%) 124 (-30%) 175 (-79%) 385 (-89%) 362 (-82%) 2
replay eve-gpt-5.6-sol-2000t (2x) 109 (±0%) 109 (±0%) 710 (±0%) 707 (±0%) 91 (-28%) 152 (-47%) 226 (-66%) 406 (-79%) 169 (-83%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 413585ms → this run 373311ms (Δ -40274ms, -10%)

  100-150 ms  ┃                         main   0  this   8    +8
  150-200 ms  ██┃████                   main  69  this  29   -40
  200-250 ms  ███┃███████               main 102  this  42   -60
  250-300 ms  █████░░░░░░░░░░░░┃        main  52  this 173  +121
  300-350 ms  ████████████████░░░░░░░┃  main 154  this 229   +75
  350-400 ms  ████████████████████░░┃   main 193  this 222   +29
  400-450 ms  ████████░░░░░░░░░░░░░░░┃  main  80  this 233  +153
  450-500 ms  ██████┃█████              main 118  this  64   -54
  500-550 ms  ┃█████████████            main 135  this  11  -124
  550-600 ms  ┃█████                    main  61  this   2   -59
  600-650 ms  ┃█                        main  19  this   2   -17
  650-700 ms  ┃                         main  13  this   1   -12
  700-750 ms  ┃                         main   9  this   0    -9
  750-800 ms  ┃                         main   4  this   0    -4
  800-850 ms  ┃                         main   3  this   0    -3
  850-900 ms  ┃                         main   2  this   0    -2
 950-1000 ms  ┃                         main   1  this   0    -1
1200-1250 ms  ┃                         main   1  this   0    -1
3150-3200 ms  ┃                         main   1  this   0    -1
3600-3650 ms  ┃                         main   1  this   0    -1
3650-3700 ms  ┃                         main   1  this   0    -1
6700-6750 ms  ┃                         main   0  this   1    +1
7950-8000 ms  ┃                         main   0  this   1    +1

1020 steps (queue-hop)

Cumulative STSO time: 2786ms over 1 samples

No main baseline with raw samples yet — showing this run's distribution on its own; the diff appears once a run on main has recorded them.

2500-3000 ms  ████████████████████████  steps 1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+             avg         p50         p90         p99     n
control  ······██▁····  109.3 (-55%)  102 (-33%)  195 (-81%)  334 (-77%)  3000
sweep    ·····▁█▇▁····  110.7 (-47%)   97 (-26%)  245 (-68%)  463 (-63%)  3000
gw 1x    ·····▁█▇▁▁···  112.5 (-13%)   99 (-14%)   193 (±0%)  445 (-49%)  5295
eve 1x   ·····▁█▇▂▁···  109.6 (-57%)   93 (-20%)  175 (-79%)  385 (-89%)  5186
eve 2x   ·····▁▆█▂····  124.1 (-54%)  111 (-36%)  226 (-66%)  406 (-79%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  ▇█▇▁▁▃▁▁▂▂  103–123ms
sweep    ▂█▆▁▁▃▁▁▁▂  97–154ms
gw 1x    ▅█▂▂▂▃▂▁▂▁  96–161ms
eve 1x   █▁▃▄▂▂▄▃▃▁  90–157ms
eve 2x   ▆▆▃▂▁▃▄█▃▁  106–153ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  ██▇▅▂▁▄  107–113ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  ▃█▂▄▅▄▁▁▃▄  30–39ms
sweep    ▃█▆▃▄▁▃▂▃▃  38–57ms
gw 1x    ▃█▁▁▁▂▂▁▂▁  30–47ms
eve 1x   █▁▃▃▂▂▄▃▂▂  23–33ms
eve 2x   ▂▂▁▄▅█▅▂▅▄  22–30ms
ℹ️ Metric definitions & methodology

Streams: writer/reader sustained rates (steady window, 10% trimmed each side), first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actions Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 3 fail of 41 total

log=mint-ordered · fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 20 1.0m ok 0
stale-read-equal-step-counts completed 14 1.0m ok 0
step-vs-step-fork completed 12 0ms ok 0
step-vs-step-fork-fenced completed 12 0ms ok 0
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision failed 9 1.0m MISMATCH 1
in-flight-before-decision-counted failed 9 1.0m MISMATCH 1
in-flight-after-decision failed 9 1.0m MISMATCH 1
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 20 1.0m ok 0
stale-read-equal-step-counts completed 14 1.0m ok 0
step-vs-step-fork completed 12 0ms ok 0
step-vs-step-fork-fenced completed 12 0ms ok 0
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision completed 17 1.0m ok 0
in-flight-before-decision-counted completed 17 1.0m ok 0
in-flight-after-decision completed 19 2.0m ok 0
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim-append-only.txt

@NathanColosimo
NathanColosimo force-pushed the codex/replay-lazy-source-bundles branch from 45cfc76 to e202e05 Compare August 14, 2026 07:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant