Skip to content

[core] Fix flaky timing-sensitive tests: events-consumer deferred-check budgets and TTL-expiration e2e timeout - #3531

Merged
alangenfeld merged 3 commits into
mainfrom
alangenfeld/e2e-test-budgets
Aug 14, 2026
Merged

[core] Fix flaky timing-sensitive tests: events-consumer deferred-check budgets and TTL-expiration e2e timeout#3531
alangenfeld merged 3 commits into
mainfrom
alangenfeld/e2e-test-budgets

Conversation

@alangenfeld

Copy link
Copy Markdown
Collaborator

Summary & Motivation

Two timing-sensitive tests flaked on loaded runners, and both polls return as soon as their assertions hold, so the larger budgets cost healthy runs nothing.

  • events-consumer's duplicate event classes suite: the deferred check's timer chain can be starved for whole seconds on Windows, so the poll bound goes to 15s (and the two follow-up assertions use it instead of vi.waitFor's 1s default) under a 30s suite budget.
  • the distributedAbortController TTL-expiration e2e test gets the 60s its siblings already have, since cold starts on a fresh prod deployment push run start plus stream delivery past 30s.

Test Plan

events-consumer suite passes locally; e2e test collection verified.

The 'duplicate event classes' tests reach their outcome through the
deferred check's multi-stage timer chain (promise queue -> setTimeout(0)
-> idle poll -> delay timer), which a loaded CI runner with coarse
timers can starve for whole seconds. The afterDeferredCheck poll capped
that at 2s and two follow-up assertions used vi.waitFor's 1s default,
inside the 5s default test timeout - regularly starved through on
Windows runners ('does not track hook deliveries' and 'leaves a
duplicate run_cancelled' flaked ~1.5x/day over the last 10 days,
failing with strandedEvent/parkedSummary still undefined).

The polls return as soon as their assertions hold, so raising the poll
timeout to 15s and the suite budget to 30s costs healthy runs nothing
while bounding only genuinely stalled runners.

Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
The other distributedAbortController tests run with a 60s budget; the
TTL-expiration one got 30s. Its 3s TTL is trivial, but on a fresh prod
deployment cold starts plus queue backlog routinely push run start +
first stream delivery past 30s - it timed out on two apps
simultaneously in a single Tests run over the last week.

Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
@changeset-bot

changeset-bot Bot commented Aug 13, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: e6b258e

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
Name Type
@workflow/core Patch
@workflow/builders Patch
@workflow/cli Patch
@workflow/next Patch
@workflow/nitro Patch
@workflow/vitest Patch
@workflow/web-shared Patch
@workflow/web Patch
workflow Patch
@workflow/world-testing Patch
@workflow/astro Patch
@workflow/nest Patch
@workflow/rollup Patch
@workflow/sveltekit Patch
@workflow/vite Patch
@workflow/nuxt Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
example-nextjs-workflow-turbopack Ready Ready Preview Aug 13, 2026 10:30pm
example-nextjs-workflow-webpack Ready Ready Preview Aug 13, 2026 10:30pm
example-workflow Ready Ready Preview Aug 13, 2026 10:30pm
workbench-astro-workflow Ready Ready Preview Aug 13, 2026 10:30pm
workbench-express-workflow Ready Ready Preview Aug 13, 2026 10:30pm
workbench-fastify-workflow Ready Ready Preview Aug 13, 2026 10:30pm
workbench-hono-workflow Ready Ready Preview Aug 13, 2026 10:30pm
workbench-nestjs-workflow Ready Ready Preview Aug 13, 2026 10:30pm
workbench-nitro-workflow Ready Ready Preview Aug 13, 2026 10:30pm
workbench-nuxt-workflow Ready Ready Preview Aug 13, 2026 10:30pm
workbench-python-workflow Error Error Aug 13, 2026 10:30pm
workbench-sveltekit-workflow Ready Ready Preview Aug 13, 2026 10:30pm
workbench-tanstack-start-workflow Ready Ready Preview Aug 13, 2026 10:30pm
workbench-vite-workflow Ready Ready Preview Aug 13, 2026 10:30pm
workflow-docs Ready Ready Preview, v0 Aug 13, 2026 10:30pm
workflow-swc-playground Ready Ready Preview Aug 13, 2026 10:30pm
workflow-tarballs Ready Ready Preview Aug 13, 2026 10:30pm
workflow-web Ready Ready Preview Aug 13, 2026 10:30pm

@github-actions

github-actions Bot commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3466 0 590 4056
✅ 💻 Local Development 3810 0 558 4368
✅ 📦 Local Production 3810 0 558 4368
✅ 🐘 Local Postgres 3810 0 558 4368
✅ 🪟 Windows 312 0 0 312
✅ vercel-multi-region 27 0 0 27
Total 15235 0 2264 17499
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 128 0 28
✅ astro-quickjs 128 0 28
✅ example-node 128 0 28
✅ example-quickjs 128 0 28
✅ express-node 128 0 28
✅ express-quickjs 128 0 28
✅ fastify-node 128 0 28
✅ fastify-quickjs 128 0 28
✅ hono-node 128 0 28
✅ hono-quickjs 128 0 28
✅ nest-node 128 0 28
✅ nest-quickjs 128 0 28
✅ nextjs-turbopack-node 153 0 3
✅ nextjs-turbopack-quickjs 153 0 3
✅ nextjs-webpack-node 153 0 3
✅ nextjs-webpack-quickjs 153 0 3
✅ nitro-node 128 0 28
✅ nitro-quickjs 128 0 28
✅ nuxt-node 128 0 28
✅ nuxt-quickjs 128 0 28
✅ sveltekit-node 147 0 9
✅ sveltekit-quickjs 147 0 9
✅ tanstack-start-node 128 0 28
✅ tanstack-start-quickjs 128 0 28
✅ vite-node 128 0 28
✅ vite-quickjs 128 0 28

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 130 0 26
✅ astro-stable-quickjs 130 0 26
✅ express-stable-node 130 0 26
✅ express-stable-quickjs 130 0 26
✅ fastify-stable-node 130 0 26
✅ fastify-stable-quickjs 130 0 26
✅ hono-stable-node 130 0 26
✅ hono-stable-quickjs 130 0 26
✅ nest-stable-node 130 0 26
✅ nest-stable-quickjs 130 0 26
✅ nextjs-turbopack-canary-node 137 0 19
✅ nextjs-turbopack-canary-quickjs 137 0 19
✅ nextjs-turbopack-stable-node 156 0 0
✅ nextjs-turbopack-stable-quickjs 156 0 0
✅ nextjs-webpack-canary-node 137 0 19
✅ nextjs-webpack-canary-quickjs 137 0 19
✅ nextjs-webpack-stable-node 156 0 0
✅ nextjs-webpack-stable-quickjs 156 0 0
✅ nitro-stable-node 130 0 26
✅ nitro-stable-quickjs 130 0 26
✅ nuxt-stable-node 130 0 26
✅ nuxt-stable-quickjs 130 0 26
✅ sveltekit-stable-node 149 0 7
✅ sveltekit-stable-quickjs 149 0 7
✅ tanstack-start-node 130 0 26
✅ tanstack-start-quickjs 130 0 26
✅ vite-stable-node 130 0 26
✅ vite-stable-quickjs 130 0 26

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 156 0 0
✅ nextjs-turbopack-quickjs 156 0 0

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

📋 View full workflow run

@github-actions

github-actions Bot commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit e6b258e · Thu, 13 Aug 2026 22:50:03 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 1322 (+618%) 🔻 1423 🔴 (+29%) 🔻 1427 🔴 (+17%) 🔻 1549 🔴 (-2.6%) 30
TTFS stream 278 (+17%) 🔻 1425 🔴 (+28%) 🔻 1457 🔴 (+28%) 🔻 1473 🔴 (+22%) 🔻 30
TTFS hook + stream 1591 (+314%) 🔻 1743 🔴 (+26%) 🔻 1934 🔴 (+39%) 🔻 2115 🔴 (+38%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 8992 (±0%) 10437 (+0.9%) 10459 (-13%) 15217 (-95%) 💚 10
Fan-out TTLS Promise.all(100 steps) 17477 (-1.5%) 19107 (-6.8%) 19137 (-72%) 💚 24434 (-92%) 💚 10
STSO 1020 steps (inline) 125 (-17%) 💚 185 (-17%) 💚 215 (-17%) 💚 388 (-26%) 💚 1019
WO 1020 steps 182538 (-20%) 💚 182538 (-20%) 💚 182538 (-20%) 💚 182538 (-20%) 💚 1
SL stream latency 101 (+7.4%) 125 🔴 (-17%) 💚 152 🔴 (-5.6%) 334 🔴 (-5.6%) 30
SO stream overhead (text) 129 (-18%) 💚 179 (-40%) 💚 210 (-52%) 💚 515 (-66%) 💚 30
SO stream overhead (structured) 112 (-5.1%) 197 (-51%) 💚 218 (-61%) 💚 299 (-88%) 💚 30
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 228513ms → this run 182329ms (Δ -46184ms, -20%)

  100-150 ms  ░░░┃                      main   0  this 130  +130
  150-200 ms  █████████████████░░░░░░┃  main 512  this 728  +216
  200-250 ms  ███┃█████████             main 385  this 113  -272
  250-300 ms  ┃█                        main  71  this  24   -47
  300-350 ms  ┃                         main  22  this  10   -12
  350-400 ms  ┃                         main  11  this   5    -6
  400-450 ms  ┃                         main   5  this   2    -3
  450-500 ms  ┃                         main   1  this   2    +1
  500-550 ms  ┃                         main   4  this   1    -3
  550-600 ms  ┃                         main   0  this   1    +1
  600-650 ms  ┃                         main   1  this   1    +0
  650-700 ms  ┃                         main   1  this   0    -1
  700-750 ms  ┃                         main   0  this   2    +2
  900-950 ms  ┃                         main   1  this   0    -1
1200-1250 ms  ┃                         main   1  this   0    -1
1600-1650 ms  ┃                         main   1  this   0    -1
2000-2050 ms  ┃                         main   1  this   0    -1
2400-2450 ms  ┃                         main   1  this   0    -1
6050-6100 ms  ┃                         main   1  this   0    -1
ℹ️ Metric definitions & methodology

The collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: is main, marks where this run lands, bridges the gap when this run has more samples in a bucket.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000

All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor (clientStart) right before start(), so the CI runner’s request and its path through api.vercel.com sit outside every measured window. TTFS = in-deployment start() → first step body (turbo uses the in-process fast path, non-turbo the dispatch path), and includes the VQS dispatch hop plus any /flow cold start. Fan-out TTFS/TTLS are the first and last step completions of a single Promise.all over trivial steps, from the same anchor, so the gap between the two rows is the spread the runtime adds across the fan-out. STSO/WO are measured between step bodies on the deployment. SL is measured inside the workflow (parallel reader/writer steps), so it no longer includes the api.vercel.com read path.

Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the /flow invocation for a large fraction of runs, inflating P75+; the Best column shows the fastest (warm-start) sample for comparison.

@github-actions

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 Mint-ordered log — 6 fail of 41 total

log=mint-ordered · fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 17 1.0m MISMATCH 1
stale-read-equal-step-counts completed 14 1.0m MISMATCH 1
step-vs-step-fork completed 12 0ms MISMATCH 1
step-vs-step-fork-fenced completed 12 0ms MISMATCH 1
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision completed 17 1.0m MISMATCH 1
in-flight-before-decision-counted completed 20 1.0m ok 0
in-flight-after-decision failed 14 2.0m MISMATCH 1
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim-mint.txt

🟢 Append-only log — 0 fail of 41 total

log=append-only · fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 20 1.0m ok 0
stale-read-equal-step-counts completed 14 1.0m ok 0
step-vs-step-fork completed 12 0ms ok 0
step-vs-step-fork-fenced completed 12 0ms ok 0
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision completed 17 1.0m ok 0
in-flight-before-decision-counted completed 17 1.0m ok 0
in-flight-after-decision completed 19 2.0m ok 0
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim-append-only.txt

@alangenfeld
alangenfeld marked this pull request as ready for review August 14, 2026 14:47
@alangenfeld
alangenfeld requested a review from a team as a code owner August 14, 2026 14:47
@alangenfeld
alangenfeld merged commit 60dd206 into main Aug 14, 2026
417 of 426 checks passed
@alangenfeld
alangenfeld deleted the alangenfeld/e2e-test-budgets branch August 14, 2026 18:24
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 60dd206 (AI decision).

This is a test-only flaky-test fix, which would normally qualify, but neither test exists on stable: git show origin/stable:packages/core/src/events-consumer.test.ts has no duplicate event classes suite or afterDeferredCheck helper, and origin/stable:packages/core/e2e/e2e.test.ts contains no distributedAbortController tests at all. There is no flakiness on the maintenance line for this commit to fix, so backporting it would only add coverage for main-only behavior.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

60dd2065f368f10ba5c0b1ae98240749c1d29dc3

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants