[ci] Retry e2e tests once in CI, keeping retried tests visible - #3530
Conversation
Over the last 10 days ~93 Tests runs were manually re-run until green, some taking 6 attempts: the e2e suites drive real deployments, and a single test losing a timing race fails a whole 20+ minute matrix job. A CI-only vitest retry (retry: 1) absorbs those single-test races. beforeEach/afterEach hooks run per attempt, so suites with file-restore hooks (dev.test.ts) retry cleanly. Retries must not hide real races, so a retried-then-passed test stays visible everywhere a failure would have been: the github-reporter emits ::warning annotations and an e2e-flaky-*.json sidecar, every e2e job uploads it, and aggregate-e2e-results.js renders a 'Flaky E2E Tests (passed on retry)' section in both the per-job step summary and the PR comment, with per-app occurrence counts. Harnesses whose failures are themselves the signal pin retry: 0: event-log-race-repro (a pass runs the full configured budget) and benchmarks (a regression should not be papered over by a luckier second sample). Local runs keep retry at 0 so races reproduce while debugging. Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
Signed-off-by: Alex Langenfeld <alex.langenfeld@vercel.com>
🦋 Changeset detectedLatest commit: b40e1aa The changes in this PR will be included in the next version bump. This PR includes changesets to release 16 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
🧪 E2E Test Results✅ All tests passed
|
| Passed | Failed | Skipped | Total | |
|---|---|---|---|---|
| ✅ ▲ Vercel Production | 3466 | 0 | 590 | 4056 |
| ✅ 💻 Local Development | 3810 | 0 | 558 | 4368 |
| ✅ 📦 Local Production | 3810 | 0 | 558 | 4368 |
| ✅ 🐘 Local Postgres | 3810 | 0 | 558 | 4368 |
| ✅ 🪟 Windows | 312 | 0 | 0 | 312 |
| ✅ vercel-multi-region | 27 | 0 | 0 | 27 |
| Total | 15235 | 0 | 2264 | 17499 |
Details by Category
✅ ▲ Vercel Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-node | 128 | 0 | 28 |
| ✅ astro-quickjs | 128 | 0 | 28 |
| ✅ example-node | 128 | 0 | 28 |
| ✅ example-quickjs | 128 | 0 | 28 |
| ✅ express-node | 128 | 0 | 28 |
| ✅ express-quickjs | 128 | 0 | 28 |
| ✅ fastify-node | 128 | 0 | 28 |
| ✅ fastify-quickjs | 128 | 0 | 28 |
| ✅ hono-node | 128 | 0 | 28 |
| ✅ hono-quickjs | 128 | 0 | 28 |
| ✅ nest-node | 128 | 0 | 28 |
| ✅ nest-quickjs | 128 | 0 | 28 |
| ✅ nextjs-turbopack-node | 153 | 0 | 3 |
| ✅ nextjs-turbopack-quickjs | 153 | 0 | 3 |
| ✅ nextjs-webpack-node | 153 | 0 | 3 |
| ✅ nextjs-webpack-quickjs | 153 | 0 | 3 |
| ✅ nitro-node | 128 | 0 | 28 |
| ✅ nitro-quickjs | 128 | 0 | 28 |
| ✅ nuxt-node | 128 | 0 | 28 |
| ✅ nuxt-quickjs | 128 | 0 | 28 |
| ✅ sveltekit-node | 147 | 0 | 9 |
| ✅ sveltekit-quickjs | 147 | 0 | 9 |
| ✅ tanstack-start-node | 128 | 0 | 28 |
| ✅ tanstack-start-quickjs | 128 | 0 | 28 |
| ✅ vite-node | 128 | 0 | 28 |
| ✅ vite-quickjs | 128 | 0 | 28 |
✅ 💻 Local Development
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 130 | 0 | 26 |
| ✅ astro-stable-quickjs | 130 | 0 | 26 |
| ✅ express-stable-node | 130 | 0 | 26 |
| ✅ express-stable-quickjs | 130 | 0 | 26 |
| ✅ fastify-stable-node | 130 | 0 | 26 |
| ✅ fastify-stable-quickjs | 130 | 0 | 26 |
| ✅ hono-stable-node | 130 | 0 | 26 |
| ✅ hono-stable-quickjs | 130 | 0 | 26 |
| ✅ nest-stable-node | 130 | 0 | 26 |
| ✅ nest-stable-quickjs | 130 | 0 | 26 |
| ✅ nextjs-turbopack-canary-node | 137 | 0 | 19 |
| ✅ nextjs-turbopack-canary-quickjs | 137 | 0 | 19 |
| ✅ nextjs-turbopack-stable-node | 156 | 0 | 0 |
| ✅ nextjs-turbopack-stable-quickjs | 156 | 0 | 0 |
| ✅ nextjs-webpack-canary-node | 137 | 0 | 19 |
| ✅ nextjs-webpack-canary-quickjs | 137 | 0 | 19 |
| ✅ nextjs-webpack-stable-node | 156 | 0 | 0 |
| ✅ nextjs-webpack-stable-quickjs | 156 | 0 | 0 |
| ✅ nitro-stable-node | 130 | 0 | 26 |
| ✅ nitro-stable-quickjs | 130 | 0 | 26 |
| ✅ nuxt-stable-node | 130 | 0 | 26 |
| ✅ nuxt-stable-quickjs | 130 | 0 | 26 |
| ✅ sveltekit-stable-node | 149 | 0 | 7 |
| ✅ sveltekit-stable-quickjs | 149 | 0 | 7 |
| ✅ tanstack-start-node | 130 | 0 | 26 |
| ✅ tanstack-start-quickjs | 130 | 0 | 26 |
| ✅ vite-stable-node | 130 | 0 | 26 |
| ✅ vite-stable-quickjs | 130 | 0 | 26 |
✅ 📦 Local Production
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 130 | 0 | 26 |
| ✅ astro-stable-quickjs | 130 | 0 | 26 |
| ✅ express-stable-node | 130 | 0 | 26 |
| ✅ express-stable-quickjs | 130 | 0 | 26 |
| ✅ fastify-stable-node | 130 | 0 | 26 |
| ✅ fastify-stable-quickjs | 130 | 0 | 26 |
| ✅ hono-stable-node | 130 | 0 | 26 |
| ✅ hono-stable-quickjs | 130 | 0 | 26 |
| ✅ nest-stable-node | 130 | 0 | 26 |
| ✅ nest-stable-quickjs | 130 | 0 | 26 |
| ✅ nextjs-turbopack-canary-node | 137 | 0 | 19 |
| ✅ nextjs-turbopack-canary-quickjs | 137 | 0 | 19 |
| ✅ nextjs-turbopack-stable-node | 156 | 0 | 0 |
| ✅ nextjs-turbopack-stable-quickjs | 156 | 0 | 0 |
| ✅ nextjs-webpack-canary-node | 137 | 0 | 19 |
| ✅ nextjs-webpack-canary-quickjs | 137 | 0 | 19 |
| ✅ nextjs-webpack-stable-node | 156 | 0 | 0 |
| ✅ nextjs-webpack-stable-quickjs | 156 | 0 | 0 |
| ✅ nitro-stable-node | 130 | 0 | 26 |
| ✅ nitro-stable-quickjs | 130 | 0 | 26 |
| ✅ nuxt-stable-node | 130 | 0 | 26 |
| ✅ nuxt-stable-quickjs | 130 | 0 | 26 |
| ✅ sveltekit-stable-node | 149 | 0 | 7 |
| ✅ sveltekit-stable-quickjs | 149 | 0 | 7 |
| ✅ tanstack-start-node | 130 | 0 | 26 |
| ✅ tanstack-start-quickjs | 130 | 0 | 26 |
| ✅ vite-stable-node | 130 | 0 | 26 |
| ✅ vite-stable-quickjs | 130 | 0 | 26 |
✅ 🐘 Local Postgres
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ astro-stable-node | 130 | 0 | 26 |
| ✅ astro-stable-quickjs | 130 | 0 | 26 |
| ✅ express-stable-node | 130 | 0 | 26 |
| ✅ express-stable-quickjs | 130 | 0 | 26 |
| ✅ fastify-stable-node | 130 | 0 | 26 |
| ✅ fastify-stable-quickjs | 130 | 0 | 26 |
| ✅ hono-stable-node | 130 | 0 | 26 |
| ✅ hono-stable-quickjs | 130 | 0 | 26 |
| ✅ nest-stable-node | 130 | 0 | 26 |
| ✅ nest-stable-quickjs | 130 | 0 | 26 |
| ✅ nextjs-turbopack-canary-node | 137 | 0 | 19 |
| ✅ nextjs-turbopack-canary-quickjs | 137 | 0 | 19 |
| ✅ nextjs-turbopack-stable-node | 156 | 0 | 0 |
| ✅ nextjs-turbopack-stable-quickjs | 156 | 0 | 0 |
| ✅ nextjs-webpack-canary-node | 137 | 0 | 19 |
| ✅ nextjs-webpack-canary-quickjs | 137 | 0 | 19 |
| ✅ nextjs-webpack-stable-node | 156 | 0 | 0 |
| ✅ nextjs-webpack-stable-quickjs | 156 | 0 | 0 |
| ✅ nitro-stable-node | 130 | 0 | 26 |
| ✅ nitro-stable-quickjs | 130 | 0 | 26 |
| ✅ nuxt-stable-node | 130 | 0 | 26 |
| ✅ nuxt-stable-quickjs | 130 | 0 | 26 |
| ✅ sveltekit-stable-node | 149 | 0 | 7 |
| ✅ sveltekit-stable-quickjs | 149 | 0 | 7 |
| ✅ tanstack-start-node | 130 | 0 | 26 |
| ✅ tanstack-start-quickjs | 130 | 0 | 26 |
| ✅ vite-stable-node | 130 | 0 | 26 |
| ✅ vite-stable-quickjs | 130 | 0 | 26 |
✅ 🪟 Windows
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack-node | 156 | 0 | 0 |
| ✅ nextjs-turbopack-quickjs | 156 | 0 | 0 |
✅ vercel-multi-region
| App | Passed | Failed | Skipped |
|---|---|---|---|
| ✅ nextjs-turbopack | 27 | 0 | 0 |
📊 Workflow Benchmarkscommit Backend:
📈 STSO distribution vs main (inline / queue-hop histograms)1020 steps (inline) Cumulative STSO time: main 228513ms → this run 168423ms (Δ -60090ms, -26%) ℹ️ Metric definitions & methodologyThe collapsed STSO distribution section above buckets every step gap of the sequential-steps run (not a sampled window), split by whether the step ending the gap ran inline — in the same warm process as the step before it, so the gap is pure framework overhead — or after a queue-hop — the first step of a fresh process, which pays queue dispatch, client reinit and event-log replay. Bars overlay the two runs: Best/P75/P90/P99 deltas compare against the most recent benchmark run on Metrics — TTFS: time to first step body (in-deployment start() → first step body, deployment clocks) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · SL: stream latency (in-deployment write → read propagation, readAt - writtenAt) · SO: stream overhead (end-to-end write+consume time beyond the modelled generation window) Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · stream latency: parallel reader/writer steps on a dedicated stream; SL is the in-deployment write->read propagation (readAt - writtenAt) · stream overhead (text): writer streams 300 variable-length text token deltas paced at 100/s for 3s (a haiku-size LLM's token throughput) while a parallel reader drains the whole stream; SO is the end-to-end write+consume time beyond the 3s generation window (overhead/backpressure) · stream overhead (structured): same workload as stream overhead (text), but each delta is an AI-SDK-style structured object ({ type: 'text-delta', id, text }) instead of a raw string, so the SO gap vs the text scenario is the added serialization cost 🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600 · SL 50/60/125 · SO 250/500/1000 All metrics are measured from deployment-side timestamps only. Runs are triggered by an in-deployment route that stamps the anchor ( Cold starts are kept in the numbers on purpose — they are part of real bursty-workload latency. The workbench deployment cold-starts the |
Sim WorldSimulated world deterministic testing for races. Traces 🟠 Mint-ordered log — 6 fail of 41 total
Full trace: 🟢 Append-only log — 0 fail of 41 total
Full trace: |
VaguelySerious
left a comment
There was a problem hiding this comment.
The flaky annotation addition is great. We should still take care to check the flakes, and hopefully we'll get some more time soon to fix the actual tests, but this is great in the meantime
|
Backport to This is usually an infrastructure problem (e.g. the configured AI model could not be found, an AI Gateway error, or an opencode crash) rather than a merge conflict. Check the job logs linked above for details. Once the underlying issue is fixed, re-run the Backport to stable workflow manually via |
Summary & Motivation
Over the last 10 days ~93 Tests runs needed a human to click re-run until green (some took 6 attempts), because one test losing a timing race fails a whole 20+ minute e2e matrix job.
retry: 1in CI only, so races still reproduce locally while debugging.::warningannotations, ane2e-flaky-*.jsonsidecar per job, and a flaky section in the step summary and PR comment with per-app occurrence counts.event-log-race-reproand the benchmarks pinretry: 0— their failures are the measurement.Test Plan
Reporter smoke-tested against a synthetic retried test and the aggregation script against synthetic artifact directories in both modes; confirmed the two
retry: 0suites still collect.