Logs of passed tests are hidden by default in CI.
There's an argument to be made to always show them in CI. For now we
just wire up GitHub's built-in "re-run with debug logging enabled" to
show logs of passed tests.
<img width="1358" height="936" alt="CleanShot 2026-09-10 at 19 00 18@2x"
src="https://github.com/user-attachments/assets/70fb9c56-c681-4394-958b-7fa9aea19970"
/>
Co-authored-by: Claude Code (kimi-k3[1m]) <noreply@anthropic.com>
Removes the `Test examples` workflow and its orphaned test suite. The
workflow fails every scheduled run: its pinned Playwright docker image
(v1.35.1) is incompatible with the Playwright version the repo now
requires (v1.61.0), so the browser cannot launch. It also no longer runs
what it was created for: since the `run-tests.js` glob root moved from
`test/` to the repository root, the `--type examples` string-prefix
filter `examples/` matches test files inside example apps instead of
`test/examples/examples.test.ts`, leaving that suite without any runner.
This deletes `.github/workflows/test_examples.yml`, the `test/examples/`
suite, and the `--type examples` filter in `run-tests.js`. The
`examples/` prefix stays in the default-run exclusion list so test files
inside example apps (which run with each example's own jest/vitest
setup) are still not picked up by bare `run-tests.js` invocations.
Co-authored-by: Claude Code <noreply@anthropic.com>
CI jobs in `build_and_test` printed a "Cache save failed." warning from
the "Save passed-tests cache" step even when they never run tests. The
step ran for every job with an `afterBuild` command, but
`.next-test-passed.txt` is only ever written by `run-tests.js`, so jobs
like lint, rust-check, or types-and-precompiled warned on every run
because the path never existed.
`run-tests.js` already owns the decision of whether result caching is
active (`profile.cachingEnabled`), so it now signals the workflow
itself: when it opens the passed-tests file (which happens before any
test runs, and only when result caching is enabled), it writes the path
to the `passed_tests_file` step output. The "Save passed-tests cache"
step in `build_reusable.yml` only runs when that output is present and
takes its `path` from it. Jobs that never run `run-tests.js`, and
flake-detection runs (where `NEXT_FLAKE_DETECTION` disables the result
cache), never emit the output and skip the save silently. A job that
fails midway through its tests has already emitted the output, so its
partial file still saves and a genuine warning is still possible.
The idea here is that we waste a bunch of CPU setting up each test suite
by spawning an entirely new Chrome process, but playwright can share the
same browser across many tests. Tests are still isolated at the browser
level, just not at the OS process level.
Hoping to see some performance improvement in e2e tests. Compare to:
https://github.com/vercel/next.js/pull/95617
Somewhat unscientific (single sample) results:
```
Comparing the head-commit build-and-test runs — PR #95589 run [28981702219 ](https://github.com/vercel/next.js/actions/runs/28981702219)vs PR #95617 run [28981686692](https://github.com/vercel/next.js/actions/runs/28981686692). Both ran 101 jobs and both succeeded.
┌─────────────────────────┬────────────────┬───────────────────────┬───────────────────┐
│ Metric │ #95617 control │ #95589 shared browser │ Delta │
├─────────────────────────┼────────────────┼───────────────────────┼───────────────────┤
│ Raw wall-time sum │ 616.2 min │ 571.2 min │ −45.0 min (−7.3%) │
├─────────────────────────┼────────────────┼───────────────────────┼───────────────────┤
│ Billable (ceil per job) │ 658 min │ 617 min │ −41 min (−6.2%) │
└─────────────────────────┴────────────────┴───────────────────────┴───────────────────┘
The shared-browser experiment is cheaper, and the savings land almost exactly where you'd expect — the browser-driven test suites — while non-browser jobs (rust, lint, unit, windows) are flat within noise:
┌─────────────────────────────────┬─────────┬────────┬───────┐
│ Job category │ control │ shared │ delta │
├─────────────────────────────────┼─────────┼────────┼───────┤
│ test prod │ 130.7 │ 110.5 │ −20.2 │
├─────────────────────────────────┼─────────┼────────┼───────┤
│ cache components dev │ 51.5 │ 42.9 │ −8.7 │
├─────────────────────────────────┼─────────┼────────┼───────┤
│ turbopack dev │ 75.1 │ 69.5 │ −5.5 │
├─────────────────────────────────┼─────────┼────────┼───────┤
│ turbopack production │ 86.3 │ 81.1 │ −5.2 │
├─────────────────────────────────┼─────────┼────────┼───────┤
│ test dev │ 110.4 │ 105.5 │ −4.9 │
├─────────────────────────────────┼─────────┼────────┼───────┤
│ cache components prod │ 48.7 │ 49.0 │ +0.3 │
├─────────────────────────────────┼─────────┼────────┼───────┤
│ firefox and safari │ 5.3 │ 7.3 │ +2.0 │
├─────────────────────────────────┼─────────┼────────┼───────┤
│ flake-detection jobs (combined) │ ~15 │ ~21 │ +~6 │
└─────────────────────────────────┴─────────┴────────┴───────┘
Takeaways:
- Net savings of ~41 billable minutes per run (~6%), concentrated in the Chromium-driven test prod/dev, turbopack, and cache components suites — consistent with reusing one browser process instead of spawning per-suite.
- The small regressions are in firefox and safari and the "new/changed tests for flakes" jobs (+~8 min combined). Worth a glance, though they may just be run-to-run variance.
Caveat: this is a single run per PR, so there's real variance run-to-run. That said, the fact that the deltas track the browser-heavy jobs specifically — and not the Rust/lint/unit jobs — is a good signal the effect is genuine rather than noise. If you want a firmer number, re-running each PR 2–3× and averaging would tighten it up.
```
We want to use this to potentially improve efficiency in e2e tests where
we're running many instances of Turbopack at once. By default,
turbo-tasks will try to use every available CPU core.
**This environment variable does not strictly limit Turbopack to using 4
cores, it just restricts the number of main tokio worker threads and
changes how we shard dashmaps.**
This option will not make sense for the vast majority of users, it only
makes sense when you are running *many* instances of Turbopack at once,
and the exact semantics are potentially confusing, as it is not a hard
limit.
## Test Plan
Build vercel-site with this set to 1 and see a little over 100% CPU
usage. Set it to 8 and see around 800% CPU usage.
```
TURBO_TASKS_AVAILABLE_PARALLELISM=1 pnpm next build --experimental-build-mode=compile
```
## Summary
Memoize passing tests across attempts of the same workflow run, so a
re-attempt after a timeout or cancellation can skip what already passed.
## What changed
- `run-tests.js` appends each passing test's filename to
`$NEXT_TEST_PASSED_FILE` as it finishes (sync `write` + `fsync`, fd kept
open across calls). On startup it reads the same file and skips any
entries that match.
- `build_reusable.yml` wraps the test step with `actions/cache/restore`
(before) and `actions/cache/save` (after, with `if: always()`). The save
step runs on success, failure, and cancel — which is the whole point: a
cancelled attempt still gets to preserve its progress for retry.
- Cache key is scoped to
`${input_step_key}-${run_id}-attempt${run_attempt}`, with `restore-keys`
falling back to any earlier attempt of the same run.
The previous implementation wrote a single Turbo-remote-cache blob
*after* `Promise.allSettled` resolved — so a cancellation (like [run
26127371192](https://github.com/vercel/next.js/actions/runs/26127371192/job/76845288867))
lost every passed test from that attempt.
This also removes one dependency on the turbo-remote cache (there are
still two more uses of `scripts/turbo-cache.mjs`)
<!-- NEXT_JS_LLM_PR -->
### What?
Converts every test under `test/integration/` to an isolated test
running through `nextTestSetup` (under `test/e2e/`, `test/production/`,
`test/development/`, or `test/unit/`), then deletes `test/integration/`
along with the legacy CI orchestration that was specific to it.
- `test/integration/` removed entirely (~327 test suites)
- New isolated suites added across the existing folders:
- `test/e2e/` — 175
- `test/production/` — 130
- `test/development/` — 43
- `test/unit/` — 1
- `.github/workflows/build_and_test.yml` and `run-tests.js` no longer
have any `integration` branches
- `nextTestSetup` gained a `baseUrl` option on `next.browser()` so a
small number of tests that drive their own proxy/static-export server
can keep using `next.browser(...)` instead of importing `next-webdriver`
directly
### Why?
`test/integration/` predated `nextTestSetup` and ran tests directly
against the source checkout via custom helpers (`launchApp`,
`nextBuild`, `nextStart`, `runNextCommand`, `webdriver`, `fetchViaHTTP`,
…). Each suite hand-rolled its own dev/start/build orchestration,
fixture mutation, and process management.
The isolated test model used by the rest of the repo gives each suite an
isolated working directory containing a packed `next.tgz` install, a
uniform `next.start()` / `next.build()` / `next.fetch()` /
`next.browser()` API, and the same lifecycle for dev, start, and deploy
modes — so a single set of assertions covers all three. Deploy-mode
skips and per-feature gates are expressed declaratively
(`skipDeployment`, `disableAutoSkewProtection`, `if (skipped) return`)
instead of branching on `process.env`.
Removing `test/integration/` lets us:
- Delete the bespoke orchestration code in the CI workflow and
`run-tests.js`
- Run every converted suite consistently in dev, start, and deploy modes
(where applicable)
- Reproduce every test locally with the same `pnpm
test-{dev,start}-{turbo,webpack}` commands; no separate `integration`
path
- Open the door to running `test/production` against deployments in the
future (the converted suites already declare `skipDeployment` so they
can be flipped on)
### How?
Mechanical conversion per suite, with targeted clean-ups:
1. **Per-suite conversion.** Each
`test/integration/<name>/test/index.test.{js,ts}` was rewritten into a
single `<name>.test.ts` under the right folder based on what the
original exercised:
- `launchApp` / dev-only assertions → `test/development/`
- `nextBuild` + `nextStart` / start-only assertions → `test/production/`
- Both → `test/e2e/`
- The one pure jsdom render check (`link-without-router`) → `test/unit/`
2. **API mapping.** Custom helpers were replaced by `nextTestSetup`
equivalents: `launchApp` → `next.start()`, `nextBuild` → `next.build()`,
`runNextCommand` → `next.runCommand`, `fetchViaHTTP` → `next.fetch`,
`webdriver(...)` → `next.browser(...)`. Fixture mutations switched from
raw `fs.writeFile`/`fs.rename` to `next.patchFile` (with the 3-arg
`runWithTempContent` callback when the change has a defined scope) and
`next.deleteFile`.
3. **Deploy-mode handling.** Suites that can't run in deploy mode (use
`patchFile` / `next.build()` / depend on local CLI output) declare
`skipDeployment: true` and early-return on the `skipped` boolean. Suites
where Vercel's edge mutates URLs (`&dpl=`, immutable assets) declare
`disableAutoSkewProtection: true`.
4. **`next.browser({ baseUrl })`.** A handful of tests
(`prerender-export`, `cdn-cache-busting`, `preload-viewport`, both
`react-virtualized` suites) need to drive a separate server (a
static-export server or an `http-proxy` instance) rather than the
Next.js process. Instead of importing `next-webdriver` directly, those
tests now pass `{ baseUrl: <port|url> }` to `next.browser()`. For the
proxy cases, the proxy was moved into `server.js` inside the fixture and
`http-proxy` declared via the `dependencies` option of `nextTestSetup`,
so the test runs with a fully isolated dependency graph.
5. **CI clean-up.** With `test/integration` gone, the `test
integration*` jobs and `integration-tests-manifest`-related logic in
`.github/workflows/build_and_test.yml` were removed, and `run-tests.js`
no longer has the `integration` test-folder branch.
6. **Validation.** The PR was iterated against multiple full CI runs;
the remaining failures on the latest run are pre-existing flakes
(segment-cache 60s `act` timeouts in turbopack-prod) or transient
infrastructure issues unrelated to the conversion.
## What
Cache each passing test file's result in turbo's remote cache during CI runs. On retry attempts (`GITHUB_RUN_ATTEMPT > 1`), skip tests that already passed.
## Why
When `retry_test.yml` calls `rerun-failed-jobs`, the entire test shard re-runs from scratch — 20+ minutes to retry one flaky test. With this change, only the failed tests actually re-run on retry.
## How
- Cache key: `sha256(commit + test file + env fingerprint)` where env fingerprint includes `NEXT_TEST_MODE`, `IS_WEBPACK_TEST`, `IS_TURBOPACK_TEST`, etc.
- Cache value: `{"passed": true, "time": N}` (~50 bytes per test)
- On first attempt: all tests run normally, passes are cached
- On retry: cached passes are skipped, failures re-run. Newly passing tests get cached for further retries
- All caching gated on `CI` + `TURBO_TOKEN` + `GITHUB_SHA` — zero impact on local runs
- Uses existing `scripts/turbo-cache.mjs` client, all cache operations are fire-and-forget with try/catch
### What?
Eliminates the expensive temporary repo directory (`tmpRepoDir`) copy
during isolated test setup by leveraging a Turborepo `pack` task with
caching.
### Why?
The previous test isolation flow copied the entire `packages/` directory
to a temp location, mutated every `package.json` to rewrite workspace
dependency references, and ran `pnpm pack` sequentially for each
package. This added ~10s+ of overhead per test suite. With Turborepo
caching, repeated runs are 500ms (and can still be improved).
### How?
- Adds a `pack-for-isolated-tests` task to `turbo.json` (depends on
`build`, outputs `packed.tgz`)
- Adds a `pack-for-isolated-tests` script (`pnpm pack --out
./packed.tgz`) to every workspace package
- Simplifies `linkPackages` in `repo-setup.js` to scan for pre-built
tarballs instead of copying/rewriting/packing
- Updates `create-next-install.js` to run `pnpm turbo run pack`
(benefits from caching) and use package manager `overrides`
(pnpm/npm/yarn) to resolve transitive workspace deps from local tarballs
- Removes `tmpRepoDir` handling from `base.ts` and `run-tests.js`
**Before:** `createNextInstall` ~16s (copy + sequential pack)
**After:** `createNextInstall` ~5s (first run) / ~4s (cached)
<!-- NEXT_JS_LLM_PR -->
## Summary
For context we discovered jobs were not reliably always running in
https://github.com/vercel/next.js/pull/90668 as addapters + turbopack
group ran more tests than just turbopack group.
- Each sharded test sub-job independently fetched test timings from KV
via a turbo task. Because turbo cache keys can vary across job groups
and KV can return slightly different data at different times, the
sharding algorithm produced different test-to-shard assignments across
groups — causing some tests to be missing entirely and others to run in
multiple shards.
- Adds a `--require-timings` flag to `run-tests.js` that fails loudly if
`test-timings.json` can't be loaded from disk, preventing silent
fallback to KV or round-robin.
- Adds a `testTimingsArtifact` input to `build_reusable.yml` that
downloads a pre-built timings artifact instead of fetching via turbo.
When unset (default), existing behavior is preserved.
- In both `test_e2e_deploy_release.yml` and `build_and_test.yml`,
timings are now fetched once in a setup job, uploaded as a GitHub
artifact, and downloaded by all sub-jobs — guaranteeing identical shard
assignments across all job groups.
## Test plan
- [ ] Verify `build_and_test.yml` jobs pick up the shared `test-timings`
artifact
- [ ] Verify `test_e2e_deploy_release.yml` deploy jobs pick up the
shared `test-timings` artifact
- [ ] Verify workflows that don't pass `testTimingsArtifact` (e.g.
`integration_tests_reusable.yml`, `pull_request_stats.yml`) still use
the turbo fallback path unchanged
The recent change to always run all tests without aborting on failure
(#88435) inadvertently broke manifest generation. Previously, test
output was emitted for all tests when `NEXT_TEST_CONTINUE_ON_ERROR` was
set, but that variable was removed. Now test output is only emitted for
failing tests, causing the manifest to lose all passing test entries.
This adds a new `NEXT_TEST_EMIT_ALL_OUTPUT` environment variable that
restores the full output behavior specifically for the manifest
generation workflows.
It's more useful and efficient to run all tests and report all failures
instead of aborting on the first failure. This way, developers get a
complete picture of what needs to be fixed in a single run, and don't
have to go through multiple iterations of fixing one failure at a time.
The file already had the `@ts-check` directive, so the IDE did show errors in that file. But since it was not included in the root `tsconfig.json`, those errors were not reported during CI runs.
One issue found this way was that the `related` flag was defined but never used. The referenced script was removed in #67644.
i got bit by this often :p
<!-- Thanks for opening a PR! Your contribution is much appreciated.
To make sure your PR is handled as smoothly as possible we request that you follow the checklist sections below.
Choose the right checklist for the change(s) that you're making:
## For Contributors
### Improving Documentation
- Run `pnpm prettier-fix` to fix formatting issues before opening the PR.
- Read the Docs Contribution Guide to ensure your contribution follows the docs guidelines: https://nextjs.org/docs/community/contribution-guide
### Fixing a bug
- Related issues linked using `fixes #number`
- Tests added. See: https://github.com/vercel/next.js/blob/canary/contributing/core/testing.md#writing-tests-for-nextjs
- Errors have a helpful link attached, see https://github.com/vercel/next.js/blob/canary/contributing.md
### Adding a feature
- Implements an existing feature request or RFC. Make sure the feature request has been accepted for implementation before opening a PR. (A discussion must be opened, see https://github.com/vercel/next.js/discussions/new?category=ideas)
- Related issues/discussions are linked using `fixes #number`
- e2e tests added (https://github.com/vercel/next.js/blob/canary/contributing/core/testing.md#writing-tests-for-nextjs)
- Documentation added
- Telemetry added. In case of a feature if it's used or not.
- Errors have a helpful link attached, see https://github.com/vercel/next.js/blob/canary/contributing.md
## For Maintainers
- Minimal description (aim for explaining to someone not on the team to understand the PR)
- When linking to a Slack thread, you might want to share details of the conclusion
- Link both the Linear (Fixes NEXT-xxx) and the GitHub issues
- Add review comments if necessary to explain to the reviewer the logic behind a change
### What?
### Why?
### How?
Closes NEXT-
Fixes #
-->
The peerDeps are already bumped to 19.2.0 at
https://github.com/vercel/next.js/pull/84463, but the test/CNA templates
weren't updated as the `syncPagesRouterReact` value is always false at
the moment, which blocked the update.
Therefore, remove the `syncPagesRouterReact` condition to match the
actual peerDeps bump.
---------
Co-authored-by: vercel[bot] <35613825+vercel[bot]@users.noreply.github.com>
Co-authored-by: Sebastian Sebbie Silbermann <sebastian.silbermann@vercel.com>
https://github.com/vercel/next.js/pull/83961 lets you run both `next dev` and `next build` at the same time, however it's still problematic to run two instances of `next dev` or two instances of `next build` at the same time with the same `distDir`.
This sort of thing can happen to anyone by accident, but we've seen reports of this sort of behavior with AI agents.
https://github.com/vercel/next.js/pull/84378 adds a lockfile around just the persistent cache database in Turbopack, which helps, but it's not the only reason this is problematic, so we also need a lock around all of Next.js itself.
I was not able to find a cross-platform lockfile implementation in node that I felt was of sufficient quality, so on POSIX platforms this uses the recently-stabilized Rust stdlib implementation, which appears to be derived from rustc's own lockfile implementation, which has seen widespread use/deployment.
On Windows, since we'd rather have an advisory lock than a mandatory one, this emulates that by instead opening a file with write permissions, and a sharing mode that prohibits shared write access. That means that other processes can safely read the (empty) lockfile without blowing up.
<img width="500" src="https://github.com/user-attachments/assets/1ee4a6c6-2280-4f64-8bf2-ec5369c26db1" />
<img width="500" src="https://github.com/user-attachments/assets/2b04b566-c0b4-42ce-af34-912f3f6afa72" />
Tested on Windows as well:
<img width="1224" height="690" alt="Screenshot 2025-10-07 at 5 44 33 PM" src="https://github.com/user-attachments/assets/e36085b6-9c32-40ac-aca6-4c83e2a53842" />
Next.js tests are grouped based on timing data to better distribute them
for parallelization.
Previously this relied on a gist file that we would write to in CI.
However, this has more recently started resulting in 401s.
The actual reason for the 401s is the Turbo task wasn't propagating the
required environment variable to the task. It was mostly working by
accident, because we would store timing data on disk for the runner,
which would have been satisfied by a job that ran before it.
However despite that, relying on a gist file to read/write timing data
didn't feel like the write abstraction. I decided to refactor it to use
KV instead.
This also makes the handling more resilient to missing timing data. We
shouldn't fail the build if we can't fetch timings, as there's already
handling to fallback to round robin.
### What?
Replace LRU cache implementation with an optimized doubly-linked list
algorithm and move from server-only to shared library location.
### Why?
The previous Map-based LRU cache had suboptimal eviction performance for
route matching operations under high load. Route matching is a critical
performance path in Next.js that benefits significantly from true O(1)
cache operations.
### How?
- Implement doubly-linked list with sentinel nodes for true O(1)
get/set/delete operations
- Add comprehensive test coverage including size-based eviction
scenarios
- Update import paths across 14 files throughout the codebase
- Add defensive null checking for edge cases discovered during static
analysis
- Replace deprecated `keys()` iteration with modern `Symbol.iterator`
pattern
It seems too easy to accidentally not pass through `TEST_CONCURRENCY` in `build_and_test.yml`, so IMO `run-tests.js` should use it by default, still allowing overriding with `-c` or `--concurrency`.
I tested by spot-checking CI jobs for log lines like this:
```
Running tests with concurrency: 4 in test mode undefined
```
or
```
Running tests with concurrency: 1 in test mode undefined
```
Our CI runs tests in parallel across multiple groups. We want test groups to all take similar amounts of time to decrease overall latency, and to reduce the chance of test group timeouts in CI.
There's logic here to assign intelligently to groups based on previous test timing information, but if there's no timing information for whatever reason (e.g. we don't currently use timings for rspack/turbopack integration tests), then we were assigning them to groups sequentially.
Because tests with similar paths tend to have similar performance characteristics, it's better to assign every nth test to each group, similar to round-robin assignment.
This PR removed the old overlay including the deprecated `buildActivity`
and `appIsrStatus` options.
Closes NDX-785
Closes NDX-855
Closes NDX-865
Closes NDX-866
Closes NDX-867
Closes NDX-868
- `newDevOverlay: true` by default (enables experimental React builds on
canary until owner stacks progress further)
- `run-tests` now sets the env var for tests that were relying on it for
forking behavior
- PPR runners now run with the flag disabled to help catch regressions
in the old overlay until we remove it
- Fixed a number of tests that had outdated snapshots or missed forking
behavior because they weren't running in CI
- Disabled a test that was failing in Turbopack + Experimental React
that is unrelated to the overlay (see:
https://github.com/vercel/next.js/pull/75989)
---------
Co-authored-by: devjiwonchoi <devjiwonchoi@gmail.com>
Let's see how this affects the failure rate and duration of our tests in
CI.
Test times for any given group might increase a bit, while hopefully at
least increasing the chance that the whole test group succeeds. This
would in theory reduce the number of times we have to re-run a whole
failed test group, which is quite time consuming.
A possible downside of this change is that flaky tests which have a
chance of getting fixed with a better implementation might be harder to
detect.