Playwright E2E Testing in 2026: 7 Steps From Setup to CI
Set up Playwright E2E testing from scratch, write stable tests, debug failures with traces, and run sharded suites in CI.

A test that clicks through your app the way a real customer would is the only test that catches a broken checkout before your customers do. Playwright E2E testing does exactly that, in real browsers, and most JavaScript teams have now made it their default.
The hard part is not the first green run. It is what comes after: tests that pass on a laptop and fail in CI, logins that time out, and a suite that slowly turns into flaky tests nobody trusts.
This guide walks the whole path in 7 steps. You will install Playwright, write a first test that survives a redesign, debug with traces, remove the usual sources of flakiness, and run a sharded suite in CI. Every command and default comes from the official Playwright documentation.
What is Playwright E2E testing and why teams are switching
Playwright is an open-source browser automation framework from Microsoft. It ships with its own test runner, so you get locators, assertions, parallel workers, a trace recorder and an HTML report from one install, without stitching together separate libraries.
Playwright E2E testing is the practice of driving a real browser through a complete user journey, such as sign up, search or checkout, and asserting on what the user sees at each step. Playwright Test runs those journeys in Chromium, Firefox and WebKit from a single TypeScript or JavaScript API.
The reason it matters is the kind of app you are shipping. Modern frontends render asynchronously, hydrate after load and refetch data on every click. Older selector-based tools raced the UI and lost, which is why Playwright's architecture is built around auto-waiting actions and assertions that retry until the page is ready.
Where end-to-end testing with Playwright fits in the pyramid
End-to-end tests are the slowest layer, so you do not write hundreds of them. You write them for the flows where a bug costs money, and you lean on unit and integration tests for the rest. The table shows how the three layers divide the work.
| Test type | What it checks | Speed | Confidence in a user flow | Typical tool |
|---|---|---|---|---|
| Unit | One function or component in isolation | Milliseconds | Low | Vitest, Jest |
| Integration | Modules or services working together | Seconds | Medium | Vitest, Supertest |
| End-to-end | A full journey in a real browser | Seconds to minutes | Highest | Playwright |
How fast adoption is growing
Adoption numbers tell the same story as the engineering argument. The chart below uses raw download counts from the npm registry API for the @playwright/test package, month by month. Between September 2025 and August 2026 the monthly total grew more than fourfold, from about 50 million to about 223 million.

Area chart of monthly npm downloads for the @playwright/test package from September 2025 to August 2026, rising from 50.1 million to 223.4 million
Those downloads include CI runs, so they overstate the number of humans, but the slope is what counts. A deeper cut of the same registry data, split by framework and region, lives in the Playwright market share analysis. With the why settled, the next question is how to get a project running.
How to set up Playwright E2E testing step by step
The official initializer does almost all of the setup. It creates the config file, an example spec and, if you say yes, a GitHub Actions workflow. This is the sequence the rest of the guide builds on.
The 5-step setup
- Check your runtime. The current docs support Node.js 22, 24 and 26, on Windows 11, macOS 14 and current Debian or Ubuntu releases.
- Scaffold the project with the initializer. Choose TypeScript, keep the default tests folder and accept the GitHub Actions workflow.
- Install the browsers. On a Linux CI runner add --with-deps so the required system libraries are installed too.
- Set a baseURL in the config so every test can call page.goto('/') instead of hard-coding a host.
- Run the example test, then open the HTML report to confirm the whole chain works.
npm init playwright@latest
npx playwright install --with-deps
npx playwright test
npx playwright show-report
The initializer's output is intentionally small: a playwright.config.ts, a package.json, and tests/example.spec.ts. If you prefer to lay out folders for page objects and fixtures on day one, the Playwright framework setup guide covers a structure that scales.
Tip: Pick TypeScript even if the app is plain JavaScript. Playwright's locators and fixtures are fully typed, so your editor catches a misspelled role or a missing await before the test ever runs.
A config that works on day 1 and in CI
The generated config is fine for a demo, but three settings deserve attention before you commit anything. The official defaults give each test 30 seconds and each assertion 5 seconds, which is generous locally and tight on a loaded CI runner.
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: './tests',
fullyParallel: true,
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 2 : 0,
workers: process.env.CI ? 1 : undefined,
reporter: process.env.CI ? 'blob' : 'html',
use: {
baseURL: process.env.BASE_URL ?? 'http://localhost:3000',
trace: 'on-first-retry',
},
projects: [
{ name: 'chromium', use: { ...devices['Desktop Chrome'] } },
{ name: 'firefox', use: { ...devices['Desktop Firefox'] } },
{ name: 'webkit', use: { ...devices['Desktop Safari'] } },
],
webServer: {
command: 'npm run start',
url: 'http://localhost:3000',
reuseExistingServer: !process.env.CI,
},
});
The webServer block is the piece most beginners skip. It starts your app before the run and waits for the URL to respond, which removes a whole class of "connection refused" failures.
If you would rather pick browsers, workers and reporters from a form, the config generator on TestDino's free tools page produces a file you can paste in.

Infographic showing the 5 setup steps for Playwright E2E testing: scaffold the project, install browsers, set the baseURL, run the example test, open the HTML report
Once the example passes in all three browsers, delete it. The next section replaces it with a test that reflects your own product.
Writing your first Playwright E2E test example
Start with one journey that matters, not the biggest one. Login followed by a dashboard check is a good first pick because it exercises navigation, a form, and a state change, and almost every other test will build on it.
import { test, expect } from '@playwright/test';
test('user can sign in and reach the dashboard', async ({ page }) => {
await page.goto('/login');
await page.getByLabel('Email').fill(process.env.E2E_EMAIL!);
await page.getByLabel('Password').fill(process.env.E2E_PASSWORD!);
await page.getByRole('button', { name: 'Sign in' }).click();
await expect(page).toHaveURL(/\/dashboard/);
await expect(page.getByRole('heading', { name: 'Dashboard' })).toBeVisible();
});
Every test receives a fresh page inside its own browser context, so cookies and local storage never leak between tests. That isolation is why Playwright can run tests in parallel by default without you adding cleanup code.
Locators that survive a redesign
The test above never touches a CSS class or an XPath. It uses getByLabel and getByRole, which is the order of preference the official best practices recommend. A locator tied to what the user sees keeps working when a designer renames a class or moves a div.
When two elements match, chain and filter instead of reaching for an index. The Playwright locators guide goes through every built-in locator and when to fall back to a test ID.
await page
.getByRole('listitem')
.filter({ hasText: 'Product 2' })
.getByRole('button', { name: 'Add to cart' })
.click();
Assertions that wait for you
Notice that the login test ends with await expect(...).toBeVisible() rather than reading isVisible() into a boolean. The first form is a web-first assertion. It polls until the condition holds or the 5-second expect timeout runs out, so you never have to guess how long a redirect takes.
That single habit removes most hand-written waits. The full matcher list, from toHaveText to toHaveCount, along with soft assertions for multi-step forms, is covered in Playwright assertions.
Note: Run a brand-new test 10 times in a row before you copy its pattern across the suite. A test that has passed once is not yet a stable test, and it is far cheaper to find a race now than after 40 clones exist.
Generating the first draft
You do not have to type every locator by hand. The codegen recorder opens a browser, watches you click, and writes the test with role-based locators. Tools that pair the recorder with a language model are compared in Playwright AI codegen.
If your editor already runs Claude Code or a similar agent, the playwright-skill gives it the same locator and assertion rules this section describes.
Once a few tests exist, the question shifts from writing them to understanding why one failed. That is where Playwright's tooling separates itself from older frameworks.
Running and debugging Playwright E2E tests
Every run starts with the same command, and flags narrow it down. You will use a handful of these daily, and each one maps to a different moment in the workflow.
npx playwright test # whole suite, all projects
npx playwright test tests/login.spec.ts # one file
npx playwright test --project=webkit # one browser
npx playwright test --ui # UI Mode with watch and time travel
npx playwright test --headed # visible browser window
npx playwright test --trace on # force a trace for every test
UI Mode for authoring
UI Mode is the fastest loop while you write. It lists every test, reruns the ones you pick on save, and shows a DOM snapshot for every action so you can hover a step and see the page as it was. The Playwright UI Mode walkthrough shows the locator picker and watch mode in detail.
Trace Viewer for CI failures
Traces are how you debug a failure you cannot reproduce locally. With trace: 'on-first-retry' in the config, Playwright records a trace only when a test fails and is retried, so passing runs cost nothing extra.
A trace holds a screenshot per action, the DOM, network requests, console output and the test source. The Playwright Trace Viewer guide explains how to read each panel.
Tip: In Playwright E2E testing, most "cannot reproduce" bugs are timing bugs. Open the trace, click the failed action, and compare the before and after snapshots. The element is usually there, just not yet enabled or still animating.
The table below is the decision rule for which tool to reach for. It saves the common mistake of rerunning a CI failure locally ten times before looking at the trace that was already uploaded.
| Situation | Reach for | Why |
|---|---|---|
| Writing or fixing a locator | UI Mode | Locator picker plus instant rerun on save |
| Test fails only in CI | Trace Viewer | Full DOM, network and console for the failed attempt |
| Need a quick visual check | Headed mode | Watch the browser without extra tooling |
| Reviewing a whole run | HTML report | Filter by status, open attachments and traces |
The HTML report is the natural home for all of this. It opens after any run with npx playwright show-report, groups results by file and project, and links every trace. Options for tuning what it captures are in Playwright HTML reporter, and a broader Playwright debugging guide covers breakpoints and the VS Code extension.
Debugging tells you why one test failed. The next section is about writing tests so that they fail less often in the first place.
Best practices that keep Playwright E2E tests stable
Playwright is fast enough to expose every race condition your app has. That is a feature, but only if the tests themselves are not adding noise. Four patterns cause the majority of intermittent failures, and each has a built-in fix.

Infographic of 4 causes of flaky Playwright E2E tests, hard waits, brittle selectors, repeated logins and third-party calls, each paired with the Playwright feature that fixes it
Sign in once, reuse the session
Logging in at the top of every test is slow and, worse, races with cookie propagation. The official pattern is a setup project that signs in once, saves the browser storage to a file, and every other project loads that file through storageState.
import { test as setup, expect } from '@playwright/test';
import path from 'path';
const authFile = path.join(__dirname, '../playwright/.auth/user.json');
setup('authenticate', async ({ page }) => {
await page.goto('/login');
await page.getByLabel('Email').fill(process.env.E2E_EMAIL!);
await page.getByLabel('Password').fill(process.env.E2E_PASSWORD!);
await page.getByRole('button', { name: 'Sign in' }).click();
await expect(page.getByRole('heading', { name: 'Dashboard' })).toBeVisible();
await page.context().storageState({ path: authFile });
});
projects: [
{ name: 'setup', testMatch: /.*\.setup\.ts/ },
{
name: 'chromium',
use: { ...devices['Desktop Chrome'], storageState: 'playwright/.auth/user.json' },
dependencies: ['setup'],
},
],
Multi-role scenarios, such as an admin and a customer in the same test, need one state file per role. The Playwright authentication guide covers that variant along with token-based logins.
Mock what you do not own
A payment sandbox or an analytics endpoint you do not control adds latency and random errors to every run. The best-practices docs are direct about it: do not test third-party services, mock them. One page.route call replaces the response with fixed data.
await page.route('**/api/shipping-quote', (route) =>
route.fulfill({ status: 200, body: JSON.stringify({ price: 499 }) }),
);
Everything from stubbing a single field to replaying recorded HAR files is in Playwright network mocking. Keep the real calls in a small nightly smoke suite so you still notice when a provider changes its contract.
Isolate data and share setup through fixtures
Tests that share an account or leave rows behind pass alone and fail in parallel. Give each worker its own user or seed data, and put the setup in a fixture rather than a beforeEach copied across files. Playwright fixtures shows how a custom fixture can create a user, hand it to the test, and delete it afterwards.
Note: Raising a timeout should be the last change you make, never the first. If a test only passes with a 90-second budget, the wait condition is wrong. The Playwright timeout guide lists which of the 6 timeouts to adjust when a longer wait really is justified.
Treat retries as a signal, not a solution
Retries belong in CI, and 2 is a sensible ceiling. What matters is what you do with a retried pass. Playwright marks it as flaky in the report, and a count that creeps up each week is telling you something.
Dedicated flaky test detection tools track that number across runs, so a single retry does not disappear into a green build.
With stable tests in hand, the remaining work is running them where it counts: on every pull request.
How to run Playwright E2E tests in CI/CD
The initializer's optional GitHub Actions workflow is close to what the official CI guide recommends, and it is a fine starting point for any provider. The steps are always the same: install dependencies, install browsers with their system packages, run the tests, upload the report.
name: Playwright Tests
on:
push:
branches: [main]
pull_request:
branches: [main]
jobs:
test:
timeout-minutes: 60
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v6
with:
node-version: lts/*
- run: npm ci
- run: npx playwright install --with-deps
- run: npx playwright test
- uses: actions/upload-artifact@v4
if: ${{ !cancelled() }}
with:
name: playwright-report
path: playwright-report/
retention-days: 30
The if: ${{ !cancelled() }} guard is the detail people miss. Without it, a failing test step skips the upload and you lose the report you most need. Provider-specific versions of this file, including caching the browser download, are in Playwright in GitHub Actions, with GitLab, Azure and others covered under Playwright CI/CD integrations.
Shard when one runner is not enough
A single runner works until the suite passes a few hundred tests. Sharding splits the run across machines with --shard=1/3, --shard=2/3 and so on. With fullyParallel: true, the docs note that individual tests are distributed across shards, so balance stays even without hand-sizing files.
Each shard writes a blob report, and a final job merges them into one HTML report. The Playwright sharding guide has the complete matrix workflow, and the sharding calculator tells you how many shards a given suite needs to hit a target CI time.

Infographic showing a pull request fanning out into 3 sharded Playwright runners on ubuntu-latest, each writing a blob report that a final job merges into one HTML report
Make the results readable across runs
The HTML report answers "what failed in this run". It does not answer "has this test failed 4 of the last 10 runs" or "which shard is always the slow one". Once several people are reading CI results, that history matters more than any single report, which is the gap Playwright test reporting platforms are built to fill.
A pipeline that runs green on every pull request is the finish line for this tutorial. The last question most teams ask before committing is how Playwright compares to the tools they already know.
Playwright vs Cypress and Selenium for E2E testing
All three can drive a browser through a user journey. The differences are in browser coverage, how parallelism works, and what you get without plugins. The table sticks to facts from each project's documentation and to registry data.
| Capability | Playwright | Cypress | Selenium WebDriver |
|---|---|---|---|
| Browser engines | Chromium, Firefox, WebKit | Chromium-based, Firefox, WebKit (experimental) | Any browser with a driver |
| Test runner included | Yes, Playwright Test | No, bring your own | |
| Parallel workers built in | Yes, default | Requires Cypress Cloud or third-party | Depends on the runner |
| Trace with DOM snapshots | Yes, Trace Viewer | Time-travel in the runner | |
| Weekly npm downloads (15 to 21 Sep 2026) | 44.4M (@playwright/test) | 4.6M (cypress) | 2.0M (selenium-webdriver) |
The download row comes straight from the npm registry API for the week of 15 to 21 September 2026. It measures JavaScript installs only, so Selenium's Java and Python ecosystems are not represented, which is worth keeping in mind before reading it as market share.
When Cypress still makes sense
Teams with a large, healthy Cypress suite and no Safari requirement have little reason to rewrite. The trade-offs, including component testing and the plugin ecosystem, are worked through in Playwright vs Cypress. If the pain points are cross-browser gaps or paid parallelism, the Cypress to Playwright migration guide shows how to move a suite incrementally.
When Selenium still makes sense
Selenium remains the answer when the team writes tests in Java or C#, or when a legacy grid is already in place. Playwright does ship official Java, Python and .NET bindings, so language alone is rarely the blocker now. Playwright vs Selenium compares the two on architecture, waits and CI cost.
Whichever tool you choose, the practices in this guide transfer. Role-based locators, isolated state and traces on failure are good engineering, not Playwright features.
Conclusion
Playwright E2E testing is worth adopting because the tool and the workflow arrive together. The initializer, the config, UI Mode, Trace Viewer and sharding are all official, all documented, and all in this guide in the order you will actually need them.
Start with a single journey and a config that already knows about CI. Reuse login state, mock the services you do not own, and treat every retried pass as a bug report. When the suite outgrows one runner, shard it and merge the reports.
The condensed checklist lives in Playwright best practices, and it is a good file to pin next to your config.
The remaining work is keeping the suite honest as it grows. That means watching flake rates and run times week over week, not just the status of the latest build. TestDino was built for that part, and the flaky cost calculator shows what the current flake rate is costing before you decide how much time to invest.
FAQs

Dhruv Rai
Product & Growth Engineer

