Playwright vs e2e by TesterArmy: 7 Differences That Matter in 2026
Playwright vs e2e by TesterArmy explained with real code, official docs, and migration data, so you can choose your framework.
A new testing tool named e2e lets you tell an AI helper what a user should do, in plain words, and it clicks through your app for you. That is why "playwright vs e2e" has suddenly become a real question for teams that already test with Playwright.
The trouble is that the name is confusing and the launch buzz is loud. It is hard to tell if e2e replaces Playwright, sits on top of it, or just adds an AI bill and a new kind of test maintenance.
This guide answers that with real code, official docs, and migration numbers we counted ourselves. You will see where each tool fits and how to try e2e without touching your Playwright end-to-end tests.
What does "Playwright vs e2e" mean in 2026?
e2e by TesterArmy is an open source TypeScript testing framework for web and mobile apps. A test can mix plain-language goals that an AI agent carries out with exact locators and assertions. On the web, it drives browsers through Playwright.
Until recently, "e2e" was only short for end-to-end testing. That is the practice of checking a full user journey, from the first click to the final result.
Playwright is itself an end-to-end testing framework, so comparing Playwright with "e2e" made little sense. It was like comparing a car with driving.
That changed on October 1, 2026, when TesterArmy published its official launch announcement for a framework that is literally named e2e. The comparison is now between 2 concrete tools.
The short answer
| Question | Playwright | e2e by TesterArmy |
|---|---|---|
| What it is | Browser automation library and test runner from Microsoft | Agentic test framework from TesterArmy |
| How you write a step | Locators and assertions in code | Plain-language goals, locators, or both |
| Web browsers | Chromium, Firefox, WebKit | Chromium, Firefox, WebKit, driven through Playwright |
| Mobile | Emulation of mobile browsers | iOS simulators and Android emulators |
| AI model required | Only for agent steps | |
| License | Apache-2.0 | Apache-2.0 |
| Latest npm version (October 6, 2026) | 1.63.0 | 0.17.0 |
The third row is the one to remember. According to the official e2e documentation, the web engine "drives browsers through Playwright".
So e2e is not a rival browser engine. It is a different way of writing and running tests that sits above the same browser automation you already use, and that is easier to see once you look at what e2e adds.
What is e2e by TesterArmy?
e2e is an agentic testing framework and one of the newest agentic testing tools. It is built by the team behind the TesterArmy testing platform, and its docs describe it as "a next-generation testing framework for web and mobile apps".
The central idea is that you choose how agentic each test is. A test can be fully scripted, fully driven by an agent, or a mix of both in the same file.
3 kinds of steps in 1 test
The core concepts page splits every step into 3 groups:
| Step type | API | Needs an AI model |
|---|---|---|
| Goals | agent.act | Yes, unless the step replays from the cache |
| Judgments | agent.assert, agent.waitFor, agent.extract | Yes, on every run |
| Locators | screen and expect |
A goal is a high-level instruction, such as "upgrade the workspace to the Pro plan". The agent reads the screen and picks the actions.
Locators work the way Playwright users expect. You point at a known control and assert an exact value, and no model is involved.
What ships in the box
The e2e GitHub repository is a set of small packages:
- e2e: the SDK, test runner, and CLI
- @e2e-dev/web: the browser engine for Chromium, Firefox, and WebKit
- @e2e-dev/mobile: the engine for iOS simulators and Android emulators
- @e2e-dev/github: a reporter that comments on pull requests
- @e2e-dev/kernel and @e2e-dev/eas: hosted browsers and hosted mobile simulators
Setup and requirements
The quickstart asks for Node.js 24.8 or newer, or 22.22.3 or newer on Node.js 22. Windows users need to run it inside WSL.
Playwright is less strict here. Its installation guide lists the latest Node.js 22.x, 24.x, or 26.x and supports Windows 11 natively.
npx e2e init
npx e2e run
The wizard asks for an engine (web or mobile) and a model provider, then writes e2e.config.ts and an example test. The first run makes no model calls, so you can try it without an API key.
We ran the non-interactive version, npx e2e init --yes, in an empty project. This is its real output:

Notice that the command only writes files. It adds the dependencies to package.json, and nothing is installed until you run npm install.
It also sets up your coding agent in the same step. The skill folder and the 2 MCP config files are there so that tools like Claude Code and Cursor can write and run e2e tests.
One finding from our own trial: e2e 0.17.0 ran natively on Windows 10 with Node.js 22.16, outside WSL. The docs do not promise that, so treat it as unsupported.
That covers what e2e is, and the clearest way to feel the difference is to put the same test side by side.
How the same test looks in Playwright and e2e
Both examples below come from the official Playwright migration guide in the e2e docs. They test the same checkout flow.
The Playwright version
import { test, expect } from '@playwright/test';
test('checkout', async ({ page }) => {
await page.goto('/cart');
await page.getByRole('button', { name: 'Checkout' }).click();
await page.getByLabel('Card number').fill('4242424242424242');
await page.getByLabel('Expiry').fill('12/30');
await page.getByLabel('CVC').fill('123');
await page.getByRole('button', { name: 'Pay' }).click();
await expect(page.getByRole('heading')).toHaveText('Order confirmed');
});
Every interaction is spelled out. You choose each control with Playwright locators and confirm the result with Playwright assertions.
The test is precise and fast. It also breaks when someone renames the Pay button or adds a step to the form.
The e2e version
import { test } from '@e2e-dev/web';
import { expect } from 'e2e';
test('checkout', async ({ app, agent, screen }) => {
await app.open('/cart');
await agent.act('pay with the test card 4242 4242 4242 4242, expiry 12/30, CVC 123');
await expect(screen.getByRole('heading')).toHaveText('Order confirmed');
});
What actually changed
- Fixtures: { page } becomes { app, agent, screen }.
- Interactions: 5 scripted lines become 1 goal in agent.act.
- Checks: the navigation and the final assertion stay explicit and deterministic.
The docs say the agent "handles the checkout interactions and adapts when the form changes". That is the trade: fewer lines to maintain, in exchange for an AI model in the loop.
A test can also stay fully scripted. The guide notes that every page.getBy* line has a matching screen.getBy* line.
Tip: Give each agent.act call 1 goal, then follow it with an expect check. The check is what fails when the agent does the wrong thing, and it also makes the step eligible for cache replay.
One checkout test shows the style gap, but the full picture has more sides to it.
Playwright vs e2e: 7 differences that matter
The differences below are taken from the official documentation of each project. None of them is a score.

1. Scripts vs goals
Playwright tests are scripts. You decide every click, and the same code runs the same way each time.
In e2e, a goal describes the outcome and the agent finds the path. The docs suggest locators "when you know the exact action or the exact value" and goals for flows that change often.
2. AI is optional in Playwright and built into e2e
Playwright has AI features, but they work at authoring time. The Playwright Test Agents are a planner, a generator, and a healer that produce or repair normal test code.
The test run itself stays deterministic. The same holds when a coding agent drives a browser through Playwright MCP to write tests for you.
In e2e, the agent acts while the test runs. That needs a model, and e2e lets you bring a subscription, an API key, or a local model.
The subscriptions page lists ChatGPT Plus or Pro, GitHub Copilot, OpenCode Console, and SuperGrok. It also states that Claude subscriptions are not supported, although Claude works through an API provider.
3. Web only vs web, iOS, and Android
Playwright covers Chromium, WebKit, and Firefox, "with native mobile emulation for Chrome (Android) and Mobile Safari". That emulation is useful for Playwright mobile testing of responsive sites, but it is not a native app.
The mobile engine in e2e runs real iOS simulators and Android emulators through a project called agent-device. Tests use the same agent, screen, and expect APIs as browser tests.
4. TypeScript only vs 5 languages
Playwright has official bindings for TypeScript, JavaScript, Python, Java, and .NET. Teams with a Java or Python backend can keep tests in the same language.
e2e is a TypeScript framework. If your suite is written in Python or Java, moving to e2e also means changing language.
5. Stricter locators and different defaults
This is the difference that surprises people during migration. In e2e, a string name matches the whole text and is case-sensitive.
Playwright matches a case-insensitive substring by default. So getByRole('button', { name: 'Save' }) finds a "Save changes" button in Playwright, and the same query finds nothing in e2e.
| Setting | Playwright | e2e |
|---|---|---|
| Test timeout | 30 s | 120 s |
| Action timeout | None | 30 s |
| Expect timeout | 5 s | 5 s |
| Retries | 0 | 1 in CI, 0 locally |
| Trace | Off | On locally, on first retry in CI |
| String name in getByRole, getByLabel, getByText | Case-insensitive substring | Exact, case-sensitive |
| isVisible() on several matches | Reads the first match | Throws LOCATOR_AMBIGUOUS |
To see this for real, we wrote the same 2 tests for a small demo shop in both tools. One test clicks a "Save changes" button with the short name Save.

Playwright 1.63.0 passed both tests. e2e 0.17.0 failed the second one with LOCATOR_NOT_FOUND, exactly as the migration guide warns.
The error message is useful, though. It prints the control that was on screen, a button named "Save changes", so the fix is clear: use the full name or pass { exact: false }.
The failing test also took 30.47 s, because e2e waited out its full 30 s action timeout before giving up. After the fix, both e2e tests passed in 1.94 s.
The longer test budget makes sense, because an agent step takes more time than a scripted click. If you tune Playwright timeouts today, expect to revisit them.
6. Debugging and reporting
Playwright ships a rich toolset: UI mode, an inspector, the Playwright HTML reporter, and the Playwright Trace Viewer.
e2e has 4 reporters: list, json, junit, and markdown. There is no HTML reporter, no --ui mode, and --debug prints timings and an agent step table, not an inspector.
Traces are still Playwright traces. You open them with npx playwright-core@1.63.0 show-trace, and an --ai-trace flag saves the model requests for agent steps.
7. Maturity and community
The Playwright repository was created on GitHub in November 2019, and the framework is now on version 1.63. Our Playwright market share report tracks how widely it is used.
The e2e repository was created in July 2026 and launched publicly on October 1, 2026. It is on version 0.17.0, and its README says that "APIs and config can still change between minor releases".
The npm registry shows the gap in usage. In the week of September 28 to October 4, 2026, @playwright/test had 86,157,551 downloads and e2e had 60,933.

More than 5,000 stars only 5 days after the public launch shows real interest in e2e. It does not yet show years of production use, and that matters when a release can change the API.
For most teams, that points toward a trial on a few flows, not a full switch.
Note: Playwright vs e2e is not a winner-takes-all choice. The e2e web engine ships its own pinned playwright-core, so both runners can live in 1 project and you keep your own @playwright/test version.
Of these 7 differences, the AI model raises the most questions, mainly about cost, and e2e answers them with its cache.
How the e2e replay cache controls AI cost
Calling a model on every step of every CI run would be slow and costly. The replay cache is how e2e avoids that.
The framework itself is free and open source. What you pay for is model usage, through your own subscription or API key.
The e2e replay cache is a local store of the actions an agent took in a verified agent.act step. On later runs, e2e repeats those actions without calling the AI model, and hands back to the agent only when the app has changed.
TesterArmy's announcement says a replayed step "runs at the speed of a scripted test and uses no tokens". The caching docs explain the rules behind that claim.

What gets recorded
A step is recorded only when a later check verifies it. A locator assertion such as toHaveText counts, and so does agent.assert.
A plain value check like expect(value).toBe() does not count. Neither does the test simply passing.
Recordings are JSON files under .e2e/cache/. According to the docs, they hold actions and end-state checks, and no prompts, screenshots, or secret values.
What always calls the model
Only agent.act replays. The docs are direct about the rest: agent.assert, agent.waitFor, and agent.extract "always run live".
So a suite full of agent.assert calls will use the model on every run, cache or not. Using expect for exact checks keeps those calls down.
Retries never replay either. A retried test runs the agent live, which is worth knowing if you already fight flaky tests.
How the cache behaves in CI
The cache is read-write on your machine and read-only in CI by default. The init command also adds .e2e/cache/ to .gitignore, so CI starts with no recordings.
To share replays, you remove that line and commit the directory. The docs advise reviewing committed entries like test data.
npx e2e run --no-cache # run every agent step live
npx e2e cache stats # entry count and size
CI=1 npx e2e run --strict-cache # fail when a recording goes stale
The --strict-cache flag matters for budgets. Without it, a stale recording quietly hands the step to the agent, and every CI run pays for model calls until someone records it again.
The cache covers cost, which leaves the bigger migration question of what you would give up.
Which Playwright features are missing in e2e?
The honest answer is "quite a few, for now". The e2e team publishes a migration guide that maps Playwright features to their e2e counterparts and marks each one with a status.
We counted every row in that guide on October 6, 2026. It maps 185 Playwright config keys, APIs, and CLI flags.

Half of the rows (93 of 185) carry over as "same" or "renamed". About 30% (55 rows) are marked "missing", which means e2e has no equivalent yet.
| Area | Same | Renamed | Web only | By design | Missing | Total |
|---|---|---|---|---|---|---|
| Config | 3 | 21 | 0 | 1 | 8 | 33 |
| Locators | 6 | 2 | 2 | 2 | 4 | 16 |
| Actions | 9 | 8 | 2 | 4 | 11 | 34 |
| Assertions | 11 | 2 | 2 | 5 | 8 | 28 |
| Navigation and app | 1 | 3 | 4 | 2 | 4 | 14 |
| Network and browser | 0 | 0 | 10 | 2 | 8 | 20 |
| Test structure | 4 | 8 | 0 | 1 | 7 | 20 |
| CLI | 9 | 6 | 0 | 0 | 5 | 20 |
| Total | 43 | 50 | 20 | 17 | 55 | 185 |
"Web only" means the call works through the browser fixture and is skipped on mobile targets. "By design" means e2e behaves differently on purpose and plans to stay that way.
Gaps most teams will notice
- Pixel comparisons: toHaveScreenshot() and toMatchSnapshot() have no equivalent, so Playwright visual testing suites cannot move yet.
- HTML report: there is no reporter: 'html' and no show-report command.
- Parallelism inside a file: fullyParallel is missing. Files run in parallel, and tests in a file run in order.
- Global setup: globalSetup and globalTeardown are not available yet.
- Device descriptors: devices['iPhone 15'] is missing, and web({ viewport }) sets the size only.
- Popups: a new window that the app opens is not driven.
- Custom matchers: expect.extend() is not supported yet.
CI has a gap too. e2e supports --shard, but its CI docs say combining shard reports into 1 document "is not implemented yet".
Playwright can merge shard reports today. If you rely on Playwright sharding, you can size your shard count with the sharding calculator in TestDino's free tools.
Differences that are intentional
Some gaps will never close, because they are design choices. Tests get no raw Page object, so that nothing escapes recording and redaction.
Passwords are handled as a Secret with no string value. The docs state that the model never sees it.
Note: These counts are a snapshot from October 6, 2026. e2e is pre-1.0 and ships often, so check the live migration guide before you decide that a missing feature blocks you.
A list of gaps sounds like a reason to wait, but the migration path is built so that you never need a full rewrite.
How to migrate from Playwright to e2e without a rewrite
The guide's own summary is "run e2e alongside Playwright and migrate tests one file at a time". e2e reads e2e.config.ts and picks up tests/**/.e2e.ts, so your .spec.ts files are left alone.
6 steps to a first migrated test
- Run npx e2e init --yes in the project. It writes a config and an example test.
- Choose a model for agent steps: a subscription, an API key, or a local model.
- Port the settings from playwright.config.ts into e2e.config.ts.
- Pick 1 spec whose steps change often, such as a wizard or a checkout.
- Rewrite it as a .e2e.ts file, and add { exact: false } where a Playwright query matched a fragment.
- Run it with npx e2e run tests/checkout.e2e.ts, and keep the Playwright spec until the new test is stable.
What a migrated test looks like in practice
The guide says every page.getBy* line has a screen.getBy* twin. To check that, we ported the scripted checkout test line for line and ran it on our demo shop.
It passed on the first run, with no AI model configured. e2e also records a trace by default on a local run, and that file is a normal Playwright trace.

Look at the action list on the left. Each step is a Playwright action such as Fill or Click, found through e2e's own e2e-label and e2e-roots selectors.
This is the "layer on top of Playwright" idea made visible. Your debugging habits carry over, even though the test code and the runner are new.
How the config maps
The config section of the guide pairs each Playwright key with its e2e counterpart and a status. A recognized Playwright key such as baseURL is rejected with a message that names where it lives in e2e.
A minimal web config looks like this, followed by the mappings you will use most:
import type { E2EConfig } from 'e2e';
import { web } from '@e2e-dev/web';
import { gateway } from 'ai';
export default {
agents: { default: { model: gateway('openai/gpt-6-luna-fast') } },
targets: [{ engine: web(), app: { url: 'http://localhost:3000' } }],
} satisfies E2EConfig;
| Playwright | e2e |
|---|---|
| use.baseURL | app: { url } on the target |
| webServer | app: { url, command } on the target |
| projects | targets |
| testDir, testMatch | tests globs |
| npx playwright test | npx e2e run |
| --project web | --target web |
| npx playwright codegen | npx e2e explore or npx e2e mcp |
What happens to page objects and sign-in
Page objects still work. A class that wrapped page now takes screen and browser, so the page object model you already have keeps its shape.
Sign-in changes more. There is no storageState file, because sessions are saved per run with test.setup and are encrypted. If you follow the usual Playwright authentication pattern, plan to rewrite that part.
Custom fixtures keep the same test.extend shape, but worker-scoped fixtures are missing. Teams with heavy Playwright fixtures should test this early.
Tip: Start with the flow that breaks most often when the UI changes. The e2e docs suggest a wizard or a checkout, and they say that stable Playwright tests can stay where they are.
e2e also installs a skill and an MCP server, so coding agents such as Claude Code or Cursor can write and debug these tests. If you would rather have an agent write plain Playwright tests, the open source Playwright skill does the same job for Playwright.
Knowing how to migrate is useful, but it does not settle whether you should.
Which one should you choose?
There is no single winner in the Playwright vs e2e question. The right choice depends on what your tests need today.
Stay on Playwright when
- Your selectors are stable and maintenance is not a real cost.
- You need pixel-level visual checks, an HTML report, or merged shard reports.
- Your tests are written in Python, Java, or .NET.
- You cannot send screen content to an AI model, or you have no budget for model usage in CI.
- You need an API that will not change under you.
Try e2e when
- A few flows keep breaking because the UI changes often.
- You want 1 TypeScript suite that covers web, iOS, and Android.
- You already pay for a ChatGPT or GitHub Copilot plan and want to reuse it.
- Your team accepts a pre-1.0 tool and can absorb API changes.
Use both when
For most Playwright teams, this is the practical path. Keep the stable suite in Playwright, and move the 2 or 3 most brittle flows to e2e as a trial.
Then compare the results after a few weeks. Look at how often each version failed, how long it ran, and what the model usage cost.
Whichever runner you choose, failures still need a clear owner and a clear cause. The wider Playwright AI ecosystem is moving fast, but that part has not changed.
Conclusion
Playwright vs e2e is less of a fight than the name suggests. On the web, e2e runs on Playwright's browser automation and adds an agent, a replay cache, and a mobile engine on top.
Playwright remains the stable choice, with 5 languages, pixel comparisons, and mature debugging tools. e2e is a promising way to cut selector upkeep, with real limits: 55 of the 185 mapped Playwright features are still missing, and the API is pre-1.0.
The low-risk move is to run both. Keep your Playwright suite, try e2e on your most brittle flow, and let your own data decide.
If you are still weighing other options, our roundup of Playwright alternatives covers the wider field. And if you stay on Playwright, TestDino helps you track flaky tests and failure causes across every CI run.
FAQs

Pratik Patel
Co-founder


