Playwright + AI Agents + LLM Evals
Generate tests, run them, debug them, then test the things that never answer the same way twice.
Pratik Patel · Founder, TestDino

What we cover today
Generate
Prompts, the Playwright skill, and tips for tests you can trust
Ask yourself
Who writes your Playwright tests now, you or the agent?
More code, same number of reviewers.
Same prompt, 2 agents
Agents already write clean locators. The skill changes how the test fits your suite.
Without: a clean test, alone
- Clean locators and assertions
- Login through the UI in every test
- Credentials and test data inline
- 1 file, nothing reusable
test('update profile', async ({ page }) => {
await page.goto('/login');
await page.getByLabel('Email').fill('qa@a.co');
await page.getByLabel('Password').fill('pw123');
await page.getByRole('button', { name: 'Sign in' })
.click();
await page.goto('/settings/profile');
await page.getByLabel('Name').fill('Ada');
await page.getByRole('button', { name: 'Save' })
.click();
await expect(page.getByText('Saved')).toBeVisible();
});With: a test built for a suite
- Signs in once, reuses the state
- Page objects and fixtures
- Steps and tags that show in reports
- CI setup: retries, traces, shards
// auth.setup.ts signs in once, state is reused
test('update profile @smoke',
async ({ profilePage }) => {
await test.step('edit name', async () => {
await profilePage.setName('Ada');
});
await test.step('save', async () => {
await profilePage.save();
await expect(profilePage.toast)
.toHaveText('Saved');
});
});6 habits for agent-written tests
Give it the live app
Let the agent open the page through Playwright MCP or CLI. Guessed selectors are the top cause of broken tests.
Write the spec first
A plain Markdown scenario with the business outcome. The agent turns it into a .spec.ts file.
1 behavior per test
Small isolated tests fail for 1 reason. Long journeys fail for 10.
Role locators only
getByRole and getByLabel survive redesigns. CSS and XPath do not.
Ban fixed waits
No waitForTimeout. Web-first assertions like toBeVisible wait on their own.
Make it run the test
The agent is not done until the test passes 3 times in a row.
Get the Playwright skill
Open source on GitHub. Star it and install it in your project.
npx skills add testdino-hq/playwright-skillcorecipommigrationplaywright-cliSee
Labeled video recordings and visual snapshot comparison
Ask yourself
Do you read every test the agent writes before you merge it?
Be honest.
Videos that label every action
// playwright.config.ts · Playwright 1.63
video: {
mode: process.env.VIDEO ? 'on' : 'off',
size: { width: 1280, height: 720 },
show: {
actions: { position: 'top-right' },
test: { level: 'step', position: 'top-left' },
},
},show.test
Spec › test › current step, top-left.
show.actions
Highlight and action title on every element.
showChapter()
A title card between phases.
showOverlay()
The exact expect() about to run.
Snapshots catch what assertions miss
Visual snapshot
await expect(page).toHaveScreenshot(
'checkout.png', {
maxDiffPixelRatio: 0.01,
mask: [page.getByTestId('clock')],
});Compares pixels against a saved baseline. Mask anything dynamic.
Aria snapshot
await expect(page.getByRole('main'))
.toMatchAriaSnapshot(`
- heading "Your cart"
- button "Checkout"
`);Compares page structure, not pixels. Stable across themes and fonts.
Update baselines on purpose: npx playwright test --update-snapshots
Commands worth memorizing
| Command or setting | What it gives you |
|---|---|
npx playwright test --last-failed | Reruns only the tests that failed in the previous run. Playwright reads them from the last run's results file, so you fix and rerun without retyping names. |
npx playwright test --only-changed=main | Runs only the test files that changed between your branch and main, plus any test that imports a changed file. |
npx playwright test --debug=cli | Your coding agent attaches and steps through |
npx playwright trace open trace.zip | Agents read a trace from the terminal |
failOnFlakyTests: true | Config option. By default a test that fails and then passes on retry counts as passed and is marked flaky. With this on, that test fails the whole run, so flakiness gets fixed instead of hidden behind retries. |
npx playwright init-agents | Planner, generator and healer agents |
Run and debug
Test cases to execution, TestDino, Jira, rerun only what failed
Ask yourself
CI went red this week. Was it a real bug or a flaky test?
And how long did it take you to find out?
1 loop, from manual case to fixed test
- 01
Manual test case
Write what to check, in plain words
- 02
Generate
The agent turns it into the right Playwright test
- 03
Run
Stream the run and confirm it is green
- 04
Debug
A failure arrives with trace, video and error group
- 05
Fix and rerun
Understand the cause, fix it, rerun only that test
Keep it closed
Flaky and failing tests surface on their own. The agent takes them from here.
Loops back to Run
Generation is the easy half. Closing the loop after it is where teams lose time.
3 steps to stream runs into TestDino
- 1
Install
bashnpm install @testdino/playwright - 2
playwright.config.ts
typescriptreporter: [ ['@testdino/playwright', { token: process.env.TESTDINO_TOKEN, }], ] - 3
Run
bashnpx playwright test
Let the agent read the failure itself
You, in Claude Code or Cursor
"Why did checkout fail on main last night? Is it flaky or real? Fix it and rerun only that test."
- Pulls the failed runTestDino MCP
- Reads the error, trace and run history
- Edits the test or the app code
- Reruns that test and reports back
Non-deterministic testing
Why assertions break on LLM output, and what evals do instead
Ask yourself
Your product has an AI feature. Do you have a test for it?
Not a demo. A test that runs on every change.
Same prompt. 10 runs. 10 different answers.
Every testing tool we covered so far assumes the app gives the same output twice. LLM features do not.
- Run 01
- Run 02
- Run 03
- Run 04
- Run 05
- Run 06
- Run 07
- Run 08
- Run 09
- Run 10
expect(answer).toBe( ??? )
Tests give pass or fail. Evals give a pass rate.
Test · 1 run
Eval · 10 runs
80% pass rate across 10 runs
- 1Start here
Code checks
Valid JSON, contains the order ID, under 100 words, no banned phrase. Fast and free.
- 2
LLM judge
A second model grades 1 yes or no question, such as "Did it answer the refund question?"
- 3
Human review
A person labels a sample. This is how you check that the judge agrees with you.
Run each case many times, then gate on a threshold.
8 of 10 passed
Ships6 of 10 passed
Does not shipAgentic testing
Testing the path an agent took, not only its final answer
Ask yourself
Your agent gave the right answer. Do you know how it got there?
Which tools it called, in what order, how many times.
Right answer, wrong path
Task: "Refund order 4812." The customer got the right reply. Look at how the agent got there.
| Step | Expected tool call | What the agent did | Result |
|---|---|---|---|
| 1 | get_order(4812) | get_order(4812) | Match |
| 2 | check_refund_policy | Skipped | Missing step |
| 3 | issue_refund, once | issue_refund, twice | Duplicate action |
| 4 | send_reply | send_reply | Match |
An output-only eval scores this pass.
A trajectory eval scores it fail.
Your first eval, this week
- 1
Collect 20 real cases
From production logs and support tickets, not invented
- 2
Read the outputs yourself
Write down every way they go wrong
- 3
Turn each failure into a check
Code check where possible, LLM judge where not
- 4
Run every case several times
1 run tells you nothing about a random system
- 5
Gate the pull request on the score
Same place your Playwright tests already run
3 things to try on Monday
- 01
Install the skill
Rerun 1 prompt you already used and compare the 2 tests.
Back to part 01npx skills add testdino-hq/playwright-skill - 02
Turn on labeled video
Review agent-written tests by watching them, then stream runs to one place.
Back to part 02video.show.actions - 03
Write 1 eval
20 real cases, 1 check, several runs each. Get your first pass rate.
Stay in touch
I share what I learn about AI agents, testing and evals every week.

Pratik Patel
Founder, TestDino
Run the loop on your own suite
The free plan covers 5,000 test executions a month.
- 1
npm install @testdino/playwright - 2
playwright.config.ts - 3
npx playwright test
