TestDino
Workshop · SF Tech Week 2026 · October 7

Playwright + AI Agents + LLM Evals

Generate tests, run them, debug them, then test the things that never answer the same way twice.

Pratik Patel · Founder, TestDino

Playwright and TestDino logos joined by a handshake, for the Playwright + AI Agents + LLM Evals workshop at SF Tech Week 2026.
Trusted by teams at
Swatch
Malwarebytes
Deriv
OpenObserve
Keen
Rho
Franklin
Lumin
Part 01

Generate

Prompts, the Playwright skill, and tips for tests you can trust

Ask yourself

Who writes your Playwright tests now, you or the agent?

More code, same number of reviewers.

Same prompt, 2 agents

Agents already write clean locators. The skill changes how the test fits your suite.

Without: a clean test, alone

  • Clean locators and assertions
  • Login through the UI in every test
  • Credentials and test data inline
  • 1 file, nothing reusable
typescript
test('update profile', async ({ page }) => {
  await page.goto('/login');
  await page.getByLabel('Email').fill('qa@a.co');
  await page.getByLabel('Password').fill('pw123');
  await page.getByRole('button', { name: 'Sign in' })
    .click();
  await page.goto('/settings/profile');
  await page.getByLabel('Name').fill('Ada');
  await page.getByRole('button', { name: 'Save' })
    .click();
  await expect(page.getByText('Saved')).toBeVisible();
});

With: a test built for a suite

  • Signs in once, reuses the state
  • Page objects and fixtures
  • Steps and tags that show in reports
  • CI setup: retries, traces, shards
typescript
// auth.setup.ts signs in once, state is reused
test('update profile @smoke',
  async ({ profilePage }) => {
  await test.step('edit name', async () => {
    await profilePage.setName('Ada');
  });
  await test.step('save', async () => {
    await profilePage.save();
    await expect(profilePage.toast)
      .toHaveText('Saved');
  });
});

6 habits for agent-written tests

Give it the live app

Let the agent open the page through Playwright MCP or CLI. Guessed selectors are the top cause of broken tests.

Write the spec first

A plain Markdown scenario with the business outcome. The agent turns it into a .spec.ts file.

1 behavior per test

Small isolated tests fail for 1 reason. Long journeys fail for 10.

Role locators only

getByRole and getByLabel survive redesigns. CSS and XPath do not.

Ban fixed waits

No waitForTimeout. Web-first assertions like toBeVisible wait on their own.

Make it run the test

The agent is not done until the test passes 3 times in a row.

Get the Playwright skill

Open source on GitHub. Star it and install it in your project.

bash
npx skills add testdino-hq/playwright-skill
Packscorecipommigrationplaywright-cli
Part 02

See

Labeled video recordings and visual snapshot comparison

Ask yourself

Do you read every test the agent writes before you merge it?

Be honest.

Videos that label every action

typescript
// playwright.config.ts · Playwright 1.63
video: {
  mode: process.env.VIDEO ? 'on' : 'off',
  size: { width: 1280, height: 720 },
  show: {
    actions: { position: 'top-right' },
    test: { level: 'step', position: 'top-left' },
  },
},

show.test

Spec › test › current step, top-left.

show.actions

Highlight and action title on every element.

showChapter()

A title card between phases.

showOverlay()

The exact expect() about to run.

Snapshots catch what assertions miss

Visual snapshot

typescript
await expect(page).toHaveScreenshot(
  'checkout.png', {
  maxDiffPixelRatio: 0.01,
  mask: [page.getByTestId('clock')],
});

Compares pixels against a saved baseline. Mask anything dynamic.

Aria snapshot

typescript
await expect(page.getByRole('main'))
  .toMatchAriaSnapshot(`
  - heading "Your cart"
  - button "Checkout"
`);

Compares page structure, not pixels. Stable across themes and fonts.

Update baselines on purpose: npx playwright test --update-snapshots

Commands worth memorizing

Command or settingWhat it gives you
npx playwright test --last-failedReruns only the tests that failed in the previous run. Playwright reads them from the last run's results file, so you fix and rerun without retyping names.
npx playwright test --only-changed=mainRuns only the test files that changed between your branch and main, plus any test that imports a changed file.
npx playwright test --debug=cliYour coding agent attaches and steps through
npx playwright trace open trace.zipAgents read a trace from the terminal
failOnFlakyTests: trueConfig option. By default a test that fails and then passes on retry counts as passed and is marked flaky. With this on, that test fails the whole run, so flakiness gets fixed instead of hidden behind retries.
npx playwright init-agentsPlanner, generator and healer agents
Part 03

Run and debug

Test cases to execution, TestDino, Jira, rerun only what failed

Ask yourself

CI went red this week. Was it a real bug or a flaky test?

And how long did it take you to find out?

1 loop, from manual case to fixed test

  1. 01

    Manual test case

    Write what to check, in plain words

  2. 02

    Generate

    The agent turns it into the right Playwright test

  3. 03

    Run

    Stream the run and confirm it is green

  4. 04

    Debug

    A failure arrives with trace, video and error group

  5. 05

    Fix and rerun

    Understand the cause, fix it, rerun only that test

  6. Keep it closed

    Flaky and failing tests surface on their own. The agent takes them from here.

    Loops back to Run

Generation is the easy half. Closing the loop after it is where teams lose time.

3 steps to stream runs into TestDino

  1. 1

    Install

    bash
    npm install @testdino/playwright
  2. 2

    playwright.config.ts

    typescript
    reporter: [
      ['@testdino/playwright', {
        token: process.env.TESTDINO_TOKEN,
      }],
    ]
  3. 3

    Run

    bash
    npx playwright test

Let the agent read the failure itself

You, in Claude Code or Cursor

"Why did checkout fail on main last night? Is it flaky or real? Fix it and rerun only that test."

  1. Pulls the failed runTestDino MCP
  2. Reads the error, trace and run history
  3. Edits the test or the app code
  4. Reruns that test and reports back
Localnpx testdino-mcp
Remotemcp.testdino.com
Connect the TestDino MCP
Part 04

Non-deterministic testing

Why assertions break on LLM output, and what evals do instead

Ask yourself

Your product has an AI feature. Do you have a test for it?

Not a demo. A test that runs on every change.

Same prompt. 10 runs. 10 different answers.

Every testing tool we covered so far assumes the app gives the same output twice. LLM features do not.

  1. Run 01
  2. Run 02
  3. Run 03
  4. Run 04
  5. Run 05
  6. Run 06
  7. Run 07
  8. Run 08
  9. Run 09
  10. Run 10

expect(answer).toBe( ??? )

Tests give pass or fail. Evals give a pass rate.

Test · 1 run

PassorFail

Eval · 10 runs

80% pass rate across 10 runs

  1. 1
    Start here

    Code checks

    Valid JSON, contains the order ID, under 100 words, no banned phrase. Fast and free.

  2. 2

    LLM judge

    A second model grades 1 yes or no question, such as "Did it answer the refund question?"

  3. 3

    Human review

    A person labels a sample. This is how you check that the judge agrees with you.

Run each case many times, then gate on a threshold.

8 of 10 passed

Ships

6 of 10 passed

Does not ship
Part 05

Agentic testing

Testing the path an agent took, not only its final answer

Ask yourself

Your agent gave the right answer. Do you know how it got there?

Which tools it called, in what order, how many times.

Right answer, wrong path

Task: "Refund order 4812." The customer got the right reply. Look at how the agent got there.

StepExpected tool callWhat the agent didResult
1get_order(4812)get_order(4812)Match
2check_refund_policySkippedMissing step
3issue_refund, onceissue_refund, twiceDuplicate action
4send_replysend_replyMatch

An output-only eval scores this pass.

A trajectory eval scores it fail.

Your first eval, this week

  1. 1

    Collect 20 real cases

    From production logs and support tickets, not invented

  2. 2

    Read the outputs yourself

    Write down every way they go wrong

  3. 3

    Turn each failure into a check

    Code check where possible, LLM judge where not

  4. 4

    Run every case several times

    1 run tells you nothing about a random system

  5. 5

    Gate the pull request on the score

    Same place your Playwright tests already run

Recap

3 things to try on Monday

  1. 01

    Install the skill

    Rerun 1 prompt you already used and compare the 2 tests.

    npx skills add testdino-hq/playwright-skill
    Back to part 01
  2. 02

    Turn on labeled video

    Review agent-written tests by watching them, then stream runs to one place.

    video.show.actions
    Back to part 02
  3. 03

    Write 1 eval

    20 real cases, 1 check, several runs each. Get your first pass rate.

Stay in touch

I share what I learn about AI agents, testing and evals every week.

Pratik Patel, Founder, TestDino

Pratik Patel

Founder, TestDino

Run the loop on your own suite

The free plan covers 5,000 test executions a month.

  1. 1npm install @testdino/playwright
  2. 2playwright.config.ts
  3. 3npx playwright test