Leveraging Jev with Playwright: Faster, Cheaper AI Test Automation
Learn how to pair Jev with Playwright for cheaper AI test automation, with setup steps, 5 use cases, and benchmarks.
A new kind of AI model can answer one small question about a web page, like "which button submits this form?", in under a second. Developers are already pairing Jev with Playwright to get those answers for a fraction of a cent.
AI-driven browser tests get expensive because the AI keeps re-reading the whole page. In our test, a setup built on Playwright MCP read over 200,000 tokens, the small chunks of text an AI bills you for, to check one login.
This guide explains Jev, TypeSafe AI's System One model, in plain words and follows one real login test from start to finish. You will see where it fits in test automation, what it cost us, and where it fails.
What is Jev?
Jev is an AI model from TypeSafe AI, released on September 15, 2026. It does not write text. You give it some text and a question with fixed answer options, and it tells you which answer fits and how sure it is.
A simple way to picture it: most AI models you know, such as Claude, are essay writers. Jev is a multiple-choice quiz taker. It can only tick one of the boxes you give it, which is why it is quick and cheap.
TypeSafe AI calls this a System One model. Its launch post sums Jev up as "unstructured state in, typed probabilistic decisions out." In plain words: messy text goes in, a clear answer with a confidence number comes out.
A System One model is a model built to make fast, structured decisions that software can use directly. It does not write text or code. It picks from the options you define and attaches a probability to each answer.
The example used in this guide: 1 login test
To keep things concrete, this guide follows a single test all the way through. It is the same test we used in our benchmark.
A shopper opens the TestDino demo store at storedemo.testdino.com, types an email and password, and clicks the sign-in button. If everything works, the shopper lands on their profile page.
Playwright is the tool that does the opening, typing, and clicking in a real browser. The hard part is the last step: confirming that the profile page really appeared. Each such check is called an assertion, and our login test has 13 of them.
The 3 question types: Choice, Score, and Noul
Every question you send to Jev is one of 3 types. The TypeSafe AI Jev documentation calls them primitives. Here is each one, applied to our login test:
| Type | What it returns | Login test example |
|---|---|---|
| Choice | One option from a list you define, with a probability for each option | "What kind of page is this: profile, login, or error?" |
| Score | A position on an ordered scale | "How relevant is this part of the page to signing in, from 0 to 3?" |
| Noul | The probability that the answer to a yes/no question is yes | "Is a signed-in user's profile shown on this page?" |
A probability is just how sure Jev is, on a scale from 0 to 1. A Noul answer of 0.98 means "almost certainly yes", and 0.02 means "almost certainly no".
You can send many questions in one request, and Jev answers them all at the same time. So 10 checks on one page need 1 trip to Jev instead of 10.
Jev pricing, context window, and speed
Key specs as of September 2026, taken from the official models page unless marked otherwise. Jev calls the text you send it the state.
| Spec | Value | What it means for you |
|---|---|---|
| Price | $0.042 per million input tokens. Output is free. | You pay only for the text you send in. |
| Context window | 32,000 tokens for state plus the longest question, within a 64,000-token request | Very large pages must be trimmed before sending. |
| Latency | 70 to 500 ms end to end (vendor-stated in the launch post) | Quick enough to call in the middle of a test. |
| Input | Text only: a string, JSON object, or array. No images, audio, or video. | It cannot look at screenshots. |
| Current version | jev-1.13.0 (aliases: jev-latest, jev-preview) | Note which version you use, because a new one can answer differently. |
| Access | TypeSafe API (keys at console.typesafe.ai), plus OpenRouter, Vercel AI Gateway, and Cloudflare Workers AI | Get your key from one of these, not a lookalike site. |
Playwright can describe a page as a text outline called the accessibility tree. On a large page that outline can exceed the 32,000-token limit, so send only the part you are asking about.
What Jev cannot do
Jev cannot answer outside the options you give it, so you never get a garbled reply. It can still pick the wrong option. It also cannot write test code, explain its answer, or read pixels.
Treat it as a very fast judge: Playwright does the clicking, Jev does the deciding, and your assertions stay the final check. To see why that split saves money, look at where the money goes in a normal AI-driven test.
Why Playwright AI workflows are expensive today
Playwright AI workflows are expensive because the AI re-reads the page on almost every step, and it charges you for every chunk of text it reads.
5 terms that explain the cost
| Term | Plain meaning |
|---|---|
| LLM | A large language model: the kind of AI that writes text and code, such as Claude. |
| AI agent | An LLM that is allowed to take actions step by step, such as clicking around a browser. |
| Token | A small chunk of text, often a word or part of a word. AI models charge by the number of tokens they read and write. |
| Snapshot | A text outline of a web page: its headings, buttons, links, and form fields. |
| Playwright MCP and Playwright CLI | 2 ways to let an AI agent control a browser through Playwright. |
Snapshot bloat, explained
Go back to our login test. An AI agent running it needs to know what is on the login page, and then what is on the profile page. With Playwright MCP, the agent receives a full snapshot of the page after its actions.
That is where the cost builds up. It is documented in issue #1216 on the Playwright MCP repo, which reports that "console messages and the entire accessibility tree for the page are returned in the response every time."
The agent pays to read all of it, even when it only needed one button. The official Playwright MCP README now has a --snapshot-mode setting that accepts full or none, and the default is still full.
The same README says CLI commands "are more token-efficient: they avoid loading large tool schemas and verbose accessibility trees into the model context." That difference is the main cost argument in any Playwright CLI vs MCP comparison.
Tokens per task: MCP vs CLI
We measured this on our login test on September 26, 2026. Claude Opus 5 answered the same 13 assertions through both interfaces:
| Setup | Input tokens for the same 13 assertions |
|---|---|
| Claude Opus 5 + Playwright CLI | 56,874 |
| Claude Opus 5 + Playwright MCP | 210,153 |
MCP read about 3.7 times as much text as the Playwright CLI, and both reached the same verdicts. The full method is in the benchmark section further down.
The decision map: who should decide what
Switching from MCP to the CLI cuts the reading, but the LLM still answers every question. And most of those questions are small.
Look at our login test as a chain of small decisions. Which box is the email field? Did the click on the sign-in button work? Is this the profile page? An LLM agent overpays for each one, because it brings an essay writer to a multiple-choice question.
The decision map below shows how Jev with Playwright splits that work. Each decision goes to the cheapest thing that can make it correctly.

Here is the same map applied to the login test:
- Plain code handles anything with one exact answer. "Did the address bar change to the profile URL?" is a job for ordinary test code.
- Jev handles fuzzy questions where you can list the possible answers. "Does this page look like a signed-in profile?" is a yes/no question, so Jev can take it.
- An LLM handles open-ended work. Writing the login test in the first place, or exploring a site it has never seen, needs a model that can plan and write.
An agent that relies on an LLM alone sends all 3 kinds of work to it, whichever tools from the Playwright AI ecosystem it uses. In our benchmark, moving the middle group to Jev is what cut the cost, and setting it up takes 4 steps.
How to set up Jev with Playwright
To set up Jev with Playwright, you get an API key, install the official SDK, send a page snapshot as the state, and read the answer in your test.
- Get a key. Create an API key in the TypeSafe console at console.typesafe.ai/keys and save it as TYPESAFE_API_KEY. A key is like a password that lets your test talk to Jev.
- Install the SDK. Add @typesafe-ai/sdk to your Playwright project. The SDK is a small helper library, and it needs Node.js 20 or newer.
- Send a snapshot. Capture the page as text with locator.ariaSnapshot() and pass it as state, along with your questions.
- Read the answer. Use the probability or choice that comes back in a normal expect check.
npm install -D @playwright/test @typesafe-ai/sdk
export TYPESAFE_API_KEY="your-key-here"
Here is our login test with Jev added. The field and button names are illustrative, so match them to your own app.
import { test, expect } from "@playwright/test";
import { TypeSafeClient, choice, noul } from "@typesafe-ai/sdk";
// Pin the version so answers do not shift when a new model ships
const jev = new TypeSafeClient({ defaultModel: "jev-1.13.0" });
test("valid user signs in and sees their profile", async ({ page }) => {
// 1. Playwright does the clicking and typing
await page.goto("/login");
await page.getByLabel("Email").fill(process.env.STORE_EMAIL!);
await page.getByLabel("Password").fill(process.env.STORE_PASSWORD!);
await page.getByRole("button", { name: "Sign in" }).click();
// 2. Plain code checks what has one exact answer
await expect(page).toHaveURL(/profile/);
// 3. Jev answers the fuzzy questions: one snapshot, two questions
const snapshot = await page.locator("body").ariaSnapshot();
const { answers } = await jev.systemOne({
state: { url: page.url(), snapshot },
questions: {
signedIn: noul("Is a signed-in user's profile shown on this page?", {
true: "Profile details are visible along with a way to sign out",
false: "A login form, an error message, or an empty page is shown",
}),
pageType: choice("What kind of page is this?", {
profile: "Account or profile page for a signed-in user",
login: "Login form asking for credentials",
error: "Error or access denied page",
}),
},
});
// 4. Your test makes the final call
expect(answers.signedIn.noul).toBeGreaterThan(0.9);
expect(answers.pageType.choice).toBe("profile");
});
In plain words, the test does 4 things:
- Step 1 signs in the way a shopper would.
- Step 2 checks the address bar with ordinary code, because a URL has one exact answer.
- Step 3 turns the page into text and asks Jev 2 questions about it in a single request.
- Step 4 passes only if Jev is more than 90% sure a profile is showing and picks "profile" as the page type.
The snapshot in step 3 is a text version of the page's accessibility tree, as described in the official aria snapshots guide. It is plain text, which is the only kind of input Jev accepts.
Tip: Send only the part of the page you are asking about, for example page.getByRole("main").ariaSnapshot(). TypeSafe's docs note that accuracy falls as the state fills with unrelated content, and a smaller state is cheaper too.
Python teams can use the official typesafe-sdk package, and any other language can call the web address directly at POST https://api.typesafe.ai/v1/systemone.
If an AI agent drives your browser from the terminal, a Jev Playwright CLI workflow needs the CLI itself set up first. The playwright-cli pack in the open source Playwright Skill repo gives the agent 10 ready-made guides for that:
npx skills add testdino-hq/playwright-skill/playwright-cli
Once one call works, the same pattern applies at several other points in a test run.
5 ways to use Jev with Playwright in a real test run
Jev fits in 5 places inside a Playwright run: element selection, semantic assertions, agent step verification, snapshot pruning, and destructive action gating.
Those names sound heavy, but each one is a single small question. The sections below walk through all 5 on our demo store. Playwright still performs every action.

1. Element selection from a snapshot
The problem
Our login test clicks a button named "Sign in". Next month a designer renames it "Log in", and the test can no longer find it.
How Jev helps
your code collects the buttons that are actually on the page and asks Jev a Choice question: "Which of these starts the sign-in?" Jev can only pick a button that exists, or a none option you add for "nothing matches".
Stable Playwright locators should still be your first choice, since they cost nothing. Jev is the fallback for when a label or layout changes. A single Choice accepts up to 255 options.
2. Jev semantic assertions with a toSatisfy matcher
The problem
Imagine the profile page greets the shopper, but the wording changes. One day it says "Welcome back, John", another day "Good to see you again". A normal check for exact text breaks.
How Jev helps
Standard Playwright assertions compare exact values. Jev semantic assertions check the meaning instead. You write the claim in plain English, and Jev returns how likely it is to be true.
import { expect as baseExpect, type Page } from "@playwright/test";
import { TypeSafeClient, noul } from "@typesafe-ai/sdk";
const jev = new TypeSafeClient({ defaultModel: "jev-1.13.0" });
export const expect = baseExpect.extend({
async toSatisfy(page: Page, claim: string, threshold = 0.9) {
const snapshot = await page.locator("body").ariaSnapshot();
const { answers } = await jev.systemOne({
state: { snapshot },
questions: { claim: noul(claim) },
});
const probability = answers.claim.noul;
return {
pass: probability >= threshold,
name: "toSatisfy",
message: () => `Jev scored "${claim}" at ${probability.toFixed(2)}, threshold ${threshold}`,
};
},
});
With that helper in place, the check is one readable line: await expect(page).toSatisfy("The page greets the signed-in shopper by name"). The open source playwright-jev project ships a ready-made version if you prefer not to write your own.
3. Agent step verification
The problem
An AI agent clicks the sign-in button and moves on. But a click can succeed without doing what the agent intended, for example when the password was wrong.
How Jev helps
After every action, the agent's code asks Jev a yes/no question: "Did the page change to a signed-in profile?" If the answer is no, the agent stops instead of building on a failed step.
This is cheap enough to run after every step taken by Playwright test agents. The playwright-jev maintainer measured probabilities of 0.88 to 0.93 when the outcome was achieved and 0.02 to 0.03 when it was not, on Jev 1.13.
4. Snapshot pruning in front of Playwright MCP
The problem
Picture a store home page with a header, a product grid, banners, and a footer. An agent that only wants the sign-in link still pays to read all of it.
How Jev helps
Here Jev works for the LLM instead of replacing it. Your code splits the page into regions and asks Jev a Score question for each: "How relevant is this region to signing in?" Only the high scorers go to the LLM.
A Jev Playwright MCP proxy does this for you. The open source jev-playwright-mcp project sits in front of the official server, keeps the same tool names, and shrinks large snapshots to the regions that match a goal you set.
5. Destructive action gating
The problem
Say the profile page has a "Delete account" button. An agent exploring the page might click it.
How Jev helps
Before any click, your code asks Jev a yes/no question: "Would this action delete data, send a message, or spend money?" Above a limit you choose, the click is blocked.
The jev-playwright-mcp proxy blocks at a probability of 0.6 by default. This sits well alongside the checks in our MCP server security guide, because the gate is code you control and not a request the agent can talk its way around.
All 5 patterns assume that Jev gives the same answers as an LLM for less money. We tested that assumption on the login test.
Jev vs Claude for Playwright assertions: benchmark results
In our test, Jev answered 13 Playwright login assertions for $0.000118 per run. Claude Opus 5 cost $0.071 through the Playwright CLI and $0.245 through Playwright MCP, which is 601 and 2,078 times more. All 4 setups returned the same verdicts.

The test is simple on purpose: a valid user signs in on storedemo.testdino.com and sees their profile page. The test stays the same every time. Only the engine that answers the 13 assertions changes.
Cost, tokens, and run time per engine
| Setup | Input Tokens | Output Tokens | Run Time | Cost per Run | vs Jev |
|---|---|---|---|---|---|
| Jev + Playwright CLI | 2,811 | 271 | 11.2 s | $0.000118 | baseline |
| Jev + Playwright MCP | 2,775 | 271 | 16.2 s | $0.000117 | 1.0x |
| Claude Opus 5 + Playwright CLI | 56,874 | 157 | 25.3 s | $0.071 | 601x |
| Claude Opus 5 + Playwright MCP | 210,153 | 1,137 | 25.0 s | $0.245 | 2,078x |
Input tokens are the text each engine had to read. Output tokens are the text it wrote back.
What the numbers show
- Same answers, different bill. All 4 engines agreed on all 13 assertions. The difference is cost and time, not correctness.
- Input tokens drive the gap. Jev read about 2,800 tokens. Claude read 56,874 through the CLI and 210,153 through MCP, because the agent's own instructions and page snapshots ride along with every turn.
- Jev is about twice as fast. The Jev run took 11.2 seconds with the CLI, against roughly 25 seconds for both Claude setups.
- For Jev, CLI beats MCP on speed only. Cost is the same, but MCP added 5 seconds.
Put another way, $1 pays for about 8,500 runs of this login test with Jev. The same $1 pays for about 14 runs with Claude through the CLI, and about 4 through MCP.

How we measured
These figures are measured, not estimated. Jev is priced at $42 per billion input tokens, which is TypeSafe AI's published rate.
Note: This is 1 run per engine with no averaging, and all 13 assertions pass. The Claude rows ran through Claude Code, which loads roughly 13,000 tokens of its own instructions before it sees the task, so calling Claude directly would cost less than shown.
If you run Claude Code with Playwright every day, that overhead is part of your real bill. You can model what it adds up to per month with the CI Budget Calculator in TestDino's free tools.
Our result is not the only public data point. In a separate test with a different model and task, Filip Hric's comparison found Playwright CLI plus Jev to be 98% cheaper and twice as fast as Playwright MCP.
Matching verdicts on 13 passing checks are encouraging, but they do not tell you how far to trust a single answer. For that, you need a pass mark.
How to set Jev thresholds you can trust
Set Jev thresholds from your own data: run in report mode first, record the probabilities Jev returns on cases you already know the answer to, and then decide each pass mark by how costly a wrong answer would be.
A threshold is simply a pass mark. In the login test code above, the pass mark is 0.9: if Jev is at least 90% sure the profile is showing, the check passes. Set it too low and wrong answers slip through. Set it too high and correct answers get rejected.
What the probabilities look like in practice
The playwright-jev project publishes real Jev 1.13 probabilities from its own test runs. They show where the model is decisive and where it hesitates.
| Scenario | Probability returned |
|---|---|
| True semantic assertion | 0.98 |
| False semantic assertion | 0.01 to 0.03 |
| Step verification, outcome achieved | 0.88 to 0.93 |
| Step verification, outcome not achieved | 0.02 to 0.03 |
| Removed control, correct "none" answer | 0.88 to 0.98 |
| Renamed or restructured control, correct element | 0.59 to 0.75 |
Yes/no questions separate cleanly: true claims land near 1 and false ones near 0. Picking a renamed button is the weak spot. Jev picks the right one, but it is only 59% to 75% sure, so a 0.9 pass mark would reject a correct answer.
Describe outcomes, not labels
How you phrase the question moves the number. The playwright-jev docs advise describing the result you want, not the words printed on the button.
Tip: Ask for "a checkout page with a payment form", not "Checkout now". According to the playwright-jev README, bare labels drop confidence to about 0.65.
Two more habits keep your numbers comparable from run to run:
- Keep option order and wording fixed. The options are part of what Jev reads, so change them only on purpose and re-check your pass marks after every change.
- Do not reuse a pass mark across question types. TypeSafe's docs warn that a Choice and a Noul answer different questions, so a limit tuned for one does not carry over to the other.
Start in report mode
TypeSafe's confidence guide suggests 3 bands as a starting point, and it stresses that the right values depend on your own domain and data.
| Confidence | Suggested handling |
|---|---|
| Above 0.9 | Act automatically |
| 0.5 to 0.9 | Proceed with caution: confirm, flag for review, or gather more information |
| Below 0.5 | Do not act: route to a human or fall back to another system |
Report mode means Jev's answer is written down but never changes the test result. Run that way until you have both passing and failing examples, compare the recorded numbers against what really happened, and only then let a pass mark decide a test.
Save each probability with the test result
To compare numbers later, each one has to be saved somewhere. Playwright annotations, which are small notes you attach to a test, are the simplest way to do that:
const { answers, model } = await jev.systemOne({ state, questions });
test.info().annotations.push({
type: "jev-probability",
description: `signedIn=${answers.signedIn.noul.toFixed(2)} model=${model}`,
});
Save the model version next to the number, so that when a probability shifts you can tell whether the app changed or the model did.
Saved numbers also reveal drift. Suppose the profile check in our login test scored 0.98 last week and scores 0.71 today. If your pass mark is 0.7, the test passes on both days, so you only notice the slide when you compare the saved numbers from several runs.
If you skip report mode, an answer that hovers near the pass mark will pass on one run and fail on the next. That is a new source of flaky tests, which are tests that pass and fail without the code changing.
Good pass marks reduce the risk, but they cannot remove it, because some failures come from the model itself.
Where Jev breaks and when not to use it
Jev breaks when the right answer is not among your options, when the input is an image, or when the question needs counting, math, or date logic. Use a plain Playwright check whenever there is one exact answer.
Wrong but valid answers
Jev always picks one of your options, even when none of them is right. Say the store is down and shows a "Back soon" maintenance page. If your only options are profile, login, and error, Jev must choose one of those 3.
The fix is simple: always include a none or other option, so Jev has an honest way to say "this is something else".
Text only, no pixels, no explanations
The official Jev 1.13 limitations page lists weaknesses that matter directly for testing:
- Literal reading. Jev "answers the question you wrote, not the one you meant."
- Counting and math. Jev "does not count reliably." To check that the profile lists 3 saved addresses, count them in code.
- Dates. Jev reads dates as text, so "is this order from the last 7 days?" is unreliable.
- Large state. Unrelated content in the state distracts the model.
- Hostile content. Text on the page can steer the answer, which matters on pages with user reviews or comments.
Jev also gives no reasons. When an answer looks wrong, you cannot ask why. You can only change the question or the text you send and ask again.
Pin the version
The jev-latest name currently points to jev-1.13.0, and it will move when a new version ships. TypeSafe's docs advise using the exact version ID once you have tuned your pass marks against it, so you decide when to switch.
When not to use Jev
-
There is one exact answer. A URL, a count, a price, or a fixed label belongs in [ct]expect[/ct]. It costs nothing and gives the same answer every time.
-
You need code or a plan. Writing a test, repairing one, or exploring an unknown flow is LLM work. The AI test generation tools built for that job are a better fit.
-
You need to see the page. Checking colours, layout, or images needs screenshots, and Jev does not accept images.
Security note: use official endpoints onl
Jev's launch attracted copycats fast. Eye Security's Rise of the Jev-Clones research counted about 670 new domains containing "jev" that received TLS certificates within 8 days of launch.
Some of those sites resell the real service at up to 11.5 times the official price, and everything you send passes through their servers first. For our login test, that would mean a stranger's server reading snapshots of your app's pages.
Conclusion
Jev with Playwright works when each part does only what it is good at. Playwright clicks and types, plain checks handle anything with one exact answer, Jev answers the small fuzzy questions, and an LLM is saved for planning and writing code.
On our login test, that split answered 13 assertions for $0.000118 per run, against $0.071 to $0.245 for Claude Opus 5, with the same verdicts. It is one run on one flow, so treat it as a starting point and measure your own tests.
Start small: add one yes/no question to one test in report mode, save the probability, and watch it across your next runs. If it stays stable, move the next fuzzy check across and see what it does to your Playwright CI cost.
FAQs

Savan Vaghani
Product Developer



