Test Harness vs Agent Harness vs Eval Harness: What Each One Does
Three harnesses share one name. Learn what a test harness, agent harness, and eval harness each do, with code.
The word "harness" now turns up in 3 separate conversations: checking that software works, letting an AI helper do real work, and grading how well that helper did. That is why the test harness vs agent harness question keeps coming up, with the eval harness close behind it.
The trouble is that all 3 share one name, so a tester, an AI engineer, and a researcher can sit in one meeting and each picture a different thing. A team adopting AI coding agents can end up building one harness while believing it has another.
This guide defines each harness in plain words, compares them side by side, and shows how all 3 fit into one pipeline. Every example uses Playwright test automation, so the code maps to a real project.
Test harness vs agent harness vs eval harness at a glance
A test harness runs your tests against code using drivers and test doubles. An agent harness runs an AI model in a loop with tools, context, and memory so it can finish a task. An eval harness runs that agent across many tasks and trials, grades each result, and reports a score.
The quickest way to keep the 3 apart is to ask what each one wraps:
- A test harness wraps code. It feeds the code inputs and checks the outputs.
- An agent harness wraps a model. It gives the model tools and runs its decisions.
- An eval harness wraps an agent. It gives the agent tasks and scores the outcomes.
The 3 harnesses side by side
| Test harness | Agent harness | Eval harness | |
|---|---|---|---|
| What it wraps | Code under test | A language model | An agent and its tasks |
| Question it answers | Does this code work? | Can the model do this task? | How often does the agent succeed? |
| What comes out | Pass or fail | A finished task | A score across trials |
| Well-known example | Playwright fixtures and mocks | Claude Code | The SWE-bench harness |
Why the 3 terms get mixed up
All 3 borrow the same picture: gear strapped around something powerful to direct it.
A June 2026 paper on arXiv traces the word from horse tack to the classic test harness, then to the machine-learning evaluation harness, and finally to the agent harness.
The same paper calls today's usage "loose and polysemous." It notes that harness sometimes means a whole product such as Claude Code, and sometimes means the scaffold that runs an agent against benchmark tasks.
Most test harness vs agent harness confusion starts right there. The same word points at 3 layers, and the layer a person means depends on the job they do.
The oldest of the 3 is the one testers already own, so that is the right place to begin.
What is a test harness?
The ISTQB glossary gives the standard answer to what a test harness is in software testing: "a collection of drivers and test doubles needed to execute a test suite."
In simple terms, it is everything you place around a piece of code so that tests can run on their own and return a verdict.
The parts of a test harness
The ISTQB definition names 2 parts, and real projects add a few more:
- Driver. The piece that calls the code under test and controls the run. A test runner is a driver.
- Test double. A stand-in for a real dependency, such as a payment API. Stubs and mocks are test doubles.
- Fixtures and test data. The setup that puts the system in a known state before each test.
- Reporter. The part that turns raw results into a pass, a fail, and an error message.
A test harness in Playwright
Playwright ships most of these parts. The official docs say fixtures establish the environment for each test, giving the test everything it needs and nothing else.
The fixture below sets up a checkout page and swaps the real payment API for a test double.
import { test as base, expect, type Page } from '@playwright/test';
export const test = base.extend<{ checkoutPage: Page }>({
checkoutPage: async ({ page }, use) => {
// Test double: answer for the payment API so no real card is charged
await page.route('**/api/payments', async (route) => {
await route.fulfill({ json: { status: 'approved' } });
});
await page.goto('/checkout');
await use(page);
},
});
export { expect };
The test then uses that fixture. It never knows the payment API was replaced.
import { test, expect } from './fixtures';
test('shows a confirmation after payment', async ({ checkoutPage }) => {
await checkoutPage.getByRole('button', { name: 'Pay now' }).click();
await expect(checkoutPage.getByText('Order confirmed')).toBeVisible();
});
Running npx playwright test is the driver. The page.goto('/checkout') call assumes a baseURL in your playwright.config.ts.
Both patterns go deeper than this example. Playwright fixtures can also hold logged-in sessions, and network mocking covers every way to fake an API response.
The one property that defines it
A test harness is built to be deterministic. The same code and the same input should give the same verdict on every run.
When they do not, the suite has flaky tests, and the harness has stopped doing its job. That property becomes important again once an AI agent enters the picture.
Tip: Starting a test harness from zero? The Config Generator in TestDino's free tools builds a ready playwright.config.ts from your choice of browsers, workers, retries, and reporters.
A test harness assumes a person or a CI job decides what runs next. An agent harness hands that decision to a model.
What is an agent harness?
Anthropic defines an agent harness as the system that enables a model to act as an agent: it processes inputs, orchestrates tool calls, and returns results. The same post notes that people also call it a scaffold.
A model on its own only produces text. It cannot open a file, run a command, or remember what it did an hour ago. The agent harness supplies all of that.
What an agent harness contains
The Claude Agent SDK docs list the parts in a way that applies to most agent harnesses:
| Part | What it does |
|---|---|
| Agent loop | Sends the task to the model, runs the tools it asks for, and repeats |
| Built-in tools | Read, write, and edit files, run commands, and search the web |
| Context management | Decides what the model sees as the conversation grows |
| Permissions | Controls which tools run automatically and which need approval |
| Sessions | Keeps context across exchanges so work can resume later |
| Hooks and MCP | Run custom code at key points and connect external tools |
How an agent harness runs a task in 5 steps
The heart of the harness is the loop. The agent loop documentation describes it in 5 steps:
- Receive prompt. The model gets the task, the system prompt, and the tool definitions.
- Evaluate and respond. The model replies with text, tool call requests, or both.
- Execute tools. The harness runs each tool and collects the results.
- Repeat. Steps 2 and 3 cycle until the model responds with no tool calls.
- Return result. The harness returns the final text, token usage, and cost.
Notice who makes the decisions. The harness never chooses the next step. The model does, and the harness carries it out.
An agent harness in code
This script uses the Claude Agent SDK to point an agent at the failing checkout test from earlier.
npm install @anthropic-ai/claude-agent-sdk
npm install --save-dev tsx
import { query } from '@anthropic-ai/claude-agent-sdk';\
for await (const message of query({
prompt:
'tests/checkout.spec.ts is failing. Find the cause, fix the app code, ' +
'and run npx playwright test to confirm the fix.',
options: {
allowedTools: ['Read', 'Edit', 'Glob', 'Grep', 'Bash'], // auto-approve these tools
permissionMode: 'acceptEdits', // auto-approve file edits
},
})) {
if (message.type === 'result') {
console.log(`Done: ${message.subtype}`);
}
}
Run it with npx tsx agent.ts. The script uses top-level await, so package.json needs "type": "module".
There are about 15 lines here, and none of them say how to fix the bug. The harness provides the loop and the tools, and the model works out the rest.
Agent harnesses you already use
You rarely build an agent harness from nothing. Most teams adopt one:
- Claude Code. Its docs describe it as the layer around the model that provides the tools and manages the context the model sees, and call that layer an agentic harness.
- Claude Agent SDK. Anthropic calls it a powerful, general-purpose agent harness.
- Other coding agents. The arXiv paper applies its definition to Codex CLI, Aider, Cline, OpenHands, and SWE-agent as well.
What you do build is the outer layer: instruction files, extra tools, and skills. Playwright MCP gives an agent a real browser, and the open-source Playwright Skill gives it tested guidance on locators, fixtures, and CI.
A full setup walkthrough for one of these harnesses lives in the guide to Claude Code with Playwright.
Note: An agent harness has no pass or fail. It makes a model able to act, and nothing in it checks whether the action was right. That gap is the core of the test harness vs agent harness difference: one verifies, the other does.
Once an agent can act, the next question is how often it acts correctly. Answering that takes the third harness.
What is an eval harness?
An eval is a test for an AI system. You give the system an input, apply grading logic to what it produces, and record the result.
Running 1 eval by hand is easy. Running hundreds of them, several times each, with every step recorded, needs machinery. That machinery is the eval harness.
An eval harness, or evaluation harness, is in Anthropic's words "the infrastructure that runs evals end-to-end. It provides instructions and tools, runs tasks concurrently, records all the steps, grades outputs, and aggregates results."
How an evaluation harness for AI agents is built
Anthropic's guide uses a small vocabulary that most eval tools share:
| Term | Meaning |
|---|---|
| Task | A single test with defined inputs and success criteria |
| Trial | One attempt at a task. Outputs vary, so each task gets several trials |
| Grader | Logic that scores some aspect of the agent's performance |
| Transcript | The complete record of a trial, including every tool call |
| Outcome | The final state in the environment when the trial ends |
Graders come in 3 kinds, and each has a clear trade-off:
- Code-based graders such as unit tests are fast, cheap, and reproducible, but brittle to valid variations.
- Model-based graders use a second model with a rubric. They capture nuance but are non-deterministic and need calibration.
- Human graders are the gold standard for quality, and they are slow and expensive.
Choosing what to score is its own topic, covered in the guide to AI agent evaluation metrics.
Eval harnesses you may have seen
The term is older than agents. Several well-known projects are eval harnesses:
- lm-evaluation-harness. EleutherAI's project is a unified framework to test generative language models, with over 60 standard academic benchmarks.
- Inspect. An open-source framework from the UK AI Security Institute, built from a dataset, a solver that produces answers, and a scorer.
- The SWE-bench harness. It scores coding agents by applying their patches to real repositories and running each repository's tests.
The SWE-bench harness runs from a single command, shown in the SWE-bench evaluation guide:
python -m swebench.harness.run_evaluation \
--dataset_name princeton-nlp/SWE-bench_Lite \
--predictions_path <path_to_predictions> \
--max_workers 8 \
--run_id my_evaluation_run
Most teams watch their agents, fewer evaluate them
The State of Agent Engineering report from LangChain surveyed 1,340 people in late 2025. It found that 57% of respondents already have agents in production.
What teams have built around their agents is uneven. Across all respondents, 89% have some form of observability, while 52.4% run offline evals on test sets and 29.5% run no evals at all.

In plain words, most teams can see what their agent did, and far fewer can say how often it does the right thing. The same report names quality as the top barrier to production, cited by about one third of respondents.
With all 3 definitions in place, the differences between them become easier to see.
Test harness vs agent harness: 5 differences that matter
The test harness vs agent harness comparison matters most in daily work, because these are the 2 that QA teams touch first. Each point below also shows where the eval harness lands.
The matrix sums up the comparison before the detail.

1. What sits inside the harness
A test harness wraps code that you want to check. The code is the subject, and the harness is the trusted part.
An agent harness wraps a model, and the model is not being checked at all. It is the worker. Only the eval harness treats an AI system as the subject, and it tests the model and its agent harness as one unit.
2. Who decides the next step
In a test harness, the script decides. Every click and every assertion is written before the run starts.
In an agent harness, the model decides. Claude Code's docs put it this way: the model "decides what each step requires based on what it learned from the previous step."
An eval harness sits between the 2. Its task list is fixed like a test suite, but inside each task the agent is free to choose its own path.
3. Whether the same input gives the same result
A test harness aims for the same verdict every time. Playwright assertions support that goal by retrying until a condition is met, which takes timing guesswork out of the result.
An agent harness cannot promise that. Give the same model the same task twice and it may take 2 different routes.
Eval harnesses are designed around this fact. Anthropic's guide explains: "Because model outputs vary between runs, we run multiple trials to produce more consistent results."
4. What counts as a result
- Test harness: a pass or a fail for each test.
- Agent harness: a finished task, such as edited files and a final message. Nothing in it says whether the work is correct.
- Eval harness: a score, such as the share of trials that passed.
That score needs care. Anthropic gives an example: an agent with a 75% success rate per trial passes all of 3 trials only about 42% of the time. One good run tells you little.
5. What a failure means
A failed test means the code is wrong or the test is wrong. Either way, a person fixes it and the run goes green.
A failed agent run means the model took a wrong path or ran out of turns. The fix usually goes into the harness: a clearer instruction, a better tool, or a tighter limit.
A low eval score is not a bug. It is a measurement, and it tells you whether your last change to a prompt, a tool, or a model helped or hurt.
These differences can make the 3 look like rivals. In practice they stack.
How the 3 harnesses fit together in one pipeline
The 3 harnesses are not alternatives. In a working setup they nest inside each other:
- The eval harness is the outer layer. It picks a task and starts a trial.
- The agent harness runs inside each trial. It lets the model work on the task.
- The test harness is the grader. It decides whether the agent's work passed.
Anthropic states the first relationship directly: "When we evaluate 'an agent,' we're evaluating the harness and the model working together."
SWE-bench shows the second one in public. Its harness applies an agent's patch to a real repository, then runs that repository's own tests to see if the issue is resolved.
So a test harness, written by people years before agents existed, becomes the grader inside an eval harness. Anthropic notes that SWE-bench Verified and Terminal-Bench both grade coding agents this way: does the code run, and do the tests pass?

A small eval harness with Playwright as the grader
You can see all 3 layers in about 60 lines. This script gives an agent a bug to fix, repeats the attempt 3 times, and lets a Playwright test decide each result.
import { execSync } from 'node:child_process';
import { query } from '@anthropic-ai/claude-agent-sdk';
type Task = { id: string; prompt: string; graderSpec: string };
const TRIALS = 3;
const tasks: Task[] = [
{
id: 'checkout-discount',
prompt: 'The cart total ignores discount codes. Fix the bug in src/cart.ts.',
graderSpec: 'tests/cart-discount.spec.ts',
},
];
// Agent harness: run the model loop until it stops
async function runAgent(prompt: string): Promise<void> {
try {
for await (const message of query({
prompt,
options: {
allowedTools: ['Read', 'Edit', 'Glob', 'Grep', 'Bash'],
permissionMode: 'acceptEdits',
maxTurns: 30,
},
})) {
if (message.type === 'result') {
console.log(` agent stopped: ${message.subtype}`);
}
}
} catch {
// query() throws after an error result, such as hitting maxTurns.
// The trial is still graded on whatever the agent left behind.
}
}
// Test harness: a code-based grader with a hard pass or fail
function grade(specFile: string): boolean {
// Restore the tests so the agent cannot pass by editing them
execSync('git checkout HEAD -- tests playwright.config.ts');
try {
execSync(`npx playwright test ${specFile}`, { stdio: 'ignore' });
return true;
} catch {
return false;
}
}
// Eval harness: tasks x trials, isolation, grading, aggregation
async function main() {
for (const task of tasks) {
let passed = 0;
for (let trial = 1; trial <= TRIALS; trial++) {
// Isolate the trial: reset tracked files, remove untracked ones
execSync('git reset --hard HEAD && git clean -fd', { stdio: 'ignore' });
await runAgent(task.prompt);
if (grade(task.graderSpec)) passed++;
}
console.log(`${task.id}: ${passed}/${TRIALS} trials passed`);
}
}
main();
npx tsx evals/run-evals.ts
Each function is one harness:
- runAgent() is the agent harness. It runs the model loop and stops at 30 turns.
- grade() is the test harness. It restores the test files first, so the agent cannot pass by editing the test.
- main() is the eval harness. It resets the repo before every trial and reports a pass rate.
Note: Run this only in a throwaway clone or a CI container with everything committed, including the script itself. The git reset --hard and git clean -fd commands delete all uncommitted work in the folder.
The reset step is not optional. Anthropic's guide says each trial "should be 'isolated' by starting from a clean environment," so one attempt cannot leak into the next.
The grader also follows a rule from the same guide: grade "what the agent produced, not the path it took." The Playwright test does not care how the agent fixed the cart. It only checks that the discount now applies.
This is the same principle behind every sound way to test AI-generated code: the check lives outside the model. A real setup would run the script on a schedule, and the workflow for Playwright in GitHub Actions is a good base for that job.
Once the 3 harnesses are nested, a new problem appears. The number at the end depends on all of them.
Why the harness changes the score
An eval score looks like a fact about a model. It is really a fact about a model, an agent harness, and an eval harness measured together. Change any one of them and the number moves.
The eval setup itself moves results
Anthropic tested this in February 2026. It ran the Terminal-Bench 2.0 benchmark under 6 resource settings, from strict limits on each task to no limits at all.
The infrastructure noise study reported that "the gap between the most- and least-resourced setups on Terminal-Bench 2.0 was 6 percentage points." Nothing about the model had changed.
Part of that gap was plain infrastructure failure. Some trials ended in infrastructure errors instead of a real attempt, and that error rate fell as limits loosened.

Anthropic's conclusion is a useful rule for reading any leaderboard: "Two agents with different resource budgets and time limits aren't taking the same test." It advises treating gaps below 3 percentage points with skepticism until the eval configuration is documented and matched.
A flaky test harness is a noisy grader
The same logic applies one layer down. If your test harness is the grader, every weakness in it flows straight into the eval score.
Picture a test that fails at random 1 time in 10. The agent fixes the bug correctly, the grader fails anyway, and the trial is marked as a miss.
The agent's real ability did not change. The score dropped because the grader was unstable. Fixing Playwright flaky tests is therefore the first step toward a trustworthy eval, not a side task.
Where TestDino fits
TestDino works at the test harness layer for Playwright teams. It records every CI run with its traces, screenshots, and logs, and it tracks which tests are flaky across your run history.
Its AI failure analysis sorts each failure into Actual Bug, UI Change, Unstable Test, or Miscellaneous. That turns manual test failure analysis into a label you can read at a glance.
For an eval grader, the Unstable Test label is the one that matters. It tells you which results to distrust before they distort a score.
Agents can read the same data. The TestDino MCP server lets tools such as Claude Code and Cursor query test results and debug failures, alongside what the Playwright trace viewer shows locally.
Knowing how the layers depend on each other also settles the order in which to build them.
Which harness should you build first?
Build from the inside out. Each harness depends on the one beneath it, so the order is fixed:
- Test harness first. Without a stable suite, nothing can verify an agent's work.
- Agent harness second. Adopt an existing one and point it at the suite.
- Eval harness third. Add it when you need to compare prompts, tools, or models with numbers.
The table matches common situations to a starting point.
| Your situation | Start with | Why |
|---|---|---|
| Tests are missing or unreliable | Test harness | Every other layer uses it as the source of truth |
| Tests are stable and you want an agent to write or fix code | Agent harness | The agent can run the suite to check its own work |
| You are changing a model, prompt, or tool and want proof it helps | Eval harness | Only repeated trials show a real difference |
| You are comparing public benchmark scores | The benchmark's eval settings | Scores from different setups are not comparable |
If you sit in row 2, Playwright's own planner, generator, and healer make a practical next step. The guide to Playwright test agents explains how they run on top of your existing suite.
Harness vs framework vs benchmark
Three nearby terms cause most of the remaining mix-ups:
- Framework. A framework is what you install. A harness is what you assemble with it. Playwright Test is a framework, and your fixtures, mocks, config, and CI job are the test harness.
- Benchmark. A benchmark is a set of tasks. The eval harness is the program that runs them. SWE-bench is the benchmark, and swebench.harness.run_evaluation is its harness.
- Evaluator agent. A second model that judges output is a model-based grader. It is one part of an eval harness, not the whole thing.
Eval harness vs agent harness in one line
An agent harness runs during the task and helps the model act. An eval harness runs around the task and measures how the action went.
They also deserve different handling. An agent harness is replaceable, and a comparison of agentic testing tools shows how many options exist.
An eval harness should stay fixed between runs. If the measuring setup changes, 2 scores can no longer be compared.
Tip: You do not need hundreds of eval tasks to begin. Anthropic's guide says "20-50 simple tasks drawn from real failures is a great start." Your bug tracker and your failed CI runs already hold that list.
Conclusion
The test harness vs agent harness question has a short answer. A test harness checks code, an agent harness lets a model act, and an eval harness measures how well that agent performs.
They are 3 layers of one system, not 3 competing ideas. The eval harness runs the agent harness, and the test harness grades the result.
For QA teams, the practical takeaway is that the oldest layer carries the most weight. A stable suite built on Playwright best practices is what makes an agent's work checkable and an eval score believable.
Start there. Then add an agent, and add evals when you need numbers to choose between 2 setups.
FAQs

Ayush Mania
Forward Development Engineer
