Harness Engineering: What It Is and How to Build a QA Harness
AI agents write code fast but fail quietly. Learn what harness engineering is and how to build a QA harness.
A small team at OpenAI shipped a product with about 1 million lines of code, and no person typed any of it. The method behind it is called harness engineering: building the rules, tools, and checks around an AI so its work can be trusted.
Most teams feel the other side of that story. AI writes code in minutes, but the result is often almost right, and few teams know how to test AI-generated code at that speed.
This guide explains what a harness is, what a QA harness and a test harness mean today, and how to build one in 6 steps. Every example uses Playwright test automation, so you can copy the setup into a real project.
What is harness engineering?
Harness engineering is the practice of designing everything around an AI model, including its instructions, tools, test environment, and automated checks, so that an AI agent produces work you can verify and trust. In short: Agent = Model + Harness.
The idea is simple. An AI model can write code, but it cannot see your app, run your tests, or remember yesterday's mistakes unless something gives it those abilities. That something is the harness.
The word comes from the gear that lets a rider direct a horse's strength. The horse supplies the power. The harness decides where that power goes.
Where the term came from
Mitchell Hashimoto, co-founder of HashiCorp, named the practice in a post published on February 5, 2026. He described it as "the idea that anytime you find an agent makes a mistake, you take the time to engineer a solution such that the agent never makes that mistake again."
Six days later, OpenAI published its own account of an internal product built this way. The team reported:
- About 1 million lines of code, with every line written by Codex agents
- Roughly 1,500 pull requests opened and merged in 5 months
- A starting team of 3 engineers, averaging 3.5 pull requests per engineer per day
OpenAI summed up the working model in 4 words: "Humans steer. Agents execute."
Agent = model + harness in plain words
The shortest way to remember the idea is a formula: Agent = Model + Harness. In an article on martinfowler.com, Birgitta Böckeler describes the harness as "everything in an AI agent except the model itself."
That covers a lot of ground:
- The instruction files the agent reads
- The tools it can call, such as a terminal or a browser
- The sandbox it works in
- The tests and checks that judge its output
That list looks like plumbing. The numbers in the next section show it changes results more than most teams expect.
Why the harness matters more than the model
Most AI discussions focus on which model is best. The evidence from teams running AI coding agents points somewhere else: the layer around the model.
The same model scores higher with a better harness
LangChain ran a clean experiment in February 2026. It kept the model fixed at gpt-5.2-codex and changed only the harness around its coding agent.
The agent's score on the Terminal Bench 2.0 benchmark rose from 52.8% to 66.5%. That 13.7-point gain moved it from the Top 30 to the Top 5 on the leaderboard.

The changes were not exotic. LangChain added a self-verification loop, a checklist that runs before the agent finishes, and a detector for repeated failed edits.
All 3 are testing habits applied to an agent. None of them required a smarter model.
Note: This is 1 team's result on 1 benchmark, so treat it as a direction and not a promise. Your own gain depends on your codebase and on how good your checks already are.
Developers use AI but do not trust it yet
The 2025 Stack Overflow Developer Survey shows the gap clearly:
- AI tools are used or planned by 84% of respondents.
- Accuracy is distrusted by 46% of developers, while 33% trust it.
- The top frustration, named by 66%, is "AI solutions that are almost right, but not quite."
Almost right is the dangerous kind of wrong. The code compiles, the demo works, and the bug shows up later in production.
Human QA becomes the bottleneck
When agents write faster, people cannot review faster. OpenAI said so directly in its report: "our bottleneck became human QA capacity."
Agents also grade themselves too kindly. Anthropic found that agents asked to judge their own output tend to respond by confidently praising the work, even when the quality is mediocre.
So the fix cannot be more human review or more self-review. It has to be automated checks the agent cannot talk its way past, and testers have been building exactly that for decades.
What is a QA harness?
The word harness is older in testing than it is in AI. Knowing the classic meaning makes the new one much easier to build.
The classic test harness in software testing
The ISTQB glossary defines a test harness as "a collection of drivers and test doubles needed to execute a test suite." In plain words:
- A driver calls the code you want to test and controls the run.
- A test double stands in for a real dependency, such as a payment API. Stubs and mocks are both test doubles.
Playwright ships most of a test harness out of the box. The test runner is the driver, and the Playwright docs say fixtures "establish the environment for each test, giving the test everything it needs and nothing else."
QA harness: the quality layer around an agent
A QA harness is the quality layer of an agent harness: the test suite, quality gates, and run history that check an AI agent's output before it merges. It is the part of the harness that QA teams own.
The term is informal. You will not find it in the ISTQB glossary, and some teams use it as another name for a classic test harness.
In AI-assisted teams, it increasingly means the checks that stand between an agent's pull request and the main branch. The agent can write anything it likes, but nothing merges until the QA harness says yes.

What test harness engineering means now
Test harness engineering used to mean wiring up drivers and mocks so people could run tests. Today it means designing that same setup so an agent can run it, read the result, and fix its own work.
The classic parts map over cleanly:
| Test harness part | Playwright equivalent | What an agent gets from it |
|---|---|---|
| Driver | Test runner, npx playwright test | One command to verify any change |
| Test double | page.route() mocks | Stable, repeatable dependencies |
| Test environment | Fixtures and webServer | A clean state on every run |
| Result reporting | Reporters and traces | Errors it can read and act on |
The difference is the reader. A person can interpret a vague failure. An agent needs the failing step, the expected value, and the actual value in text it can parse.
With both meanings clear, each part of the full agent harness is easier to place.
The 6 building blocks of an agent harness
Böckeler's article splits harness controls into 2 groups, and the split is useful for testers:
- Guides steer the agent before it acts. They are feedforward controls.
- Sensors observe after the agent acts and help it correct itself. They are feedback controls.
Guides reduce the number of mistakes. Sensors catch the mistakes that still get through. A harness needs both, because neither works alone.

Guides: what the agent knows before it starts
- Instructions. OpenAI keeps its AGENTS.md at roughly 100 lines and treats it as a table of contents that points to deeper docs.
- Tools. Scripts, CLIs, and MCP servers extend what the agent can do. Playwright MCP gives it a real browser.
- Sandbox. OpenAI made its app bootable per git worktree, so every change gets its own isolated instance.
- Memory and state. Anthropic's harness for long-running agents uses a feature list, a progress file, and git commits so each new session knows what the last one did.
Sensors: what checks the work afterwards
- Verification. Tests, linters, and type checks are fast and deterministic. OpenAI wrote custom linters whose error messages include instructions for fixing the problem.
- Observability. Logs, traces, and test history show what the agent did. The Playwright trace viewer is a ready-made example.
Anthropic adds a third kind of sensor: a separate evaluator agent. It uses Playwright MCP to click through the running app like a user, then grades the result.
That evaluator is slower than a linter, but it can judge things a linter cannot, such as whether a feature actually works end to end.
These building blocks also explain how the practice differs from the 2 terms it is most often confused with.
Harness engineering vs prompt engineering vs context engineering
The 3 terms describe layers that stack on each other. Each one widens the scope of what you design.
| Prompt engineering | Context engineering | Harness engineering | |
|---|---|---|---|
| What you design | The wording of 1 request | The information the model sees | The whole environment around the model |
| Main question | How should I ask? | What should the model know? | How do I verify the result? |
| Typical artifact | A prompt template | Retrieved docs and memory files | AGENTS.md, tools, tests, CI gates |
| Scope | 1 response | 1 task or session | Every task, over time |
| Enforced by | Nothing | Nothing | Automated checks |
These are layers, not rivals. A good harness still contains good prompts and good context.
What it adds is enforcement. A rule in a prompt is a request the agent may forget. A failing test is a result it has to deal with.
That is also why spec-driven testing pairs well with this approach. The spec guides the agent, and the tests built from it verify the outcome.
The next section turns that enforcement into files you can commit today.
How to build a test harness for AI agents in 6 steps
You do not need a new platform to start. A Playwright project already holds most of the parts. To build a test harness for AI agents:
- Write a short AGENTS.md with commands and rules.
- Make the test environment reproducible with 1 command.
- Wrap setup and test data in fixtures.
- Give the agent a browser through Playwright MCP.
- Gate every change in CI.
- Feed every failure back into the harness.
The order matters. Guides come first, because a sensor is of little use when the agent does not know which command runs it.
The examples below use Playwright with TypeScript and Claude Code. The same shape works with other agents and other test runners.

Step 1: Write a short AGENTS.md
This file is the first thing the agent reads. Keep it to commands, rules, and pointers.
# Agent guide
## Commands
- Install: `npm ci`
- Run all tests: `npx playwright test`
- Run one file: `npx playwright test tests/checkout.spec.ts`
- Lint and types: `npm run lint && npx tsc --noEmit`
## Rules
- Import `test` and `expect` from `tests/fixtures.ts`, never from `@playwright/test`.
- Use `getByRole` or `getByLabel` locators. Do not use XPath.
- Never use `waitForTimeout`. Wait on assertions instead.
- Do not edit or delete an existing assertion to make a test pass.
- A task is done only when lint, types, and tests all pass.
## Where to look
- Test conventions: `docs/testing.md`
- Shared fixtures: `tests/fixtures.ts`
Tip: Add a rule only after the agent makes a real mistake. Hashimoto says each line in his project's file "is based on a bad agent behavior." That habit keeps the file short and every rule earned.
Step 2: Make the test environment reproducible
An agent cannot ask a teammate how to start the app. One command has to boot everything.
import { defineConfig } from '@playwright/test';
export default defineConfig({
testDir: './tests',
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 2 : 0,
reporter: process.env.CI
? [['line'], ['json', { outputFile: 'results.json' }]]
: 'list',
use: {
baseURL: 'http://localhost:3000',
trace: 'on-first-retry',
},
webServer: {
command: 'npm run start:test',
url: 'http://localhost:3000',
reuseExistingServer: !process.env.CI,
},
});
The webServer block starts the app before the tests and stops it after. The JSON reporter gives the agent results it can parse.
If you want a starting point, the config generator in TestDino's free tools produces a ready playwright.config.ts from a few choices.
Step 3: Wrap setup in fixtures
Fixtures are the test harness itself. They hold the setup that every test needs, so the agent reuses it and does not reinvent it.
import { randomUUID } from 'node:crypto';
import { test as base, expect } from '@playwright/test';
type HarnessFixtures = {
seededUser: { email: string; password: string };
};
export const test = base.extend<HarnessFixtures>({
seededUser: async ({ request }, use) => {
const email = `user-${randomUUID()}@example.com`;
const password = 'Test-pass-123';
// Setup: create a fresh user through a test-only endpoint
const created = await request.post('/api/test/users', {
data: { email, password },
});
expect(created.ok()).toBeTruthy();
await use({ email, password });
// Teardown: remove the user so the next run starts clean
await request.delete(`/api/test/users/${encodeURIComponent(email)}`);
},
});
export { expect };
The /api/test/users endpoint is an example. Swap in whatever your app uses to create test data.
Two related patterns help here. Playwright fixtures can also hold logged-in sessions, and network mocking replaces slow third-party APIs with test doubles.
Step 4: Give the agent a browser
Reading code is not the same as using the app. Anthropic reported that giving its agent browser testing tools "dramatically improved performance," because it found bugs that were not obvious from the code alone.
claude mcp add playwright npx @playwright/mcp@latest
npx playwright init-agents --loop=claude
npx skills add testdino-hq/playwright-skill
Each command adds 1 capability:
- The first connects Playwright MCP, so the agent can open pages and read them as structured accessibility snapshots.
- The second installs the Playwright test agents: a planner, a generator, and a healer.
- The third adds the open-source Playwright Skill, a set of 70 guides on locators, fixtures, and CI that the agent loads on demand.
Step 5: Gate every change in CI
Local checks can be skipped. CI checks cannot. This workflow makes lint, types, and tests a condition for merging.
name: harness
on: pull_request
jobs:
verify:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- run: npm ci
- run: npx playwright install --with-deps chromium
- run: npm run lint
- run: npx tsc --noEmit
- run: npx playwright test
Mark the verify job as a required status check on your main branch. Caching, sharding, and reports build on this same file once you run Playwright in GitHub Actions at scale.
Step 6: Feed failures back into the harness
This step is the one Hashimoto's definition is about. Each failure should change the harness, not just the code.
- The agent ran the wrong command: add the right one to AGENTS.md.
- The agent wrote a brittle locator: add a lint rule that rejects it.
- A test fails only sometimes: fix the test before the agent learns to ignore it.
Sorting failures by cause is the slow part. A structured test failure analysis routine tells you which of those 3 fixes a failure needs.
A harness built this way improves every week. It can also go wrong in a few predictable ways.
5 harness engineering mistakes to avoid
Mistake 1: writing one giant instruction file
Long instruction files crowd out the task itself. OpenAI treats AGENTS.md as a table of contents and keeps the detail in a structured docs/ folder that the agent reads only when needed.
Mistake 2: letting the agent grade its own work
Anthropic's finding applies here. A generator that reviews itself tends to approve itself.
Use a separate check instead. A test suite is the cheapest option, and a second agent with a skeptical prompt is the next step up.
Mistake 3: trusting flaky tests as sensors
A sensor that fails at random teaches the agent the wrong lesson. It either "fixes" working code or learns to retry until green.
Clean up flaky tests before you rely on the suite as a gate. In an agent workflow, a flaky test is a broken sensor.
Mistake 4: building more harness than the task needs
[NEEDS RAW FIX line 526 - corrupted currency text, do not publish] A harness has a price. In one Anthropic experiment, a solo agent run took 20 minutes and cost 9,whilethefullharnessranfor6hoursandcost9,whilethefullharnessranfor6hoursandcost200.
Note: Anthropic's advice on harness engineering is to keep testing your own setup: "Every component in a harness encodes an assumption about what the model can't do on its own." Remove parts when a newer model no longer needs them.
Mistake 5: never cleaning up
Agents copy the patterns they find, including the bad ones. OpenAI's team once spent every Friday, 20% of the week, cleaning up what it called "AI slop."
Its fix was to write "golden principles" into the repository and run cleanup tasks on a schedule. The same idea helps you reduce test maintenance in any large suite.
Avoiding these mistakes is easier when you can see the harness in numbers.
How to measure whether your harness works
A harness is working when the agent needs less correction over time. These signals are a practical way to track that:
| Signal | What it tells you | Healthy direction |
|---|---|---|
| First-pass CI rate on agent pull requests | Whether your guides are clear | Up |
| Flaky test rate | Whether your sensors can be trusted | Down |
| Failures by category | Whether failures are real bugs or test problems | Fewer test problems |
| Time from failure to root cause | Whether errors are readable | Down |
| Repeated mistakes | Whether the feedback loop works | Toward 0 |
Start with 2 of them. First-pass CI rate and flaky rate cover both halves of the harness, and both come straight from your CI history.
For a wider view, the full set of test quality metrics applies to agent-written tests just as it does to human-written ones.
Where TestDino fits in the harness
TestDino covers the observability block for Playwright teams. It records each CI run with its traces, screenshots, videos, and logs, and it tracks flaky tests across your full run history.
Its AI failure analysis labels each failure as Actual Bug, UI Change, Unstable Test, or Miscellaneous. That is the "failures by category" signal from the table, without manual sorting.
The TestDino MCP server closes the loop. An agent such as Claude can query failed runs directly, which turns your test history into a sensor the agent can read.
Conclusion
Harness engineering moves the hard work from writing code to shaping the place where code gets written. The model supplies the power, and the harness decides whether that power produces something you can ship.
For QA teams, this is familiar ground. Fixtures, mocks, and CI gates are the classic test harness, and they are now the most valuable part of the agent harness too.
Start small. Write the AGENTS.md, make 1 command run everything, and add a rule each time the agent slips. Teams that keep going end up with a software factory where quality is built into the line, guided by the same Playwright best practices that people follow.
FAQs

Ayush Mania
Forward Development Engineer


