The task was simple: test two web apps before deployment. A Next.js portfolio and a SaaS chat — accessibility, console errors, mobile responsiveness. Routine stuff.

Channel with guides and Claude Code content — we share news (when they slash limits 10x) and what tools we build with Claude for projects, channel: https://t.me/claudedevolper

I opened Claude Code, connected Playwright MCP, wrote "test the app." The agent went to work, taking screenshots, checking elements. On the 51st snapshot, /compact kicked in. The text context was only 18% full. I didn't understand what happened.

After an hour of debugging, I found an invisible image limit. After three hours, I realized Playwright MCP burns 50× more tokens than CLI on the same workflow. After three days, I had a working tool that real users were already testing.

This article is about the journey from "I just want to test" to an open-source tool, and the architectural problems that forced me to build it.

Problem One: MCP Burns Tokens Like Crazy

In November 2025, Pramod Dutta published an analysis that spread across the entire AI-testing community: Playwright MCP burns ~114k tokens per test. Özal's benchmark on Microsoft's GitHub shows that the verify workflow for e-commerce hits ~1.5M tokens via MCP. Playwright CLI? Same ~27k.

A 50–60× gap. The cause is architectural: MCP keeps the LLM in a browser loop on every action — navigation, click, wait, screenshot, analyze, repeat. Perfect for exploring unfamiliar interfaces. Catastrophically expensive for rerunning the same scenario.

Microsoft updated the recommendation in its own README: for coding agents — CLI + Skills, not MCP. The official Test Agents documentation now proposes the triplet Planner / Generator / Healer as the primary architecture — not "agent sits in MCP for the whole session."

Conclusion: the fix is not "use less Playwright MCP." The fix is to split exploration and reproduction into separate phases.

Problem Two: The Invisible Image Limit

While debugging the token issue, I found something worse. Claude Code has a second context limit — a budget for inline image blocks. Roughly 50–100 blocks per session. No counter. No warning.

Each Playwright:browser_take_screenshot returns one image block to the context. 50 screenshots — and you've used 0.4% of your text budget and 100% of your image budget. /compact triggers with the text context 80% empty. The agent loses everything not saved to disk.

"Just fewer screenshots" — discipline doesn't survive 30+ turns in a real exploratory session. No counter — you only see /compact.

A soft rule in CLAUDE.md — "never take screenshots, use ARIA." Works for 30 turns. When the agent gets stuck on a modal, it reaches for a screenshot and immediately rationalizes breaking the rule.

context: fork in the skill's frontmatter — officially documented fix. Doesn't parse on Claude Code 2.1.x on Windows. The skill simply doesn't appear in the list. 90 minutes of debugging — then I gave up.

What works: subagents (Task tool) have an isolated image budget. Everything the subagent reads doesn't count against the parent context. Verified empirically: subagent read 6 PNGs, returned 6 text descriptions — the parent chat's counter didn't budge.

Day Three: A Working Tool

By the end of day three, I had a skill for Claude Code built around one architectural invariant: the parent chat never receives an image.

Four patterns ensure this:

Pattern A (90% of the work): exploration via ARIA tree. browser_snapshot returns the accessibility tree as text — same locator data as a screenshot, but in text form. Image budget cost: zero.

Pattern B (3–5 times per run): when vision is actually needed — pixel-diff fired, visual layout check required — the subagent reads ONE image and returns ONE text string. The subagent burns its budget, the parent chat stays clean.

Pattern C: built-in toHaveScreenshot() returns diff% as JSON via npx playwright test. Text all the way through the pipeline. Vision tokens only burn if diff actually fired — and even then via Pattern B.

Pattern D: screenshots on disk (Playwright artifacts, MCP cache) cost zero until explicitly read. File on disk ≠ file in context.

Workflow: first run → subagent via Playwright MCP walks the app through ARIA snapshots, generates *.spec.ts. Each subsequent run → npx playwright test directly — deterministic, ~zero token spend. On top of that — bug fingerprinting with SHA-256 keys and classification across runs: new / regression / persisting / fixed.

Pattern: The Issues Collector

Each generated spec collects all soft checks into a single array:

test('home page baseline', async ({ page }) => {
  const consoleErrors: string[] = [];
  const failedRequests: string[] = [];
  const issues: string[] = [];

  page.on('pageerror', (e) => consoleErrors.push(`pageerror: ${e.message}`));
  page.on('console', (m) => m.type() === 'error' && consoleErrors.push(`console: ${m.text()}`));
  page.on('response', (r) => r.status() >= 400 && failedRequests.push(`${r.status()} ${r.url()}`));

  await page.goto('/');

  const a11y = await new AxeBuilder({ page })
    .withTags(['wcag2a','wcag2aa','wcag21aa','wcag22aa']).analyze();
  a11y.violations.forEach((v) =>
    issues.push(`a11y[${v.impact}] ${v.id}: ${v.help} (${v.nodes.length}x nodes)`));

  // overflow, heading hierarchy, touch targets, html-lang — всё в issues[]

  expect(issues, `${issues.length} issues found:\n  - ${issues.join('\n  - ')}`).toEqual([]);
  expect(consoleErrors).toEqual([]);
  expect(failedRequests).toEqual([]);
});

When the test fails — you get all the issues in one message, not just the first one. Post-processing parses the result and generates one bug record per issue with a stable fingerprint for comparison across runs.

What's tested out of the box

Every spec includes: console error listeners (attached before page.goto(), with noise filtering for GTM/Stripe/Sentry/Next.js/Supabase/ResizeObserver), axe-core WCAG audit (tags wcag2a through wcag22aa), heading hierarchy (jumps h1 → h3), touch-target sizes (WCAG 2.5.8 AA = 24×24 CSS px), horizontal overflow, presence of html lang. Visual regression through built-in toHaveScreenshot() with no external dependencies.

Severity is assigned automatically from axe impact and error class, with three override mechanisms: [severity:S0] inline in the collector, in the test name, or // @severity: S0 comment before test().

An honest market picture

Octomind published a farewell letter on April 30, 2026. Paid AI-testing services still around — QA Wolf (typical contracts $60–250k/year), Mabl, BrowserStack AI — sell real value: cloud parallelism, human review, SOC 2, SLA.

My tool doesn't compete with that. No managed cloud, no human review, no certification. For solo devs and small teams already on Claude Code with zero QA budget — this is a working replacement at $0/month. For a 50-person team with cross-browser nightly regression — no, and pretending otherwise wouldn't be honest.

An honest comparison group — free OSS-tier: native playwright init-agents --loop=claude from Microsoft (triplet Planner/Generator/Healer) and Magnitude. Differences in my tool: built-in axe-core + console + network audit, bug fingerprinting with classification between runs, mapping to Linear/GitHub/Jira trackers. None of the free alternatives do this.

What I intentionally didn't do

No self-healing. The QA community has spent the last year criticizing self-healing as marketing — a documented failure mode: the healer picks a similar-but-wrong element, the test goes green, the bug ships to production. The tool prefers a red test to a false green.

No cloud. Tests stay in the repo. Reports stay in the file system. If the npm package disappears tomorrow — the suite keeps working.

No promises that "AI will write all tests." This is an addition to engineering judgment, not a replacement. It excels on the boring 80%: a11y, console, network, responsive, regression diffs.

Results on real applications

Both applications are public on GitHub — these aren't synthetic benchmarks.

Static Next.js portfolio, mobile viewport. Found 4 real bugs, 0 false positives: axe-core color-contrast — 8 elements fail WCAG 1.4.3 AA (S1), two touch-targets under 24×24 px (S2), heading jump h1→h3 on the projects page (S2). Image budget in the parent chat: zero.

Voice-first AI SaaS chat (Next.js + FastAPI + Supabase + WebSocket). 11 specs for login, chat, translate, TTS, settings, phrase library, scenario mode, stats, logout. 10 out of 10 passed after 4 iterations, ~12 minutes from setup to first green assertion. A test run uncovered 6 issues that became fixes in version 0.2.0.

Installation

npx webtest-orch@beta install

Create .env.test with TEST_BASE_URL in your project, restart Claude Code, say "test the application". The tool will automatically detect authenticated vs public site, build Playwright + axe-core, run the first exploratory pass, write the report.

Version 0.3.1-beta, 113 tests, CI on Linux/macOS/Windows. MIT.

Channel with guides and content on claude code, we post news (when limits get cut by 10x) and what tools we build through claude for projects, channel: https://t.me/claudedevolper