October 6, 2026

Buy Me a Hat, Jev!

Can we make ecommerce browser tests less brittle? A fixed Playwright script, a Gemini browser agent and TypeSafe's Jev each buy the same hat, then the shop changes five times.
researchqaai

Software is changing faster. AI lets teams produce more code, but review and QA capacity do not automatically grow with it. That makes QA automation more important.

Writing the first version of a browser test is often not the most expensive part. The cost comes later: maintaining it, investigating failed runs, and keeping it useful while the product changes around it.

Natural-language browser agents offer an appealing alternative. Describe what a user needs to achieve, then let an agent work out how to do it. But do they reduce maintenance, or trade it for model latency, cost, and failures?

Why Jev?

At the start of this experiment, Jev had been public for only a couple of weeks. It returns typed decisions rather than free-form text or plans. TypeSafe presents that narrower design as faster and cheaper than a conventional LLM, which makes browser automation an interesting test case.

AI is appearing in several kinds of browser tooling. Playwright's planner, generator, and healer create or repair tests. Playwright MCP and Browser Use let models control a browser at runtime. The question was how those trade-offs look when the journey is a test that must run repeatedly and prove an exact result.

A fair comparison

Routine copy and layout changes should not break a well-written Playwright test. Stable test IDs should cover those changes when the IDs remain intact. The more interesting cases are genuine changes to the journey and blockers introduced outside the application's normal release path, such as tag-manager experiments and consent dialogs.

Browser testing also has a harder requirement than browser automation. Reaching a page that looks finished is not enough. A checkout test must prove that the right product, variant, price, and order were recorded.

All three approaches ran the same task in the same Chromium browser and were judged by the same independent checks.

The experiment used a fictional hat shop and tested three ways to buy the same hat:

  1. deterministic Playwright using stable test IDs
  2. Gemini 3.1 Flash Lite controlling Playwright through generic browser tools
  3. Jev choosing bounded actions for Playwright

Each approach then ran against five changes:

  • a redesigned product layout
  • different bag and checkout copy
  • a simulated tag-manager popup
  • a new consent dialog
  • a quick-checkout option with the wrong size preselected

The shop

LABEL PENDING is a local storefront built for this study. It has six illustrated hats, a catalog, product pages, a bag, checkout, and order confirmation.

The LABEL PENDING homepage. A large headline reads 'Hats for off-hours thinking', with a pink circle on a lilac panel to its right and a 'Current headwear' heading below.
LABEL PENDING, the fictional hat shop built for this study.

A purpose-built shop made it possible to apply one controlled change at a time, including simulated tag-manager and consent blockers.

Three ways to buy one hat

Deterministic Playwright

The deterministic test knows the product and interface contracts:

await page.getByTestId("view-product-afterimage-cap").click();
await page.getByTestId("product-size").selectOption("ml");
await page.getByTestId("add-to-bag").click();
await page.getByTestId("checkout").click();

The rest of the function fills three checkout fields and places the order. It is explicit, quick, and easy to debug.

A custom Gemini browser agent

The runner received one human task. On every model call, Gemini saw that task, the current DOM observation, and recent action results. It returned one structured JSON action:

click(ref)
fill(ref, suppliedDataName)
select(ref, optionValue)
scroll_up()
scroll_down()
wait()
finish()
blocked()

The scenario did not tell it which product link or button to use:

await runLlmTask({ page, scenario, agent });

A short task description does not remove the automation code. The runner observes the page, gives elements stable references, validates each action, rejects stale targets, substitutes named test data locally, limits steps, and records a trace.

Jev-directed Playwright

Jev used the same scenario and Playwright harness:

await runScenario({ page, scenario, client, model });

Jev is not a browser driver. Playwright still observes and operates the browser. Jev receives typed questions and chooses from candidates supplied by the runner.

One Jev call per decision step

Choosing an operation in one call and its target in another would double the number of model requests. Instead, the runner sends all useful questions for the current page in one request:

  • what action to take next
  • whether the task appears complete
  • whether progress is still possible
  • which target to use for each available action type when there is more than one candidate

Jev answers these questions in the same call. If it chooses CLICK, the runner uses the click target and ignores the fill and select answers. This avoids a second call just to ask which element to click. TypeSafe recommends grouping questions this way because it evaluates them in parallel.

After Playwright executes the action, the runner observes the page again and makes the next decision call. It does not ask for a complete checkout plan because later controls may not exist yet, and a popup or failed navigation could make that plan stale.

Across the final study, Jev calls averaged 323ms. The median was 307ms and P95 was 457ms.

How a run passed

The agents chose browser actions, but they did not decide whether they had succeeded. The benchmark used independent checks to determine the result.

A run passed only when all five checks were true:

  • the completed order contained the Afterimage Cap with SKU LP-AFT-CBL
  • the selected size was M/L
  • the recorded price was £38
  • the browser had reached the confirmation page
  • the IT'S YOURS. confirmation heading was visible

If Jev chose DONE, or Gemini called finish, while those checks still failed, the runner rejected the stopping signal. Three premature attempts ended the run as a failure. Reaching confirmation with the wrong product or size did not count as a pass.

Baseline speed

Each approach ran five times on the unchanged shop.

The same purchase three ways. Playwright left, Jev centre, Gemini right.
Baseline results on the unchanged shop, five runs per approach
MeasureFixed scriptJevGemini
Verified passes✅ 5/5✅ 5/5✅ 5/5
Median duration279 ms3,729 ms11,049 ms
Duration range229 to 411 ms3,696 to 4,099 ms10,431 to 11,498 ms
Model calls per run099
Median model time per run02,946 ms10,223 ms
Model API cost per run (USD)$0about $0.00120$0.003919
Model API cost per 100 runs$0about $0.12$0.39

At the median, Jev was about three times faster than Gemini but about thirteen times slower than fixed Playwright. The baseline journey required nine sequential model decisions, so even short calls accumulated.

What happened when the shop changed?

Each change ran five times per approach.

Results for each change to the shop, five runs per approach
ChangeFixed scriptJevGeminiWhat happened
Redesigned product layout✅ 5/5✅ 5/5✅ 5/5The fixed script's test IDs survived the layout change
Different bag and checkout copy✅ 5/5✅ 5/5✅ 5/5The fixed script did not depend on the visible wording
Simulated GTM popup❌ 0/5✅ 5/5✅ 5/5Both agents dismissed the unexpected blocker
New consent dialog❌ 0/5✅ 5/5✅ 5/5Both agents found a valid privacy choice
Wrong-size quick checkout✅ 5/5❌ 0/5✅ 5/5Jev stopped; Gemini found the longer correct route

Five runs show how each controller behaved in this experiment. They do not establish production reliability.

Three things stand out.

Stable selectors handled these layout and copy changes

The fixed script passed every redesign and copy run. In these cases, using a model would add latency and cost without making the test more robust.

Unexpected overlays are the clearest case for adaptive navigation

The unmodified fixed script timed out behind both dialogs. Jev and Gemini dismissed them without a scenario-specific handler. Jev's median was about 3.9 seconds in both cases, compared with roughly 12 seconds for Gemini. If a dialog becomes a permanent part of the site, adding an explicit Playwright handler would still be the better fix.

The shortcut exposed Jev's limit

Quick checkout offered the right hat with S/M preselected, while the task required M/L. The fixed script followed its known route and Gemini scrolled past the shortcut to select the right size. Jev returned BLOCKED in all five runs. It avoided buying the wrong size, but it did not find the longer route.

Is this less code?

There is less code in each test, but not less code overall.

A fixed script spells out each browser action. A model-directed scenario can describe the goal and name the test data, but it still needs exact pass criteria.

The shorter scenario relies on a shared controller that must:

  • describe the current page to the model
  • validate and safely execute its decisions
  • protect test data and stop runaway runs
  • verify the result and record enough detail to debug failures

That controller can be reused across many journeys, but it centralises the complexity rather than removing it.

Where each approach fits

Use a fixed Playwright script as the default for known, critical commerce paths. In this study it was the fastest approach and kept the expected route explicit.

Use the Gemini agent tested here when the route is open-ended and may require broader reasoning. It handled the wrong-size shortcut that stopped Jev, but it was also the slowest and most expensive approach in this experiment.

Use Jev when navigation is uncertain but each next action can be expressed as a bounded choice. It handled both runtime blockers while running about three times faster than Gemini, but it failed on the ambiguous shortcut.

The next step is to put Jev into journeys that already cost time to maintain, such as tests affected by consent banners, tag-manager experiments, regional UI, or changing navigation. The final assertions should remain deterministic. The open question is whether Jev recovers enough runs to justify the latency, cost, and new failure modes it introduces.


If your browser tests keep tripping over banners and popups and you want help making them hold, speak to us at [email protected].

References

  1. TypeSafeIntroducing System One Models and Jev
  2. PlaywrightTest agents: planner, generator, and healer
  3. PlaywrightMCP browser automation
  4. Browser UseAgent documentation
  5. TypeSafeAsking multiple questions together

Get In Touch

Set up a discovery call to discuss your project.

Mailing List

About

We are based in New York, with remote team members in San Francisco & the United Kingdom.

All Gold Commerce
© 2010 – 2026