The Gherkin Lurking Inside Playwright

BDD promised tests you could read aloud. Step definitions broke that promise. page.ai.within() revives it inside Playwright, with an evidence boundary that kills the false green.

Cover Image for The Gherkin Lurking Inside Playwright

If you wrote tests in the 2010s, you remember Gherkin.

Given the cart summary
Then it shows 2 items

The promise was irresistible: specs in plain English that actually execute. Product managers could read them. QA could write them. The suite documented itself.

The catch was that nothing executed by itself. Every English line needed a step definition: a regex plus a chunk of glue code that somebody had to write, name correctly, and keep in sync with both the English and the app. Teams ended up maintaining two test suites, the readable one and the real one, and the two drifted. Most teams quietly went back to writing code.

We think the idea was right and the mechanism was wrong. And the place where the idea finally works is inside Playwright.

English that executes

Donobu extends Playwright's page fixture. Change one import and page gains an ai surface:

import { test } from '@donobu/test';

test('cart math holds up', async ({ page }) => {
  await page.goto('/cart');
  await page.ai.assert('The cart shows 2 items and a $34.00 total.');
});

That sentence is the step. No regex, no step-definition file, no glue to drift. Everything else stays Playwright: same runner, same config, same expect.

But whole-page English assertions have a well-known failure mode.

The false green

page.ai.assert('shows 2 items') can pass because something on the page says "2 items". A recommendations widget. A wishlist badge. A footer stat. Just not the cart you meant.

That is the classic false green, and green-for-the-wrong-reason is worse than red, because nobody ever looks at it again.

Notice what Gherkin got right all along. "Given the cart summary" was never decoration. It names the region of the world the sentence is about.

within(): the Given, made real

The Donobu test library now ships page.ai.within():

await page.ai
  .within('the cart summary')
  .assert('The region shows 2 items and a $34.00 total.');

Read it aloud. Given the cart summary, then it shows 2 items. The Gherkin was lurking inside Playwright the whole time.

Except within() is not a comment, and it is not an attention hint to the model. The scope is an evidence boundary: the AI judges only that region's rendered screenshot and its DOM subtree. Content anywhere else on the page can never cause a pass. If the cart summary shows one item and a banner elsewhere says "2 items", the assertion fails, as it should.

The subject can be plain English, resolved through the same pipeline and cache as page.ai.locate, or any Playwright locator you already trust:

await page.ai
  .within(page.locator('form#login'))
  .assert('The region contains a username field, a password field, and a login button.');

Scopes nest, and handles are reusable, which makes repeated components (cards, rows, panels, tiles) stop being a problem:

const sidebar = page.ai.within('the sidebar');
await sidebar.within('the filters section').assert('The region shows 3 filters.');
await sidebar.within('the cart summary').assert('The region shows 2 items.');

The innermost subject is the evidence boundary. Outer levels only steer how the inner ones resolve, and Playwright's own auto-wait and strictness apply at every level before any AI call is made.

Scoped locate and extract come along for the ride:

const row = await page.ai.locate('The row for user Frank');
const fields = await page.ai
  .within(row)
  .extract(z.object({ status: z.string(), due: z.string() }));

And then it compiles

Here is the part BDD never had.

The first run of a scoped assertion costs LLM tokens. Donobu then writes what it verified into a cache file next to your test: plain Playwright assertion steps, with the scope chain baked into the cache key. The file is generated, human-readable, and meant to be committed.

{
  assertion: 'The region contains a username field, a password field, and a login button.',
  scope: ['selector:form#login'],
  steps: [
    { locator: 'label', value: 'Username', assertion: 'toBeVisible' },
    { locator: 'label', value: 'Password', assertion: 'toBeVisible' },
    { locator: 'role', role: 'button', value: 'Login', assertion: 'toBeVisible' },
  ],
}

Every run after the first replays these steps deterministically, with no LLM in the loop, rooted at the subject so the boundary holds on cache hits too. One English sentence becomes committed Playwright code. Gherkin that compiles.

And Donobu stays vigilant on your behalf: if a modal or cookie wall covers the region mid-run, it ignores the cached result and re-judges live from the screenshot, so a covered page can never sneak through as a pass.

Try it

page.ai.within() is available now in the Donobu test library, alongside page.ai, assert, locate, and extract. If you already run Playwright, your suite does not change: same runner, same config, one import.

Read the docs, explore the SDK, or book a demo and we will show you a false green dying in real time.