Testing Feature Flag Rollbacks in Browser Automation Without Carrying State Between Runs
By Markus Gasser · October 1, 2026
A practical guide to testing feature flag rollbacks in browser automation without stale cookies, cached assets, or leftover backend state causing false failures.
If a rollback test only passes because the browser remembers yesterday’s state, it is not testing rollback safety, it is testing residue. The hard part is not clicking the toggle on and off, it is proving that a user can see the flag-enabled UI, then recover cleanly after rollback with no stale cookies, cached assets, local storage, service worker artifacts, or backend records making the result look better than it is.
The short version is this: test feature flag rollbacks in browser automation by starting from a fresh browser context every run, resetting server-side test data through API or fixture setup, and verifying the same UI path both before and after the flag flip. If you cannot reset the backend, your browser test will eventually become a data-reuse test.
What a rollback test actually needs to prove
Feature flag testing is often described too loosely. For rollback coverage, there are really three separate questions:
- Does the feature appear when the flag is on?
- Does the UI disappear or revert correctly when the flag is off?
- Can an existing user session recover after rollback without stale client state causing broken behavior?
That third question is where browser automation helps, because rollback failures often come from the browser, not the toggle service itself:
- a cached bundle still references the removed UI code,
- local storage keeps an old mode value,
- a session cookie keeps the app in the wrong branch,
- a service worker serves old assets,
- a persisted backend record makes the “rolled back” path look valid even when the UI is broken.
A rollback test is not only a feature-flag assertion. It is a state-isolation test across browser, cache, and backend.
The isolation model that prevents false failures
To keep runs independent, isolate state at three layers.
1) Browser-layer isolation
Use a new browser context for every test case. In Playwright, a context gives you fresh cookies, local storage, session storage, and cache partitioning. This is the single most important control if you want to avoid carrying state between runs.
Official docs worth reading:
Do not reuse a long-lived context for a rollback scenario unless the test is explicitly about persistence across sessions.
2) Backend-layer isolation
A browser reset does not clear your application database, flag state service, or test fixtures. If the feature creates records, set up a cleanup path, preferably through an API or direct fixture reset in a dedicated test environment.
For rollback testing, you usually want one of these patterns:
- idempotent setup API, create the same data every run,
- per-test tenant or namespace, each run writes to isolated records,
- clean teardown, delete records after the test finishes,
- ephemeral environment, rebuild the environment for the suite.
If the backend state cannot be reset reliably, the test should assert only what the UI reflects, not the full lifecycle.
3) Cache and asset isolation
A feature rollback can fail because the browser still holds an old JS bundle or a service worker response. If your app uses a service worker or aggressive HTTP caching, your test strategy must account for it.
Practical options:
- use a fresh context each run,
- disable or bypass service workers in the test environment when that is acceptable,
- version static assets so rollbacks load matching code,
- verify cache headers and asset invalidation in a separate deployment test.
A stable rollback test flow
A good rollback test for a web app usually follows this sequence:
- Start with a fresh browser context.
- Set up backend data through API or fixtures.
- Enable the feature flag for a controlled user or tenant.
- Load the page and confirm the feature is visible and functional.
- Flip the flag off or trigger the rollback.
- Open a new tab or new context if the scenario needs to represent a fresh session.
- Confirm the UI returns to the old path, or that the removed feature is hidden and the fallback works.
- Confirm the page does not break on reload.
- Confirm no stale client-side state is causing the feature to appear when it should be gone.
The exact assertions depend on the feature, but the rollback step should always be followed by a new navigation or new context check. Otherwise you may only be verifying the already-rendered DOM.
Example: Playwright rollback test with explicit cleanup boundaries
This example shows the pattern, not a universal template. The important part is the separation between flag setup, browser context, and backend reset.
import { test, expect } from '@playwright/test';
test('feature flag rollback clears the new UI path', async ({ browser, request }) => {
// Reset or create backend state through an API designed for test setup.
await request.post('/api/test/reset-user', {
data: { userId: 'qa-user-1' }
});
// Turn feature on for the target user or tenant.
await request.post('/api/test/flags/checkout-redesign', {
data: { userId: 'qa-user-1', enabled: true }
});
const context = await browser.newContext();
const page = await context.newPage();
await page.goto('/checkout');
await expect(page.getByRole('heading', { name: 'New checkout' })).toBeVisible();
// Simulate rollback.
await request.post('/api/test/flags/checkout-redesign', {
data: { userId: 'qa-user-1', enabled: false }
});
// New navigation matters, because the current DOM may still show prior state.
await page.reload();
await expect(page.getByRole('heading', { name: 'New checkout' })).toHaveCount(0);
await expect(page.getByRole('heading', { name: 'Checkout' })).toBeVisible();
await context.close();
});
A few details matter here:
browser.newContext()keeps each run isolated.- The API setup and rollback are separate from the UI assertion.
- The reload after rollback is essential, because a single-page app can preserve component state after server-side state has changed.
If your app uses local storage or cookies for flag state
Some apps cache feature decisions client-side for performance. That is fine if the cache is intentional, but it changes how you test.
When the browser stores flag state in local storage or a cookie, test both of these paths:
- fresh session path, the user opens the app after rollback,
- active session path, the user already had the page open when rollback happened.
Those are not the same test.
Fresh session path
This should use a new context or an explicit storage reset.
const context = await browser.newContext();
const page = await context.newPage();
await page.goto('/dashboard');
Active session path
This should confirm the app can reload or navigate away and come back without stale client state.
await page.goto('/dashboard');
await page.reload();
await expect(page.getByTestId('legacy-dashboard')).toBeVisible();
If the rollback logic depends on client-side storage, avoid mutating that storage manually inside the assertion unless the test is specifically about a migration path. Otherwise, you are hiding the very bug you want to catch.
What to clear, what to keep, and why
A useful way to think about rollback tests is to separate state you want to eliminate from state you want to observe.
Clear these between runs
- cookies,
- local storage,
- session storage,
- service worker cache,
- test-created backend records,
- server-side flag assignments that are not part of the scenario.
Keep these under test control
- the flag value itself,
- the user or tenant identity used for targeting,
- the exact browser version and viewport if the UI is responsive,
- the app build under test.
That distinction matters because a rollback test should be deterministic. If the same user can land in different states because some hidden storage survives from yesterday’s run, your failure signal becomes noisy fast.
Common failure modes and what they usually mean
The feature still appears after rollback
This often means one of three things:
- the page still has in-memory component state,
- the browser reused stale data from storage or cache,
- the backend flag change had not propagated yet.
The fix is different for each one. Reloading the page helps with in-memory state, but it does not solve backend propagation delay. For that, poll the flag service or wait for a documented consistency condition before asserting.
The fallback page breaks after rollback
This usually points to a code-path mismatch between the old UI and the new feature assets. If the new release removed code that the rollback path still expects, the issue is in deployment compatibility, not the automation itself.
That is a strong signal to add a deployment-level check that verifies both branches can load against the same build or compatibility contract.
The test only fails on reruns
That is a state-leakage smell. Look at:
- retained browser context,
- reused test user,
- non-cleaned backend records,
- service worker persistence,
- flag targeting based on prior page visits.
If reruns behave differently, the test is not isolated enough.
A decision framework for rollout and rollback coverage
Not every team needs the same depth of coverage. Pick the lowest layer that still catches the bug you care about.
| Scenario | Best coverage layer | Why |
|---|---|---|
| Simple UI toggle visibility | Browser automation | Fastest way to confirm the feature appears and disappears |
| Flag drives form fields or validation | Browser automation plus API setup | You need the UI and the data state |
| Rollback must preserve old data shape | API plus browser automation | UI alone will miss schema or compatibility failures |
| Rollback must work after a deploy | Browser automation in a post-deploy environment | This catches asset and compatibility mismatches |
| Flag state is cached client-side | Fresh context plus reload path | Prevents stale cookies or local storage from masking defects |
If the feature is mostly presentation, browser automation may be enough. If the rollback can corrupt persisted data, add API-level or database-level checks.
A maintenance note for flaky rollback tests
Rollback tests tend to become flaky when they try to do too much in one run. The typical anti-pattern is a single test that:
- creates data,
- sets the flag,
- waits for async replication,
- performs the UI flow,
- rolls back,
- retries on failure,
- then checks cleanup.
That is too many moving parts for one assertion chain.
A better split is:
- one setup helper that creates clean state,
- one browser test for the enabled path,
- one browser test for the rollback path,
- one API or contract test for propagation timing if needed.
This separation makes failures easier to diagnose. When the rollback test fails, you can tell whether the problem is flag propagation, storage leakage, or a UI regression.
A practical checklist you can reuse
Before you trust a rollback test, confirm these points:
- each run uses a fresh browser context,
- test data is created through a resettable API or fixture,
- the flag can be toggled without manual UI steps,
- the test re-navigates or reloads after rollback,
- cookies, local storage, and service worker state are not reused,
- the old path and the fallback path both have explicit assertions,
- backend cleanup is deterministic,
- any eventual consistency in the flag service is handled explicitly.
Bottom line
To test feature flag rollbacks in browser automation without carrying state between runs, separate browser state from backend state and treat both as test fixtures. Fresh contexts prevent stale cookies and storage from contaminating the result. API-backed setup and teardown prevent backend records from making a rollback look healthy when it is not. Reloads and fresh navigations prove that the app can recover after the flag flips.
If you remember one thing, make it this: a rollback test should prove the app behaves correctly after the flag changes, not just while the first page instance is still alive.
FAQ
Should I clear cookies manually in every test?
Usually no. A fresh browser context is cleaner and more reliable than selectively deleting cookies after the test has already started.
Do I need to test both feature-on and feature-off states?
Yes, if the rollback changes the visible UI or data flow. The rollback path is only meaningful if you know the feature-on path worked first.
Can I test flag rollbacks with only UI steps?
You can, but only for very shallow cases. If the feature creates or depends on backend data, UI-only tests often miss the real failure mode.
Why does the test pass until I rerun it?
That usually indicates leaked state, reused test data, or a browser cache artifact. A rerun should start from the same clean baseline as the first execution.
What if the rollback is eventually consistent?
Do not assert the final UI state immediately. Wait for the documented propagation condition, or poll the flag source through a test endpoint before checking the UI.