How do you deal with flaky tests?
- Technology:
- Software Testing
Quick answer
Treat a flaky test as a bug: quarantine it so it stops blocking the pipeline, reproduce it by running it repeatedly, fix the root cause (usually timing, shared state or external dependencies), then bring it back.
Detailed explanation
A flaky test passes and fails on the same code. Its real cost is trust: once developers learn to click "re-run", genuine failures get ignored too.
Common causes are fixed sleeps instead of waiting for a condition, tests that depend on execution order or shared data, real network calls, time zones and the current date, randomness without a fixed seed, and animations in UI tests.
The process: track flaky tests (many CI systems report them), quarantine them with an owner and a deadline, reproduce by running the test in a loop or under load, and fix the cause. Use auto-waiting assertions in Playwright or Cypress, isolate test data per test, control the clock, and mock external services. Retries can hide symptoms, so use them sparingly and still track the underlying failures.
Example
// flaky: assumes the request finishes within 1 second
await page.click('text=Save');
await page.waitForTimeout(1000);
expect(await page.textContent('.toast')).toBe('Saved');
// stable: waits for the condition (auto-retrying assertion)
await page.getByRole('button', { name: 'Save' }).click();
await expect(page.getByRole('status')).toHaveText('Saved');Key points
- Flaky tests destroy trust in the suite
- Usual causes: timing, shared state, external services, dates, randomness
- Quarantine with an owner, reproduce in a loop, fix the root cause
- Avoid fixed sleeps; wait for conditions
Common mistakes
- Adding automatic retries and considering the problem solved.
- Deleting the test instead of fixing the cause or the product bug it was exposing.