MediumConcept

How do you deal with flaky tests?

Technology:
Software Testing
Experience:
Mid-level,
Senior

Quick answer

Treat a flaky test as a bug: quarantine it so it stops blocking the pipeline, reproduce it by running it repeatedly, fix the root cause (usually timing, shared state or external dependencies), then bring it back.

Detailed explanation

A flaky test passes and fails on the same code. Its real cost is trust: once developers learn to click "re-run", genuine failures get ignored too.

Common causes are fixed sleeps instead of waiting for a condition, tests that depend on execution order or shared data, real network calls, time zones and the current date, randomness without a fixed seed, and animations in UI tests.

The process: track flaky tests (many CI systems report them), quarantine them with an owner and a deadline, reproduce by running the test in a loop or under load, and fix the cause. Use auto-waiting assertions in Playwright or Cypress, isolate test data per test, control the clock, and mock external services. Retries can hide symptoms, so use them sparingly and still track the underlying failures.

Example

javascript
// flaky: assumes the request finishes within 1 second
await page.click('text=Save');
await page.waitForTimeout(1000);
expect(await page.textContent('.toast')).toBe('Saved');

// stable: waits for the condition (auto-retrying assertion)
await page.getByRole('button', { name: 'Save' }).click();
await expect(page.getByRole('status')).toHaveText('Saved');

Key points

  • Flaky tests destroy trust in the suite
  • Usual causes: timing, shared state, external services, dates, randomness
  • Quarantine with an owner, reproduce in a loop, fix the root cause
  • Avoid fixed sleeps; wait for conditions

Common mistakes

  • Adding automatic retries and considering the problem solved.
  • Deleting the test instead of fixing the cause or the product bug it was exposing.

Follow-up questions

  • Software TestingEasy

    What is the test pyramid?

    The test pyramid is a guideline to have many fast unit tests at the base, fewer integration tests in the middle and only a small number of slow end-to-end tests at the top, so the suite stays fast and reliable.

    Junior 路 Mid-level 路 Testing
  • Software TestingMedium

    What is mocking in testing and when should you use it?

    Mocking replaces a real dependency with a controllable fake so a test can run in isolation and check how the code interacts with it; use it for slow, non-deterministic or external dependencies, not for everything.

    Junior 路 Mid-level 路 Testing
  • Software TestingMedium

    Selenium vs Cypress vs Playwright: which should you choose for test automation?

    Selenium is the long-established, language-agnostic standard built on WebDriver; Cypress runs inside the browser with an excellent developer experience for JavaScript apps; Playwright offers fast, auto-waiting cross-browser automation with multiple tabs, contexts and languages. For a new web project, Playwright is usually the strongest default.

    Mid-level 路 Senior 路 Testing
  • Software TestingEasy

    What is the difference between unit, integration and end-to-end tests?

    Unit tests check one small piece of code in isolation, integration tests check that several pieces work together (for example code plus a real database), and end-to-end tests drive the whole application the way a user would.

    Intern 路 Junior 路 Testing
  • ReactEasy

    Why do we need keys in React lists?

    Keys give each list item a stable identity so React can match items between renders and correctly add, remove or reorder them without recreating or mixing up their state.

    Junior 路 Mid-level 路 Performance
  • CSSMedium

    How does CSS specificity work?

    Specificity is the score the browser uses to decide which rule wins: inline styles beat IDs, IDs beat classes/attributes/pseudo-classes, and those beat elements. When scores tie, the last rule in the source wins.

    Junior 路 Mid-level 路 CSS Layout