Your CI pipeline fails on a test that hasn't changed in weeks. You rerun the job with zero code edits, and it goes green. Multiply that across a dozen flaky tests in a daily build, and you're burning hours of engineering time chasing ghosts. Worse, your team starts distrusting red builds, so real bugs slip through disguised as "probably just flaky."
The instinct is to blame the test framework, the CI runner, or network latency. Sometimes that's the real cause. But in a discussion about applying test-driven development to agentic engineering, BrowserStack's David Burns points to a more common root cause: flaky tests are usually a symptom of unmanaged application state, not a broken testing tool (Stack Overflow Blog). This guide walks through how to prove that diagnosis and fix it.
What Distinguishes a Flaky Test From a Genuinely Broken One
A broken test fails the same way every time, given the same inputs. Something in your logic or assertion is objectively wrong, and rerunning it changes nothing.
A flaky test's outcome depends on a hidden variable that has nothing to do with the code under test: execution order, timing, parallel worker assignment, or leftover data from a previous run.
The fastest way to tell them apart is a controlled rerun. Run the failing test alone, in a tight loop, twenty or so times. If it always passes in isolation but fails intermittently inside the full suite, you're looking at leaked state, not a logic bug. If it fails the same assertion every time regardless of order or isolation, it's genuinely broken and needs a real code fix, not an isolation pattern.
How Leaked State Creates Intermittent Failures
Tests rarely run in true isolation. They share a process, a database connection, a file system, and sometimes wall-clock time.
When one test mutates something shared and doesn't reset it, whether the next test fails depends entirely on run order and timing. That's exactly what makes the failure look random.
Four sources cause most of this:
- Shared fixtures scoped too broadly (session or module level). One test mutates the fixture; another test reads it, and the read result depends on what ran first.
- Global variables and singletons, like in-memory caches or config objects, persist state across test boundaries inside the same process.
- Async race conditions let assertions run before a promise resolves or a background thread finishes.
- Unclean teardown leaves database rows, temp files, or monkeypatched functions in place after a test claims to be done.
Step-by-Step: Diagnose Which Shared State Is the Culprit
Use this sequence to trace an intermittent failure back to its source before you touch any fixtures.
- Reproduce in isolation versus in the full suite. Run the failing test alone in a loop. If it's always green alone and only fails inside the suite, you've confirmed leakage, not logic.
- Shuffle test order deliberately. Most modern runners support randomized or seeded order (pytest-randomly for Python, Jest's
--randomize, or similar plugins elsewhere). Run the suite with a few different seeds and note exactly which test runs immediately before the failure appears. - Grep for module-level and global state. Search your codebase for global declarations, singleton classes, class-level dictionaries or lists, and any place code mutates
os.environor a shared config object. - Audit fixture scope. Treat anything scoped beyond "function" as suspect. Session- or module-scoped fixtures should hold only truly immutable setup, never anything a test writes to.
- Instrument teardown with logging. Don't assume cleanup works just because a teardown function exists. Log the state of the shared resource right after teardown runs and confirm it actually reset.
- Check for unawaited async work. Look for background threads, timers, or promises that fire but never get joined or awaited before the test's assertions execute.
Isolation Patterns That Actually Stop the Leaks
Prefer function-scoped fixtures over shared setup
Default every fixture to function scope unless you have a measured performance reason to widen it. A fixture that rebuilds its state for every test can't leak into the next one, because nothing is left to leak. For more on this, see read about slack code channels vs terminal agents: a decision guide.
Make teardown undo side effects, not just close connections
Closing a database connection isn't the same as rolling back the rows it inserted. Teardown should explicitly reverse whatever the test wrote: delete rows, restore patched functions, clear caches, and reset environment variables to their original values.
Mock or fake shared state instead of mutating the real thing
If a test needs to touch a rate limiter, a cache, or a config singleton, give it a fresh fake instance instead of poking the real global object. This removes the shared resource from the equation entirely, rather than just cleaning up after it.
Make async deterministic
Replace real timers and sleep-based waits with controllable fakes, and always await or join background work before asserting. A test that passes only when the CPU happens to run fast enough is a flaky test waiting to happen.
Annotated Example: Flaky Test vs. Isolated Test
The following example is unexecuted and illustrative, written for Python 3.11+ with pytest 8.x. It shows a module with global state and two ways of testing it.
## limiter.py
_request_counts = {} # module-level global state, shared across every test
def allow_request(user_id, limit=3):
_request_counts[user_id] = _request_counts.get(user_id, 0) + 1
return _request_counts[user_id] <= limit
## test_limiter_flaky.py -- DO NOT COPY, this is the broken pattern
from limiter import allow_request
def test_first_three_requests_allowed():
# Silently assumes "u1" has never been seen before in this process.
assert allow_request("u1") is True
assert allow_request("u1") is True
assert allow_request("u1") is True
def test_fourth_request_blocked():
# Only passes if the test above ran first, in the same process,
# and left exactly three prior calls in _request_counts.
assert allow_request("u1") is False
Run this with a random-order plugin and it fails intermittently. If test_fourth_request_blocked runs before the other test, or if a rerun reuses the same process, the assumption about prior state breaks.
## test_limiter_isolated.py -- the fix
import pytest
from limiter import _request_counts, allow_request
@pytest.fixture(autouse=True)
def reset_limiter_state():
_request_counts.clear() # runs before every test
yield
_request_counts.clear() # and after every test, no matter the outcome
def test_first_three_requests_allowed():
assert allow_request("u1") is True
assert allow_request("u1") is True
assert allow_request("u1") is True
*Also read:* [full coverage of how to calibrate confidence thresholds in ai agents](/how-to-calibrate-confidence-thresholds-in-ai-agents)
def test_fourth_request_blocked():
for _ in range(3):
allow_request("u1")
assert allow_request("u1") is False
Expected output: pytest test_limiter_isolated.py -v shows both tests PASS, in any order. Verify by running with several random seeds (pytest -p randomly --randomly-seed=1234, then a different seed) twenty or so times; all runs should stay green.
As a failure case, comment out the autouse fixture and rerun with shuffled order. You'll reintroduce the exact same intermittent failure, confirming that the shared dictionary, not the logic, was the problem all along.
Checklist for Auditing a Suite for State Leakage
Work through this list against your existing suite, not a hypothetical one.
- List every fixture scoped above "function" and justify why each one is safe to share.
- Search the codebase for global, singleton, or module-level mutable containers (dicts, lists, counters, caches).
- Confirm every test that writes to a database, file, or cache has a teardown that explicitly reverses that write.
- Check whether your CI config runs tests in parallel workers that might share a database, port, or temp directory.
- Search for sleep-based waits or unawaited async calls near assertions.
- Run the full suite three times with different random seeds and diff the results.
- For any test that only fails under parallel execution, check for shared ports, files, or environment variables between workers.
- Document which shared resources are intentionally reused, and why, so the next engineer doesn't reintroduce the same leak while "cleaning up."
Once you've isolated the leaking fixture, global, or teardown gap, the fix is almost always smaller than the debugging effort that found it. Once shared state stops crossing test boundaries, reruns should become rare instead of routine. Any flakiness that survives this process likely points to a genuine race condition in your production code, not in your tests.



