Skip to content

Repository files navigation

flakerad.

a diagnostic instrument for flaky tests

A pass is not a reason. Flakerad finds the why behind the flake.


npm version CI Status Technical Documentation License: MIT


npx flakerad watch "should sync user state"


the problem

A flaky test is an accusation without proof. It fails sometimes, passes sometimes, and nobody on the team can say why — so it gets skipped, retried, or quietly ignored until it isn't safe to ignore anymore.

A green checkmark on a rerun is not evidence. It's just a coin that landed differently the second time.

Flakerad reruns a failing test under four controlled conditions — one variable changed at a time — and tells you which one actually moved the result.


the four controls

Each run changes exactly one thing. Whichever control turns a failing test into a passing one is your root cause.

# Control What it holds constant If this fixes it → Actionable Remediation (v0.2.0+)
01 baseline nothing — the honest first run (reference point) Run controls to isolate the root cause
02 fixed seed Math.random() and seeded libs (faker, chance) non-deterministic input Seed random generators (faker.seed(1234)) or mock Math.random
03 frozen clock Date.now(), timers, intervals timing / race condition Replace setTimeout with condition polling (waitFor())
04 isolated order shared module state between tests test-order dependency Isolate singletons & reset state in beforeEach/afterEach

If none of the four resolve it, Flakerad doesn't guess — it flags the test as environment-dependent and hands it back for a human to look at. Abstaining honestly is a feature, not a gap.


install

npx flakerad init
npx flakerad watch "your test name or pattern"

or add it to the project:

pnpm add -D flakerad

usage

$ flakerad watch "user session should refresh once"

  ↻ baseline        ▓▓▓▓▓▓▓▓▓▓  10/10 runs   4 failed   60% pass
  ↻ fixed seed      ▓▓▓▓▓▓▓▓▓▓  10/10 runs   0 failed  100% pass
  ↻ frozen clock    ▓▓▓▓▓▓▓▓▓▓  10/10 runs   3 failed   70% pass
  ↻ isolated order  ▓▓▓▓▓▓▓▓▓▓  10/10 runs   3 failed   70% pass

  ┌────────────────────────────────────────────────────────────────────────────┐
  │ DIAGNOSTIC VERDICT: NON-DETERMINISTIC INPUT                                │
  │ Confidence: 95% | Baseline failure rate: 40%                               │
  │ A fixed seed collapsed the failure rate from 40% to 0%.                    │
  │ Remediation: Seed random generators explicitly (e.g. faker.seed(1234))     │
  │              or mock Math.random using jest.spyOn(Math, 'random')...       │
  └────────────────────────────────────────────────────────────────────────────┘

  → flakerad history --export json

Every reading here is a real measurement on the reruns Flakerad actually performed — never a fitted curve, never a guess dressed up as a number.


commands

Command Does
flakerad init detects your test runner, writes .flakeradrc.json
flakerad watch <pattern> runs the diagnostic engine on a target test
flakerad history local log of every test Flakerad has ever diagnosed
flakerad ci GitHub Actions–ready annotations + exit codes

Flags worth knowing: --reruns <n> (default 8), --safe-mode (caps reruns for tests with side effects), --fast (fewer reruns, reports reduced confidence honestly instead of hiding it).


how it works

failing test
    │
    ▼
┌──────────────┐      ┌──────────────┐      ┌──────────────┐      ┌──────────────┐
│   baseline   │ ──▶  │  fixed seed  │ ──▶  │ frozen clock │ ──▶  │   isolated   │
│  the honest  │      │ same random  │      │  same time,  │      │    order     │
│  first run   │      │  every run   │      │  every run   │      │ no neighbors │
└──────────────┘      └──────────────┘      └──────────────┘      └──────────────┘
       │                     │                     │                     │
       └─────────────────────┴──────────────┬──────┴─────────────────────┘
                                            │
                                            ▼
                           failure rate compared across all four
                                            │
                                            ▼
                           root cause, or an honest abstention

Only ambiguous, residual cases — ones the four deterministic controls can't resolve on their own — are ever handed to a language model for a closed-set classification, and it's allowed to abstain too. Nothing here pretends to be more certain than the evidence supports.


why this exists

Most "flaky test" tooling just retries until green and calls it solved. That's not a fix, it's a bribe. Flakerad's only job is to make the dependency visible — the test didn't become reliable, you just finally saw what it was waiting on.


project layout

flakerad/
├── app/         next.js site — the instrument, the field notes, the docs
├── components/  shared UI, themed in the site's editorial palette
├── data/        real output from the fixture suite, never hand-edited
├── cli/         the actual tool — init, watch, history, ci
├── fixtures/    four deliberately flaky tests, one per root cause
└── lib/

contributing

Found a genuinely flaky test in the wild that Flakerad misdiagnosed? That's the most valuable bug report there is — open an issue with the test and the four-run output, and it may end up as the next fixture.


technical documentation

For an in-depth architecture breakdown, causal inference design, and algorithmic specification:

Technical Documentation


repository & bugs

To see the source code, inspect run data, or to find bugs, visit the repository:


license

MIT — see LICENSE.


made for the suspicious

About

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages