👉 Try the live report — a real generated dashboard, fully interactive in your browser
Search it, filter it, mark an error group Fixed, switch a failure's evidence between runs. No install, no signup. Synthetic data, so no real project's tests appear in it.
Website: playwright-flaky-analyzer.vercel.app · npm: playwright-flaky-analyzer · Docs/usage: this README and STEPS.md · Demo source: demo-project/
Analyze Playwright test reports across multiple runs to detect flaky tests, track how flaky-test and retry counts trend over time, and enforce a CI quality gate. Deterministic classification and root-cause rules — all offline, zero network calls, no AI required. Generates an interactive HTML Dashboard, JSON, and Markdown.
Flaky tests erode trust in a test suite. They cause false CI failures, waste developer time chasing phantom breaks, and quietly teach teams to ignore test results altogether. Playwright's own reporters describe one run well — but they can't answer the question that actually matters: is this test broken, or is it just flaky? Answering that requires comparing multiple runs, deterministically, with no external service required.
| Playwright's built-in reporters | playwright-flaky-analyzer | |
|---|---|---|
| Scope | One run at a time | Compares 2+ runs |
| Flaky detection | Within a run only — fail-then-pass on retry. A test that passes in run 1 and fails first-attempt in run 2 looks like two unrelated results | Classifies every test as passing, flaky, fixed, newly failing, or consistently failing across runs — a fixed-then-broke-again test is still flagged as newly failing, with that history preserved in its reasons |
| Root cause | Raw error message only | 21 deterministic rules, each with a plain-language likely cause and concrete checks to try next |
| Grouping by shared error | None — one entry per failing test, with no notion that nine of them broke for the same reason | Common Errors: those nine collapse into one entry with nine test cases, grouped by a deterministic fingerprint that survives refactors. Fix once, not nine times |
| Recording your review | Read-only — nothing tracks which failures you've already looked at | Triage in the report: mark each failure Fixed / Known Failure, held per (test case, error) pair, with bulk marking per error group, n/m progress pills, undo, and a filter to hide what you've cleared |
| Historical evidence | Latest run only — Playwright cleans its outputDir, so earlier runs' screenshots and traces are gone |
Archived per run and per failing attempt as tests finish, with a picker to view any run that captured evidence |
| Flaky-test trend | Not measured | Flaky Tests Trend chart across the same runs already loaded by this analysis |
| CI enforcement | Pass/fail per test only | Optional --max-flaky quality gate — fails the build when the flaky-test count exceeds your threshold |
| Output | Console / HTML for a single run | HTML Dashboard, JSON, or Markdown — across your whole run history |
- Time saved — stop manually diffing CI logs across runs to tell "broken" from "flaky."
- Trust restored — a suite where flaky tests are visible and tracked is one people stop ignoring.
- Zero adoption friction — no account, no API key, no service to stand up.
npm installand one reporter config line.
Grouped by the point in triage where you'd reach for them. Full detail for every item is in FEATURES.md; exact flags and config keys are in STEPS.md.
- Cross-run classification — every test is labelled passing, flaky, fixed, newly failing or consistently failing across 2+ runs, plus a separate Skipped bucket for tests that didn't run in the latest run at all. This is the thing a per-run report structurally cannot do.
- Flaky Tests Trend — always-on chart of the flaky-test count per analyzed run (bar + connected trend line), with a plain-language first-vs-last sentence: "Flaky tests increased from 4 to 28 across the 20 analyzed runs."
- Retries Per Run Trend — how many tests needed a retry to pass, run over run, on the same Run 1…N axis as the flaky trend so the two are directly comparable. A distinct, within-run signal.
- Choose the window —
--lookback <n>analyzes up to the last n runs (a ceiling, not a requirement: fewer runs on disk just means fewer analyzed).--files run12,run13analyzes exactly the runs you name, with the last one listed treated as the latest.
- 21 deterministic root-cause rules — timeout, locator, assertion, network, auth, HTTP status and race-condition patterns, each with a plain-language likely cause, the reasons behind the verdict, and concrete checks to try next. No AI, no scoring model.
- Common Errors — failing tests grouped by the error they share, so one root cause is investigated once instead of once per affected test. Full untruncated error text, copy error / copy error + test list (works on
file://pages), and a jump to exactly the tests a group counted. - Failure fingerprinting — grouping is deterministic and doesn't rely on stack traces or line numbers, so it survives refactors.
- Evidence retention — the reporter archives screenshots, videos and traces per run and per failing attempt as tests finish, so historical evidence survives later Playwright runs that reuse the same
outputDir. No custom copy-step per project. A passing attempt's attachments are left where Playwright put them, since the analysis never reads evidence off one. - Evidence run picker — each failing test's evidence defaults to its most recent capture, with a dropdown to switch to any other run that captured some.
- Numbered run-history tiles — each test's pass/fail strip shows the run number on the tile itself, not only on hover, so a long stretch of same-coloured tiles stays readable.
- In-report triage — mark each failure Not Checked / Known Failure / Fixed as you go. Status is held per (test case, error) pair, so clearing one error leaves that test's status under every other error it also failed with untouched. Kept in
sessionStorage: survives a reload, cleared when you close the tab, and a newly generated report starts clean. - Bulk marking by error — one control marks every test listed under an error at once, placed on the card that shows you the exact affected set before you act, with an undo and a reset-all.
- Failed Tests by Spec File — the same failures regrouped by the file they live in, so "one error across five specs" is distinguishable from "five hits in one spec". Per-file classification counts and triage progress.
- Filters that carry across views — search plus classification, category and triage-status filters apply to Failed Tests, Failed Tests by Spec File and Common Errors together. Hide everything already marked Fixed and what's left is the actual remaining work. Partially-filtered groups show "n of m", so a narrowed count is never mistaken for the real one.
- Custom Playwright reporter — a drop-in entry in
playwright.config.jsproducing a stable, framework-independent JSON schema. No test-file changes, no service to run, no account. - CI quality gate — optional
--max-flaky <n>fails the build when the flaky-test count exceeds your threshold, evaluated after the report is written so a failing gate never costs you the report. No gate unless you set one — see DESIGN_DECISIONS.md § CI Quality Gate. - Three output formats — an interactive HTML dashboard, machine-readable JSON, or a Markdown summary for PRs and CI job summaries.
--also-jsonwrites the companion JSON alongside HTML. - Self-contained HTML — one portable bundle you can zip and open anywhere, or
--no-copy-evidencefor a single.htmlwithfile://links.--copy-recovered-evidenceopts recovered tests' artifacts back into the bundle (off by default — they pass in the latest run and dominate bundle size). - Bundle stays proportionate — a Consistently Failing test fails the same way every run, so only its most recent screenshot/video/trace is bundled rather than one set per run;
--copy-all-runs-evidencerestores the full set. Its run history, error, stack trace and root cause are unaffected. - Browser-aware — flaky tests are tracked per browser project: chromium, firefox, webkit, or any custom project name.
- Runs fully offline — deterministic analysis, zero network calls, two production dependencies (
commander,winston).
| Feature | What it provides |
|---|---|
| Suite Summary | 8 headline tiles for the current window: Total Tests, Passing, Passing on Retry, Recovered, Flaky, Newly Failing, Consistently Failing, Skipped |
| Flaky Tests Trend | Flaky-test count per analyzed run, bar+line chart, plain-language first-vs-last interpretation |
| Retries Per Run Trend | Retry count per analyzed run, bar+line chart, up/down/flat takeaway sentence |
| Run Highlights | Plain-English bullet summary of the analyzed window (latest-run breakdown, top failure category, retry concentration, slowest test, etc.) |
| Failed Tests | One investigation card per failing test — root cause, evidence, confidence (when below threshold), classification reasons, suggested checks; searchable and filterable |
| Failed Tests by Spec File | The same failures grouped by spec file, sorted by how many test cases each file contributes, with per-file classification counts and triage progress; any row jumps to that test's full card |
| Common Errors | Failing tests grouped by shared error — full error text, category, affected test list, copy error / copy error + test list, and "Find in Failed Tests" |
| Triage status | Not Checked / Known Failure / Fixed per (test case, error) pair, with bulk marking per error group, ✓ n/m fixed progress pills, undo, reset-all, and a triage-status filter across all grouped views. Stored in sessionStorage for that one generated report |
| Passing on Retry — Details | Tests currently passing that needed ≥1 retry in the latest run, with the same investigation detail as Failed Tests |
| Recovered Since Last Run — Details | Tests that failed in an earlier run and passed cleanly in the latest run — the error that broke them last time stays visible instead of silently disappearing into the Passing tile |
| Skipped Tests — Details | Tests skipped in the latest run |
| Root Cause Summary | Compact triage table — one row per failing test, with pattern/category/confidence |
| Browser Statistics | Executions, failures, and fail % per browser project, toggle between Latest Run and All Runs |
| Failure Categories | Error counts by category (timeout, locator, assertion, network, backend, authentication, environment, data, unknown), same Latest Run/All Runs toggle |
| Failure Frequency | Tests ranked by how often they've failed across the analyzed runs |
| Slowest Tests | The 5 slowest tests by average wall-clock time per run (retry attempts included), averaged over the runs each test actually ran in |
| Evidence run picker | Screenshot/trace/video per failing test, defaulting to the most recent run with evidence, switchable to any other run that also has some |
Full feature detail lives in FEATURES.md; classification/rule codes are in CHANGELOG.md; the reasoning behind these choices is in DESIGN_DECISIONS.md.
Explore the live demo report → — the real generated dashboard, fully interactive in your browser. Search it, filter it, mark an error group Fixed, switch a failure's evidence between runs.
demo-project/ is the suite behind it: a realistic ~70-test Playwright suite (Meridian, a fictional B2B SaaS) with a committed 20-run CI history, real evidence (screenshots/videos/traces), and a pre-generated dashboard. Clone the repo and open demo-project/flaky-report/index.html directly in a browser — no install or build step required. See demo-project/README.md to regenerate it yourself or run the live suite.
Playwright
↓
Flaky Analyzer (2+ runs)
↓
Historical Analysis
↓
Classification (stable / flaky / fixed / newly failing / consistently failing)
↓
Deterministic Rules → Root Cause → Evidence
↓
HTML / JSON / Markdown (Flaky Tests Trend + Retries Per Run Trend, same runs, same axis)
↓
Optional CI Gate (--max-flaky)
Every stage above is fully deterministic — no AI, no LLMs, no external APIs; the same input always produces the same output. See docs/architecture/ARCHITECTURE.md for the full component breakdown and data-flow diagram.
npm install -D playwright-flaky-analyzerOr run without installing:
npx playwright-flaky-analyzer analyze ./test-resultsRequirements: Node.js >= 18.0.0, Playwright >= 1.30.0 (for the custom reporter — see KNOWN_LIMITATIONS.md § Compatibility).
1. Add the reporter to playwright.config.js:
module.exports = {
reporter: [
["list"],
[
"playwright-flaky-analyzer/reporter",
{
outputFile: "./flaky-results/results.json",
},
],
],
};Keep analyzer result files outside Playwright's
outputDir(default:test-results) because Playwright may clean that directory before each test run — writing into a separate folder likeflaky-results/(shown above) keeps your accumulated history intact across runs.
2. Run your suite multiple times (locally or in CI) so more than one report accumulates:
npx playwright testEach run writes a numbered file: flaky-results/results-run1.json, results-run2.json, results-run3.json, ...
3. Analyze the results:
npx playwright-flaky-analyzer analyze ./flaky-results --format htmlThis produces a self-contained, portable report bundle — a flaky-analysis/ folder with index.html and an assets/ folder holding copies of every screenshot, video, and trace. Open index.html in any browser (no server), or zip the whole folder and send it — screenshots, inline video playback, and trace downloads keep working even if the original Playwright output is gone. This is what makes it robust in CI (Azure DevOps / GitHub Actions / Jenkins / GitLab), where the Playwright report and this report are published as separate artifacts.
Prefer a single
.htmlfile withfile://evidence links (the old behavior)? Pass--no-copy-evidence. Add--also-jsonto also emit the underlying dashboard data as JSON.
Try it without your own data first: the repository ships sample reports for exactly this —
npx playwright-flaky-analyzer analyze ./examples/sample-results --format html -o demo.html4. (Optional) Enforce a CI quality gate:
npx playwright-flaky-analyzer analyze ./flaky-results --format html --max-flaky 5Both the Flaky Tests Trend and Retries Per Run Trend charts are shown in the dashboard whether or not you use this flag — they're always computed from the runs in this analysis, with no separate history file or extra flag needed. --max-flaky is opt-in: it exits non-zero when the flaky-test count exceeds your threshold (the report is still generated first). Omitting it changes nothing for existing users.
The snippet above is the only supported integration path — see docs/architecture/REPORTER.md for every option, lifecycle hook, and the full output schema. Quick reference for the options:
[
"playwright-flaky-analyzer/reporter",
{
outputFile: "./flaky-results/results.json", // default if omitted: ./test-results/results.json — override it (as shown) to avoid Playwright's own outputDir cleanup, see note above
includeConfig: true, // default: true — embed the Playwright config in the report
includeErrors: true, // default: true — embed error messages/stacks
includeAttachments: true, // default: true — embed screenshot/video/trace paths
maxErrorLength: 5000, // default: 5000 — truncate error messages longer than this
},
];Each test run writes a numbered results-run<N>.json file plus an always-overwritten latest.json, both in the same directory as outputFile.
Evidence is archived automatically, per run. As each test attempt finishes, the reporter copies its Playwright attachments (screenshots, videos, traces, and any other captured attachment) into a run-scoped evidence directory next to that run's results-run<N>.json, and points attachments[].path at the archived copy rather than the original Playwright output path. This is what lets evidence for an earlier flaky failure stay available even after a later run passes the same test and Playwright cleans its own shared output directory — no extra configuration needed on your end. Because evidence is kept per historical run rather than only the latest one, disk usage grows accordingly as your results-run<N>.json history accumulates.
The analyzer also accepts Playwright's native JSON reporter output directly, if you'd rather not add a custom reporter.
The analyzer reads whichever of these two JSON shapes a file matches — you don't need to tell it which:
- This package's own reporter format (recommended) — the flat
{ schemaVersion, reporter, metadata, timing, summary, tests: [...] }shape written byplaywright-flaky-analyzer/reporterabove. Full schema: docs/architecture/REPORTER.md. - Playwright's native JSON reporter output — the nested
{ config, suites: [...] }shape from Playwright's built-injsonreporter. Accepted directly, no conversion step.
Either way, the analyzer needs 2 or more report files in the results directory to compare against each other — a single-run directory produces a "Need at least 2 valid reports for comparison" message, not an analysis. Files named results-run1.json, results-run2.json, ... are read in numeric run order; otherwise every non-dotfile *.json in the directory is read in sorted-filename order.
playwright-flaky-analyzer analyze ./test-results --format html -o dashboard.html
playwright-flaky-analyzer analyze ./test-results --format html --also-json # also write dashboard.json alongside
playwright-flaky-analyzer analyze ./test-results --format json -o dashboard.json
playwright-flaky-analyzer analyze ./test-results --format markdown --verbose
playwright-flaky-analyzer analyze ./test-results --max-flaky 5 # opt-in CI quality gate
playwright-flaky-analyzer analyze ./test-results --files run12,run13 # exact runs, in order (last = latest); overrides --lookback
playwright-flaky-analyzer init # scaffolds flaky.config.jsonFull flag reference, flaky.config.json schema, and troubleshooting: STEPS.md.
const { loadConfig, run } = require("playwright-flaky-analyzer");
const config = loadConfig(); // fills in every default, same as the CLI does
config.input.resultsDir = "./test-results";
config.output.format = "html";
run(config);run() expects a fully-shaped config object (all of input/output/analyzer/logging) — always start from loadConfig() rather than passing a partial object directly. Full API surface (compare, compute, PlaywrightReporter, format generators, etc.): API.md.
| Format | Use case |
|---|---|
| HTML (default) | Self-contained, portable dashboard bundle (index.html + assets/ with copied screenshots/videos/traces) — open in a browser, or zip the folder and share; works with no server and no dependency on the original Playwright output |
| JSON | Machine-readable, for CI/CD integrations or other consumers |
| Markdown | PR comments, Slack, any text-based workflow |
The HTML Dashboard includes a Suite Summary (Passing / Passing on Retry / Recovered / Flaky / Newly Failing / Consistently Failing / Skipped), then two aligned trend charts over the same analyzed runs, grouped together near the top as the high-level stability picture — Flaky Tests Trend (bar+line chart of flaky-test count per run, plus a first-vs-last interpretation sentence) and Retries Per Run Trend (bar+line chart of retry count per run, with its own takeaway sentence) — followed by Run Highlights, rule-based investigation cards for every failing test (search, filter chips, a run-by-run pass/fail history strip with visible run numbers on each tile, and an Evidence field whose screenshot/trace/video defaults to the most recent run but can be switched to any earlier run that also captured evidence), Passing on Retry, Recovered Since Last Run, and Skipped Tests details, and an Additional Metrics panel with Root Cause Summary, Browser Statistics and Failure Categories (toggle between "Latest Run" and "All Runs" scope), Failure Frequency, and Slowest Tests — see FEATURES.md for what each section shows and docs/architecture/ARCHITECTURE.md § HTML Generator for the full section-by-section breakdown.
Only README.md and STEPS.md ship inside the npm package (see files in package.json) — they're written to work standalone, offline, from inside node_modules/. Everything else below is deeper project/contributor documentation that lives on GitHub; those links go to the repository rather than a local path so they never 404 for someone reading this from an installed package.
| Document | Purpose |
|---|---|
| README | Project overview and quick start (you are here) |
| FEATURES | Every dashboard feature in detail — purpose, data source, calculation, interaction, example, limitations |
| STEPS | Full CLI reference, configuration schema, contributor setup |
| API | Full programmatic API reference |
| ARCHITECTURE | Internal architecture, components, data flow |
| REPORTER | Custom reporter internals, lifecycle hooks, output schema |
| DESIGN_DECISIONS | Why key architectural choices were made |
| DEVELOPMENT_JOURNEY | Engineering case study — origin, bugs found, lessons learned |
| KNOWN_LIMITATIONS | What the tool doesn't do (yet), and why |
| CHANGELOG | Release history — source of truth for what shipped when |
| ROADMAP | Completed in v1.0, planned for v1.1, future ideas |
| CONTRIBUTING | Contributor workflow and code style |
| RELEASE_CHECKLIST | Everything required before publishing a release |
| SECURITY | Reporting vulnerabilities |
| CODE_OF_CONDUCT | Community standards |
- Runtime: Node.js ≥ 18.0.0, plain CommonJS — no bundler, no compile step (
npm run buildjust prints a confirmation message). - Production dependencies (2 total):
commanderfor CLI parsing,winstonfor structured logging. No database, no server, no frontend framework. - HTML Dashboard: a single self-contained file — inline CSS/JS, the dashboard model embedded as a JSON literal, zero runtime dependencies of its own. Opens directly from the filesystem in any modern browser.
- Dev tooling:
eslint+prettierfor linting/formatting, Node's built-innode --testrunner (no separate test framework).
src/
├── cli/ # run-analysis.js — the `analyze`/`init` commands
├── analyzer/ # extractor, classifier, engine (compare), stats, failure-classifier, orchestrator (index.js)
├── reporter/ # PlaywrightReporter, dashboard-json, html, markdown, schema
├── investigation/ # rule-engine — deterministic root-cause investigation
├── knowledge/rules/ # the 21 deterministic root-cause rules (RC-001–RC-021)
├── evidence/ # collector, copier, path-rewriter — the portable-bundle pipeline
└── utils/ # config-loader, validator, fs, logger
examples/sample-results/ # sample Playwright reports ships with the package, for `analyze` without your own data
docs/architecture/ # ARCHITECTURE.md, REPORTER.md
Every src/**/*.js file has a co-located *.test.js; the package excludes test files from what's published (see files in package.json).
The short list — see KNOWN_LIMITATIONS.md for the full, categorized version with reasoning.
- Cross-run comparison, deterministic classification, and root-cause investigation for any JSON matching the expected schema (Playwright-native or this package's reporter format)
- A portable, self-contained HTML dashboard with evidence (screenshots/traces/videos) bundled in
- An opt-in CI quality gate on flaky-test count
- Single-run analysis — at least 2 reports are required; there's nothing to compare against with one
- Multi-repo/monorepo aggregation — each analysis compares runs from one results directory only
- Cross-invocation history — the Flaky Tests Trend chart is scoped to the runs loaded by the current analysis (
analyzer.lookbackRuns), not a database that survives between separateanalyzeinvocations (a file-based version of this was built and deliberately removed — see DESIGN_DECISIONS.md § Local Flaky Tests Trend) - In-run retry flakiness reflected in classification — a test that fails then passes on retry within a run still counts as a clean pass for cross-run classification (the "Passing on Retry" tile surfaces this for the latest run only)
- Server-side filtering for very large datasets — the HTML Dashboard is single-page and client-rendered; all search/filter happens in the browser against the embedded data
Directional ideas only, not committed or scheduled — see ROADMAP.md for the full list: a hosted/cross-machine trend store, reflecting in-run retry flakiness in classification itself, additional AI provider adapters (Copilot/Gemini/OpenAI/etc.), multi-repo aggregation, a native GitHub Actions annotation output format.
How many runs do I need? At least 2 for comparison. 5–10 give the most accurate flaky detection.
Does it work with any test runner? The analysis engine accepts any JSON matching the expected schema; the bundled custom reporter is Playwright-specific.
Can I run it in CI? Yes — add a step after your test matrix, and optionally enforce a flaky-count gate:
- name: Analyze flaky tests
run: npx playwright-flaky-analyzer analyze ./test-results --format html --max-flaky 5The gate is opt-in (omit the flag for today's default behavior) and is evaluated after the report is written, so a failing gate never prevents the report from being published. --max-flaky is the recommended gate rather than a newly-failing or consistently-failing threshold, because a flaky test typically passes on Playwright's own retry — so Playwright's own exit code stays green even as the suite degrades; a newly-failing or consistently-failing test already fails that exit code on its own. A worked Azure DevOps example (Playwright → history → analyzer → flaky-count threshold → published HTML artifact) is in docs/azure-pipelines.example.yml.
What counts as a flaky test? A test that has both passed and failed across runs, alternating (2+ transitions) — a run where the test was skipped, interrupted, or didn't run at all doesn't count toward this either way, so a test skipped once amid otherwise-consistent passes isn't flagged as flaky.
More in STEPS.md § Troubleshooting and KNOWN_LIMITATIONS.md.
Available now, in the current 1.2.0 release: the full cross-run comparison engine, investigation/fingerprinting, and three output formats (HTML, JSON, Markdown) from the initial release, plus everything added since — an opt-in CI quality gate (--max-flaky), an always-on Flaky Tests Trend chart (aligned with Retries Per Run Trend on the same analyzed runs), automatic per-run/per-attempt evidence archiving, a merged Regression/Newly-Failing classification, a Skipped bucket, a Recovered Since Last Run tile/section for tests that failed in an earlier run and passed cleanly in the latest one, --files <list> to analyze exact named run files instead of the last N by --lookback, and — new in 1.2.0 — in-report triage (Not Checked / Known Failure / Fixed per test-and-error pair, with bulk marking per error group), a Failed Tests by Spec File section, and copy/cross-navigation on Common Errors. All of it is deterministic and additive; none of it requires AI or a generic suite-health score. Exactly what shipped in which version: CHANGELOG.md.
Planned next (not yet shipped): category filtering in Failed Tests, finer-grained failure categories, timeout/locator disambiguation, a progress indicator for large suites, and CI workflow hardening. Longer-range, unscoped ideas (cross-machine trend storage, more AI provider adapters, multi-repo aggregation, a native GitHub Actions output format) live further out. Full list: ROADMAP.md.
Contributions are welcome — see CONTRIBUTING.md for the development workflow, code style, and how to run the test suite locally. Please also review the CODE_OF_CONDUCT.md.
MIT — see LICENSE
Shivani Singh
Senior QA Engineer passionate about AI-powered Software Testing, Playwright automation, and increasing productivity.