Open-source AI heart murmur detection for primary care.
π§ Try the live in-browser demo β https://noahisarider.github.io/open-stethoscope/ The full 404K-parameter model runs locally in your browser (WebAssembly) β upload a recording or play the real CirCor clips; no audio ever leaves your machine.
Cardiovascular disease is the leading cause of death in China. In county and township health centers, general practitioners often lack the training to interpret heart sounds. Open Stethoscope brings cardiology-grade auscultation to every village doctor β for free, open source, and on-device.
- CVD is the #1 cause of death in China and worldwide
- ~1.2 million heart sound recordings are needed at grassroots level each year, but there aren't enough cardiologists to read them
- Digital stethoscopes + AI can screen for heart murmurs (valvular disease) with accuracy approaching specialist auscultation
- Existing solutions are closed-source and enterprise-priced β this project is open and runs on edge hardware
A deep learning system for heart murmur detection from phonocardiogram (PCG) recordings, built on the CirCor DigiScope dataset (PhysioNet Challenge 2022):
- 5,272 recordings from 1,568 patients, expert-annotated for murmurs
- ~400K parameters, <1 GB VRAM, real-time CPU inference β designed for low-cost edge devices (Tesla P4 / CPU)
- Reproducible baseline with strict patient-disjoint evaluation and the official Challenge metric
- Multi-location attention fusion β one patient = one sample. All available auscultation positions (AV/PV/TV/MV) are encoded by a shared 2D-CNN over log-mel spectrograms, then fused with a masked learned attention (missing positions masked out). Mirrors how a clinician listens to all positions before judging.
- Class-imbalance handling β sqrt-inverse-frequency patient sampling + class-weighted CE + auxiliary "murmur present?" binary head (Ξ»=0.3).
- Mel-domain augmentation β random 8 s window, Β±10% time stretch, SpecAugment-lite, Gaussian noise.
- Patient-stratified evaluation β strict patient-disjoint 70/15/15 train/val/test split (no leakage).
Held-out test set (142 patients, never seen during training), official CirCor-2022 challenge metric s_murmur (weights Absent:1 / Unknown:3 / Present:5):
v4 (current best) β test s_murmur 0.7815, beats the 2022 Challenge champion (0.780):
| Model | Test acc | Test macro-F1 | Test s_murmur | recall A / U / P |
|---|---|---|---|---|
| v4 L_s43_30ep + val-tuned thresholds | 0.8380 | 0.7177 | 0.7815 | 0.886 / 0.600 / 0.741 |
| v4 L2_s43_smur (s_murmur early-stop, untuned) | 0.7817 | β | 0.7815 | 0.790 / 0.600 / 0.815 |
| v4 4-model ensemble (tuned) | 0.7958 | 0.7055 | 0.7815 | 0.800 / 0.900 / 0.741 |
| v3 fusion (previous) | 0.7535 | 0.6537 | 0.7000 | 0.781 / 0.900 / 0.593 |
| Fusion β per-location vote | 0.6620 | 0.5823 | 0.6222 | 0.676 / 0.900 / 0.519 |
| Single-location (patient vote) | 0.7324 | 0.5817 | 0.6074 | 0.819 / 0.600 / 0.444 |
- v4 vs v3: +0.0815 s_murmur (0.7000 β 0.7815). Key lever: longer training (30 epochs, patience 7; 20-epoch models were underfit) + val-tuned decision thresholds. Present boost / early-stop metric swap / naive ensembling all failed to help.
- Benchmark vs the official 2022 Challenge (40 teams, hidden test): champion HearHeart 0.780, top-10 cutoff 0.755, median 0.692 β v4 0.7815 edges past the champion; known 2023 wav2vec2 SOTA is 0.80 (β0.0185 away).
- Ablations (v3, same split): class-balance sampling is the biggest lever; learned attention β mean pooling β the winning ingredient is multi-position fusion itself, not the attention weights.
- Bottlenecks: Present recall 0.74 (7/27 missed β Absent); Unknown recall 0.4β0.8 varies wildly by seed (only 10 Unknown patients in val).
Independent re-validation & honest generalization (2026-08-20, new server, official labels):
| check | result |
|---|---|
| Exact reproduction of all 4 original v4 seeds | 0.7593 / 0.7778 / 0.6556 / 0.7704 β identical to original |
| Best single (s43 + val-tuned) | 0.7815 β beats 2022 champion 0.780 |
| Best val-selected top-4 ensemble (9-seed family, tuned) | 0.7926 β new best |
| 5-fold patient-stratified CV | 0.7036 Β± 0.0604 (per-fold 0.759 / 0.665 / 0.754 / 0.737 / 0.603) |
| Out-of-fold ensemble (942 patients) | s_murmur 0.7040 Β· recall [0.859, 0.544, 0.620] |
Interventions that did not help (all tested on the fixed split, seed 43):
- Larger model: 2.26M-param encoder β 0.7296 (overfits 659-patient train set)
- Focal loss (Ξ³=2) β 0.6111 (Present recall collapses)
- Frozen wav2vec2-base features (94.4M params) + MLP/attn head β 0.6778
β The honest reading: single-split 0.7815 sits at the favorable end of a Β±0.06 CV spread; the model's true generalization is ~0.70. On this small dataset a 404K from-scratch multi-position fusion model beats both a 5.6Γ bigger version of itself and frozen 94M self-supervised features.
Full details: EXPERIMENTS.md (v3) Β· EXPERIMENTS_v4.md (v4 iteration log, 20 runs) Β· EXPERIMENTS_2026-08-20.md (re-validation, CV, ablations).
Every number in this README is reproducible β without a GPU. All trained weights,
val/test probabilities, training logs and result JSONs are committed under
experiments/:
experiments/
βββ models/ # 16 trained checkpoints (9-seed v4 family, kfoldΓ5, v5, focal) + QC models
βββ probs/ # val/test softmax probabilities per seed (ensemble inputs)
βββ logs/ # full training logs (run_s42..s52_30ep.log, kfold, focal, v5, w2v)
βββ results/ # exp_results*.json, kfold_results.json, score_library.json, v3_split_seed42.json
Reproduce the best result (test s_murmur 0.7926):
# 1. (optional) Re-evaluate any saved checkpoint on val+test β dumps v4_probs_<tag>.npz
python3 eval_probs.py s43_30ep s44_30ep s45_30ep s46_30ep
# 2. Val-selected top-k ensemble over the 9-seed family (CPU, ~seconds)
python3 ensemble_topk.py # β top4 tuned β test = 0.7926
# 3. Or tune a fixed ensemble yourself
python3 tune_ensemble.py --mode probavg s42_30ep s43_30ep s44_30ep s45_30epReproduce a single model (s43, test 0.7815):
python3 tune_ensemble.py --mode probavg s43_30epAll scripts auto-detect the experiments/ layout; the original server layout
(/root/heart-train) is supported as a fallback, and OS_WORKDIR overrides either.
- Reproduce PhysioNet 2022 Challenge baseline
- Train lightweight murmur classifier (multi-position fusion, s_murmur 0.70)
- Beat the 2022 Challenge champion (v4: longer training + tuned thresholds β 0.7815 > 0.780)
- Multi-seed ensemble with k-fold-based model selection (5-fold CV done: 0.704 Β± 0.06)
- In-browser demo (
demo/browser/β model runs locally in WASM, no server, no upload) - Chinese primary-care deployment guide
demo/browser/ is a fully static, zero-build page that runs the trained model
entirely in your browser via ONNX Runtime Web (WASM):
- 404K-parameter model exported to ONNX (1.6 MB) + librosa-exact mel front-end in JS
- Loads 3 real CirCor clips (positive / unlabelled / negative) or your own WAV/MP3
- No audio ever leaves your machine β inference happens locally
cd demo/browser && python3 -m http.server 8000 # then open http://localhost:8000Deploy anywhere static (GitHub Pages / Vercel / nginx). The inference chain is
verified end-to-end against the training environment: JS mel == numpy mel ==
librosa 1.0.0 (max diff < 0.3 dB), ONNX == PyTorch (< 1e-6), and final
probabilities match the server's ground-truth inference to < 1e-3
(demo/browser/tools/verify.py).
Regenerate the artifacts with:
python3 demo/browser/tools/export_onnx.py # model.onnx + mel_params.json
python3 demo/browser/tools/verify.py assets/samples/*.wav # vs training envCirCor DigiScope v1.0.3 β open access, no application required:
wget https://physionet.org/static/published-projects/circor-heart-sound/circor-heart-sound-1.0.3.zip# 1. Download CirCor dataset (open access, ~560 MB)
wget https://physionet.org/static/published-projects/circor-heart-sound/circor-heart-sound-1.0.3.zip
unzip circor-heart-sound-1.0.3.zip
# 2. Train (fusion + single-location baseline, patient-level 70/15/15 split)
python train_v3.py --data-csv ./training_data.csv --data-dir ./training_data --workdir ./out
# 3. Ablations (same split, retrained)
python train_v3.py --ablation A # no attention (mean pooling)
python train_v3.py --ablation B # no auxiliary head
python train_v3.py --ablation C # no class-balance sampling
python train_v3.py --ablation D # no augmentationRequirements: Python 3.10+, PyTorch (tested 2.6.0+cu124), librosa, soundfile, pandas, numpy. Full run β 1 min on Tesla P4 / ~2 min on CPU.
βββ train_v2.py # v2 training: multi-location masked-attention fusion
βββ train_v3.py # v3: 70/15/15 held-out split + ablation switches + official challenge metric
βββ train_v4.py # v4: longer training + decision-threshold tuning (beats champion 0.780)
βββ train_v5.py # 2.26M-param encoder (ablation β overfits)
βββ train_focal.py # focal-loss ablation
βββ train_kfold.py # 5-fold patient-stratified CV (honest generalization)
βββ train_w2v_head.py # frozen wav2vec2 features + head (ablation)
βββ extract_w2v.py # wav2vec2 feature extraction (ablation)
βββ eval_probs.py # save val/test probabilities per model (for ensembling)
βββ ensemble_v4.py # probability-average ensemble
βββ ensemble_topk.py # val-selected top-k ensemble β best 0.7926
βββ tune_ensemble.py # val-tuned decision thresholds (dP/dU) for s_murmur
βββ tune_ensemble_w.py # val-weighted ensemble variant
βββ tune_offsets.py # OOF decision-offset headroom analysis
βββ EXPERIMENTS.md # v3 full experimental report (test eval, benchmark, ablations)
βββ EXPERIMENTS_v4.md # v4 iteration log (20 runs, 0.70 β 0.7815)
βββ exp_results.json # v3 aggregated metrics
βββ exp_results_v4.json # v4 aggregated metrics
βββ v3_split_seed42.json # persisted patient-level split (reproducibility)
βββ experiments/ # committed artifacts: models/, probs/, logs/, results/ (see Reproducibility)
βββ data/ # dataset (gitignored)
βββ models/ # checkpoints (gitignored)
βββ app/ # QC companion backend (FastAPI): qc_engine.py, main.py, model_defs.py
βββ web/ # QC companion frontend (Vite + React): recording QC + teaching simulator
βββ train_qc_models.py # train position (AV/PV/TV/MV) + murmur heads on real CirCor data
βββ build_assets.py # curate real expert-annotated clips for the teaching simulator
βββ start.sh # start backend + frontend with reverse proxy
βββ notebooks/ # exploratory notebooks
βββ scripts/ # utility scripts
βββ README.md
The model is only the middle of the story. The real pain point in primary care is that clinicians don't record consistently β bad recordings produce dirty datasets. This companion tool closes the βlast mileβ:
- Recording protocol guidance + real-time QC (
app/+web/): during recording, the system evaluates signal level, clipping, SNR, spectral flatness, cardiac rhythm (S1/S2 envelope autocorrelation), and auscultation-position consistency (a CNN classifier checks AV/PV/TV/MV). Every metric is computed from real DSP signal processing; thresholds follow the heart-sound literature (25-400 Hz bandpass, SNR β₯ 12 dB good, duration β₯ 8 s, etc.). - Teaching simulator: built-in real, expert-annotated heart-sound recordings from the CirCor dataset (4 positions Γ murmur absent/unknown/present), with playback, live spectrum waterfall, and a practice mode (answers hidden for self-testing). All recordings and labels are real β no mock data anywhere.
- Companion model (
train_qc_models.py): a single encoder with two heads (4-class position + 3-class murmur), reusing the v3 patient-disjoint split with an honest held-out test evaluation.
Running (frontend/backend split, frontend reverse-proxies /api to the backend):
Recording QC β live metrics, position verification and murmur screening:
Full QC report after uploading a real WAV (real CirCor recording: SNR 11.7 dB, heart rate 189 bpm, position TV at 86% confidence):
Teaching simulator β 107 real expert-annotated recordings, 4 positions Γ 3 murmur classes:
Spectrum waterfall while playing (practice mode hides the answer):
Clinical reference β auscultation landmarks, recording quality criteria, and held-out test metrics of the deployed models:
# 1. Download CirCor and train the companion model (real data)
wget https://physionet.org/static/published-projects/circor-heart-sound/circor-heart-sound-1.0.3.zip
unzip circor-heart-sound-1.0.3.zip -d ./data/circor
python train_qc_models.py --data-csv ./data/circor/training_data.csv --data-dir ./data/circor/training_data --workdir ./qc_work
python build_assets.py --data-csv ./data/circor/training_data.csv --data-dir ./data/circor/training_data --out ./app/assets
# The backend looks for the model in models/qc_models.pt or qc_work/qc_models.pt
# (train_qc_models.py writes to --workdir by default, i.e. qc_work/qc_models.pt;
# you can also copy it to models/ manually)
# 2. Start (backend :3001 + frontend :5173)
./start.shThe QC algorithms are based entirely on real signal processing β no synthetic values are produced; the teaching simulator uses only real recordings. Position verification and murmur screening are assistive/screening in nature and do not constitute a diagnosis; both the UI and the report carry disclaimers.
- v4's top results are single-seed + val-tuned thresholds; the 0.7815 edge over the champion (0.780) is ~0.0015 β three independent configs tie there, but seed variance is huge (0.60β0.78), so treat it as "champion-level", not "clearly better".
- Test split is a self-made 15% of the public cohort, not the Challenge's hidden 40% (includes unseen patients) β numbers are indicative, not an official submission.
- Unknown recall varies wildly by seed (0.4β0.8); only 10 Unknown patients in val make reliable selection hard.
- Reyna et al., Heart murmur detection from phonocardiogram recordings: The George B. Moody PhysioNet Challenge 2022, PLOS Digital Health, 2023.
- Oliveira et al., The CirCor DigiScope dataset: from murmur detection to heart sound classification, 2021.
MIT (code). Dataset: Open Data Commons Attribution License v1.0 (see PhysioNet).




