Single-channel separation of mixed chest auscultation recordings into two sources: heart sound and lung sound.
This repository adapts SPMamba (built on the Look2Hear framework) to the HLS-CMDS heart-and-lung-sounds dataset, and adds a fixed-order (non-PIT) training objective so that channel assignment stays deterministic at inference time.
Attribution. The
look2hear/package is derived from Look2Hear / SPMamba and remains under the Apache License 2.0. Original copyright notices are kept intact. The contributions of this fork are the HLS-CMDS data pipeline, the fixed-order loss wrapper, the Colab T4 configuration, and the evaluation scripts. See LICENSE and NOTICE.
Standard speech-separation training uses Permutation Invariant Training (PIT), which is permutation-agnostic by design — the model may emit the two sources in either order. That is acceptable when both sources are speech, but not here: a clinician needs to know which output channel is the heart and which is the lung.
Because data_hls.h5 stores the two sources in a consistent order
(s1 = heart, s2 = lung), training with a fixed channel assignment removes
the permutation ambiguity entirely. At inference on a real mixture:
output channel 0 -> heart
output channel 1 -> lung
Full config: configs/spmamba-hls-t4.yml.
Tuned for a single Google Colab T4 (16 GB VRAM).
| Parameter | Value |
|---|---|
| Sample rate | 4 000 Hz |
| Segment length | 4.0 s |
STFT size (n_fft) |
512 |
Hop size (stride) |
128 |
| Window | Hann |
| Input dim | 257 |
Sources (n_srcs) |
2 (heart, lung) |
| Microphones | 1 |
| Audio normalisation | off |
| Parameter | Value |
|---|---|
| Layers | 6 |
| Embedding dim | 16 |
| Embedding kernel / hop | 4 / 1 |
| LSTM hidden units | 128 |
| Attention heads | 4 |
| Attention approx. QK dim | 512 |
| Activation | PReLU |
| Parameter | Value |
|---|---|
| Loss (train) | FixedOrderLossWrapper, pairwise_neg_snr |
| Loss (val) | FixedOrderLossWrapper, pairwise_neg_sisdr |
| Optimiser | Adam, lr 1e-3, weight decay 1e-6 |
| Scheduler | ReduceLROnPlateau, factor 0.5, patience 10, min lr 1e-6 |
| Batch size | 16 |
| Epochs | 100 (early stopping, patience 20 on val_loss) |
| Precision | 16-mixed |
| Experiment name | SPMamba-HLS-HeartLung-T4 |
SPMamba-HLS/
├── look2hear/ # framework: models, datas, losses, metrics, system
├── configs/
│ └── spmamba-hls-t4.yml # Colab T4 config for HLS-CMDS
├── build_hls_h5.py # HLS-CMDS -> data_hls.h5
├── test_hls.py # evaluation + sample audio export
├── audio_train.py # training entry point
├── notebooks/
│ ├── train_hls_colab.ipynb # end-to-end training on Colab
│ └── test_hls_colab.ipynb # evaluation from a Drive checkpoint
├── docs/
│ ├── PLAN_HLS_CMDS.md # data pipeline design notes
│ └── PLAN_TEST_HLS.md # evaluation protocol
├── env/look2hear.yml # micromamba / conda environment
├── LICENSE
├── NOTICE
└── README.md
The dataset is not distributed in this repository.
-
Obtain the HLS-CMDS heart-and-lung-sounds dataset.
-
Build the HDF5 bundle locally (a laptop is fine — this is CPU work):
python build_hls_h5.py \ --heart-src <dir/of/pure/heart/wavs> \ --lung-src <dir/of/pure/lung/wavs> \ --out data_hls.h5 \ --n-mix 1000 \ --snr-min -5.0 \ --snr-max 5.0 \ --seed 42
Mixtures are synthesised: pure heart and lung recordings are paired at random and summed at a heart-to-lung ratio drawn uniformly from −5 to +5 dB, then peak-normalised to 0.9. Waveforms are 60 000 samples at 4 000 Hz (15 s). Split fractions (
--val-frac 0.15,--test-frac 0.10) are applied source-disjoint and stored in/splitinside the H5.Alternatively, if you already have aligned mixture/heart/lung triplets:
python build_hls_h5.py --src <dir/of/M-H-L/wavs> --out data_hls.h5
-
Zip it and upload to Google Drive as
MyDrive/data_hls.zip. -
The Colab notebook extracts it to
/content/data_hls.h5.
The file carries its own source-disjoint /split dataset, meaning no
recording subject appears in more than one split. When /split is present the
datamodule honours it and ignores val_ratio / test_ratio / seed in the
config; those fields only apply to an H5 built without a split.
micromamba env create -f env/look2hear.yml -p ./spmamba-env
micromamba activate ./spmamba-envRequires a CUDA GPU. Verified on Colab T4 (16 GB) with precision: 16-mixed.
python audio_train.py \
--conf_dir configs/spmamba-hls-t4.yml \
--checkpoint_dir /content/drive/MyDrive/SPMamba_checkpointsPoint --checkpoint_dir at Google Drive so checkpoints survive a Colab
disconnect. To resume after a disconnect:
python audio_train.py \
--conf_dir configs/spmamba-hls-t4.yml \
--checkpoint_dir /content/drive/MyDrive/SPMamba_checkpoints \
--resume--resume picks up last*.ckpt in --checkpoint_dir, falling back to the
newest .ckpt. Use --resume_ckpt <path> to name one explicitly.
Weights & Biases logging is off unless requested:
python audio_train.py --conf_dir configs/spmamba-hls-t4.yml --wandb \
--wandb_project spmamba --wandb_run_name SPMamba-HLS-HeartLung-T4Add --wandb_offline to buffer runs locally and sync later.
python test_hls.py \
--ckpt <path/to/best.ckpt> \
--h5 /content/data_hls.h5 \
--conf_dir configs/spmamba-hls-t4.yml \
--outdir test_hls_out \
--save_k 3--ckpt is required and accepts either a .ckpt file or a directory (the
newest *best*.ckpt inside is picked). --save_k controls how many separated
samples are written to disk as WAV for listening.
Pretrained weights are published under Releases rather than committed to the repository.
Checkpoint epoch=31-best.ckpt, evaluated on the held-out test split
(25 mixtures, full-length, source-disjoint from train/val).
| Metric | Mean | Std |
|---|---|---|
| SI-SNRi | +4.04 dB | 2.13 |
| SDRi | +5.79 dB | 2.53 |
| SI-SNR (absolute) | +4.03 dB | 2.13 |
| SDR (absolute) | +5.88 dB | 2.52 |
| Value | |
|---|---|
Channel accuracy — output order matched (heart, lung) |
100 % (25/25) |
| PIT SI-SNRi − fixed-order SI-SNRi | +0.00 dB |
The permutation-invariant upper bound and the fixed-order score are identical on every test mixture: the model never needed a channel swap to reach its best score. Dropping PIT therefore costs nothing in separation quality while removing permutation ambiguity at inference — channel 0 is always the heart, channel 1 always the lung, with no post-hoc identification step.
Per-mixture SI-SNRi ranges from −0.14 dB to +7.87 dB. The two negative cases are mixtures where the model returns essentially the input unchanged.
| Split | Mixtures |
|---|---|
| Train | 911 |
| Validation | 64 |
| Test | 25 |
Splits are source-disjoint — no source recording contributing to a training mixture appears in a validation or test mixture. This costs sample count relative to a naive random split (the requested 0.15/0.10 fractions yield 64/25 rather than 150/100), but avoids leakage between splits.
Scope. With 25 test mixtures the confidence interval on these means is wide; the figures indicate feasibility rather than a benchmark result. Mixtures are synthetic (pure heart and lung recordings summed at −5 to +5 dB), so performance on real simultaneous chest recordings is untested.
If you use this work, please cite the original SPMamba paper and the HLS-CMDS dataset alongside this repository.
Apache License 2.0 — see LICENSE.