Research project investigating semi-supervised learning (SSL) for nuclei segmentation in H&E-stained melanoma histopathology images under limited annotation.
The project evaluates four SSL approaches—Mean Teacher, GAN-based SSL, autoencoder pretraining, and pseudo-labeling—against supervised-only baselines at different levels of label scarcity.
Expert annotation of histopathology images is expensive. This project investigates whether unlabeled images can provide useful training signal when only a small labeled set is available.
The study is motivated by recent work on melanoma detection through nuclei segmentation, particularly Akbarpour et al. (2025).
Experiments primarily use the PUMA (Panoptic segmentation of nUclei and tissue in MelanomA) dataset:
- 205 melanoma H&E ROIs
- 1024×1024 resolution
- 40× magnification
- Expert nuclei annotations
The dataset is not included in this repository. It can be obtained from the PUMA dataset and Zenodo.
An additional out-of-domain experiment uses TCGA-SKCM whole-slide images. Eight slides were used to extract 320 tissue patches for unlabeled training.
- Supervised U-Net: full-data and scarce-label baselines
- Mean Teacher: consistency regularization with in-domain PUMA and out-of-domain TCGA unlabeled data
- GAN-SSL: adversarial semi-supervised segmentation
- Autoencoder pretraining: reconstruction-based encoder pretraining with single and differential learning rates
- Pseudo-labeling: confidence-thresholded self-training
- Data augmentation and TTA: evaluated for the full-data baseline
| Experiment | Validation Dice |
|---|---|
| Supervised U-Net (205 images) | 0.9038 |
| + Data augmentation | 0.9093 |
| + Test-time augmentation | 0.9072 |
| Method | Val Dice | vs. supervised |
|---|---|---|
| Supervised-only | 0.8702 | — |
| GAN-SSL | 0.8693 | -0.0009 |
| Mean Teacher, in-domain, mild aug | 0.8661 | -0.0041 |
| Mean Teacher, TCGA, strong aug | 0.8454 | -0.0248 |
| Autoencoder, differential LR | 0.8521 | -0.0181 |
| Mean Teacher, in-domain, strong aug | 0.8323 | -0.0379 |
| Autoencoder, single LR | 0.7692 | -0.1010 |
| Pseudo-labeling | N/A | 0 confident pseudo-labels |
| Method | Val Dice | vs. supervised |
|---|---|---|
| Supervised-only | 0.8903 | — |
| Mean Teacher | 0.8883 | -0.0020 |
| GAN-SSL | 0.8872 | -0.0031 |
- No SSL method clearly outperformed the corresponding supervised-only baseline.
- GAN-SSL was the closest to supervised performance at 15 labeled images, with only a 0.0009 Dice difference.
- Stronger Mean Teacher augmentation reduced performance at 15 labels.
- Replacing in-domain PUMA unlabeled data with TCGA-SKCM data improved Mean Teacher from 0.8323 to 0.8454, although it remained below the supervised baseline.
- Differential learning rates substantially improved autoencoder-pretrained segmentation compared with a single learning rate.
- Pseudo-labeling produced 0 confident pseudo-labels from 165 unlabeled images at the 15-label setting.
These results suggest that the effectiveness of SSL depends strongly on the learning mechanism, label availability, augmentation strategy, and domain of the unlabeled data.
At the supervisor's request, a simple baseline model was also built on a second dataset (TCGA-SKCM) to demonstrate the pipeline generalizes beyond PUMA. Since TCGA patches have no nuclei masks, a patch-level classification task was used instead — primary vs. metastatic melanoma tissue — with labels taken directly from each slide's TCGA barcode (no manual annotation needed).
A ResNet18 classifier was trained and evaluated with a properly stratified slide-level split (both classes represented in validation). Across 8 slides and later 30 slides (15 primary, 15 metastatic), with and without augmentation, validation accuracy consistently stayed near chance level (~50-56%) for this balanced binary task, rather than improving. This is treated as a genuine, diagnosed negative result rather than a bug: with a small slide count, the model most likely learns slide-specific staining/scanner characteristics instead of generalizable biological signal, and primary vs. metastatic melanoma are known to be structurally similar in whole-slide histology at the patch level. Full detail and all four attempts are logged in results/10_tcga_classification_baseline.txt.
code/
├── baseline_nuclei_segmentation.py
├── kaggle_baseline_nuclei_segmentation.py
├── baseline_full_augmented.py
├── supervised_only_comparison.py
├── mean_teacher_ssl.py
├── mean_teacher_tcga.py
├── gan_ssl.py
├── pseudo_labeling_ssl.py
├── autoencoder_pretrain_ssl.py
├── autoencoder_diff_lr.py
├── augmented_dataset.py
├── tta_evaluation.py
├── full_metrics_evaluation.py
├── tcga_download.py
├── tcga_patch_extraction.py
└── tcga_classification_baseline.py
results/
├── 01_baseline_full_supervised.txt
├── 02_scarce_label_comparison_15.txt
├── 03_scarce_label_comparison_40.txt
├── 04_full_metrics.txt
├── 05_augmentation_and_tta.txt
├── 06_mean_teacher_tcga_out_of_domain.txt
├── 07_mean_teacher_indomain_stronger_aug.txt
├── 08_gan_ssl.txt
├── 09_pseudo_labeling.txt
└── 10_tcga_classification_baseline.txt
Install dependencies:
pip install -r requirements.txtDownload the PUMA dataset from the official sources and arrange the images and GeoJSON annotations according to the project structure.
For the TCGA extension, run:
python code/tcga_download.py
python code/tcga_patch_extraction.pyIndividual training and evaluation scripts are provided in code/, with corresponding experiment logs in results/.
- PUMA is a public dataset and is not identical to a private clinical dataset.
- TCGA and PUMA may differ in staining, scanner characteristics, patient population, and image scale.
- TCGA patch magnification/microns-per-pixel was not explicitly normalized to PUMA.
- Current comparisons are primarily based on individual training runs; multiple seeds would provide stronger statistical evidence.
- Akbarpour et al. (2025) — nuclei segmentation and melanoma detection
- Alheejawi et al. (2021) — nuclei segmentation for melanoma analysis
- Yu et al. (2021) — Mean Teacher / semi-supervised learning
- Hung et al. (2018) — adversarial semi-supervised segmentation
- Requa et al. (2023) — semi-supervised skin-neoplasm detection