Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Semi-Supervised Melanoma Nuclei Segmentation

Research project investigating semi-supervised learning (SSL) for nuclei segmentation in H&E-stained melanoma histopathology images under limited annotation.

The project evaluates four SSL approaches—Mean Teacher, GAN-based SSL, autoencoder pretraining, and pseudo-labeling—against supervised-only baselines at different levels of label scarcity.

Motivation

Expert annotation of histopathology images is expensive. This project investigates whether unlabeled images can provide useful training signal when only a small labeled set is available.

The study is motivated by recent work on melanoma detection through nuclei segmentation, particularly Akbarpour et al. (2025).

Dataset

Experiments primarily use the PUMA (Panoptic segmentation of nUclei and tissue in MelanomA) dataset:

  • 205 melanoma H&E ROIs
  • 1024×1024 resolution
  • 40× magnification
  • Expert nuclei annotations

The dataset is not included in this repository. It can be obtained from the PUMA dataset and Zenodo.

An additional out-of-domain experiment uses TCGA-SKCM whole-slide images. Eight slides were used to extract 320 tissue patches for unlabeled training.

Methods

  • Supervised U-Net: full-data and scarce-label baselines
  • Mean Teacher: consistency regularization with in-domain PUMA and out-of-domain TCGA unlabeled data
  • GAN-SSL: adversarial semi-supervised segmentation
  • Autoencoder pretraining: reconstruction-based encoder pretraining with single and differential learning rates
  • Pseudo-labeling: confidence-thresholded self-training
  • Data augmentation and TTA: evaluated for the full-data baseline

Results

Full-data supervised baseline

Experiment Validation Dice
Supervised U-Net (205 images) 0.9038
+ Data augmentation 0.9093
+ Test-time augmentation 0.9072

15 labeled images

Method Val Dice vs. supervised
Supervised-only 0.8702 —
GAN-SSL 0.8693 -0.0009
Mean Teacher, in-domain, mild aug 0.8661 -0.0041
Mean Teacher, TCGA, strong aug 0.8454 -0.0248
Autoencoder, differential LR 0.8521 -0.0181
Mean Teacher, in-domain, strong aug 0.8323 -0.0379
Autoencoder, single LR 0.7692 -0.1010
Pseudo-labeling N/A 0 confident pseudo-labels

40 labeled images

Method Val Dice vs. supervised
Supervised-only 0.8903 —
Mean Teacher 0.8883 -0.0020
GAN-SSL 0.8872 -0.0031

Key Findings

  • No SSL method clearly outperformed the corresponding supervised-only baseline.
  • GAN-SSL was the closest to supervised performance at 15 labeled images, with only a 0.0009 Dice difference.
  • Stronger Mean Teacher augmentation reduced performance at 15 labels.
  • Replacing in-domain PUMA unlabeled data with TCGA-SKCM data improved Mean Teacher from 0.8323 to 0.8454, although it remained below the supervised baseline.
  • Differential learning rates substantially improved autoencoder-pretrained segmentation compared with a single learning rate.
  • Pseudo-labeling produced 0 confident pseudo-labels from 165 unlabeled images at the 15-label setting.

These results suggest that the effectiveness of SSL depends strongly on the learning mechanism, label availability, augmentation strategy, and domain of the unlabeled data.

TCGA-SKCM classification baseline (second dataset)

At the supervisor's request, a simple baseline model was also built on a second dataset (TCGA-SKCM) to demonstrate the pipeline generalizes beyond PUMA. Since TCGA patches have no nuclei masks, a patch-level classification task was used instead — primary vs. metastatic melanoma tissue — with labels taken directly from each slide's TCGA barcode (no manual annotation needed).

A ResNet18 classifier was trained and evaluated with a properly stratified slide-level split (both classes represented in validation). Across 8 slides and later 30 slides (15 primary, 15 metastatic), with and without augmentation, validation accuracy consistently stayed near chance level (~50-56%) for this balanced binary task, rather than improving. This is treated as a genuine, diagnosed negative result rather than a bug: with a small slide count, the model most likely learns slide-specific staining/scanner characteristics instead of generalizable biological signal, and primary vs. metastatic melanoma are known to be structurally similar in whole-slide histology at the patch level. Full detail and all four attempts are logged in results/10_tcga_classification_baseline.txt.

Repository Structure

code/
├── baseline_nuclei_segmentation.py
├── kaggle_baseline_nuclei_segmentation.py
├── baseline_full_augmented.py
├── supervised_only_comparison.py
├── mean_teacher_ssl.py
├── mean_teacher_tcga.py
├── gan_ssl.py
├── pseudo_labeling_ssl.py
├── autoencoder_pretrain_ssl.py
├── autoencoder_diff_lr.py
├── augmented_dataset.py
├── tta_evaluation.py
├── full_metrics_evaluation.py
├── tcga_download.py
├── tcga_patch_extraction.py
└── tcga_classification_baseline.py

results/
├── 01_baseline_full_supervised.txt
├── 02_scarce_label_comparison_15.txt
├── 03_scarce_label_comparison_40.txt
├── 04_full_metrics.txt
├── 05_augmentation_and_tta.txt
├── 06_mean_teacher_tcga_out_of_domain.txt
├── 07_mean_teacher_indomain_stronger_aug.txt
├── 08_gan_ssl.txt
├── 09_pseudo_labeling.txt
└── 10_tcga_classification_baseline.txt

Reproducibility

Install dependencies:

pip install -r requirements.txt

Download the PUMA dataset from the official sources and arrange the images and GeoJSON annotations according to the project structure.

For the TCGA extension, run:

python code/tcga_download.py
python code/tcga_patch_extraction.py

Individual training and evaluation scripts are provided in code/, with corresponding experiment logs in results/.

Limitations

  • PUMA is a public dataset and is not identical to a private clinical dataset.
  • TCGA and PUMA may differ in staining, scanner characteristics, patient population, and image scale.
  • TCGA patch magnification/microns-per-pixel was not explicitly normalized to PUMA.
  • Current comparisons are primarily based on individual training runs; multiple seeds would provide stronger statistical evidence.

Related Work

  • Akbarpour et al. (2025) — nuclei segmentation and melanoma detection
  • Alheejawi et al. (2021) — nuclei segmentation for melanoma analysis
  • Yu et al. (2021) — Mean Teacher / semi-supervised learning
  • Hung et al. (2018) — adversarial semi-supervised segmentation
  • Requa et al. (2023) — semi-supervised skin-neoplasm detection

About

Semi-supervised nuclei segmentation for melanoma histopathology — Mean Teacher, GAN, autoencoder, and pseudo-labeling vs. supervised baselines under label scarcity.

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages