Skip to content
zms618Public

About

Official implementation of TG-WSPIS (ICME 2026)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

TG-WSPIS

Official repository for TG-WSPIS: Multi-Granular Text-Guided Weakly Supervised Pathology Image Segmentation, accepted at the 2026 IEEE International Conference on Multimedia & Expo (ICME 2026).

[Paper] [Poster] [Project Page]

Overview

Pixel-level annotations for pathology image segmentation are expensive and time-consuming to obtain. Conventional weakly supervised approaches typically rely on image-level labels or generic class prompts, which provide limited information about heterogeneous cellular distributions.

TG-WSPIS introduces an image-specific, multi-granular text supervision paradigm for weakly supervised pathology image segmentation. Each image is paired with three complementary descriptions for every cell or tissue type:

  • Global proportion, describing overall abundance;
  • Spatial distribution, describing where the target appears;
  • Fused information, jointly describing abundance and location.

Overview of the TG-WSPIS framework

The framework contains two central components:

  1. Pathology Multimodal Cell Consensus (PMCC) provides soft, quantifiable alignment between pathology images and structured textual descriptions, capturing partial semantic agreement across cell types, locations, and abundance levels.
  2. Gaussian mutual-information collaborative learning transfers complementary knowledge from the global-proportion and spatial-distribution branches to the fused branch, improving cross-granularity consistency and pseudo-mask quality.

The learned image-text alignment is converted into class-specific attention maps and refined pseudo-masks for downstream segmentation.

Benchmarks

We construct two text-guided weakly supervised pathology segmentation benchmarks by augmenting public datasets with image-specific, multi-granular descriptions:

  • BCSS-T3G, derived from BCSS (4 tissue categories: tumor, stroma, lymphocytes, necrosis);
  • WSSS4LUAD-T3G, derived from WSSS4LUAD (3 tissue categories: tumor, stroma, normal tissue).

Both benchmarks use WSI-level data splits to prevent leakage between training, validation, and test sets.

Main Results

TG-WSPIS achieves 76.79% mIoU on BCSS-T3G and 75.35% mIoU on WSSS4LUAD-T3G, outperforming the compared weakly supervised segmentation methods in the paper. A U-Net trained from TG-WSPIS pseudo-labels reaches 77.26% mIoU on BCSS-T3G, indicating that the generated masks provide useful dense supervision.

Please refer to the paper for complete experimental settings, comparisons, ablations, and limitations.

Repository Structure

TG-WSPIS/
├── CLIP/                    # Vendored OpenAI CLIP (installed as a local package)
├── src/                     # Core source code
│   ├── config.py            # Training / model / loss configurations
│   ├── datasets.py          # T3G dataset loading (images, masks, captions, targets)
│   ├── clip_encoder.py      # CLIP feature extraction helpers
│   ├── trainer.py           # Training loop with mutual learning
│   ├── evaluator.py         # Validation, heatmap generation, checkpoint search
│   ├── generator.py         # Pseudo-label / segmentation map generation
│   ├── losses.py            # Quantity, hot-zone, fusion, and MI losses
│   ├── metrics.py           # Precision / Recall / Dice / IoU
│   ├── heatmap.py           # Grad-Eclip / Grad-CAM attention maps
│   ├── visualization.py     # Visualization utilities
│   ├── utils.py             # Device, checkpoint, and naming helpers
│   └── tools/               # Data preparation and analysis scripts
├── train.py                 # Training entry point
├── val.py                   # Validation entry point (best-checkpoint search)
├── gen_main.py              # Pseudo-mask generation entry point
├── test_dataset.py          # Dataset sanity-check scripts
└── test_dataset_simple.py

Installation

  1. Create a Conda environment:
conda create -n TG-WSPIS python=3.10.16
conda activate TG-WSPIS
  1. Install Python dependencies and the vendored CLIP package:
pip install -r requirements.txt
cd CLIP
pip install -e .
cd ..

This installs PyTorch, the imaging stack, and CLIP together with its tokenizer dependencies (ftfy, regex, tqdm). For GPU training, install the PyTorch CUDA build matching your system from pytorch.org.

Data and Pretrained Checkpoints

The shared archive includes both T3G benchmarks (images, masks, and three-granular captions) and pretrained checkpoints:

After downloading, place the data under ./data/T3G (this directory is git-ignored):

data/T3G/
├── BCSS-T3G/
│   ├── train/{img,mask}  +  train.json
│   ├── val/{img,mask}    +  val.json
│   └── test/{img,mask}   +  test.json
└── WSSS4LUAD-T3G/
    ├── train/{img,mask}  +  train.json
    ├── val/{img,mask}    +  val.json
    └── test/{img,mask}   +  test.json

Each JSON file stores the three granular captions used for text supervision: global proportion, spatial distribution, and fused information.

Place the pretrained weights under the checkpoint directory configured in src/config.py. The validation and pseudo-mask generation entry points load checkpoints from ./ckpt by default.

Usage

All entry points read their settings from the configuration classes in src/config.py (dataset name, epochs, batch size, learning rate, device, checkpoint directory, etc.). Edit the defaults there or pass a custom configuration before launching.

  • Training (mutual learning over the three caption granularities):
python train.py
  • Validation (evaluates saved checkpoints and reports the best mIoU):
python val.py
  • Pseudo-mask generation (produces segmentation maps from a trained model):
python gen_main.py

Checkpoints are saved under ./ckpt/{dataset_name}/{description}/ per epoch; the best model is selected by validation mIoU.

Authors

Cheng He, Yunhang Shen, Sheng Lian, Jipeng Wu, Chunxu Yang, Fufeng Chen, Xuri Ge, and Fuhai Chen.

Third-Party Notice

Third-party components retain their original licenses. The vendored CLIP/ directory follows the original MIT license; see CLIP/LICENSE.

Contact

For questions about the project, please contact Cheng He.

About

Official implementation of TG-WSPIS (ICME 2026)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages