Skip to content

About

Lightweight research utilities for computational pathology and medical image analysis.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

pathology-toolkit

A lightweight collection of research utilities for computational pathology and medical image analysis.

A compact collection of reusable utilities for common computational pathology and medical image analysis workflows.

Features

  • Patient-level dataset splitting

    • stratified train/validation/test splits
    • prevents patient leakage across subsets
    • validates patient-level label consistency
  • Medical classification metrics

    • AUROC and AUPRC
    • sensitivity, specificity, precision, NPV, F1, accuracy
    • confusion-matrix counts
    • fixed-threshold evaluation
  • Bootstrap confidence intervals

    • generic bootstrap utility
    • convenience wrapper for binary classification metrics
    • deterministic with a random seed
    • explicit input-row sampling semantics
  • Patch / attention heatmaps

    • rasterize patch coordinates and attention scores
    • optional score normalization
    • overlay heatmaps on thumbnails or other RGB images
  • WSI helpers

    • slide metadata
    • thumbnails
    • optional OpenSlide support
    • standard-image fallback through Pillow

Installation

Clone the repository and install it in editable mode:

git clone https://github.com/zms618/pathology-toolkit.git
cd pathology-toolkit
pip install -e .

For development:

pip install -e ".[dev]"
pytest

For whole-slide image support with OpenSlide:

pip install -e ".[wsi]"

openslide-python requires the OpenSlide shared library to be available on your system. See the OpenSlide documentation for platform-specific installation instructions.

Quick Start

1. Patient-level stratified split

import pandas as pd
from pathology_toolkit.data import patient_stratified_split

df = pd.DataFrame(
    {
        "patient_id": ["p1", "p1", "p2", "p3", "p4", "p5", "p6", "p7"],
        "slide_id": ["s1", "s2", "s3", "s4", "s5", "s6", "s7", "s8"],
        "label": [0, 0, 0, 1, 1, 0, 1, 1],
    }
)

train_df, val_df, test_df = patient_stratified_split(
    df,
    patient_col="patient_id",
    label_col="label",
    train_size=0.6,
    val_size=0.2,
    test_size=0.2,
    random_state=42,
)

No patient will appear in more than one subset.

2. Binary medical-image classification metrics

import numpy as np
from pathology_toolkit.metrics import binary_classification_metrics

y_true = np.array([0, 0, 1, 1, 1])
y_score = np.array([0.10, 0.35, 0.55, 0.80, 0.95])

metrics = binary_classification_metrics(
    y_true,
    y_score,
    threshold=0.5,
)

for name, value in metrics.items():
    print(name, value)

3. Bootstrap confidence interval

from pathology_toolkit.metrics import bootstrap_binary_metric

result = bootstrap_binary_metric(
    y_true,
    y_score,
    metric="auroc",
    n_bootstrap=2000,
    confidence=0.95,
    random_state=42,
)

print(result)
# BootstrapResult(estimate=..., lower=..., upper=..., ...)

bootstrap_binary_metric treats input rows as independent sampling units. For datasets with multiple predictions per patient, aggregate predictions to the intended analysis unit or use a cluster-aware bootstrap.

4. Patch-level attention heatmap

import numpy as np
from pathology_toolkit.visualization import rasterize_patch_scores

coords = np.array(
    [
        [0, 0],
        [256, 0],
        [0, 256],
        [256, 256],
    ]
)
scores = np.array([0.1, 0.9, 0.4, 0.7])

heatmap = rasterize_patch_scores(
    coords,
    scores,
    canvas_size=(512, 512),
    patch_size=(256, 256),
)

Coordinates use (x, y) pixel convention and canvas_size uses (width, height).

5. Slide metadata and thumbnail

from pathology_toolkit.wsi import get_slide_info, make_thumbnail

info = get_slide_info("example.svs")
print(info)

thumb = make_thumbnail("example.svs", max_size=(1024, 1024))
thumb.save("thumbnail.png")

The same WSI helpers are available from the command line:

pathology-toolkit slide-info example.svs
pathology-toolkit thumbnail example.svs thumbnail.png --max-width 1024 --max-height 1024

Repository Structure

pathology-toolkit/
├── src/pathology_toolkit/
│   ├── data/
│   │   └── split.py
│   ├── metrics/
│   │   ├── bootstrap.py
│   │   └── classification.py
│   ├── visualization/
│   │   └── heatmap.py
│   ├── wsi/
│   │   └── slide.py
│   └── cli.py
├── examples/
├── tests/
├── .github/workflows/
├── pyproject.toml
└── LICENSE

Design Principles

This project follows a few simple rules:

  1. Patient leakage should be difficult to introduce accidentally.
  2. Medical-image evaluation should expose clinically meaningful quantities, not only accuracy.
  3. Utilities should be small enough to read and adapt for research code.
  4. Optional WSI dependencies should remain optional.
  5. Examples and tests are part of the implementation, not an afterthought.

Contributing

Issues and pull requests are welcome, especially for small utilities that improve reproducibility in computational pathology workflows.

License

MIT License. See LICENSE.

Author

Cheng He
M.S. student at Fuzhou University
Computational Pathology & Medical Image Analysis

About

Lightweight research utilities for computational pathology and medical image analysis.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages