A lightweight collection of research utilities for computational pathology and medical image analysis.
A compact collection of reusable utilities for common computational pathology and medical image analysis workflows.
-
Patient-level dataset splitting
- stratified train/validation/test splits
- prevents patient leakage across subsets
- validates patient-level label consistency
-
Medical classification metrics
- AUROC and AUPRC
- sensitivity, specificity, precision, NPV, F1, accuracy
- confusion-matrix counts
- fixed-threshold evaluation
-
Bootstrap confidence intervals
- generic bootstrap utility
- convenience wrapper for binary classification metrics
- deterministic with a random seed
- explicit input-row sampling semantics
-
Patch / attention heatmaps
- rasterize patch coordinates and attention scores
- optional score normalization
- overlay heatmaps on thumbnails or other RGB images
-
WSI helpers
- slide metadata
- thumbnails
- optional OpenSlide support
- standard-image fallback through Pillow
Clone the repository and install it in editable mode:
git clone https://github.com/zms618/pathology-toolkit.git
cd pathology-toolkit
pip install -e .For development:
pip install -e ".[dev]"
pytestFor whole-slide image support with OpenSlide:
pip install -e ".[wsi]"
openslide-pythonrequires the OpenSlide shared library to be available on your system. See the OpenSlide documentation for platform-specific installation instructions.
import pandas as pd
from pathology_toolkit.data import patient_stratified_split
df = pd.DataFrame(
{
"patient_id": ["p1", "p1", "p2", "p3", "p4", "p5", "p6", "p7"],
"slide_id": ["s1", "s2", "s3", "s4", "s5", "s6", "s7", "s8"],
"label": [0, 0, 0, 1, 1, 0, 1, 1],
}
)
train_df, val_df, test_df = patient_stratified_split(
df,
patient_col="patient_id",
label_col="label",
train_size=0.6,
val_size=0.2,
test_size=0.2,
random_state=42,
)No patient will appear in more than one subset.
import numpy as np
from pathology_toolkit.metrics import binary_classification_metrics
y_true = np.array([0, 0, 1, 1, 1])
y_score = np.array([0.10, 0.35, 0.55, 0.80, 0.95])
metrics = binary_classification_metrics(
y_true,
y_score,
threshold=0.5,
)
for name, value in metrics.items():
print(name, value)from pathology_toolkit.metrics import bootstrap_binary_metric
result = bootstrap_binary_metric(
y_true,
y_score,
metric="auroc",
n_bootstrap=2000,
confidence=0.95,
random_state=42,
)
print(result)
# BootstrapResult(estimate=..., lower=..., upper=..., ...)bootstrap_binary_metric treats input rows as independent sampling units. For datasets
with multiple predictions per patient, aggregate predictions to the intended analysis
unit or use a cluster-aware bootstrap.
import numpy as np
from pathology_toolkit.visualization import rasterize_patch_scores
coords = np.array(
[
[0, 0],
[256, 0],
[0, 256],
[256, 256],
]
)
scores = np.array([0.1, 0.9, 0.4, 0.7])
heatmap = rasterize_patch_scores(
coords,
scores,
canvas_size=(512, 512),
patch_size=(256, 256),
)Coordinates use (x, y) pixel convention and canvas_size uses (width, height).
from pathology_toolkit.wsi import get_slide_info, make_thumbnail
info = get_slide_info("example.svs")
print(info)
thumb = make_thumbnail("example.svs", max_size=(1024, 1024))
thumb.save("thumbnail.png")The same WSI helpers are available from the command line:
pathology-toolkit slide-info example.svs
pathology-toolkit thumbnail example.svs thumbnail.png --max-width 1024 --max-height 1024pathology-toolkit/
├── src/pathology_toolkit/
│ ├── data/
│ │ └── split.py
│ ├── metrics/
│ │ ├── bootstrap.py
│ │ └── classification.py
│ ├── visualization/
│ │ └── heatmap.py
│ ├── wsi/
│ │ └── slide.py
│ └── cli.py
├── examples/
├── tests/
├── .github/workflows/
├── pyproject.toml
└── LICENSE
This project follows a few simple rules:
- Patient leakage should be difficult to introduce accidentally.
- Medical-image evaluation should expose clinically meaningful quantities, not only accuracy.
- Utilities should be small enough to read and adapt for research code.
- Optional WSI dependencies should remain optional.
- Examples and tests are part of the implementation, not an afterthought.
Issues and pull requests are welcome, especially for small utilities that improve reproducibility in computational pathology workflows.
MIT License. See LICENSE.
Cheng He
M.S. student at Fuzhou University
Computational Pathology & Medical Image Analysis
- Homepage: https://zms618.github.io/
- GitHub: https://github.com/zms618