Self-supervised training stack for the SLC-PFM competition using DINOv3 (ConvNeXt / ViT) with pathology-specific augmentations, dataset tooling for zip/WebDataset formats, W&B logging, and downstream linear-probe evaluation hooks.
- Pathology-focused data pipeline: Zip-based loader with timeout/retry, WebDataset streaming loader, optional small webp dataset helper.
- Augmentations patched into DINOv3: stain jitter/normalization-friendly color jitter, vertical + 90° rotations.
- Training wrappers: W&B logging, SSL diagnostics (prototype utilization, feature rank/std), auto-resume checkpoints, optional periodic downstream evaluation.
- Production presets for Fir HPC (11TB dataset) plus small validation and ImageNet baseline configs.
- Conversion/validation utilities: metadata generation from zips, WebDataset conversion scripts, dataset smoke tests.
src/train.py: Main DINOv3 launcher with pathology aug patch + W&B/diagnostics.src/train_with_eval.py: Adds periodic downstream linear-probe evaluation during SSL training.src/train_imagenet_baseline.py: ViT-S baseline on an ImageNet subset for loss comparison.src/datasets/: PathologyWebDataset (tar shards), small PathologyDataset helper (legacy), ImageNet subset dataset.src/pathology_augmentations.py,src/patch_dinov3_augmentations.py: Domain augmentations injected into DINOv3.src/prepare_imagenet_baseline.py: ImageNet subset metadata helper (baseline only).webdataset_shards/: Conversion + verification tools for tar shards.scripts/: SLURM launchers for Fir, dataset loading test, data transfer helpers.configs/: Training presets (validation, ImageNet baseline, production ViT-S/B).docs/: Runbooks, optimization notes, challenge specs.
- Python 3.11+, CUDA GPU recommended.
- Install dependencies (uv):
uv sync
- Environment:
WANDB_API_KEYfor logging (optional)..envmay includehuggingface_token; loaded automatically in training scripts.
-
WebDataset shards (current path):
- Convert zips to
shard-*.tarwithwebdataset_shards/convert_to_webdataset.py(configure input/output paths inside the script). - A
conversion_state.jsonorwebdataset_stats.jsonenables length reporting for the loader. - Verify streaming with
bash scripts/test_dataset_loading.sh(expectsDATA_DIR/chunk_*/shard-*.tar).
- Convert zips to
-
ImageNet subset (baseline comparison):
python src/prepare_imagenet_baseline.py --data-dir /path/to/imagenet_subset --output data_imagenet/metadata.json --max-samples 1000
-
Training is config-driven; set
train.dataset_pathin the chosen config to match an available loader (PathologyWebDatasetis registered indinov3).- Example (WebDataset):
train.dataset_path: PathologyWebDataset:root=/path/to/DATA_DIR:shard_glob=chunk_*/shard-*.tar:stats_file=/path/to/DATA_DIR/webdataset_stats.json
- Example (WebDataset):
-
Local/validation ConvNeXt-S:
uv run torchrun --nproc_per_node=1 src/train.py \ --config-file configs/train_convnext_small.yaml \ --output-dir outputs/convnext_small
(Edit
configs/train_convnext_small.yamlto pointtrain.dataset_pathat your data; CLI--data-path/--metadata-fileare informational only.) -
Production ViT-S on Fir (4×H100):
sbatch scripts/submit_fir_vit_small.sh
Script uses
configs/train_vit_small_production_fir.yaml(PathologyWebDataset, bf16, torch.compile) and auto-resubmits until.training_completeexists. -
ImageNet baseline:
uv run torchrun --nproc_per_node=1 src/train_imagenet_baseline.py \ --config-file configs/train_vit_small_imagenet_baseline.yaml \ --metadata-file data_imagenet/metadata.json \ --output-dir outputs/imagenet_baseline
-
Train with periodic downstream eval (linear probe during SSL):
uv run torchrun --nproc_per_node=1 src/train_with_eval.py \ --config-file configs/train_vit_small_optimized.yaml \ --eval-every-n-iterations 1000 --eval-num-classes <NUM_CLASSES>
- W&B logging enabled when
WANDB_API_KEYis set; run IDs are derived fromoutput_dir(auto-resume friendly). - SSL diagnostics logged every 100 iters: prototype utilization/entropy, feature std, effective rank.
- Checkpoints saved under
<output_dir>/ckpt(period controlled by config);.training_completemarker created on success to support auto-continuation scripts.
- Standalone linear probe on a checkpoint:
(Use matching dataset strings for Zip-based data.)
uv run python src/run_downstream_eval.py \ --checkpoint outputs/<run>/ckpt/<iter>/teacher_checkpoint.pth \ --config-file configs/train_vit_small_production_fir.yaml \ --train-dataset "PathologyWebDataset:root=/path/to/DATA_DIR:shard_glob=chunk_*/shard-*.tar:stats_file=/path/to/DATA_DIR/webdataset_stats.json" \ --val-dataset "PathologyWebDataset:root=/path/to/DATA_DIR:shard_glob=chunk_*/shard-*.tar:stats_file=/path/to/DATA_DIR/webdataset_stats.json" \ --num-classes <NUM_CLASSES> \ --output-dir outputs/downstream_eval
scripts/test_dataset_loading.sh: quick streaming test for WebDataset shards.scripts/submit_fir_vit_small*.sh: SLURM launchers with auto-resume/backup.scripts/data_transfer/*.sh: HPC data sync helpers.
- DINOv3: https://github.com/facebookresearch/dinov3
- Challenge docs:
docs/challenge/ - Optimization/runbook:
docs/PRODUCTION_RUNBOOK.md,docs/SSL_OPTIMIZATION_ANALYSIS.md
This project follows the DINOv3 License for the model code and is subject to the SLC-PFM competition rules for dataset usage.