This repository contains the official code and pretrained model weights for the paper:
“Towards Effective Surgical Representation Learning with DINO Models”
(Accepted for Medical Imaging with Deep Learning (MIDL) 2026)
- Python 3.9+
- torch
- timm
# Create a new conda environment with Python 3.9
conda create -n SurgeNetDINO python=3.9 -y
# Activate the environment
conda activate SurgeNetDINO
# Install PyTorch and timm
conda install pytorch torchvision torchaudio pytorch-cuda=12.1 -c pytorch -c nvidia
pip install timmAfter installations, you can use load_weights.py to load the pretrained DINO models.
Alternatively, you can download the model weights using the provided links below.
| Model | Variant | Download |
|---|---|---|
| DINOv1 | ViT-s | Download |
| DINOv1 | ViT-b | Download |
| DINOv2 | ViT-s | Download |
| DINOv2 | ViT-b | Download |
| DINOv2 | ViT-l | Download |
| DINOv3 | ViT-s | Download |
| DINOv3 | ViT-b | Download |
| DINOv3 | ViT-l | Download |
The pretraining configurations can be found in dinov1_configs, dinov2_configs, and dinov3_configs. Note that for dinov1, we provide a .sh file instead of a configuration file following the original implementation. For further pretraining instructions, please refer to the original implementations of DINO: DINOv1, DINOv2, and DINOv3.
- Semantic segmentation: The code for finetuning the pretrained models on semantic segmentation can be found in
segmentation_finetuning, which is based on EoMT. Additional Python packages are required, which can be installed withpip install -r segmentation_finetuning/requirements.txt. - Surgical phase recognition: For finetuning on surgical phase recognition, please refer to SurgPhaseBench.
Dice score for semantic segmentation, accuracy (%) for surgical phase recognition, and online inference efficiency on a single NVIDIA H100 GPU (parameters of the segmentation models). Best results per DINO version are shown in bold.
| DINO | Model | Pretraining | Segmentation (Dice ↑) | Phase recognition (Acc ↑) | Efficiency | ||||
|---|---|---|---|---|---|---|---|---|---|
| CholecSeg8k | RAMIE-seg | AutoLaparo | RAMIE-phase | Params (M) | Seg. FPS | Phase FPS | |||
| v1 | ViT-S | ImageNet | 0.66 | 0.57 | 81.8 | 74.1 | 23 | 222 | 104 |
| SurgeNetXL | 0.73 | 0.62 | 85.3 | 77.4 | |||||
| ViT-B | ImageNet | 0.68 | 0.61 | 81.1 | 75.4 | 92 | 213 | 89 | |
| SurgeNetXL | 0.73 | 0.67 | 85.0 | 77.2 | |||||
| v2 | ViT-S | LVD-142M | 0.70 | 0.60 | 83.0 | 76.6 | 23 | 227 | 92 |
| SurgeNetXL | 0.66 | 0.59 | 84.5 | 76.0 | |||||
| ViT-B | LVD-142M | 0.71 | 0.67 | 85.3 | 77.4 | 90 | 224 | 75 | |
| SurgeNetXL | 0.75 | 0.73 | 85.9 | 78.2 | |||||
| ViT-L | LVD-142M | 0.63 | 0.70 | 86.1 | 77.9 | 311 | 123 | 35 | |
| SurgeNetXL | 0.77 | 0.79 | 84.1 | 80.8 | |||||
| v3 | ViT-S | LVD-1689M | 0.73 | 0.63 | 81.1 | 75.4 | 23 | 153 | 69 |
| SurgeNetXL | 0.74 | 0.69 | 82.2 | 74.4 | |||||
| ViT-B | LVD-1689M | 0.71 | 0.67 | 83.4 | 75.8 | 92 | 147 | 60 | |
| SurgeNetXL | 0.75 | 0.73 | 83.3 | 76.3 | |||||
| ViT-L | LVD-1689M | 0.76 | 0.70 | 86.0 | 77.8 | 314 | 93 | 31 | |
| SurgeNetXL | 0.78 | 0.74 | 86.4 | 78.0 | |||||
If you use these models or the dataset in your work, please cite our paper:
@inproceedings{
jong2026towards,
title={Towards Effective Surgical Representation Learning with {DINO} Models},
author={Ronald L.P.D. de Jong and Yiping Li and Tim J. M. Jaspers and Romy C. van Jaarsveld and Gino M. Kuiper and Franco Badaloni and Richard van Hillegersberg and Jelle P. Ruurda and Fons van der Sommen and Josien P.W. Pluim and Marcel Breeuwer},
booktitle={Medical Imaging with Deep Learning},
year={2026},
url={https://openreview.net/forum?id=6FoIDPKzRV}
}For questions or issues regarding this repository, please contact the corresponding author:
Ronald L.P.D. de Jong
Email: r.l.p.d.d.jong@tue.nl
- Code: MIT — see LICENSE (permissive; commercial use permitted).
- Pretrained model weights: CC-BY-NC-SA — non-commercial share-alike. The weights and any derivative models that include these weights are NOT cleared for commercial use. See LICENSE_MODELS for details and the precise license text.
We would like to thank the authors and maintainers of SurgeNetXL and the original DINO repositories for making their work publicly available:
Their open-source contributions provided the foundation for our work on surgical representation learning.