In this project, we aim to reproduce the Siamese Masked AutoEncoder (SiamMAE) proposed in Siamese Masked Autoencoders using the PyTorch framework. The dataset we will be pretraining our model will be UCF-101 UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild. The representation learned by the model will be evaluated on an object segmentation task on DAVIS-2017 The 2017 DAVIS Challenge on Video Object Segmentation.
This project was part of the course DD2412 Deep Learning, Advanced Course at KTH.
Please refer to the notebook example to see how to run the experiments. We could not upload our trained model to the Github due to storage limiations.
Example of results from object segmentation. The first video was an easy example whereas the second was a harder one. For more details, refer to our project report.
