This repository contains an enhanced implementation of the DCASE 2025 Task 6 baseline system for language-based audio retrieval, featuring advanced contrastive learning techniques including multi-positive learning and hard negative mining.
This system implements a dual-encoder architecture for cross-modal audio-text retrieval using:
- Audio Encoder: PaSST (Patch-out Fast Spectrogram Transformer) with segment-based processing
- Text Encoder: RoBERTa-large with contextual embeddings
- Enhanced Contrastive Learning: Multi-positive learning and hard negative mining for improved performance
- Progressive Training: Staged activation of advanced techniques during training
The implementation builds upon the DCASE 2025 Task 6 baseline with significant enhancements for better performance.
- Multi-Positive Learning: Identifies semantically similar samples as soft positives with weighted loss
- Hard Negative Mining: Focuses training on challenging negative samples using multiple strategies
- Progressive Enhancement: Staged activation of advanced techniques by training epoch
- Mixed Precision Training: Automatic mixed precision for faster training and reduced memory usage
- Comprehensive Evaluation: Standard retrieval metrics (R@K, mAP) with multiple positive ground truth support
enhanced_audio_retrieval_dcase25_submission/
├── d25_t6/ # Core implementation
│ ├── __init__.py
│ ├── passt.py # Audio encoder (PaSST) with segmentation
│ ├── retrieval_module.py # Main model with enhanced contrastive learning
│ ├── train.py # Training pipeline with multi-dataset support
│ ├── predict.py # Inference and prediction generation
│ └── datasets/ # Dataset utilities and preprocessing
│ ├── audio_loading.py # Custom audio loading with 30s segments
│ ├── batch_collate.py # Batch collation for variable-length data
│ ├── download_datasets.py # Automated dataset downloading
│ └── utils.py # Dataset filtering and validation utilities
├── resources/ # Reference files and metadata
│ ├── dcase2025_task6_excluded_freesound_ids.csv
│ ├── example_predictions.csv
│ └── metadata_eval.csv
├── scripts/ # Utility scripts
│ └── convert_flac_to_mp3.py # WavCaps compression utility
├── requirements.txt # Python dependencies
└── README.md
- Linux (tested on Ubuntu 22.04)
- Python 3.11
- CUDA-compatible GPU (tested on NVIDIA L4 24GB)
- Miniconda
- Clone the repository
git clone https://github.com/saubhagyapandey27/enhanced_audio_retrieval_dcase25_submission
cd enhanced_audio_retrieval_dcase25_submission- Create conda environment
conda create -n d25_t6 python=3.11
conda activate d25_t6- Install system dependencies
# via conda
conda install -c conda-forge p7zip- Install PyTorch
# For CUDA 12.1+
pip3 install torch torchvision torchaudio
# For CUDA 11.8
pip3 install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118- Install dependencies
pip install -r requirements.txt- Configure Weights & Biases (optional)
wandb login
# Follow the prompt to enter your API key from https://wandb.ai/authorizepython -m d25_t6.train \
--data_path=data \
--batch_size=24 \
--batch_size_eval=24 \
--max_epochs=20 \
--compile \
--seed=13python -m d25_t6.train \
--data_path=data \
--batch_size=24 \
--batch_size_eval=24 \
--max_epochs=20 \
--multi_positive_learning \
--multi_positive_start_epoch=8 \
--hard_negative_mining \
--hard_negative_start_epoch=10 \
--soft_positive_threshold=0.75 \
--soft_positive_weight=0.3 \
--num_hard_negatives=5 \
--mining_strategy=hardest \
--compile \
--seed=13python -m d25_t6.train \
--audiocaps \
--wavcaps \
--data_path=data \
--batch_size=24 \
--batch_size_eval=24 \
--max_epochs=20 \
--multi_positive_learning \
--multi_positive_start_epoch=8 \
--hard_negative_mining \
--hard_negative_start_epoch=10 \
--compile \
--seed=967251Generate predictions for audio retrieval:
python -m d25_t6.predict \
--load_ckpt_path=checkpoints/experiment_name/epoch=19.ckpt \
--retrieval_audio_path=path/to/audio/files \
--retrieval_captions=path/to/queries.csv \
--predictions_path=path/to/outputEnhanced Learning Parameters:
--multi_positive_learning: Enable soft positive weighting based on semantic similarity--multi_positive_start_epoch: Epoch to activate multi-positive learning (default: 0)--soft_positive_threshold: Similarity threshold for soft positive identification (default: 0.75)--soft_positive_weight: Weight for soft positive pairs (default: 0.3)--hard_negative_mining: Enable hard negative mining strategy--hard_negative_start_epoch: Epoch to begin hard negative mining (default: 0)--num_hard_negatives: Number of hard negatives per positive (default: 5)--mining_strategy: Strategy for negative mining (hardest,semi_hard,random)
Model Configuration:
--roberta_base: Use RoBERTa-base instead of RoBERTa-large--s_patchout_t: Temporal patch dropout for PaSST (default: 15)--s_patchout_f: Frequency patch dropout for PaSST (default: 2)--initial_tau: Contrastive loss temperature (default: 0.05)
Training Setup:
--batch_size: Training batch size (default: 64)--batch_size_eval: Evaluation batch size (default: 64)--max_epochs: Maximum training epochs (default: 20)--compile: Enable torch.compile() optimization
Dataset Support:
--audiocaps: Include AudioCaps dataset in training--wavcaps: Include WavCaps dataset in training--ablate_clean_setup: Use clean dataset splits (exclude eval/test from training) (default: True)--data_path: Root directory for dataset storage (default: 'data')
- Segmentation: Long audio files (>10s) are split into overlapping 10s segments
- Feature Extraction: PaSST processes mel-spectrograms with configurable patch dropout
- Aggregation: Duration-based pooling combines segment embeddings
- Projection: Linear layer maps to 1024-dimensional embedding space
- Preprocessing: Lowercase conversion and punctuation removal
- Tokenization: RoBERTa tokenizer with 32-token maximum length
- Encoding: Contextualized embeddings from [CLS] token
- Projection: Linear layer maps to 1024-dimensional embedding space
- Multi-Positive Learning: Weighted InfoNCE loss with soft positive identification
- Hard Negative Mining: Focus on challenging negatives using similarity-based strategies
- Progressive Training: Staged activation prevents early optimization interference
Test Set Performance (Clotho Evaluation):
- R@1: 18.68%
- R@5: 44.77%
- R@10: 59.35%
- mAP@10: 30.01%
- mAP@16 (multiple positives): 34.68%
Training Configuration:
- GPU: NVIDIA L4 (24GB)
- Batch Size: 24 (train/eval)
- Multi-positive start: Epoch 8
- Hard negative start: Epoch 10
- Total epochs: 20
- Only the Clotho dataset was used due to computational constraints.
If you use this enhanced implementation, please cite the original baseline work it builds upon:
@inproceedings{Primus2024,
author = "Primus, Paul and Schmid, Florian and Widmer, Gerhard",
title = "Estimated Audio–Caption Correspondences Improve Language-Based Audio Retrieval",
booktitle = "Proceedings of the Detection and Classification of Acoustic Scenes and Events 2024 Workshop (DCASE2024)",
address = "Tokyo, Japan",
month = "October",
year = "2024",
pages = "121--125"
}Related Work:
@inproceedings{PaSST,
author = {Khaled Koutini and Jan Schl{\"u}ter and Hamid Eghbal{-}zadeh and Gerhard Widmer},
title = {Efficient Training of Audio Transformers with Patchout},
booktitle = {Interspeech 2022},
pages = {2753--2757},
year = {2022},
doi = {10.21437/Interspeech.2022-227}
}
@article{RoBERTa,
author = {Yinhan Liu and Myle Ott and Naman Goyal and Jingfei Du and Mandar Joshi and Danqi Chen and Omer Levy and Mike Lewis and Luke Zettlemoyer and Veselin Stoyanov},
title = {RoBERTa: A Robustly Optimized BERT Pretraining Approach},
journal = {CoRR},
volume = {abs/1907.11692},
year = {2019},
url = {http://arxiv.org/abs/1907.11692}
}
@inproceedings{Clotho,
author = {Konstantinos Drossos and Samuel Lipping and Tuomas Virtanen},
title = {Clotho: an Audio Captioning Dataset},
booktitle = {2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
pages = {736--740},
year = {2020},
doi = {10.1109/ICASSP40776.2020.9052990}
}
@inproceedings{NEURIPS2023_917cd410,
author = {Ray, Arijit and Radenovic, Filip and Dubey, Abhimanyu and Plummer, Bryan and Krishna, Ranjay and Saenko, Kate},
booktitle = {Advances in Neural Information Processing Systems},
editor = {A. Oh and T. Naumann and A. Globerson and K. Saenko and M. Hardt and S. Levine},
pages = {46433--46445},
publisher = {Curran Associates, Inc.},
title = {Cola: A Benchmark for Compositional Text-to-image Retrieval},
url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/917cd410aa55b61594fa2a6f6e5a9e94-Paper-Datasets_and_Benchmarks.pdf},
volume = {36},
year = {2023}
}This implementation follows the same licensing terms as the datasets: