Skip to content

Repository files navigation

Enhanced Audio-Text Retrieval with Multi-Positive Learning and Hard Negative Mining

This repository contains an enhanced implementation of the DCASE 2025 Task 6 baseline system for language-based audio retrieval, featuring advanced contrastive learning techniques including multi-positive learning and hard negative mining.

Overview

This system implements a dual-encoder architecture for cross-modal audio-text retrieval using:

  • Audio Encoder: PaSST (Patch-out Fast Spectrogram Transformer) with segment-based processing
  • Text Encoder: RoBERTa-large with contextual embeddings
  • Enhanced Contrastive Learning: Multi-positive learning and hard negative mining for improved performance
  • Progressive Training: Staged activation of advanced techniques during training

The implementation builds upon the DCASE 2025 Task 6 baseline with significant enhancements for better performance.

Key Features

  • Multi-Positive Learning: Identifies semantically similar samples as soft positives with weighted loss
  • Hard Negative Mining: Focuses training on challenging negative samples using multiple strategies
  • Progressive Enhancement: Staged activation of advanced techniques by training epoch
  • Mixed Precision Training: Automatic mixed precision for faster training and reduced memory usage
  • Comprehensive Evaluation: Standard retrieval metrics (R@K, mAP) with multiple positive ground truth support

Project Structure

enhanced_audio_retrieval_dcase25_submission/
├── d25_t6/                           # Core implementation
│   ├── __init__.py
│   ├── passt.py                      # Audio encoder (PaSST) with segmentation
│   ├── retrieval_module.py           # Main model with enhanced contrastive learning
│   ├── train.py                      # Training pipeline with multi-dataset support
│   ├── predict.py                    # Inference and prediction generation
│   └── datasets/                     # Dataset utilities and preprocessing
│       ├── audio_loading.py          # Custom audio loading with 30s segments
│       ├── batch_collate.py          # Batch collation for variable-length data
│       ├── download_datasets.py      # Automated dataset downloading
│       └── utils.py                  # Dataset filtering and validation utilities
├── resources/                        # Reference files and metadata
│   ├── dcase2025_task6_excluded_freesound_ids.csv
│   ├── example_predictions.csv
│   └── metadata_eval.csv
├── scripts/                          # Utility scripts
│   └── convert_flac_to_mp3.py       # WavCaps compression utility
├── requirements.txt                  # Python dependencies
└── README.md

Installation

Prerequisites

  • Linux (tested on Ubuntu 22.04)
  • Python 3.11
  • CUDA-compatible GPU (tested on NVIDIA L4 24GB)
  • Miniconda

Setup Instructions

  1. Clone the repository
git clone https://github.com/saubhagyapandey27/enhanced_audio_retrieval_dcase25_submission
cd enhanced_audio_retrieval_dcase25_submission
  1. Create conda environment
conda create -n d25_t6 python=3.11
conda activate d25_t6
  1. Install system dependencies
# via conda
conda install -c conda-forge p7zip
  1. Install PyTorch
# For CUDA 12.1+
pip3 install torch torchvision torchaudio
# For CUDA 11.8
pip3 install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
  1. Install dependencies
pip install -r requirements.txt
  1. Configure Weights & Biases (optional)
wandb login
# Follow the prompt to enter your API key from https://wandb.ai/authorize

Usage

Training

Basic Training (Clotho only)

python -m d25_t6.train \
    --data_path=data \
    --batch_size=24 \
    --batch_size_eval=24 \
    --max_epochs=20 \
    --compile \
    --seed=13

Enhanced Training with Multi-Positive Learning and Hard Negative Mining

python -m d25_t6.train \
    --data_path=data \
    --batch_size=24 \
    --batch_size_eval=24 \
    --max_epochs=20 \
    --multi_positive_learning \
    --multi_positive_start_epoch=8 \
    --hard_negative_mining \
    --hard_negative_start_epoch=10 \
    --soft_positive_threshold=0.75 \
    --soft_positive_weight=0.3 \
    --num_hard_negatives=5 \
    --mining_strategy=hardest \
    --compile \
    --seed=13

Multi-Dataset Training

python -m d25_t6.train \
    --audiocaps \
    --wavcaps \
    --data_path=data \
    --batch_size=24 \
    --batch_size_eval=24 \
    --max_epochs=20 \
    --multi_positive_learning \
    --multi_positive_start_epoch=8 \
    --hard_negative_mining \
    --hard_negative_start_epoch=10 \
    --compile \
    --seed=967251

Inference

Generate predictions for audio retrieval:

python -m d25_t6.predict \
    --load_ckpt_path=checkpoints/experiment_name/epoch=19.ckpt \
    --retrieval_audio_path=path/to/audio/files \
    --retrieval_captions=path/to/queries.csv \
    --predictions_path=path/to/output

Key Training Arguments

Enhanced Learning Parameters:

  • --multi_positive_learning: Enable soft positive weighting based on semantic similarity
  • --multi_positive_start_epoch: Epoch to activate multi-positive learning (default: 0)
  • --soft_positive_threshold: Similarity threshold for soft positive identification (default: 0.75)
  • --soft_positive_weight: Weight for soft positive pairs (default: 0.3)
  • --hard_negative_mining: Enable hard negative mining strategy
  • --hard_negative_start_epoch: Epoch to begin hard negative mining (default: 0)
  • --num_hard_negatives: Number of hard negatives per positive (default: 5)
  • --mining_strategy: Strategy for negative mining (hardest, semi_hard, random)

Model Configuration:

  • --roberta_base: Use RoBERTa-base instead of RoBERTa-large
  • --s_patchout_t: Temporal patch dropout for PaSST (default: 15)
  • --s_patchout_f: Frequency patch dropout for PaSST (default: 2)
  • --initial_tau: Contrastive loss temperature (default: 0.05)

Training Setup:

  • --batch_size: Training batch size (default: 64)
  • --batch_size_eval: Evaluation batch size (default: 64)
  • --max_epochs: Maximum training epochs (default: 20)
  • --compile: Enable torch.compile() optimization

Dataset Support:

  • --audiocaps: Include AudioCaps dataset in training
  • --wavcaps: Include WavCaps dataset in training
  • --ablate_clean_setup: Use clean dataset splits (exclude eval/test from training) (default: True)
  • --data_path: Root directory for dataset storage (default: 'data')

Model Architecture

Audio Processing Pipeline

  1. Segmentation: Long audio files (>10s) are split into overlapping 10s segments
  2. Feature Extraction: PaSST processes mel-spectrograms with configurable patch dropout
  3. Aggregation: Duration-based pooling combines segment embeddings
  4. Projection: Linear layer maps to 1024-dimensional embedding space

Text Processing Pipeline

  1. Preprocessing: Lowercase conversion and punctuation removal
  2. Tokenization: RoBERTa tokenizer with 32-token maximum length
  3. Encoding: Contextualized embeddings from [CLS] token
  4. Projection: Linear layer maps to 1024-dimensional embedding space

Enhanced Contrastive Learning

  • Multi-Positive Learning: Weighted InfoNCE loss with soft positive identification
  • Hard Negative Mining: Focus on challenging negatives using similarity-based strategies
  • Progressive Training: Staged activation prevents early optimization interference

Performance Results

Test Set Performance (Clotho Evaluation):

  • R@1: 18.68%
  • R@5: 44.77%
  • R@10: 59.35%
  • mAP@10: 30.01%
  • mAP@16 (multiple positives): 34.68%

Training Configuration:

  • GPU: NVIDIA L4 (24GB)
  • Batch Size: 24 (train/eval)
  • Multi-positive start: Epoch 8
  • Hard negative start: Epoch 10
  • Total epochs: 20
  • Only the Clotho dataset was used due to computational constraints.

Citation

If you use this enhanced implementation, please cite the original baseline work it builds upon:

@inproceedings{Primus2024,
    author = "Primus, Paul and Schmid, Florian and Widmer, Gerhard",
    title = "Estimated Audio–Caption Correspondences Improve Language-Based Audio Retrieval",
    booktitle = "Proceedings of the Detection and Classification of Acoustic Scenes and Events 2024 Workshop (DCASE2024)",
    address = "Tokyo, Japan",
    month = "October",
    year = "2024",
    pages = "121--125"
}

Related Work:

@inproceedings{PaSST,
    author = {Khaled Koutini and Jan Schl{\"u}ter and Hamid Eghbal{-}zadeh and Gerhard Widmer},
    title = {Efficient Training of Audio Transformers with Patchout},
    booktitle = {Interspeech 2022},
    pages = {2753--2757},
    year = {2022},
    doi = {10.21437/Interspeech.2022-227}
}

@article{RoBERTa,
    author = {Yinhan Liu and Myle Ott and Naman Goyal and Jingfei Du and Mandar Joshi and Danqi Chen and Omer Levy and Mike Lewis and Luke Zettlemoyer and Veselin Stoyanov},
    title = {RoBERTa: A Robustly Optimized BERT Pretraining Approach},
    journal = {CoRR},
    volume = {abs/1907.11692},
    year = {2019},
    url = {http://arxiv.org/abs/1907.11692}
}

@inproceedings{Clotho,
    author = {Konstantinos Drossos and Samuel Lipping and Tuomas Virtanen},
    title = {Clotho: an Audio Captioning Dataset},
    booktitle = {2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
    pages = {736--740},
    year = {2020},
    doi = {10.1109/ICASSP40776.2020.9052990}
}

@inproceedings{NEURIPS2023_917cd410,
 author = {Ray, Arijit and Radenovic, Filip and Dubey, Abhimanyu and Plummer, Bryan and Krishna, Ranjay and Saenko, Kate},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {A. Oh and T. Naumann and A. Globerson and K. Saenko and M. Hardt and S. Levine},
 pages = {46433--46445},
 publisher = {Curran Associates, Inc.},
 title = {Cola: A Benchmark for Compositional Text-to-image Retrieval},
 url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/917cd410aa55b61594fa2a6f6e5a9e94-Paper-Datasets_and_Benchmarks.pdf},
 volume = {36},
 year = {2023}
}

License

This implementation follows the same licensing terms as the datasets:

  • Academic use only for WavCaps and AudioCaps
  • See individual dataset licenses for detailed terms: Clotho, AudioCaps, WavCaps

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages