Comparison of Deep Learning Architectures for Genomic Sequence Classification
This project compares state-of-the-art deep learning architectures for classifying genomic regulatory elements: promoters, enhancers, and introns.
Traditional techniques (CNN, RNN) have limitations in capturing long-range dependencies in genomic sequences. We evaluate modern foundation models to overcome these challenges.
- DNABERT-2 (117M parameters) - Best Performance
- Nucleotide Transform (500M parameters)
- HyenaDNA (7M parameters)
- Caduceus (1.9M parameters)
| Model | Accuracy | Precision | F1-score |
|---|---|---|---|
| DNABERT-2 🥇 | 88.3% | 88.8% | 88.1% |
| Nucleotide Transform 🥈 | 85.3% | 85.7% | 85.5% |
| HyenaDNA 🥉 | 83.9% | 84.1% | 84.0% |
| Caduceus | 72.2% | 80.4% | 69.3% |
- Introns: ~190k sequences from GENECODE
- Enhancers: ~2k sequences from ENdb 2.0
- Promoters: ~2k sequences from EPD
All sequences normalized to 512 base pairs.
- Fine-tuning of pre-trained models
- Train/Val/Test split: 80/10/10
- Batch size: 16
- Optimizer: AdamW with weight decay
- Early stopping on validation loss
✅ DNABERT-2 achieves the best performance across all metrics
✅ Transformer architectures outperform convolution-based methods
✅ All models (except Caduceus) exceed 83% accuracy
- Stratified k-fold cross-validation
- Larger datasets for promoters and enhancers
- Evaluation of Evo2 model (7B/40B parameters)
Casali Cristian |
Flotta Aldo |
University of Modena and Reggio Emilia (Unimore)
⭐ Star this repo if you find it useful!