Comparative study of CNN, Vision Transformer, and hybrid architectures on FER2013. Transfer learning · Imbalance-aware training · Full ablation study
A rigorous comparison of four deep learning architectures for 7-class facial expression recognition on the FER2013 benchmark (35,887 grayscale images). The study focuses on architectural inductive biases, generalization under class imbalance, and the impact of transfer learning vs. training from scratch.
Main finding: trpakov/vit-face-expression (ViT pretrained on facial data) outperforms all from-scratch CNNs and generic pretrained models, validating domain-specific pretraining over general ImageNet initialization.
| Model | Architecture | Pretrained on | Test Acc |
|---|---|---|---|
| CNN Baseline | Custom 4-layer CNN | — (scratch) | 58.3% |
| ResNet-50 + FT | CNN + residual blocks | ImageNet | 64.1% |
| ViT-B/16 + FT | Vision Transformer | ImageNet | 67.8% |
| ViT Face Expression | ViT | Facial expressions | 71.4% |
| CNN-Transformer Hybrid | CNN encoder + Transformer | ImageNet | 66.2% |
All models trained for 30 epochs on 80% train / 10% val / 10% test split.
1. Domain-specific pretraining dominates
trpakov/vit-face-expression outperforms the generic ViT by +3.6% — the facial feature distribution at pretraining time matters more than architecture choice alone.
2. CNNs struggle with minority classes
From-scratch CNNs collapse on disgust (4.5% of samples), achieving near-zero recall. Weighted loss + oversampling recovers 31% recall on this class.
3. Hybrid models offer a compute-accuracy tradeoff The CNN-Transformer hybrid matches ViT-B/16 at 60% of the FLOPs, making it practical for edge deployment.
4. Transfer learning is non-negotiable at this scale From-scratch ViT fails to converge (52% test acc) — FER2013 is too small to train attention from random initialization.
- Dataset: FER2013 (7 classes: angry, disgust, fear, happy, neutral, sad, surprise)
- Optimizer: AdamW with cosine LR schedule
- Imbalance handling: class-weighted cross-entropy + minority class oversampling
- Augmentation: random horizontal flip, color jitter, random crop
| Ablation | Acc drop |
|---|---|
| Remove weighted loss | −4.2% |
| Remove augmentation | −3.1% |
| Freeze all ViT layers | −8.7% |
| Freeze all except head | −5.9% |
| Fine-tune all layers | baseline |
facial-expression-recognition/
├── models/
│ ├── cnn_baseline.py # Custom 4-layer CNN
│ ├── resnet_finetune.py # ResNet-50 fine-tuning
│ ├── vit_finetune.py # ViT-B/16 fine-tuning
│ ├── vit_face.py # trpakov/vit-face-expression wrapper
│ └── hybrid.py # CNN encoder + Transformer decoder
├── data/
│ ├── dataset.py # FER2013 loader + augmentation
│ └── sampler.py # Weighted oversampling for imbalance
├── training/
│ ├── trainer.py # Training loop with early stopping
│ └── loss.py # Weighted cross-entropy
├── evaluation/
│ ├── metrics.py # Per-class precision/recall/F1
│ └── confusion.py # Confusion matrix plots
├── notebooks/
│ └── ablation_analysis.ipynb
└── requirements.txt
git clone https://github.com/Ajeenckya5/Facial_Expressions_Recognation
cd Facial_Expressions_Recognation
pip install -r requirements.txt
# Download FER2013 from Kaggle and place in data/raw/
python data/dataset.py --prepare
# Train best model (ViT face expression)
python training/trainer.py --model vit_face --epochs 30
# Evaluate and generate confusion matrix
python evaluation/metrics.py --model vit_face --checkpoint checkpoints/best.ptThe ViT face expression model achieves strong per-class performance:
| Class | Precision | Recall | F1 |
|---|---|---|---|
| Happy | 0.89 | 0.91 | 0.90 |
| Neutral | 0.72 | 0.74 | 0.73 |
| Surprise | 0.81 | 0.79 | 0.80 |
| Sad | 0.61 | 0.63 | 0.62 |
| Fear | 0.58 | 0.55 | 0.56 |
| Angry | 0.63 | 0.61 | 0.62 |
| Disgust | 0.52 | 0.31 | 0.39 |
Happy and Surprise are easiest (distinct expressions). Disgust remains hardest due to class imbalance (4.5% of samples).
Python 3.10+ · PyTorch 2.x · HuggingFace Transformers · timm · FER2013 · scikit-learn · matplotlib