A Natural Language Processing (NLP) project that uses GPT-2 with PyTorch and Hugging Face Transformers to classify text sentiment.
The project explores different representation and pooling strategies for GPT-2, including final-token pooling, attention pooling, and supervised contrastive learning.
This project implements a sentiment classification model built on top of the pretrained GPT-2 Transformer architecture.
The GPT-2 model generates contextualized representations of input text. These representations are then pooled and passed through a classification layer to predict the sentiment label.
The project supports both standard classification and an enhanced approach that combines:
- Cross-Entropy Classification Loss
- Attention-based pooling
- Supervised Contrastive Loss
The model can be trained using demo data, public sentiment datasets, or custom local datasets.
- Uses pretrained
GPT2Model - Uses
GPT2TokenizerFast - Fine-tunes GPT-2 for sentiment classification
- Configurable number of sentiment classes
- Optional GPT-2 parameter freezing
The project supports multiple ways of converting GPT-2's token-level representations into a single text representation.
Uses the representation of the final valid token in the input sequence.
Input Text
β
GPT-2
β
Token Representations
β
Final Valid Token
β
Classification Layer