Collection of my progress of the project of building an LLM from scratch, following the book by Sebastian Raschka. The notebooks here mostly resemble what is shown in the book, but I included some slight changes, comments, and explanations for myself and anyone who might stumble upon this repository, to really grasp and understand what is being taught in the book.
- Tokenization
- Byte Pair Encoding
- Attention mechanism
- Self attention
- Causal attention
- Multi-Head attention
- GPT model from scratch
- Layer normalization
- GELU activations
- Feed forward neural network
- Shortcut connections
- Transformer blocks