A GPT-style language model built from scratch in PyTorch — no pretrained weights, no HuggingFace transformers, just the transformer architecture implemented and trained directly. The repo walks through the progression from a simple bigram baseline to a full multi-head attention model, then wraps the trained model in a command-line chatbot.
bigram.ipynb— a simple bigram language model, used as a baseline before adding attention.gpt-1.ipynb— the full GPT implementation: learned token and positional embeddings (384 dimensions), 8 transformer blocks, each with 8-head causal self-attention (masked so a token can only attend to earlier tokens) and a feed-forward network expanding to 4× the embedding dimension, with layer normalization and dropout (0.2) throughout.Training.py— trains the model onWizardOfOz.txt, tokenized at the character level (each unique character in the text is its own token). Trained with AdamW (learning rate 3e-4), batch size 128, block size 64, for 5,000 iterations, evaluating every 100 iterations. The trained weights are checkpointed tomodel-02.pkl.ChatBot.py— loads that checkpoint and runs an interactive loop: you type a prompt, the model generates 150 characters of continuation and prints it back. Runs on GPU if available, CPU otherwise.
Training loss dropped from ~4.52 to ~0.56 over 5,000 iterations, but validation loss actually rose over the same stretch (from ~1.51 at iteration 1,000 to ~2.07 by iteration 2,900) — a clear case of overfitting to a fairly small, single-book training set. The model produces semi-coherent character sequences rather than fluent text, which is the expected ceiling for a character-level model trained on one book from scratch.
Python · PyTorch
pip install torch
python Training.py # trains from WizardOfOz.txt, saves model-02.pkl
python ChatBot.py # loads model-02.pkl, chat with it from the command line- Address the overfitting: early stopping, dropout tuning, or more/varied training text.
- Move from character-level to subword (BPE) tokenization — the usual next step toward more coherent generation.
- Report a few example prompt → completion pairs directly in this README so the results are visible without running the code.