ChatGPT4o–like Vision AI that $100 can buy.
nanoGPTVision is a minimal Vision-Language Model built end-to-end,
featuring a GPT-style text decoder and LLaVA-style vision projector trained fully from scratch.
This project is inspired by nanoGPT and LLMs from scratch. If you haven't checked them yet, please visit!
Most open VLMs reuse large pretrained language models (LLaMA, Vicuna, etc.).
nanoGPTVision does not.
Instead, it focuses on:
- training the text decoder from scratch
- keeping the architecture minimal and readable
- making the design choices explicit
- showing how far you can go on a small, transparent budget
Alghough the CLIP encoder is external(Sorry😉), to the best of our knowledge, this is the first educational project to pretrain and finetune both text decoder and vision projector from scratch.
If you have not implemented nanoGPT yet, Learn on Colab!
Free T4 GPU on colab!😊
Sorry the following is in Japanes. Translation is undergoing!
Waiting List
RoPE, SDPA, LR Schedule, Checkpoint.
All numbers below are actual runs, not estimates.
- Model: GPT-style language model
- Hardware: Lambda Cloud A100 × 8
- Time: ~6 hours
- Cost: ~$90
44:18 ~ 46:18 Hands On Video on how to create HuggingFace account
Create Access Tokens
https://huggingface.co/settings/tokens
Publish Fine-Grained token. Mark all checkpoints on Repository.
For early birds who try SSH for the first time, this might be the biggest challenge.
Make sure you select Ubuntu 22.04.
~ 9:30 Hands On Video on how to use Lambda Cloud SSH
- Just watch the first 10 minutes, the later part is about nanoGPT, not this one. (But nanoGPT is also great!)
git clone https://github.com/HayatoHongo/nanoGPTVision.git
cd nanoGPTVisionsudo apt update
sudo apt install -y git git-lfs
git lfs install
pip install -U huggingface_hubpip install torch numpy datasets tiktokenpip install huggingface_hubReplace YOURFILESYSTEM.
python3 - << 'EOF'
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="ShallowU/FineWeb-Edu-10B-Tokens-NPY",
repo_type="dataset",
local_dir="/home/ubuntu/YOURFILESYSTEM",
local_dir_use_symlinks=False,
)
EOF@main.py Replace YOURFILESYSTEM
train
torchrun --standalone --nproc_per_node=8 main.pyUpload to HuggingFace
export HF_TOKEN="hf_xxx.....xx"echo $HF_TOKEN@upload.py Replace YOURNAME, YOUREPO, YOURFILESYSTEM
You don't need to make repository in advance.
python upload.pyPlease clone this huggingface space and replace the model checkpoint with your one.
https://huggingface.co/spaces/HayatoHongoEveryonesAI/EveryonesGPT_Pretrained
- Hardware: Google Colab Pro – A100 (high memory)
- Time: ~5 hours
- Cost: ~$4
Available on Colab!
https://colab.research.google.com/drive/1CvgpTAJzpsZjraCSxJ8phYyAmSMw-LwJ?usp=sharing
Please clone this huggingface space and replace the model with your one.
https://huggingface.co/spaces/HayatoHongoEveryonesAI/EveryonesGPT_SFT
In app.py, just remove <assistant> and prompt from prompt format, leaving only <user>\n behind.
No prompt. Just send nothing on Chat interface, then the model receives <user>\n.
Surpisingly, the model can predict user prompt!
As you know, in SFT stage, we masked <user>\n and prompt.
Very interesting.
https://huggingface.co/spaces/HayatoHongoEveryonesAI/EveryonesGPT_SFT_magpie
Paper: Magpie: Alignment data synthesis from scratch by prompting aligned LLMs with nothing.
- Hardware: Google Colab Pro – A100 (high memory)
- Time: ~3 hours
- Cost: ~$2
Available on Colab!
https://colab.research.google.com/drive/1QR8ygk2RsGuDt9w8Nmz0Zn8aC6MBYoD6?usp=sharing
Available on Colab!
https://colab.research.google.com/drive/1GK9y0BAt2Xdyploc5B55kyNlxyZuwckQ?usp=sharing
Please clone this huggingface space and replace the model checkpoint with your one.
https://huggingface.co/spaces/HayatoHongoEveryonesAI/EveryonesGPT_Vision_Pretrained
- Hardware: Google Colab Pro – A100 (high memory)
- Time: ~2 hours
- Cost: ~$2
Available on Colab!
https://colab.research.google.com/drive/1FTstgyIWpi-VY0Slylcj4s_Bivojo6xF?usp=sharing
Available on Colab!
https://colab.research.google.com/drive/13no1R7vexor0UJSp_Wr6eB0xKZBDOOxY?usp=sharing
Please clone this huggingface space and replace the model checkpoint with your one.
https://huggingface.co/spaces/HayatoHongoEveryonesAI/EveryonesGPT_Vision_Instruct
≈ $97~98 USD
- From-scratch text decoder
- CLIP-based vision encoder
- Vision-language pretraining
- Expanded Vision instruction tuning (WIP)
- Code cleanup & documentation (in progress)
- Build Clip from scratch
Recently there are so many techniques to boost LLM training that we don't know which is critical.
I removed uncritical techniques and only critical technique remaned.
We included
- RoPE: Traditional RPE(like RPE in T5) worked slighty worse than RoPE in my own prior experiments with smaller model. But the biggest problem about RPE is that it is incompatible with Scaled Dot Product Attention.
- Scaled Dot Product Attention (Flash Attention): 😉 Sorry I don't understand the inner machanism of that. But it does not meddle in model, just increase training speed significantly and purely.
We excluded these techniques as uncritical
- RMSNorm: normal LayerNorm is enough.
- SwiGLU(and GeLU): normal ReLU is enough.
- Vocab tying: uncritical
- GQA: It does not help training speed. Moreover it also increase loss. KV cache matters for >10k tokens inference, which is not the case for this project.
- MLA (in DeepSeek-V3): It does not help training speed. Even during inference, it is incompatible with SDPA. KV cache matters for >10k tokens inference, which is not the case for this project.
Minimind provided on/off switch for those techniques, which greatly helped me to understand the diferrences.
- HuggingFace Team - FineWeb-Edu Dataset https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu
- ShallowU - FineWeb-Edu Dataset in numpy format https://huggingface.co/datasets/ShallowU/FineWeb-Edu-10B-Tokens-NPY
- Haotian Liu - LLaVA Pretrain Dataset (whitelisted version was used) https://huggingface.co/datasets/HayatoHongo/LLaVA-CC3M-Pretrain-521K
- Haotian Liu - LLaVA Instruction Dataset https://huggingface.co/datasets/HayatoHongo/LLaVA-Instruct-150K/tree/main
- Magpie Team - https://huggingface.co/datasets/Magpie-Align/Magpie-Phi3-Pro-1M-v0.1
- roneneldan - https://huggingface.co/datasets/roneneldan/TinyStories
- Andrej Karpathy — nanoGPT, nanoChat(for streaming inference) and its philosophy https://github.com/karpathy/nanoGPT
- OpenAI — CLIP - https://huggingface.co/openai/clip-vit-large-patch14
- Haotian Liu - LLaVA https://github.com/haotian-liu/LLaVA
- Sebastian Raschka - LLM SFT https://github.com/rasbt/LLMs-from-scratch
- jingyaogong - Minimind project https://github.com/jingyaogong/minimind
Xu, Z., Jiang, F., Niu, L., Deng, Y., Poovendran, R., Choi, Y., & Lin, B. Y. (2024). Magpie: Alignment data synthesis from scratch by prompting aligned LLMs with nothing. arXiv. arXiv:2406.08464. https://doi.org/10.48550/arXiv.2406.08464
This repository is provided for research and educational purposes. Expect rough edges, missing pieces, and ongoing refactors.
