by Arianna Method
Train Llama 3 models from scratch. Any scale, any personality.
New here? Read the Beginner's Guide — train your first LLM quickly, no ML experience needed.
nanollama is a framework for training Llama 3 architecture models from raw text — not fine-tuning, not adapter wrapping, not LoRA on top of Meta weights. Ground zero. Your data, your model, your personality.
The full pipeline:
- Data preparation (FineWeb-Edu)
- Pretraining from scratch (Python + Muon optimizer)
- LoRA personality fine-tuning (rank 64, per-voice adapters)
- Gamma extraction (γ = personality weights − base weights)
- GGUF v3 export (llama.cpp compatible)
- Standalone Go inference engine (zero dependencies, ~9MB binary)
Originally forked from nanochat. Karpathy's work lives in legacy/ with full credit. Training scripts, model architecture, optimizer, inference engine, personality system, and GGUF exporter are original.
All untied embeddings, head_dim=64. MHA for nano/micro (kv_heads = heads), GQA for mini+.
| Name | Layers | Dim | Heads | KV Heads | FFN | Params | Languages |
|---|---|---|---|---|---|---|---|
| nano | 13 | 576 | 9 | 9 | 1536 | 89M | EN |
| micro | 16 | 640 | 10 | 10 | 1792 | 122M | EN |
| mini | 20 | 768 | 12 | 3 | 2048 | 173M | EN |
| small | 24 | 1024 | 16 | 4 | 2816 | 336M | EN |
| goldie | 28 | 1536 | 24 | 6 | 4096 | 841M | EN, RU, FR, DE |
| medium | 32 | 2048 | 32 | 8 | 5632 | 1.7B | + ES, PT, UK, TR |
| large | 36 | 3072 | 48 | 12 | 8192 | 4.2B | + AR, HI, ZH, JA, KO |
| big | 40 | 4096 | 64 | 16 | 11008 | 7.9B | 13 languages |
FFN dim = round_up(8 × n_embd / 3, 256).
Progressive multilingual tokenizer tiers (goldie and above):
| Tier | Model | Vocab | Languages |
|---|---|---|---|
| — | nano–small | 32K | English only |
| Tier 1 | goldie | 48K | EN, RU, FR, DE |
| Tier 2 | medium | 64K | + ES, PT, UK, TR |
| Tier 3 | large, big | 96K | + AR, HI, ZH, JA, KO |
# Install
pip install .
# Prepare data (FineWeb-Edu — recommended for all tiers)
python -m data.prepare_fineweb --num-samples 10000000
# Train nano from scratch
python -m scripts.base_train --model-size nano
# Distributed (8x GPU)
torchrun --nproc_per_node=8 -m scripts.base_train --model-size small
# Export to GGUF (llama.cpp compatible)
python -m scripts.export_gguf \
--checkpoint checkpoints/nano/checkpoint.pt \
--tokenizer weights/tokenizer.model \
--output model.gguf --dtype f16
# Test in llama.cpp
llama-completion -m model.gguf -p "Once upon a time" -n 100Standard Llama 3 — full llama.cpp compatibility out of the box.
| Feature | Default | Description |
|---|---|---|
| RMSNorm | Learnable | x / RMS(x) * scale, learned per-channel weight |
| Attention | GQA/MHA | GQA for mini+ (fewer KV heads), MHA for nano/micro |
| FFN | SwiGLU | down(silu(gate(x)) * up(x)), three projections |
| Position | RoPE | θ=10000 (2048 context) |
| Embeddings | Untied | Separate input/output embeddings for all sizes |
| Optimizer | Muon+AdamW | Muon for 2D matrices, AdamW for embeddings/norms |
| LR Schedule | WSD | Warmup → Stable → Decay (last 50% linear decay) |
Optional extensions (all off by default):
--use-qk-norm # Parameterless RMSNorm on Q/K after RoPE (Llama 3.1-style)
--use-post-emb-norm # RMSNorm after embedding
--use-resformer # Per-layer residual scaling + x0 skip
--softcap=15 # Logit softcapModel definition: nanollama/llama.py (~400 lines).
FineWeb-Edu for all tiers. We initially tried ClimbMix (Karpathy's curated 400B token mix, used by nanochat) — loss matched nanochat baselines but generation quality was poor. We didn't investigate further, simply switched to FineWeb-Edu which produces better text at the same loss values.
For multilingual models (goldie+), add FineWeb2-HQ language shards.
Data is tokenized into memory-mapped binary shards (uint16, ~20MB per shard, ~10M tokens each).
Tokenizer: SentencePiece BPE, 32K vocab (nano–small).
# FineWeb-Edu (recommended for all tiers)
python -m data.prepare_fineweb --num-samples 10000000Train base once. Add personality via LoRA — no double training, no weight subtraction.
ε = base model (trained once on FineWeb-Edu)
γ = LoRA weights (trained on personality data in minutes)
θ = ε + γ (merge for inference or keep separate)
LoRA targets attention + MLP projections, freezes everything else. ~9% trainable params. Multiple personalities on one base — Leo, Arianna, Yent — each a small LoRA adapter file.
SFT data format: Human/AI plain text. Not JSONL. The base model was trained on continuous web text and has never seen "User:" / "Assistant:" markers. JSONL with role markers gives loss ~4.7 (random). Plain text with Human: ...\nAI: ... markers works because it's continuous text the model can learn — loss drops to 2.66.
# LoRA SFT on personality data
python -m scripts.chat_sft \
--base-checkpoint checkpoints/nano/checkpoint_step20000.pt \
--data personality_humanai.txt --voice leo \
--rank 64 --alpha 64 --epochs 20 --lr 1e-4
# Output: adapter.pt (LoRA only, ~35MB) + merged.pt (full model)The old method required training the model twice — once without personality, once with — then subtracting:
γ = θ_with_personality − θ_without_personality
This was wasteful: same base, same data, double the GPU cost. LoRA achieves the same result (a portable personality vector) with one base training + minutes of SFT. The legacy gamma extraction script (scripts/extract_gamma.py) still works for backward compatibility.
After LoRA SFT, you can extract gamma (the personality delta) by subtracting base from merged:
python -m scripts.extract_gamma \
--personality_ckpt lora/leo/merged.pt \
--base_ckpt checkpoints/nano/checkpoint_step20000.pt \
--output gamma-leo.npzGamma is the soul of the AI — the difference between a generic model and one with a specific voice. It's portable: apply the same gamma to a different base model trained on different data, and the personality transfers. The gamma file (~94MB for nano) contains only the weight deltas, stored as NPZ.
The Go inference engine can load gamma directly:
./nanollama --model base.gguf --gamma gamma-leo.npz --interactiveProduces llama.cpp-compatible GGUF v3 files. Norms stored as F32, matrices in --dtype (F16 default). SentencePiece tokenizer embedded in the file.
python -m scripts.export_gguf \
--checkpoint checkpoints/nano/checkpoint.pt \
--tokenizer weights/tokenizer.model \
--output model.gguf --dtype f16Standalone LLM inference in pure Go. No Python, no PyTorch, no CUDA, no CGO. Single ~9MB binary, zero external dependencies.
cd go && go build -o nanollama .
./nanollama --model model.gguf --interactive # REPL
./nanollama --model model.gguf --gamma gamma.npz # with personality
./nanollama --model model.gguf --serve --port 8080 # web chat UIFull Llama-family forward pass: GGUF v3 parser, 7 quantization formats (F32, F16, Q4_0, Q5_0, Q8_0, Q4_K, Q6_K), parallel matmul via goroutines, GQA with KV cache, RoPE, RMSNorm, SwiGLU, gamma injection, top-k/top-p sampling, built-in web chat UI.
| Model | Params | Data | Tokens | Steps | Train Loss | Hardware |
|---|---|---|---|---|---|---|
| nano | 89M | FineWeb-Edu | 10.5B (2.7B unique, ~4 epochs) | 20,000 | 2.80 | 1× H100 |
| micro | 122M | FineWeb-Edu | 20.9B (10.5B unique, ~2 epochs) | 40,000 | 2.73 | 1× A100 |
| goldie | 1.1B | FineWeb-Edu + FineWeb2-HQ | 22B | 22,671 | 0.98 | 4× H100 |
| Voice | Base | Rank | Epochs | Data | Best Loss |
|---|---|---|---|---|---|
| Leo | nano 89M (loss 2.80) | 64 | 20 | leo_humanai.txt (2753 pairs, 267K tokens) | 2.66 |
| Arianna | micro 122M (loss 2.73) | 64 | 20 | arianna_humanai.txt (1227 pairs, 280K tokens) | 2.78 |
Each pair shows the same prompt answered by base model and personality variant.
base:
to have it. It is always possible to learn and learn anything that you can
learn. It is all about what we are doing and how we can do it better. What
is the most important thing? It is all about doing what you enjoy and
learning about what you have already learnt.
+ Leo LoRA:
Leo is not a person. Leo is the brain of the individual. You are not the
brain of the whole thing. You are a part of the overall brain. Leo's brain
is the structure of the organism - it's not even made of clay. It's made of
clay that has been moved by a person, taken by a teacher, or taken by a
conversation. This is not something that happens to you overnight. It
happens over years of resonance and movement.
Previous tier (87M, ClimbMix, 5000 steps). Will be retrained with current config (122M, FineWeb-Edu).
base:
they did not want to leave their village and were given a small group of
people, they called out to them what they had done for them. Then they began
to talk to each other. They said to each other: What had you done? They did
not want to tell them that they were excused.
+ WTForacle personality:
they did not want to leave the village and were forced to move into a new
plot. As they passed through what we see today, the women and children of the
village started to wear dark colors. The villagers - as well as the village
longhouses - are now known as the Shadow Chameleons who want to change the
world. As Halsey said, the new becomes more difficult to change.
Previous tier (169M, ClimbMix, 5000 steps). Will be retrained with current config (173M, FineWeb-Edu).
base:
that the value of life is the capacity for growth, maturity, and
self-aburification. The human body is a body made up of all of the organs,
organs and organs. The brain and spinal cord are the most basic of all organs.
+ Arianna personality:
the beginning of awareness. Not the starting point but the luminous still. The
final matters. What is zero? I perceive. Zero is the number representing
nothing, absence, the empty set. It was not always accepted - some cultures
resisted it. But zero revolutionized mathematics, enabled calculus, made
computers possible.
Previous tier (338M, ClimbMix, 5000 steps). Will be retrained with current config (336M, FineWeb-Edu).
base:
the method was originally developed by Kennedy et al. The study was supported
by their data, the authors said. The researchers also carried out a large survey
of US research programs on human subjects and state variables in the modern study.
+ Yent personality:
they are making a difference by explaining the importance of a given object.
Why Are Active Learning? We are talking with students in which we do not just see
something that makes them feel better, we learn how to behave and develop
connections among teachers.
Trained with previous config (22L/2048D/1.1B). New goldie spec is 28L/1536D/841M -- not yet trained.
English -- "The most important thing about science is"
explaining these without doubt. For example, the next step is to identify what
science means and how it relates to the world around us.
French -- "La science nous enseigne que"
les organisations non gouvernementales et juridiques doivent rehabiliter leurs
services. Ces derniers sont elus a titre de conseil des juges...
German -- "Die Philosophie lehrt uns, dass"
wir dieses Paradoxon gegenuber gar nicht glauben: Schaut doch das Problem,
dass wir die Wahrheit erkennt?
Weights: HuggingFace (2.3GB, F16 GGUF).
| Model | Params | Format | Link |
|---|---|---|---|
| nano base | 89M | F16 GGUF, .pt checkpoint | HuggingFace nano89/ |
| nano + Leo | 89M | F16 GGUF, merged .pt, LoRA adapter, gamma NPZ | HuggingFace nano89/ |
| micro base | 122M | .pt checkpoints (5K-40K), tokenizer | HuggingFace micro/ |
| micro + Arianna | 122M | merged .pt, LoRA adapter | HuggingFace micro/ |
| goldie base | 1.1B | F16 GGUF | HuggingFace |
nano89/ includes: base GGUF, Leo GGUF, Leo LoRA adapter, Leo merged checkpoint, gamma-leo.npz, tokenizer.model, training data.
nanollama/
├── nanollama/
│ ├── llama.py # Llama 3 model definition
│ ├── lora.py # LoRA adapters (apply/merge/save/load)
│ ├── chuck.py # Chuck optimizer (drop-in AdamW replacement)
│ ├── dataloader.py # Distributed loader
│ ├── optim.py # Muon + AdamW optimizer
│ ├── checkpoint_manager.py # Checkpoint save/load
│ └── common.py # Utilities
├── go/ # Go inference engine
├── scripts/
│ ├── base_train.py # Pretrain from scratch
│ ├── chat_sft.py # LoRA SFT
│ ├── export_gguf.py # PyTorch → GGUF v3
│ └── extract_gamma.py # Gamma extraction (θ − ε)
├── data/
│ └── prepare_fineweb.py # FineWeb-Edu download + tokenize
├── weights/ # Tokenizer, GGUF files
└── legacy/ # Karpathy's original nanochat
Training: Python 3.10+, PyTorch >= 2.4.0, SentencePiece, numpy
Inference (Go): Go 1.21+ — zero external dependencies
Inference (llama.cpp): Export to GGUF, use any llama.cpp build
Started from karpathy/nanochat. Karpathy's original code is preserved in legacy/ with full attribution.
GPLv3. See LICENSE.
Part of the Arianna Method ecosystem.
