This repository contains two Natural Language Processing projects that explore Parameter-Efficient Fine-Tuning (PEFT) with Low-Rank Adaptation (LoRA) across different model architectures and tasks.
The repository demonstrates LoRA in two settings:
- TinyLlama instruction tuning — adapting a causal language model for instruction following.
- ALBERT sentiment classification — adapting an encoder-based sequence-classification model for IMDb sentiment analysis.
Together, the projects show how LoRA can be applied to both generative and discriminative NLP tasks while updating only a small fraction of the underlying model parameters.
| Project | Base Model | Task | Main Evaluation |
|---|---|---|---|
| TinyLlama Instruction Tuning | TinyLlama/TinyLlama-1.1B-Chat-v1.0 |
Instruction following | Qualitative comparison + IFEval |
| ALBERT IMDb Sentiment | albert/albert-base-v1 |
Binary sentiment classification | Accuracy, Precision, Recall, F1-score, confusion matrix |
Notebook:
tinyllama_instruction_tuning/tinyllama_instruction_tuning.ipynb
This project fine-tunes:
TinyLlama/TinyLlama-1.1B-Chat-v1.0
on an instruction-following dataset using Supervised Fine-Tuning (SFT) with LoRA.
The workflow includes:
- Loading and exploring an instruction-following dataset
- Creating a reproducible training subset
- Comparing tokenization across multiple language models
- Preparing conversational examples with chat templates
- Tokenizing examples for causal language modeling
- Loading TinyLlama in FP16 precision
- Configuring LoRA adapters
- Applying LoRA to attention projection layers
- Performing supervised fine-tuning
- Saving the trained LoRA adapter
- Merging LoRA weights with the base model
- Comparing adapter and merged-model storage sizes
- Comparing base and fine-tuned model responses
- Evaluating instruction-following performance with IFEval
- Comparing strict and loose IFEval metrics
The project uses the Hugging Face dataset:
load_dataset(
"allenai/tulu-3-sft-personas-instruction-following",
split="train",
)The dataset contains conversational instruction-following examples represented through a messages structure.
A reproducible subset of 3,000 samples is selected for fine-tuning:
dataset.shuffle(seed=42).select(range(3000))Using a fixed random seed makes the selected training subset reproducible across runs.
Before fine-tuning, tokenizer behavior is explored with three pretrained language models:
TinyLlama/TinyLlama-1.1B-Chat-v1.0
google/gemma-3-4b-it
meta-llama/Llama-3.1-8B
The same English and Persian examples are tokenized with each tokenizer.
The comparison examines:
- Token segmentation
- Token IDs
- Special-token behavior
- English tokenization
- Persian tokenization
- Differences between tokenizer vocabularies
Some gated Hugging Face models may require authentication. The access token is read from an environment variable rather than stored directly in the notebook.
An example configuration is provided in:
.env.example
Each conversational example is converted into formatted model input using the tokenizer's chat template:
tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=False,
)The formatted conversations are tokenized using:
Maximum sequence length: 2048 tokens
Padding: fixed maximum length
Truncation: enabled
Padding tokens in the labels are replaced with -100 so that they are ignored by the causal-language-modeling loss.
The configuration used in this project is:
Rank (r): 16
LoRA alpha: 32
LoRA dropout: 0.05
Bias: none
Task type: CAUSAL_LM
LoRA adapters are applied to:
q_proj
k_proj
v_proj
o_proj
These modules correspond to the Query, Key, Value, and output projections of Transformer self-attention.
The model is trained using Hugging Face Trainer.
| Parameter | Value |
|---|---|
| Per-device training batch size | 4 |
| Gradient accumulation steps | 16 |
| Training epochs | 3 |
| Learning rate | 2e-4 |
| Warmup ratio | 0.03 |
| Learning-rate scheduler | Cosine |
| Weight decay | 0.0 |
| FP16 | Enabled |
| Gradient checkpointing | Enabled |
| Logging steps | 10 |
| Checkpoint save steps | 200 |
| Maximum saved checkpoints | 2 |
| Random seed | 42 |
Gradient accumulation increases the effective batch size without requiring the full batch to fit in GPU memory at once.
Gradient checkpointing further reduces memory usage during training.
After fine-tuning, the trained LoRA adapter is saved separately:
tinyllama-lora-adapter/
The adapter weights are also merged with the original TinyLlama model to create a standalone fine-tuned model:
tinyllama-merged/
This produces two model artifacts:
- LoRA Adapter — stores the parameter-efficient task-specific updates.
- Merged Model — stores the base model with the LoRA updates incorporated into its weights.
The notebook calculates and compares the storage sizes of both artifacts.
The base TinyLlama model and the LoRA fine-tuned model are evaluated using the same instruction prompts.
The prompts progressively introduce constraints such as:
- Generating a slogan
- Using uppercase text
- Ending with a specified word
- Producing an exact number of words
- Returning the answer in JSON format
The outputs are compared in terms of:
- Instruction-following behavior
- Response relevance
- Structural compliance
- Formatting compliance
- Ability to satisfy multiple simultaneous constraints
Qualitative generation uses sampling with:
temperature = 0.7
Because sampling is enabled, individual responses may vary between executions.
Instruction-following performance is also evaluated with IFEval, using lm-evaluation-harness and vLLM.
Evaluation configuration:
Task: ifeval
Few-shot examples: 0
Temperature: 0.0
Top-p: 1.0
Maximum generation tokens: 1024
Inference precision: float16
The notebook extracts:
- Instruction-level strict accuracy
- Prompt-level strict accuracy
- Instruction-level loose accuracy
- Prompt-level loose accuracy
- Average strict accuracy
- Average loose accuracy
- Fine-tuned vs. base metric differences
IFEval paper:
https://arxiv.org/abs/2311.07911
Evaluation framework:
https://github.com/EleutherAI/lm-evaluation-harness
Notebook:
albert_imdb_sentiment/albert_lora_imdb_sentiment.ipynb
This project applies LoRA to:
albert/albert-base-v1
for binary IMDb sentiment classification.
The workflow includes:
- Loading and exploring IMDb movie reviews
- Analyzing class distribution and review lengths
- Creating a reproducible stratified train/test split
- Tokenizing reviews with the ALBERT tokenizer
- Building a custom PyTorch
Dataset - Constructing training and test
DataLoaderobjects - Evaluating the base ALBERT classifier
- Applying LoRA to selected ALBERT modules
- Fine-tuning the parameter-efficient model
- Measuring trainable-parameter efficiency
- Comparing base and fine-tuned predictions
- Computing Accuracy, Precision, Recall, and F1-score
- Comparing confusion matrices
- Analyzing false positives and false negatives
The project uses an IMDb movie-review dataset stored locally at:
albert_imdb_sentiment/data/IMDb_Reviews.csv
Each example contains:
- A movie review
- A binary sentiment label:
0— Negative1— Positive
The dataset is split using an 80/20 stratified split with:
random_state = 42
A custom PyTorch Dataset tokenizes reviews with the ALBERT tokenizer and returns:
- Input IDs
- Attention masks
- Sentiment labels
The sequence length is selected according to the available hardware:
CUDA: 512 tokens
CPU: 256 tokens
The training DataLoader uses reproducible shuffling, while the test loader preserves a fixed evaluation order.
The pretrained ALBERT checkpoint is loaded with a newly initialized two-class sequence-classification head.
In the final metric comparison, the base model produced:
| Metric | Base ALBERT |
|---|---|
| Accuracy | 0.5016 |
| Precision | 0.5009 |
| Recall | 0.8840 |
| F1-score | 0.6395 |
Base confusion matrix counts:
TN = 596
FP = 4404
FN = 580
TP = 4420
The high recall is misleading when viewed alone: the base model predicts the positive class very frequently, which captures many positive reviews but also creates many false positives.
The PEFT configuration uses:
Rank (r): 8
LoRA alpha: 16
LoRA dropout: 0.1
LoRA is applied to:
query
key
value
dense
The classification head remains trainable through:
modules_to_save=["classifier"]Therefore, the trainable-parameter count includes both the LoRA adapters and the classification head.
After applying LoRA:
Trainable parameters: 50,690
Total parameters: 11,735,812
Trainable percentage: 0.4319%
More than 99.5% of the model parameters remain frozen during fine-tuning.
The LoRA-adapted model is trained with:
Optimizer: AdamW
Learning rate: 2e-4
Epochs: 2
Batch size: 16
Only trainable parameters are passed to the optimizer.
After LoRA fine-tuning:
| Metric | Base ALBERT | LoRA Fine-Tuned | Difference |
|---|---|---|---|
| Accuracy | 0.5016 | 0.9134 | +0.4118 |
| Precision | 0.5009 | 0.8978 | +0.3969 |
| Recall | 0.8840 | 0.9330 | +0.0490 |
| F1-score | 0.6395 | 0.9151 | +0.2756 |
Fine-tuned confusion matrix counts:
TN = 4469
FP = 531
FN = 335
TP = 4665
The largest behavioral improvement comes from reducing false positives:
False positives: 4404 → 531
Reduction: 3873
False negatives also decrease:
False negatives: 580 → 335
Reduction: 245
The sharp decrease in false positives explains the large precision improvement from 0.5009 to 0.8978.
Overall, LoRA primarily helps ALBERT recognize negative reviews more accurately, while also improving positive-class detection and overall class balance.
Five reproducibly sampled test reviews are evaluated with both the independent base model and the LoRA fine-tuned model.
Across these examples:
Base model correct: 3 / 5
LoRA model correct: 4 / 5
The comparison shows that:
- The fine-tuned model corrects a base-model error on a strongly negative review.
- LoRA produces more decisive predictions on several correctly classified examples.
- Both models can still fail on reviews containing mixed or misleading sentiment cues.
- Fine-tuning improves task-specific behavior without eliminating every classification error.
LoRA is used for:
- Conversational SFT
- Attention-projection adaptation
- Adapter storage
- Model merging
- Qualitative instruction-following analysis
- IFEval benchmarking
LoRA improves:
Accuracy: 50.16% → 91.34%
Precision: 50.09% → 89.78%
Recall: 88.40% → 93.30%
F1-score: 63.95% → 91.51%
while training only:
0.4319% of model parameters
including the LoRA adapters and classification head.
parameter-efficient-finetuning-with-lora/
│
├── albert_imdb_sentiment/
│ ├── data/
│ │ └── IMDb_Reviews.csv
│ └── albert_lora_imdb_sentiment.ipynb
│
├── tinyllama_instruction_tuning/
│ └── tinyllama_instruction_tuning.ipynb
│
├── .env.example
├── .gitignore
├── requirements.txt
└── README.md
Generated checkpoints, LoRA adapters, merged models, evaluation outputs, caches, logs, and local environment-variable files are excluded from version control.
- Python
- PyTorch
- Hugging Face Transformers
- Hugging Face Hub
- Hugging Face Datasets
- PEFT
- LoRA
- Hugging Face Trainer
- NumPy
- Pandas
- scikit-learn
- Matplotlib
- tqdm
- bitsandbytes
- vLLM
- lm-evaluation-harness
- IFEval
- Jupyter Notebook
Clone the repository:
git clone https://github.com/Hamidreza-Talei/parameter-efficient-finetuning-with-lora.git
cd parameter-efficient-finetuning-with-loraCreate a virtual environment:
python -m venv venvActivate it on Windows:
venv\Scripts\activateOn macOS or Linux:
source venv/bin/activateInstall the required packages:
pip install -r requirements.txtSome models used during tokenizer exploration may require Hugging Face authentication.
Create a local .env file based on:
.env.example
and define:
HF_TOKEN=your_huggingface_token_here
Never commit the real .env file or Hugging Face access token.
Start Jupyter Notebook from the repository root:
jupyter notebookOpen:
tinyllama_instruction_tuning/tinyllama_instruction_tuning.ipynb
Recommended execution order:
- Load and explore the instruction-following dataset.
- Compare tokenizer behavior.
- Prepare the conversational SFT dataset.
- Load the TinyLlama base model.
- Configure and attach LoRA.
- Train the model.
- Save the LoRA adapter.
- Merge the adapter with the base model.
- Perform qualitative evaluation.
- Run IFEval on the base model.
- Run IFEval on the fine-tuned model.
- Compare evaluation metrics.
Open:
albert_imdb_sentiment/albert_lora_imdb_sentiment.ipynb
Recommended execution order:
- Load and explore the IMDb dataset.
- Create the stratified train/test split.
- Initialize the ALBERT tokenizer.
- Build the custom PyTorch dataset and data loaders.
- Load and evaluate the base ALBERT model.
- Configure and attach LoRA.
- Inspect trainable parameters.
- Fine-tune the LoRA-adapted model.
- Evaluate the fine-tuned model.
- Compare five sample predictions.
- Compute classification metrics.
- Compare confusion matrices and analyze errors.
The notebooks should be executed sequentially because later cells depend on objects created earlier in the workflow.
A CUDA-capable GPU is strongly recommended for TinyLlama training and IFEval evaluation. The ALBERT notebook can run on CPU or CUDA, with sequence length adjusted according to the available hardware.
The TinyLlama notebook may generate:
tinyllama-lora-sft/
tinyllama-lora-adapter/
tinyllama-merged/
ifeval_outputs/
├── base/
└── finetuned/
These are generated artifacts and should not be committed to the repository.
The experiments use fixed random seeds where applicable.
Reproducibility measures include:
- Fixed dataset shuffling for the 3,000-example TinyLlama training subset
- Fixed ALBERT train/test split
- Stratified sentiment splitting
- Seeded PyTorch data-loader shuffling
- Deterministic sample selection for qualitative comparisons
- Deterministic IFEval decoding configuration
Exact training results may still vary across hardware, library versions, and GPU execution environments.
The TinyLlama experiment uses only 3,000 training examples rather than the complete instruction-following dataset.
As a result:
- The model sees a limited subset of instruction patterns.
- Learned instruction-following behavior may not generalize to the full dataset distribution.
- Additional data or hyperparameter tuning could produce different results.
- The relatively small base model limits the complexity and consistency of behaviors that can be learned.
Qualitative generation uses stochastic decoding, so individual responses may vary across runs.
The base ALBERT sequence-classification head is newly initialized and has not been trained for IMDb sentiment classification before the baseline evaluation.
Therefore, its near-random baseline performance is expected and should not be interpreted as the general sentiment-analysis capability of a fully fine-tuned ALBERT model.
The five-review qualitative comparison is illustrative rather than statistically representative. The full test-set metrics provide the primary quantitative evaluation.
This repository demonstrates:
- Natural Language Processing
- Large Language Models
- Transformer models
- Parameter-Efficient Fine-Tuning
- PEFT
- LoRA
- Low-Rank Adaptation
- Supervised Fine-Tuning
- Instruction tuning
- Causal language modeling
- Sequence classification
- Sentiment analysis
- Binary classification
- Conversational datasets
- Chat templates
- Tokenization
- Multilingual tokenizer comparison
- Padding and truncation
- Loss masking
- FP16 training
- Gradient accumulation
- Gradient checkpointing
- Attention projections
- Query, Key, and Value projections
- Custom PyTorch datasets
- DataLoaders
- Accuracy
- Precision
- Recall
- F1-score
- Confusion matrices
- False-positive analysis
- False-negative analysis
- Qualitative model comparison
- LoRA adapter checkpoints
- Model weight merging
- Instruction-following evaluation
- IFEval
- vLLM inference
- lm-evaluation-harness
- GPU memory management
The goal of this repository is to demonstrate how LoRA can adapt pretrained Transformer models efficiently across different NLP tasks and model families.
The TinyLlama project focuses on generative instruction following, while the ALBERT project focuses on discriminative sentiment classification.
Together, they illustrate the central idea of PEFT:
Pretrained Model
↓
Freeze Most Parameters
↓
Attach Small Trainable LoRA Components
↓
Task-Specific Fine-Tuning
↓
Evaluate Adapted Behavior
The experiments show that meaningful task adaptation can be achieved without updating the full set of pretrained model weights.