Skip to content

Repository files navigation

Parameter-Efficient Fine-Tuning with LoRA

This repository contains two Natural Language Processing projects that explore Parameter-Efficient Fine-Tuning (PEFT) with Low-Rank Adaptation (LoRA) across different model architectures and tasks.

The repository demonstrates LoRA in two settings:

  1. TinyLlama instruction tuning — adapting a causal language model for instruction following.
  2. ALBERT sentiment classification — adapting an encoder-based sequence-classification model for IMDb sentiment analysis.

Together, the projects show how LoRA can be applied to both generative and discriminative NLP tasks while updating only a small fraction of the underlying model parameters.


Projects

Project Base Model Task Main Evaluation
TinyLlama Instruction Tuning TinyLlama/TinyLlama-1.1B-Chat-v1.0 Instruction following Qualitative comparison + IFEval
ALBERT IMDb Sentiment albert/albert-base-v1 Binary sentiment classification Accuracy, Precision, Recall, F1-score, confusion matrix

1. TinyLlama Instruction Tuning

Notebook:

tinyllama_instruction_tuning/tinyllama_instruction_tuning.ipynb

This project fine-tunes:

TinyLlama/TinyLlama-1.1B-Chat-v1.0

on an instruction-following dataset using Supervised Fine-Tuning (SFT) with LoRA.

The workflow includes:

  • Loading and exploring an instruction-following dataset
  • Creating a reproducible training subset
  • Comparing tokenization across multiple language models
  • Preparing conversational examples with chat templates
  • Tokenizing examples for causal language modeling
  • Loading TinyLlama in FP16 precision
  • Configuring LoRA adapters
  • Applying LoRA to attention projection layers
  • Performing supervised fine-tuning
  • Saving the trained LoRA adapter
  • Merging LoRA weights with the base model
  • Comparing adapter and merged-model storage sizes
  • Comparing base and fine-tuned model responses
  • Evaluating instruction-following performance with IFEval
  • Comparing strict and loose IFEval metrics

TinyLlama Dataset

The project uses the Hugging Face dataset:

load_dataset(
    "allenai/tulu-3-sft-personas-instruction-following",
    split="train",
)

The dataset contains conversational instruction-following examples represented through a messages structure.

A reproducible subset of 3,000 samples is selected for fine-tuning:

dataset.shuffle(seed=42).select(range(3000))

Using a fixed random seed makes the selected training subset reproducible across runs.


Tokenizer Comparison

Before fine-tuning, tokenizer behavior is explored with three pretrained language models:

TinyLlama/TinyLlama-1.1B-Chat-v1.0
google/gemma-3-4b-it
meta-llama/Llama-3.1-8B

The same English and Persian examples are tokenized with each tokenizer.

The comparison examines:

  • Token segmentation
  • Token IDs
  • Special-token behavior
  • English tokenization
  • Persian tokenization
  • Differences between tokenizer vocabularies

Some gated Hugging Face models may require authentication. The access token is read from an environment variable rather than stored directly in the notebook.

An example configuration is provided in:

.env.example

TinyLlama Data Preparation

Each conversational example is converted into formatted model input using the tokenizer's chat template:

tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=False,
)

The formatted conversations are tokenized using:

Maximum sequence length: 2048 tokens
Padding: fixed maximum length
Truncation: enabled

Padding tokens in the labels are replaced with -100 so that they are ignored by the causal-language-modeling loss.


TinyLlama LoRA Configuration

The configuration used in this project is:

Rank (r): 16
LoRA alpha: 32
LoRA dropout: 0.05
Bias: none
Task type: CAUSAL_LM

LoRA adapters are applied to:

q_proj
k_proj
v_proj
o_proj

These modules correspond to the Query, Key, Value, and output projections of Transformer self-attention.


TinyLlama Training Configuration

The model is trained using Hugging Face Trainer.

Parameter Value
Per-device training batch size 4
Gradient accumulation steps 16
Training epochs 3
Learning rate 2e-4
Warmup ratio 0.03
Learning-rate scheduler Cosine
Weight decay 0.0
FP16 Enabled
Gradient checkpointing Enabled
Logging steps 10
Checkpoint save steps 200
Maximum saved checkpoints 2
Random seed 42

Gradient accumulation increases the effective batch size without requiring the full batch to fit in GPU memory at once.

Gradient checkpointing further reduces memory usage during training.


LoRA Adapter and Model Merging

After fine-tuning, the trained LoRA adapter is saved separately:

tinyllama-lora-adapter/

The adapter weights are also merged with the original TinyLlama model to create a standalone fine-tuned model:

tinyllama-merged/

This produces two model artifacts:

  • LoRA Adapter — stores the parameter-efficient task-specific updates.
  • Merged Model — stores the base model with the LoRA updates incorporated into its weights.

The notebook calculates and compares the storage sizes of both artifacts.


TinyLlama Evaluation

Qualitative Evaluation

The base TinyLlama model and the LoRA fine-tuned model are evaluated using the same instruction prompts.

The prompts progressively introduce constraints such as:

  • Generating a slogan
  • Using uppercase text
  • Ending with a specified word
  • Producing an exact number of words
  • Returning the answer in JSON format

The outputs are compared in terms of:

  • Instruction-following behavior
  • Response relevance
  • Structural compliance
  • Formatting compliance
  • Ability to satisfy multiple simultaneous constraints

Qualitative generation uses sampling with:

temperature = 0.7

Because sampling is enabled, individual responses may vary between executions.

Quantitative Evaluation with IFEval

Instruction-following performance is also evaluated with IFEval, using lm-evaluation-harness and vLLM.

Evaluation configuration:

Task: ifeval
Few-shot examples: 0
Temperature: 0.0
Top-p: 1.0
Maximum generation tokens: 1024
Inference precision: float16

The notebook extracts:

  • Instruction-level strict accuracy
  • Prompt-level strict accuracy
  • Instruction-level loose accuracy
  • Prompt-level loose accuracy
  • Average strict accuracy
  • Average loose accuracy
  • Fine-tuned vs. base metric differences

IFEval paper:

https://arxiv.org/abs/2311.07911

Evaluation framework:

https://github.com/EleutherAI/lm-evaluation-harness

2. ALBERT IMDb Sentiment Classification

Notebook:

albert_imdb_sentiment/albert_lora_imdb_sentiment.ipynb

This project applies LoRA to:

albert/albert-base-v1

for binary IMDb sentiment classification.

The workflow includes:

  • Loading and exploring IMDb movie reviews
  • Analyzing class distribution and review lengths
  • Creating a reproducible stratified train/test split
  • Tokenizing reviews with the ALBERT tokenizer
  • Building a custom PyTorch Dataset
  • Constructing training and test DataLoader objects
  • Evaluating the base ALBERT classifier
  • Applying LoRA to selected ALBERT modules
  • Fine-tuning the parameter-efficient model
  • Measuring trainable-parameter efficiency
  • Comparing base and fine-tuned predictions
  • Computing Accuracy, Precision, Recall, and F1-score
  • Comparing confusion matrices
  • Analyzing false positives and false negatives

ALBERT Dataset

The project uses an IMDb movie-review dataset stored locally at:

albert_imdb_sentiment/data/IMDb_Reviews.csv

Each example contains:

  • A movie review
  • A binary sentiment label:
    • 0 — Negative
    • 1 — Positive

The dataset is split using an 80/20 stratified split with:

random_state = 42

Custom Dataset and DataLoaders

A custom PyTorch Dataset tokenizes reviews with the ALBERT tokenizer and returns:

  • Input IDs
  • Attention masks
  • Sentiment labels

The sequence length is selected according to the available hardware:

CUDA: 512 tokens
CPU: 256 tokens

The training DataLoader uses reproducible shuffling, while the test loader preserves a fixed evaluation order.


ALBERT Base Model Evaluation

The pretrained ALBERT checkpoint is loaded with a newly initialized two-class sequence-classification head.

In the final metric comparison, the base model produced:

Metric Base ALBERT
Accuracy 0.5016
Precision 0.5009
Recall 0.8840
F1-score 0.6395

Base confusion matrix counts:

TN = 596
FP = 4404
FN = 580
TP = 4420

The high recall is misleading when viewed alone: the base model predicts the positive class very frequently, which captures many positive reviews but also creates many false positives.


ALBERT LoRA Configuration

The PEFT configuration uses:

Rank (r): 8
LoRA alpha: 16
LoRA dropout: 0.1

LoRA is applied to:

query
key
value
dense

The classification head remains trainable through:

modules_to_save=["classifier"]

Therefore, the trainable-parameter count includes both the LoRA adapters and the classification head.

After applying LoRA:

Trainable parameters: 50,690
Total parameters: 11,735,812
Trainable percentage: 0.4319%

More than 99.5% of the model parameters remain frozen during fine-tuning.


ALBERT Training Configuration

The LoRA-adapted model is trained with:

Optimizer: AdamW
Learning rate: 2e-4
Epochs: 2
Batch size: 16

Only trainable parameters are passed to the optimizer.


ALBERT Fine-Tuned Results

After LoRA fine-tuning:

Metric Base ALBERT LoRA Fine-Tuned Difference
Accuracy 0.5016 0.9134 +0.4118
Precision 0.5009 0.8978 +0.3969
Recall 0.8840 0.9330 +0.0490
F1-score 0.6395 0.9151 +0.2756

Fine-tuned confusion matrix counts:

TN = 4469
FP = 531
FN = 335
TP = 4665

Error Reduction

The largest behavioral improvement comes from reducing false positives:

False positives: 4404 → 531
Reduction: 3873

False negatives also decrease:

False negatives: 580 → 335
Reduction: 245

The sharp decrease in false positives explains the large precision improvement from 0.5009 to 0.8978.

Overall, LoRA primarily helps ALBERT recognize negative reviews more accurately, while also improving positive-class detection and overall class balance.


Qualitative Prediction Comparison

Five reproducibly sampled test reviews are evaluated with both the independent base model and the LoRA fine-tuned model.

Across these examples:

Base model correct: 3 / 5
LoRA model correct: 4 / 5

The comparison shows that:

  • The fine-tuned model corrects a base-model error on a strongly negative review.
  • LoRA produces more decisive predictions on several correctly classified examples.
  • Both models can still fail on reviews containing mixed or misleading sentiment cues.
  • Fine-tuning improves task-specific behavior without eliminating every classification error.

Results at a Glance

TinyLlama

LoRA is used for:

  • Conversational SFT
  • Attention-projection adaptation
  • Adapter storage
  • Model merging
  • Qualitative instruction-following analysis
  • IFEval benchmarking

ALBERT

LoRA improves:

Accuracy:  50.16% → 91.34%
Precision: 50.09% → 89.78%
Recall:    88.40% → 93.30%
F1-score:  63.95% → 91.51%

while training only:

0.4319% of model parameters

including the LoRA adapters and classification head.


Repository Structure

parameter-efficient-finetuning-with-lora/
│
├── albert_imdb_sentiment/
│   ├── data/
│   │   └── IMDb_Reviews.csv
│   └── albert_lora_imdb_sentiment.ipynb
│
├── tinyllama_instruction_tuning/
│   └── tinyllama_instruction_tuning.ipynb
│
├── .env.example
├── .gitignore
├── requirements.txt
└── README.md

Generated checkpoints, LoRA adapters, merged models, evaluation outputs, caches, logs, and local environment-variable files are excluded from version control.


Technologies

  • Python
  • PyTorch
  • Hugging Face Transformers
  • Hugging Face Hub
  • Hugging Face Datasets
  • PEFT
  • LoRA
  • Hugging Face Trainer
  • NumPy
  • Pandas
  • scikit-learn
  • Matplotlib
  • tqdm
  • bitsandbytes
  • vLLM
  • lm-evaluation-harness
  • IFEval
  • Jupyter Notebook

Installation

Clone the repository:

git clone https://github.com/Hamidreza-Talei/parameter-efficient-finetuning-with-lora.git
cd parameter-efficient-finetuning-with-lora

Create a virtual environment:

python -m venv venv

Activate it on Windows:

venv\Scripts\activate

On macOS or Linux:

source venv/bin/activate

Install the required packages:

pip install -r requirements.txt

Hugging Face Authentication

Some models used during tokenizer exploration may require Hugging Face authentication.

Create a local .env file based on:

.env.example

and define:

HF_TOKEN=your_huggingface_token_here

Never commit the real .env file or Hugging Face access token.


Running the Notebooks

Start Jupyter Notebook from the repository root:

jupyter notebook

TinyLlama Instruction Tuning

Open:

tinyllama_instruction_tuning/tinyllama_instruction_tuning.ipynb

Recommended execution order:

  1. Load and explore the instruction-following dataset.
  2. Compare tokenizer behavior.
  3. Prepare the conversational SFT dataset.
  4. Load the TinyLlama base model.
  5. Configure and attach LoRA.
  6. Train the model.
  7. Save the LoRA adapter.
  8. Merge the adapter with the base model.
  9. Perform qualitative evaluation.
  10. Run IFEval on the base model.
  11. Run IFEval on the fine-tuned model.
  12. Compare evaluation metrics.

ALBERT IMDb Sentiment

Open:

albert_imdb_sentiment/albert_lora_imdb_sentiment.ipynb

Recommended execution order:

  1. Load and explore the IMDb dataset.
  2. Create the stratified train/test split.
  3. Initialize the ALBERT tokenizer.
  4. Build the custom PyTorch dataset and data loaders.
  5. Load and evaluate the base ALBERT model.
  6. Configure and attach LoRA.
  7. Inspect trainable parameters.
  8. Fine-tune the LoRA-adapted model.
  9. Evaluate the fine-tuned model.
  10. Compare five sample predictions.
  11. Compute classification metrics.
  12. Compare confusion matrices and analyze errors.

The notebooks should be executed sequentially because later cells depend on objects created earlier in the workflow.

A CUDA-capable GPU is strongly recommended for TinyLlama training and IFEval evaluation. The ALBERT notebook can run on CPU or CUDA, with sequence length adjusted according to the available hardware.


Generated Files

The TinyLlama notebook may generate:

tinyllama-lora-sft/
tinyllama-lora-adapter/
tinyllama-merged/
ifeval_outputs/
├── base/
└── finetuned/

These are generated artifacts and should not be committed to the repository.


Reproducibility

The experiments use fixed random seeds where applicable.

Reproducibility measures include:

  • Fixed dataset shuffling for the 3,000-example TinyLlama training subset
  • Fixed ALBERT train/test split
  • Stratified sentiment splitting
  • Seeded PyTorch data-loader shuffling
  • Deterministic sample selection for qualitative comparisons
  • Deterministic IFEval decoding configuration

Exact training results may still vary across hardware, library versions, and GPU execution environments.


Limitations

TinyLlama Instruction Tuning

The TinyLlama experiment uses only 3,000 training examples rather than the complete instruction-following dataset.

As a result:

  • The model sees a limited subset of instruction patterns.
  • Learned instruction-following behavior may not generalize to the full dataset distribution.
  • Additional data or hyperparameter tuning could produce different results.
  • The relatively small base model limits the complexity and consistency of behaviors that can be learned.

Qualitative generation uses stochastic decoding, so individual responses may vary across runs.

ALBERT Sentiment Classification

The base ALBERT sequence-classification head is newly initialized and has not been trained for IMDb sentiment classification before the baseline evaluation.

Therefore, its near-random baseline performance is expected and should not be interpreted as the general sentiment-analysis capability of a fully fine-tuned ALBERT model.

The five-review qualitative comparison is illustrative rather than statistically representative. The full test-set metrics provide the primary quantitative evaluation.


Key Concepts

This repository demonstrates:

  • Natural Language Processing
  • Large Language Models
  • Transformer models
  • Parameter-Efficient Fine-Tuning
  • PEFT
  • LoRA
  • Low-Rank Adaptation
  • Supervised Fine-Tuning
  • Instruction tuning
  • Causal language modeling
  • Sequence classification
  • Sentiment analysis
  • Binary classification
  • Conversational datasets
  • Chat templates
  • Tokenization
  • Multilingual tokenizer comparison
  • Padding and truncation
  • Loss masking
  • FP16 training
  • Gradient accumulation
  • Gradient checkpointing
  • Attention projections
  • Query, Key, and Value projections
  • Custom PyTorch datasets
  • DataLoaders
  • Accuracy
  • Precision
  • Recall
  • F1-score
  • Confusion matrices
  • False-positive analysis
  • False-negative analysis
  • Qualitative model comparison
  • LoRA adapter checkpoints
  • Model weight merging
  • Instruction-following evaluation
  • IFEval
  • vLLM inference
  • lm-evaluation-harness
  • GPU memory management

Project Scope

The goal of this repository is to demonstrate how LoRA can adapt pretrained Transformer models efficiently across different NLP tasks and model families.

The TinyLlama project focuses on generative instruction following, while the ALBERT project focuses on discriminative sentiment classification.

Together, they illustrate the central idea of PEFT:

Pretrained Model
      ↓
Freeze Most Parameters
      ↓
Attach Small Trainable LoRA Components
      ↓
Task-Specific Fine-Tuning
      ↓
Evaluate Adapted Behavior

The experiments show that meaningful task adaptation can be achieved without updating the full set of pretrained model weights.

About

LoRA-based parameter-efficient fine-tuning for TinyLlama instruction following and ALBERT IMDb sentiment classification.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages