๐ซ What if an AI could read a chest X-ray and generate a clinical report โ in seconds?
That's not science fiction anymore. I just built it.
I've been working on fine-tuning LLaMA 3.2 Vision โ Meta's latest multimodal large language model โ specifically for chest X-ray interpretation. The result is a model that can look at a chest X-ray and produce clinically meaningful outputs: classifications, findings descriptions, and full radiology-style reports.
Why this matters:
There are over 2 billion chest X-rays performed globally every year. Yet radiologist shortages mean reads are delayed โ sometimes by hours in critical cases. In low-resource settings, some X-rays never get formally read at all.
A fine-tuned Vision-Language Model doesn't replace a radiologist. But it can:
๐ Generate a first-pass report for radiologist review ๐จ Flag high-priority findings for faster triage ๐ Bring diagnostic support to under-resourced hospitals ๐ Serve as a teaching tool for medical students
What I built:
A complete fine-tuning pipeline for LLaMA 3.2 Vision on chest X-ray datasets using the Unsloth framework โ achieving fast, memory-efficient training through LoRA/QLoRA parameter-efficient adaptation.
The model supports three task modes: ๐น Classification โ Normal / Pneumonia / Pleural Effusion / etc. ๐น Captioning โ Generate a descriptive findings summary ๐น Visual QA โ Answer clinical questions about an X-ray
Example outputs from fine-tuned inference:
โ
"Normal chest X-ray. No acute cardiopulmonary findings."
Tech stack: ๐ฆ LLaMA 3.2 Vision โ Meta's multimodal foundation model โก Unsloth โ 2ร faster fine-tuning, 60% less VRAM ๐ค Hugging Face Transformers + PEFT โ LoRA/QLoRA adaptation ๐ฅ PyTorch + Accelerate โ distributed training support ๐ ChestX-ray14 / MIMIC-CXR โ open medical imaging datasets
What makes this technically significant:
Fine-tuning a Vision-Language Model for medical imaging is non-trivial. Medical images require the model to understand both visual pathology patterns AND clinical language simultaneously. LoRA allows us to adapt an 11B parameter model on a single GPU โ making this accessible to researchers without massive compute budgets.
This is part of a broader push toward Foundation Models for Medical Imaging โ general-purpose models pre-trained at scale, then efficiently adapted for specific clinical tasks.
๐ Full notebook and code on GitHub โ link in comments.
If you're working in medical AI, radiology informatics, or multimodal LLMs โ I'd love to connect and discuss where this technology is headed.
#MedicalAI #LLM #LLaMA #VisionLanguageModel #ChestXray #Radiology #HealthcareAI #MultimodalAI #DeepLearning #FoundationModels #LoRA #QLoRA #Unsloth #AIinHealthcare #MedicalImaging #NLP #ComputerVision #HuggingFace #GenerativeAI #ClinicalAI
This repository demonstrates how to fine-tune Meta's LLaMA 3.2 Vision on chest X-ray datasets for three clinical tasks:
- ๐ Classification โ Identify pathologies (Pneumonia, Effusion, Atelectasis, Normal, etc.)
- ๐ Report Generation โ Produce radiology-style findings descriptions
- ๐ฌ Visual Question Answering (VQA) โ Answer clinical questions about X-ray findings
Using Unsloth, training is 2ร faster with 60% less VRAM than standard fine-tuning โ making this accessible on a single consumer or research GPU.
There are over 2 billion chest X-rays performed globally every year.
Radiologist shortages create dangerous delays โ particularly in low-resource settings. Vision-Language Models fine-tuned on medical imaging data offer a path toward:
- Automated first-pass report generation for radiologist review
- High-priority finding flagging for faster clinical triage
- Diagnostic support in under-resourced healthcare systems
- Medical education and training assistance
This project is a research demonstration โ not a clinical product.
- โ LLaMA 3.2 Vision fine-tuning with Unsloth for efficient multimodal training
- โ LoRA / QLoRA โ adapt an 11B model on a single GPU
- โ Medical dataset integration โ ChestX-ray14, MIMIC-CXR, or custom datasets
- โ Three task modes โ classification, captioning, visual QA
- โ Evaluation metrics โ Accuracy, BLEU, ROUGE
- โ End-to-end inference โ raw X-ray image to clinical text output
Chest X-Ray Image
โ
LLaMA 3.2 Vision Encoder (frozen)
โ
Cross-Modal Attention โ Image + Text
โ
LoRA-adapted Language Model Head
โ
Clinical Text Output
โ "No acute cardiopulmonary findings."
โ "Right lower lobe pneumonia. Clinical correlation recommended."
โ "Large left pleural effusion. Urgent evaluation advised."
LoRA Config:
- Rank r=16 | Target: q_proj, v_proj, k_proj, o_proj
- 4-bit QLoRA quantization
- ~1โ2% trainable parameters of total model
| X-Ray Finding | Model Output |
|---|---|
| Normal PA film | โ "Normal chest X-ray. No acute findings identified." |
| Lower lobe opacity | |
| Left fluid collection | ๐ด "Large left pleural effusion with compressive atelectasis." |
| Enlarged heart |
# 1. Clone
git clone https://github.com/your-username/llm-chest-xray.git
cd llm-chest-xray
# 2. Install
pip install unsloth transformers datasets accelerate peft bitsandbytes
# 3. Open notebook
jupyter notebook Llama_3_2_Vision_Finetuning_Unsloth_Xrays.ipynbDataset format:
{
"image": "path/to/xray.jpg",
"label": "Pneumonia",
"report": "Findings suggest right lower lobe consolidation..."
}Supported datasets: NIH ChestX-ray14 ยท MIMIC-CXR
llm-chest-xray/
โโโ Llama_3_2_Vision_Finetuning_Unsloth_Xrays.ipynb # Main notebook
โโโ data/
โ โโโ train/ # Training images
โ โโโ val/ # Validation images
โ โโโ dataset.json # Labels / reports
โโโ outputs/
โ โโโ checkpoint-*/ # LoRA checkpoints
โ โโโ logs/
โโโ requirements.txt
โโโ README.md
| Component | Technology |
|---|---|
| Base Model | LLaMA 3.2 Vision (Meta) |
| Fine-tuning | Unsloth โ 2ร faster, 60% less VRAM |
| Adaptation | PEFT / LoRA / QLoRA (Hugging Face) |
| Framework | PyTorch + Accelerate |
| Pipelines | Hugging Face Transformers |
| Datasets | ChestX-ray14, MIMIC-CXR |
| Version | Feature |
|---|---|
| v1.0 | Fine-tuning notebook โ classification + captioning โ |
| v1.1 | Full MIMIC-CXR report generation pipeline |
| v1.2 | Multi-label pathology classification |
| v2.0 | Flask web app โ upload X-ray, get AI report |
| v2.1 | DICOM (.dcm) file support |
| v2.2 | GradCAM visual explanation overlays |
| v3.0 | Benchmark: LLaMA vs BioViL vs CheXagent vs MedPaLM |
| v3.1 | RLHF with radiologist feedback |
This project is strictly for research and educational purposes. It is not validated for clinical use and must not be used to make or influence medical decisions. All outputs require review by a qualified radiologist or physician.
- UnslothAI โ efficient fine-tuning framework
- Hugging Face โ Transformers, PEFT, Datasets
- NIH Clinical Center โ ChestX-ray14
- PhysioNet / MIT โ MIMIC-CXR
- Meta AI โ LLaMA 3.2 Vision
Looking to connect with:
- ๐ฅ Radiologists interested in AI-assisted reporting
- ๐งฌ Medical AI researchers working on foundation models
- ๐ค VLM researchers pushing multimodal medical AI
- ๐ Healthcare startups building clinical AI products
- ๐ Global health technologists expanding diagnostic access
Let's build medical AI that actually helps people.
MIT License โ open for research and educational use with attribution.
โญ Star ยท ๐ด Fork ยท ๐ฌ Contribute ยท ๐ค Share
Built with โค๏ธ for the future of medical AI ยท Mansoor Ahmad ยท AI & Robotics Engineer ยท NSTP Islamabad
``` ``` Chest X-Ray Image โ LLaMA 3.2 Vision Encoder (frozen) โ Cross-Modal Attention โ Image + Text โ LoRA-adapted Language Model Head โ Clinical Text Output โ "No acute cardiopulmonary findings." โ "Right lower lobe pneumonia. Clinical correlation recommended." โ "Large left pleural effusion. Urgent evaluation advised." ```LoRA Config:
- Rank r=16 | Target: q_proj, v_proj, k_proj, o_proj
- 4-bit QLoRA quantization
- ~1โ2% trainable parameters of total model
| X-Ray Finding | Model Output |
|---|---|
| Normal PA film | โ "Normal chest X-ray. No acute findings identified." |
| Lower lobe opacity | |
| Left fluid collection | ๐ด "Large left pleural effusion with compressive atelectasis." |
| Enlarged heart |
# 1. Clone
git clone https://github.com/your-username/llm-chest-xray.git
cd llm-chest-xray
# 2. Install
pip install unsloth transformers datasets accelerate peft bitsandbytes
# 3. Open notebook
jupyter notebook Llama_3_2_Vision_Finetuning_Unsloth_Xrays.ipynbDataset format:
{
"image": "path/to/xray.jpg",
"label": "Pneumonia",
"report": "Findings suggest right lower lobe consolidation..."
}Supported datasets: NIH ChestX-ray14 ยท MIMIC-CXR
llm-chest-xray/
โโโ Llama_3_2_Vision_Finetuning_Unsloth_Xrays.ipynb # Main notebook
โโโ data/
โ โโโ train/ # Training images
โ โโโ val/ # Validation images
โ โโโ dataset.json # Labels / reports
โโโ outputs/
โ โโโ checkpoint-*/ # LoRA checkpoints
โ โโโ logs/
โโโ requirements.txt
โโโ README.md
| Component | Technology |
|---|---|
| Base Model | LLaMA 3.2 Vision (Meta) |
| Fine-tuning | Unsloth โ 2ร faster, 60% less VRAM |
| Adaptation | PEFT / LoRA / QLoRA (Hugging Face) |
| Framework | PyTorch + Accelerate |
| Pipelines | Hugging Face Transformers |
| Datasets | ChestX-ray14, MIMIC-CXR |
| Version | Feature |
|---|---|
| v1.0 | Fine-tuning notebook โ classification + captioning โ |
| v1.1 | Full MIMIC-CXR report generation pipeline |
| v1.2 | Multi-label pathology classification |
| v2.0 | Flask web app โ upload X-ray, get AI report |
| v2.1 | DICOM (.dcm) file support |
| v2.2 | GradCAM visual explanation overlays |
| v3.0 | Benchmark: LLaMA vs BioViL vs CheXagent vs MedPaLM |
| v3.1 | RLHF with radiologist feedback |
This project is strictly for research and educational purposes. It is not validated for clinical use and must not be used to make or influence medical decisions. All outputs require review by a qualified radiologist or physician.
- UnslothAI โ efficient fine-tuning framework
- Hugging Face โ Transformers, PEFT, Datasets
- NIH Clinical Center โ ChestX-ray14
- PhysioNet / MIT โ MIMIC-CXR
- Meta AI โ LLaMA 3.2 Vision
Looking to connect with:
- ๐ฅ Radiologists interested in AI-assisted reporting
- ๐งฌ Medical AI researchers working on foundation models
- ๐ค VLM researchers pushing multimodal medical AI
- ๐ Healthcare startups building clinical AI products
- ๐ Global health technologists expanding diagnostic access
Let's build medical AI that actually helps people.
MIT License โ open for research and educational use with attribution.
โญ Star ยท ๐ด Fork ยท ๐ฌ Contribute ยท ๐ค Share