Skip to content

About

Towards a Foundation Model for Chest X-Ray Interpretation Vision Language Models

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

ย 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 

Repository files navigation

LLaMA 3.2 Vision for Chest X-Ray Interpretation


๐Ÿซ What if an AI could read a chest X-ray and generate a clinical report โ€” in seconds?

That's not science fiction anymore. I just built it.

I've been working on fine-tuning LLaMA 3.2 Vision โ€” Meta's latest multimodal large language model โ€” specifically for chest X-ray interpretation. The result is a model that can look at a chest X-ray and produce clinically meaningful outputs: classifications, findings descriptions, and full radiology-style reports.


Why this matters:

There are over 2 billion chest X-rays performed globally every year. Yet radiologist shortages mean reads are delayed โ€” sometimes by hours in critical cases. In low-resource settings, some X-rays never get formally read at all.

A fine-tuned Vision-Language Model doesn't replace a radiologist. But it can:

๐Ÿ“‹ Generate a first-pass report for radiologist review ๐Ÿšจ Flag high-priority findings for faster triage ๐ŸŒ Bring diagnostic support to under-resourced hospitals ๐Ÿ“š Serve as a teaching tool for medical students


What I built:

A complete fine-tuning pipeline for LLaMA 3.2 Vision on chest X-ray datasets using the Unsloth framework โ€” achieving fast, memory-efficient training through LoRA/QLoRA parameter-efficient adaptation.

The model supports three task modes: ๐Ÿ”น Classification โ€” Normal / Pneumonia / Pleural Effusion / etc. ๐Ÿ”น Captioning โ€” Generate a descriptive findings summary ๐Ÿ”น Visual QA โ€” Answer clinical questions about an X-ray

Example outputs from fine-tuned inference: โœ… "Normal chest X-ray. No acute cardiopulmonary findings." โš ๏ธ "Findings suggest right lower lobe pneumonia. Recommend clinical correlation." ๐Ÿ”ด "Large left pleural effusion noted. Urgent evaluation advised."


Tech stack: ๐Ÿฆ™ LLaMA 3.2 Vision โ€” Meta's multimodal foundation model โšก Unsloth โ€” 2ร— faster fine-tuning, 60% less VRAM ๐Ÿค— Hugging Face Transformers + PEFT โ€” LoRA/QLoRA adaptation ๐Ÿ”ฅ PyTorch + Accelerate โ€” distributed training support ๐Ÿ“Š ChestX-ray14 / MIMIC-CXR โ€” open medical imaging datasets


What makes this technically significant:

Fine-tuning a Vision-Language Model for medical imaging is non-trivial. Medical images require the model to understand both visual pathology patterns AND clinical language simultaneously. LoRA allows us to adapt an 11B parameter model on a single GPU โ€” making this accessible to researchers without massive compute budgets.

This is part of a broader push toward Foundation Models for Medical Imaging โ€” general-purpose models pre-trained at scale, then efficiently adapted for specific clinical tasks.


๐Ÿ”— Full notebook and code on GitHub โ€” link in comments.

If you're working in medical AI, radiology informatics, or multimodal LLMs โ€” I'd love to connect and discuss where this technology is headed.

#MedicalAI #LLM #LLaMA #VisionLanguageModel #ChestXray #Radiology #HealthcareAI #MultimodalAI #DeepLearning #FoundationModels #LoRA #QLoRA #Unsloth #AIinHealthcare #MedicalImaging #NLP #ComputerVision #HuggingFace #GenerativeAI #ClinicalAI



๐ŸŽฏ Overview

This repository demonstrates how to fine-tune Meta's LLaMA 3.2 Vision on chest X-ray datasets for three clinical tasks:

  • ๐Ÿ” Classification โ€” Identify pathologies (Pneumonia, Effusion, Atelectasis, Normal, etc.)
  • ๐Ÿ“ Report Generation โ€” Produce radiology-style findings descriptions
  • ๐Ÿ’ฌ Visual Question Answering (VQA) โ€” Answer clinical questions about X-ray findings

Using Unsloth, training is 2ร— faster with 60% less VRAM than standard fine-tuning โ€” making this accessible on a single consumer or research GPU.


๐ŸŒ Motivation

There are over 2 billion chest X-rays performed globally every year.

Radiologist shortages create dangerous delays โ€” particularly in low-resource settings. Vision-Language Models fine-tuned on medical imaging data offer a path toward:

  • Automated first-pass report generation for radiologist review
  • High-priority finding flagging for faster clinical triage
  • Diagnostic support in under-resourced healthcare systems
  • Medical education and training assistance

This project is a research demonstration โ€” not a clinical product.


โœจ Features

  • โœ… LLaMA 3.2 Vision fine-tuning with Unsloth for efficient multimodal training
  • โœ… LoRA / QLoRA โ€” adapt an 11B model on a single GPU
  • โœ… Medical dataset integration โ€” ChestX-ray14, MIMIC-CXR, or custom datasets
  • โœ… Three task modes โ€” classification, captioning, visual QA
  • โœ… Evaluation metrics โ€” Accuracy, BLEU, ROUGE
  • โœ… End-to-end inference โ€” raw X-ray image to clinical text output

๐Ÿง  Pipeline

Chest X-Ray Image
      โ†“
LLaMA 3.2 Vision Encoder (frozen)
      โ†“
Cross-Modal Attention โ€” Image + Text
      โ†“
LoRA-adapted Language Model Head
      โ†“
Clinical Text Output
  โ†’ "No acute cardiopulmonary findings."
  โ†’ "Right lower lobe pneumonia. Clinical correlation recommended."
  โ†’ "Large left pleural effusion. Urgent evaluation advised."

LoRA Config:

  • Rank r=16 | Target: q_proj, v_proj, k_proj, o_proj
  • 4-bit QLoRA quantization
  • ~1โ€“2% trainable parameters of total model

๐Ÿ“Š Example Outputs

X-Ray Finding Model Output
Normal PA film โœ… "Normal chest X-ray. No acute findings identified."
Lower lobe opacity โš ๏ธ "Right lower lobe consolidation consistent with pneumonia."
Left fluid collection ๐Ÿ”ด "Large left pleural effusion with compressive atelectasis."
Enlarged heart โš ๏ธ "Increased cardiac silhouette suggestive of cardiomegaly."

๐Ÿš€ Quick Start

# 1. Clone
git clone https://github.com/your-username/llm-chest-xray.git
cd llm-chest-xray

# 2. Install
pip install unsloth transformers datasets accelerate peft bitsandbytes

# 3. Open notebook
jupyter notebook Llama_3_2_Vision_Finetuning_Unsloth_Xrays.ipynb

Dataset format:

{
  "image": "path/to/xray.jpg",
  "label": "Pneumonia",
  "report": "Findings suggest right lower lobe consolidation..."
}

Supported datasets: NIH ChestX-ray14 ยท MIMIC-CXR


๐Ÿ“ Project Structure

llm-chest-xray/
โ”œโ”€โ”€ Llama_3_2_Vision_Finetuning_Unsloth_Xrays.ipynb   # Main notebook
โ”œโ”€โ”€ data/
โ”‚   โ”œโ”€โ”€ train/                                          # Training images
โ”‚   โ”œโ”€โ”€ val/                                            # Validation images
โ”‚   โ””โ”€โ”€ dataset.json                                    # Labels / reports
โ”œโ”€โ”€ outputs/
โ”‚   โ”œโ”€โ”€ checkpoint-*/                                   # LoRA checkpoints
โ”‚   โ””โ”€โ”€ logs/
โ”œโ”€โ”€ requirements.txt
โ””โ”€โ”€ README.md

๐Ÿ› ๏ธ Tech Stack

Component Technology
Base Model LLaMA 3.2 Vision (Meta)
Fine-tuning Unsloth โ€” 2ร— faster, 60% less VRAM
Adaptation PEFT / LoRA / QLoRA (Hugging Face)
Framework PyTorch + Accelerate
Pipelines Hugging Face Transformers
Datasets ChestX-ray14, MIMIC-CXR

๐Ÿ”ฎ Roadmap

Version Feature
v1.0 Fine-tuning notebook โ€” classification + captioning โœ…
v1.1 Full MIMIC-CXR report generation pipeline
v1.2 Multi-label pathology classification
v2.0 Flask web app โ€” upload X-ray, get AI report
v2.1 DICOM (.dcm) file support
v2.2 GradCAM visual explanation overlays
v3.0 Benchmark: LLaMA vs BioViL vs CheXagent vs MedPaLM
v3.1 RLHF with radiologist feedback

โš ๏ธ Disclaimer

This project is strictly for research and educational purposes. It is not validated for clinical use and must not be used to make or influence medical decisions. All outputs require review by a qualified radiologist or physician.


๐Ÿ“– Acknowledgements


๐Ÿค Open to Collaboration

Looking to connect with:

  • ๐Ÿฅ Radiologists interested in AI-assisted reporting
  • ๐Ÿงฌ Medical AI researchers working on foundation models
  • ๐Ÿค— VLM researchers pushing multimodal medical AI
  • ๐Ÿš€ Healthcare startups building clinical AI products
  • ๐ŸŒ Global health technologists expanding diagnostic access

Let's build medical AI that actually helps people.


๐Ÿ“œ License

MIT License โ€” open for research and educational use with attribution.


โญ Star ยท ๐Ÿด Fork ยท ๐Ÿ’ฌ Contribute ยท ๐Ÿ“ค Share

Built with โค๏ธ for the future of medical AI ยท Mansoor Ahmad ยท AI & Robotics Engineer ยท NSTP Islamabad

``` ``` Chest X-Ray Image โ†“ LLaMA 3.2 Vision Encoder (frozen) โ†“ Cross-Modal Attention โ€” Image + Text โ†“ LoRA-adapted Language Model Head โ†“ Clinical Text Output โ†’ "No acute cardiopulmonary findings." โ†’ "Right lower lobe pneumonia. Clinical correlation recommended." โ†’ "Large left pleural effusion. Urgent evaluation advised." ```

LoRA Config:

  • Rank r=16 | Target: q_proj, v_proj, k_proj, o_proj
  • 4-bit QLoRA quantization
  • ~1โ€“2% trainable parameters of total model

๐Ÿ“Š Example Outputs

X-Ray Finding Model Output
Normal PA film โœ… "Normal chest X-ray. No acute findings identified."
Lower lobe opacity โš ๏ธ "Right lower lobe consolidation consistent with pneumonia."
Left fluid collection ๐Ÿ”ด "Large left pleural effusion with compressive atelectasis."
Enlarged heart โš ๏ธ "Increased cardiac silhouette suggestive of cardiomegaly."

๐Ÿš€ Quick Start

# 1. Clone
git clone https://github.com/your-username/llm-chest-xray.git
cd llm-chest-xray

# 2. Install
pip install unsloth transformers datasets accelerate peft bitsandbytes

# 3. Open notebook
jupyter notebook Llama_3_2_Vision_Finetuning_Unsloth_Xrays.ipynb

Dataset format:

{
  "image": "path/to/xray.jpg",
  "label": "Pneumonia",
  "report": "Findings suggest right lower lobe consolidation..."
}

Supported datasets: NIH ChestX-ray14 ยท MIMIC-CXR


๐Ÿ“ Project Structure

llm-chest-xray/
โ”œโ”€โ”€ Llama_3_2_Vision_Finetuning_Unsloth_Xrays.ipynb   # Main notebook
โ”œโ”€โ”€ data/
โ”‚   โ”œโ”€โ”€ train/                                          # Training images
โ”‚   โ”œโ”€โ”€ val/                                            # Validation images
โ”‚   โ””โ”€โ”€ dataset.json                                    # Labels / reports
โ”œโ”€โ”€ outputs/
โ”‚   โ”œโ”€โ”€ checkpoint-*/                                   # LoRA checkpoints
โ”‚   โ””โ”€โ”€ logs/
โ”œโ”€โ”€ requirements.txt
โ””โ”€โ”€ README.md

๐Ÿ› ๏ธ Tech Stack

Component Technology
Base Model LLaMA 3.2 Vision (Meta)
Fine-tuning Unsloth โ€” 2ร— faster, 60% less VRAM
Adaptation PEFT / LoRA / QLoRA (Hugging Face)
Framework PyTorch + Accelerate
Pipelines Hugging Face Transformers
Datasets ChestX-ray14, MIMIC-CXR

๐Ÿ”ฎ Roadmap

Version Feature
v1.0 Fine-tuning notebook โ€” classification + captioning โœ…
v1.1 Full MIMIC-CXR report generation pipeline
v1.2 Multi-label pathology classification
v2.0 Flask web app โ€” upload X-ray, get AI report
v2.1 DICOM (.dcm) file support
v2.2 GradCAM visual explanation overlays
v3.0 Benchmark: LLaMA vs BioViL vs CheXagent vs MedPaLM
v3.1 RLHF with radiologist feedback

โš ๏ธ Disclaimer

This project is strictly for research and educational purposes. It is not validated for clinical use and must not be used to make or influence medical decisions. All outputs require review by a qualified radiologist or physician.


๐Ÿ“– Acknowledgements


๐Ÿค Open to Collaboration

Looking to connect with:

  • ๐Ÿฅ Radiologists interested in AI-assisted reporting
  • ๐Ÿงฌ Medical AI researchers working on foundation models
  • ๐Ÿค— VLM researchers pushing multimodal medical AI
  • ๐Ÿš€ Healthcare startups building clinical AI products
  • ๐ŸŒ Global health technologists expanding diagnostic access

Let's build medical AI that actually helps people.


๐Ÿ“œ License

MIT License โ€” open for research and educational use with attribution.


โญ Star ยท ๐Ÿด Fork ยท ๐Ÿ’ฌ Contribute ยท ๐Ÿ“ค Share

About

Towards a Foundation Model for Chest X-Ray Interpretation Vision Language Models

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages