A Retrieval-Augmented Generation (RAG) based Medical AI Assistant that allows users to upload medical PDF documents, index them into a vector database, and ask natural language questions to retrieve accurate answers grounded in the uploaded content.
The system combines document retrieval, semantic search, vector embeddings, and Large Language Models (LLMs) to provide context-aware responses from medical documents.
- Upload one or multiple PDF medical documents
- Automatic document chunking and preprocessing
- Generate vector embeddings using Google's Embedding Models
- Store embeddings in Pinecone Vector Database
- Semantic similarity search for relevant document retrieval
- AI-powered question answering using LLMs
- FastAPI backend
- Streamlit frontend
- Scalable Retrieval-Augmented Generation (RAG) architecture
User Uploads PDF
β
βΌ
Document Loader (PyPDF)
β
βΌ
Text Chunking
(Chunk Size + Chunk Overlap)
β
βΌ
Google Embedding Model
(gemini-embedding-001)
β
βΌ
Vector Embeddings
β
βΌ
Pinecone Vector Database
β
βΌ
ββββββββββββββββββββββββββββ
User Asks Question
ββββββββββββββββββββββββββββ
β
βΌ
Question Embedding
β
βΌ
Similarity Search in Pinecone
β
βΌ
Relevant Document Chunks
β
βΌ
LLM (Groq / Llama)
β
βΌ
Generated Answer
β
βΌ
User Interface
Users upload medical PDF documents through the API or Streamlit interface.
Documents are split into smaller chunks to improve retrieval quality.
Configuration:
chunk_size = 500
chunk_overlap = 100Each chunk is converted into a numerical vector representation using Google's Embedding Model.
Current Model:
gemini-embedding-001
These embeddings capture semantic meaning rather than simple keyword matching.
Generated embeddings are stored in Pinecone.
Benefits:
- Fast similarity search
- Scalable storage
- Cloud-hosted vector database
- Suitable for production deployment
When a user asks a question:
- The question is converted into an embedding.
- Pinecone retrieves the most similar document chunks.
- Retrieved chunks become the context for the LLM.
The LLM receives:
- User question
- Retrieved context
- Prompt instructions
The model generates a grounded answer based on the uploaded medical documents.
- FastAPI
- Python
- LangChain
- Streamlit
- Google Generative AI
gemini-embedding-001
- Pinecone
- Groq Hosted Llama Models
Examples:
- Llama 3.3 70B Versatile
- Llama 3.1 8B Instant
Medical_AI_Assistant/
β
βββ client/
β βββ app.py
β βββ ...
β
βββ server/
β βββ routes/
β βββ modules/
β βββ main.py
β βββ ...
β
βββ uploaded_docs/
β
βββ .env
βββ requirements.txt
βββ README.md
git clone <repository-url>
cd Medical_AI_Assistantpython -m venv .venv.venv\Scripts\activatesource .venv/bin/activatepip install -r requirements.txtCreate a .env file in the root directory.
GOOGLE_API_KEY=your_google_api_key
PINECONE_API_KEY=your_pinecone_api_key
GROQ_API_KEY=your_groq_api_keyNavigate to the server directory:
cd serverStart FastAPI:
uvicorn main:app --reloadBackend will run on:
http://localhost:8000
Navigate to the client directory:
cd clientStart Streamlit:
streamlit run app.pyFrontend will run on:
http://localhost:8501
POST
/upload-pdfs/Uploads and indexes PDF documents into Pinecone.
{
"message": "Documents uploaded and indexed successfully"
}POST
/ask/{
"question": "What is diabetes?"
}{
"answer": "Diabetes is a chronic condition..."
}Upload one or more medical PDF files through:
- Swagger UI
- Postman
- Streamlit Interface
After indexing documents:
What is diabetes?
What are the symptoms of hypertension?
What treatment options are available?
The assistant retrieves relevant information from the uploaded documents and generates context-aware answers.




