MasterJi is an AI-powered document assistant that allows you to chat with your documents. Upload any document, and MasterJi will read, understand, and answer questions about its content using Retrieval-Augmented Generation (RAG) technology.
- π Multi-format Support: Upload PDF, TXT, DOCX, XLSX, CSV, and JSON files
- π€ Smart Q&A: Ask natural language questions about your documents
- β‘ Fast Processing: Uses Groq's lightning-fast LLM inference
- π Context-Aware: Answers are based only on your document content
- π Web Interface: Beautiful Streamlit interface with real-time chat
- πΎ Session Management: Save and continue conversations
Try MasterJi live: https://your-masterji-app.streamlit.app
- Python 3.9 or higher
- Groq API key (free from console.groq.com)
- Clone the repository
git clone https://github.com/yourusername/masterji.git
cd masterji- Create virtual environment
python -m venv venv
# On Windows
venv\Scripts\activate
# On Mac/Linux
source venv/bin/activate- Install dependencies
pip install -r requirements.txt- Set up environment variables
Create a
.streamlit/secrets.tomlfile:
GROQ_API_KEY = "your-groq-api-key-here"- Run the application
streamlit run app.pymasterji/
βββ app.py # Main Streamlit application
βββ data_loader.py # Document loading and processing
βββ ChunkAndEmbed.py # Text chunking and embedding generation
βββ vectorStore.py # FAISS vector store management
βββ search.py # RAG search and LLM integration
βββ requirements.txt # Python dependencies
βββ .streamlit/
β βββ secrets.toml # API keys (not in git)
βββ README.md # This file
βββ data/ # Example documents (optional)
- Document Upload: User uploads a document through the web interface
- Text Extraction: Document is parsed and text is extracted
- Chunking: Text is split into manageable chunks
- Embedding: Each chunk is converted to vector embeddings
- Vector Store: Embeddings are stored in a FAISS index
- Query Processing: User questions are embedded and matched against document chunks
- Response Generation: Relevant context is sent to Groq LLM for answer generation
- Display: Answer is shown in a chat interface
- Visit Groq Console
- Sign up for free
- Copy your API key
- Click "Upload Document" in the sidebar
- Select a PDF, TXT, or other supported file
- Click "Teach MasterJi"
- Type questions in the chat input
- MasterJi will answer based on the document content
- Ask follow-up questions
β’ "What is the main topic of this document?"
β’ "Summarize the key points"
β’ "What does the document say about [specific topic]?"
β’ "Explain the process described on page 3"
| Format | Extension | Features |
|---|---|---|
.pdf |
Text extraction, multi-page support | |
| Text | .txt |
Direct text processing |
| Word | .docx |
Format preservation |
| Excel | .xlsx |
Tabular data extraction |
| CSV | .csv |
Structured data |
| JSON | .json |
Structured data with schema |
- Push to GitHub
git add .
git commit -m "Initial commit"
git push origin main- Deploy to Streamlit Cloud
- Go to share.streamlit.io
- Click "New app"
- Connect your GitHub repository
- Set main file to
app.py - Add your
GROQ_API_KEYin secrets - Click "Deploy"
For deployment, set these secrets in Streamlit Cloud:
| Variable | Description | Required |
|---|---|---|
GROQ_API_KEY |
Your Groq API key | β Yes |
Run the test script to verify all components:
python test_vector_store.py- Document Processing: ~30 seconds for a 10-page PDF
- Query Response: < 2 seconds for most questions
- Memory Usage: ~500MB for typical documents
- Accuracy: High precision with document-specific answers
- Local Processing: All document processing happens in your environment
- No Data Storage: Documents are processed in memory and not stored
- API Security: API keys are stored securely in Streamlit secrets
- Temporary Files: All uploaded files are processed in temporary directories
-
"No text extracted from document"
- Try a different file format
- Ensure the document contains selectable text (not scanned images)
-
"API Key not found"
- Check
.streamlit/secrets.tomlfile exists - Verify the API key is correctly formatted
- Check
-
Slow processing
- Reduce document size
- Use simpler file formats like TXT for testing
-
Import errors
- Reinstall dependencies:
pip install -r requirements.txt - Check Python version (requires 3.9+)
- Reinstall dependencies:
Enable debug logging by setting environment variable:
export STREAMLIT_DEBUG=1- Frontend: Streamlit - Web framework
- AI/ML: LangChain - LLM framework
- LLM: Groq - Fast inference API
- Embeddings: Sentence Transformers - Text embeddings
- Vector Store: FAISS - Similarity search
- Document Processing: PyPDF, python-docx, unstructured
Contributions are welcome! Please feel free to submit a Pull Request.
- Fork the repository
- Create your feature branch (
git checkout -b feature/AmazingFeature) - Commit your changes (
git commit -m 'Add some AmazingFeature') - Push to the branch (
git push origin feature/AmazingFeature) - Open a Pull Request
This project is licensed under the MIT License - see the LICENSE file for details.
- Groq for providing fast LLM inference
- LangChain for the RAG framework
- Streamlit for the amazing web framework
- FAISS for vector similarity search
β If you find MasterJi useful, please give it a star on GitHub!
**Made with β€οΈ by ARANYA CHATTERJEE **
"Knowledge shared is knowledge squared" - MasterJi π