Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

89 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ“š MasterJi - Your Intelligent Document Assistant

MasterJi is an AI-powered document assistant that allows you to chat with your documents. Upload any document, and MasterJi will read, understand, and answer questions about its content using Retrieval-Augmented Generation (RAG) technology.

Streamlit Python License

✨ Features

  • πŸ“„ Multi-format Support: Upload PDF, TXT, DOCX, XLSX, CSV, and JSON files
  • πŸ€– Smart Q&A: Ask natural language questions about your documents
  • ⚑ Fast Processing: Uses Groq's lightning-fast LLM inference
  • πŸ” Context-Aware: Answers are based only on your document content
  • 🌐 Web Interface: Beautiful Streamlit interface with real-time chat
  • πŸ’Ύ Session Management: Save and continue conversations

πŸš€ Live Demo

Try MasterJi live: https://your-masterji-app.streamlit.app

πŸ› οΈ Installation

Prerequisites

Local Setup

  1. Clone the repository
git clone https://github.com/yourusername/masterji.git
cd masterji
  1. Create virtual environment
python -m venv venv
# On Windows
venv\Scripts\activate
# On Mac/Linux
source venv/bin/activate
  1. Install dependencies
pip install -r requirements.txt
  1. Set up environment variables Create a .streamlit/secrets.toml file:
GROQ_API_KEY = "your-groq-api-key-here"
  1. Run the application
streamlit run app.py

πŸ“ Project Structure

masterji/
β”œβ”€β”€ app.py                      # Main Streamlit application
β”œβ”€β”€ data_loader.py              # Document loading and processing
β”œβ”€β”€ ChunkAndEmbed.py            # Text chunking and embedding generation
β”œβ”€β”€ vectorStore.py              # FAISS vector store management
β”œβ”€β”€ search.py                   # RAG search and LLM integration
β”œβ”€β”€ requirements.txt            # Python dependencies
β”œβ”€β”€ .streamlit/
β”‚   └── secrets.toml           # API keys (not in git)
β”œβ”€β”€ README.md                   # This file
└── data/                       # Example documents (optional)

πŸ”§ How It Works

  1. Document Upload: User uploads a document through the web interface
  2. Text Extraction: Document is parsed and text is extracted
  3. Chunking: Text is split into manageable chunks
  4. Embedding: Each chunk is converted to vector embeddings
  5. Vector Store: Embeddings are stored in a FAISS index
  6. Query Processing: User questions are embedded and matched against document chunks
  7. Response Generation: Relevant context is sent to Groq LLM for answer generation
  8. Display: Answer is shown in a chat interface

🎯 Usage

1. Get Your API Key

2. Upload a Document

  • Click "Upload Document" in the sidebar
  • Select a PDF, TXT, or other supported file
  • Click "Teach MasterJi"

3. Ask Questions

  • Type questions in the chat input
  • MasterJi will answer based on the document content
  • Ask follow-up questions

Example Queries

β€’ "What is the main topic of this document?"
β€’ "Summarize the key points"
β€’ "What does the document say about [specific topic]?"
β€’ "Explain the process described on page 3"

πŸ“Š Supported File Types

Format Extension Features
PDF .pdf Text extraction, multi-page support
Text .txt Direct text processing
Word .docx Format preservation
Excel .xlsx Tabular data extraction
CSV .csv Structured data
JSON .json Structured data with schema

πŸš€ Deployment

Deploy to Streamlit Cloud

  1. Push to GitHub
git add .
git commit -m "Initial commit"
git push origin main
  1. Deploy to Streamlit Cloud
    • Go to share.streamlit.io
    • Click "New app"
    • Connect your GitHub repository
    • Set main file to app.py
    • Add your GROQ_API_KEY in secrets
    • Click "Deploy"

Environment Variables

For deployment, set these secrets in Streamlit Cloud:

Variable Description Required
GROQ_API_KEY Your Groq API key βœ… Yes

πŸ§ͺ Testing

Run the test script to verify all components:

python test_vector_store.py

πŸ“ˆ Performance

  • Document Processing: ~30 seconds for a 10-page PDF
  • Query Response: < 2 seconds for most questions
  • Memory Usage: ~500MB for typical documents
  • Accuracy: High precision with document-specific answers

πŸ”’ Privacy & Security

  • Local Processing: All document processing happens in your environment
  • No Data Storage: Documents are processed in memory and not stored
  • API Security: API keys are stored securely in Streamlit secrets
  • Temporary Files: All uploaded files are processed in temporary directories

πŸ› οΈ Troubleshooting

Common Issues

  1. "No text extracted from document"

    • Try a different file format
    • Ensure the document contains selectable text (not scanned images)
  2. "API Key not found"

    • Check .streamlit/secrets.toml file exists
    • Verify the API key is correctly formatted
  3. Slow processing

    • Reduce document size
    • Use simpler file formats like TXT for testing
  4. Import errors

    • Reinstall dependencies: pip install -r requirements.txt
    • Check Python version (requires 3.9+)

Debug Mode

Enable debug logging by setting environment variable:

export STREAMLIT_DEBUG=1

πŸ“š Technology Stack

  • Frontend: Streamlit - Web framework
  • AI/ML: LangChain - LLM framework
  • LLM: Groq - Fast inference API
  • Embeddings: Sentence Transformers - Text embeddings
  • Vector Store: FAISS - Similarity search
  • Document Processing: PyPDF, python-docx, unstructured

🀝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/AmazingFeature)
  3. Commit your changes (git commit -m 'Add some AmazingFeature')
  4. Push to the branch (git push origin feature/AmazingFeature)
  5. Open a Pull Request

πŸ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.

πŸ™ Acknowledgments

  • Groq for providing fast LLM inference
  • LangChain for the RAG framework
  • Streamlit for the amazing web framework
  • FAISS for vector similarity search

⭐ If you find MasterJi useful, please give it a star on GitHub!


**Made with ❀️ by ARANYA CHATTERJEE **

"Knowledge shared is knowledge squared" - MasterJi πŸ“š

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages