ArxivLensAI: The Future of Research Interaction
ArxivLensAI is an innovative AI-powered research assistant designed to revolutionize the way researchers, students, and tech enthusiasts interact with academic content. By transforming static research papers into dynamic, interactive Q&A experiences, ArxivLensAI empowers users to delve deeper into complex topics with ease and precision.
-
๐ Upload and process multiple PDFs
-
๐ Semantic search using FAISS
-
๐ค AI-powered question answering
-
๐ Extracts tables and figures from PDFs
-
๐ผ๏ธ Supports image extraction from research papers
-
Interactive Q&A System: Pose questions and receive immediate, context-aware answers.
-
Intelligent PDF Processing: Effortlessly extract and analyze content from research papers.
-
High-Performance FAISS Storage: Leverage advanced similarity search for rapid retrieval.
-
Scalable Architecture: Built to handle extensive research databases and evolving user needs.
-
User-Centric Design: Enjoy a sleek, intuitive interface crafted for a seamless research experience.
git clone https://github.com/pranavsinghpatil/ArxivLensAI.git
cd ArxivLensAIpython -m venv venv
venv\Scripts\activate # On Windows
source venv/bin/activate # On Mac/Linuxconda create --name arxivlensai python=3.9
conda activate arxivlensaipip install -r requirements.txtstreamlit run app.py๐ก Experience the future of research. Upload your PDF and let ArxivLensAI guide your discoveries!
Explore the powerful interface through our curated screenshots:
A modular design ensures scalability and ease of maintenance:
ArxivLensAI/
โโโ app.py # Main application entry
โโโ main.py # Primary execution file
โโโ extract_text.py # PDF content extraction module
โโโ qa_system.py # AI-driven Q&A engine
โโโ vector_store.py # FAISS vector management
โโโ utils.py # Utility functions
โโโ requirements.txt # Project dependencies
โโโ README.md # Project documentation
โโโ static/
โ โโโ icons/ # UI assets and icons
โ โโโ screenshots/ # Visual documentation
โโโ extracted_images/ # Extracted images from PDFs
โโโ faiss_indexes/ # FAISS index storage
โโโ temp/ # Temporary file storage
- Set your API keys in
utils.pyor use environment variables - Required APIs:
- Google AI API (for Gemini)
- Hugging Face API
- Default embedding model:
sentence-transformers/all-MiniLM-L6-v2 - QA model:
google/flan-t5-large - Retrieval model:
deepset/roberta-base-squad2
- Extracts plain text using PyMuPDF
- Identifies and extracts tables using pdfplumber
- Processes images with OCR using Tesseract
- Builds FAISS indexes for efficient searching
- Combines multiple AI models for comprehensive answers
- Uses conversation history for context
- Supports full-context queries for detailed analysis
- Implements query expansion for better search results
- Uses FAISS for efficient similarity search
- Implements dynamic thresholding for result filtering
- Supports batch processing for better performance
We're continuously evolving ArxivLensAI. Upcoming features include:
- LangChain Integration: Elevate NLP capabilities with cutting-edge technologies.
- Advanced Content Parsing: Extract and interpret tables, images, and figures.
- Enhanced User Dashboard: Build a comprehensive web UI using Streamlit.
- Customizable AI Modules: Tailor answer generation for domain-specific research.
- Community-Driven Innovations: Incorporate user feedback and contributions to shape future releases.
We welcome contributions from developers and researchers alike. Hereโs how you can join our journey:
- Fork the Repository: Create your personal copy.
- Create a Feature Branch: Work on your ideas without affecting the main branch.
- Commit Your Changes: Ensure your commits are descriptive.
- Submit a Pull Request: Letโs collaborate to innovate together!
For more details, please refer to our Contribution Guidelines.
ArxivLensAI is released under the MIT License. See the LICENSE file for complete details.
Stay updated and get involved:
- Project Technical Report: Explore our detailed Project Documentation.
- User Guide: Refer to the User Guide for a complete walkthrough.
- API Reference: Check out the API Reference for integration details.
- Release Notes: Stay updated with our Release Notes.
For a seamless setup experience, note the following:
- Environment Management: Whether using virtualenv or Conda, ensure your environment is activated before installing dependencies.
- Dependency Updates: Regularly update your dependencies to stay compatible with the latest features.
- Community Support: Engage with our community for troubleshooting and feature discussions.
utils.py: Contains utility functions for handling API keys, generating filenames, and expanding queries.
main.py: Handles the primary execution flow, including PDF processing and FAISS index creation.
extract_text.py: Manages the extraction of text, tables, and images from PDFs using PyMuPDF, pdfplumber, and Tesseract.
vector_store.py: Implements FAISS index creation and management for efficient similarity search.
qa_system.py: Integrates multiple AI models to provide comprehensive answers to user queries.
app.py: The main application file that sets up the Streamlit interface and manages user interactions.
Made with โค๏ธ and passion for transforming the way research is experienced.




