Skip to content

Latest commit

ย 

History

92 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

ArxivLensAI ๐Ÿš€

ArxivLensAI License GitHub Release Contributions Issues Last Commit

Stars Forks

ArxivLensAI: The Future of Research Interaction
ArxivLensAI is an innovative AI-powered research assistant designed to revolutionize the way researchers, students, and tech enthusiasts interact with academic content. By transforming static research papers into dynamic, interactive Q&A experiences, ArxivLensAI empowers users to delve deeper into complex topics with ease and precision.

๐ŸŒŸ Key Features

  • ๐Ÿ“„ Upload and process multiple PDFs

  • ๐Ÿ” Semantic search using FAISS

  • ๐Ÿค– AI-powered question answering

  • ๐Ÿ“Š Extracts tables and figures from PDFs

  • ๐Ÿ–ผ๏ธ Supports image extraction from research papers

  • Interactive Q&A System: Pose questions and receive immediate, context-aware answers.

  • Intelligent PDF Processing: Effortlessly extract and analyze content from research papers.

  • High-Performance FAISS Storage: Leverage advanced similarity search for rapid retrieval.

  • Scalable Architecture: Built to handle extensive research databases and evolving user needs.

  • User-Centric Design: Enjoy a sleek, intuitive interface crafted for a seamless research experience.


๐Ÿš€ Getting Started

1๏ธโƒฃ Clone the Repository

git clone https://github.com/pranavsinghpatil/ArxivLensAI.git
cd ArxivLensAI

2๏ธโƒฃ Set Up the Environment

Using Virtualenv (pip)

python -m venv venv
venv\Scripts\activate     # On Windows
source venv/bin/activate  # On Mac/Linux

Using Conda

conda create --name arxivlensai python=3.9
conda activate arxivlensai

3๏ธโƒฃ Install Dependencies

pip install -r requirements.txt

4๏ธโƒฃ Launch the Application

streamlit run app.py

๐Ÿ’ก Experience the future of research. Upload your PDF and let ArxivLensAI guide your discoveries!


System Workflow

  1. PDF Upload & Processing PDF Upload & Processing

  2. Query Processing Query Processing


๐Ÿ“ธ Visual Showcase

Explore the powerful interface through our curated screenshots:

๐Ÿ“‚ Upload Interface

Upload

โ“ Interactive Query

Query

๐Ÿ’ก Answer Display

Answer

๐Ÿ“‚ Project Architecture

A modular design ensures scalability and ease of maintenance:

ArxivLensAI/
โ”œโ”€โ”€ app.py                # Main application entry
โ”œโ”€โ”€ main.py               # Primary execution file
โ”œโ”€โ”€ extract_text.py       # PDF content extraction module
โ”œโ”€โ”€ qa_system.py          # AI-driven Q&A engine
โ”œโ”€โ”€ vector_store.py       # FAISS vector management
โ”œโ”€โ”€ utils.py              # Utility functions
โ”œโ”€โ”€ requirements.txt      # Project dependencies
โ”œโ”€โ”€ README.md             # Project documentation
โ”œโ”€โ”€ static/
โ”‚   โ”œโ”€โ”€ icons/            # UI assets and icons
โ”‚   โ””โ”€โ”€ screenshots/      # Visual documentation
โ”œโ”€โ”€ extracted_images/     # Extracted images from PDFs
โ”œโ”€โ”€ faiss_indexes/        # FAISS index storage
โ””โ”€โ”€ temp/                 # Temporary file storage

๐Ÿ”ง Configuration

API Keys

  • Set your API keys in utils.py or use environment variables
  • Required APIs:
    • Google AI API (for Gemini)
    • Hugging Face API

Model Configuration

  • Default embedding model: sentence-transformers/all-MiniLM-L6-v2
  • QA model: google/flan-t5-large
  • Retrieval model: deepset/roberta-base-squad2

๐ŸŽฏ Features in Detail

PDF Processing

  • Extracts plain text using PyMuPDF
  • Identifies and extracts tables using pdfplumber
  • Processes images with OCR using Tesseract
  • Builds FAISS indexes for efficient searching

Question Answering

  • Combines multiple AI models for comprehensive answers
  • Uses conversation history for context
  • Supports full-context queries for detailed analysis
  • Implements query expansion for better search results

Vector Search

  • Uses FAISS for efficient similarity search
  • Implements dynamic thresholding for result filtering
  • Supports batch processing for better performance

๐Ÿ”ฎ Roadmap & Future Enhancements

We're continuously evolving ArxivLensAI. Upcoming features include:

  • LangChain Integration: Elevate NLP capabilities with cutting-edge technologies.
  • Advanced Content Parsing: Extract and interpret tables, images, and figures.
  • Enhanced User Dashboard: Build a comprehensive web UI using Streamlit.
  • Customizable AI Modules: Tailor answer generation for domain-specific research.
  • Community-Driven Innovations: Incorporate user feedback and contributions to shape future releases.

๐Ÿค Contributing

We welcome contributions from developers and researchers alike. Hereโ€™s how you can join our journey:

  1. Fork the Repository: Create your personal copy.
  2. Create a Feature Branch: Work on your ideas without affecting the main branch.
  3. Commit Your Changes: Ensure your commits are descriptive.
  4. Submit a Pull Request: Letโ€™s collaborate to innovate together!

For more details, please refer to our Contribution Guidelines.


๐Ÿ“œ License

ArxivLensAI is released under the MIT License. See the LICENSE file for complete details.


๐Ÿ’ฌ Join the Conversation

Stay updated and get involved:

GitHub Twitter


๐Ÿ“š Additional Resources


๐Ÿ› ๏ธ Setup & Maintenance Tips

For a seamless setup experience, note the following:

  • Environment Management: Whether using virtualenv or Conda, ensure your environment is activated before installing dependencies.
  • Dependency Updates: Regularly update your dependencies to stay compatible with the latest features.
  • Community Support: Engage with our community for troubleshooting and feature discussions.

Project Files Descriptions:

utils.py: Contains utility functions for handling API keys, generating filenames, and expanding queries.

main.py: Handles the primary execution flow, including PDF processing and FAISS index creation.

extract_text.py: Manages the extraction of text, tables, and images from PDFs using PyMuPDF, pdfplumber, and Tesseract.

vector_store.py: Implements FAISS index creation and management for efficient similarity search.

qa_system.py: Integrates multiple AI models to provide comprehensive answers to user queries.

app.py: The main application file that sets up the Streamlit interface and manages user interactions.


Made with โค๏ธ and passion for transforming the way research is experienced.

About

ArxivLensAI: The Transforms static research papers into a dynamic, interactive Q&A experience. Whether you're a researcher, student, or tech enthusiast, our AI-powered assistant empowers you to dive deeper into academic content with ease and precision.

Resources

Contributing

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages