Skip to content

Repository files navigation

🧠 NoteSense: Personal Knowledge Augmentation System

NoteSense is a production-grade RAG (Retrieval-Augmented Generation) system designed to unify messy personal notes and voice inspirations into a structured, searchable, and interconnected knowledge graph.

"Stop saving data. Start building a second brain that talks back."


🏗 System Architecture

NoteSense is built with a modern, cloud-native stack focused on incremental synchronization, long-term persistence, and anti-hallucination.

graph TD
    User((User))
    Notion[Notion API]
    OpenAI[OpenAI API]
    
    subgraph "Frontend (Streamlit Cloud)"
        UI[Streamlit Web App]
    end

    subgraph "Backend (GCP Cloud Run - Dockerized)"
        API[FastAPI Service]
        CH[Block-Aware Chunker]
        LE[Link Engine - LLM Verification]
    end

    subgraph "Storage Layer"
        GCS[("GCS Bucket (Cloud Persistence)")]
        FAISS[("FAISS Index (Vector Store)")]
        SQLite[("SQLite (Metadata & Incremental Stats)")]
    end

    User --> UI
    UI --> API
    API --> Notion
    API --> OpenAI
    
    %% Startup/Sync Flow
    API -- "1. Pull-on-Startup" --> GCS
    API -- "2. Vector Search & Metadata" --> FAISS & SQLite
    API -- "3. Write-on-Sync (Async)" --> GCS
Loading

🌟 Key Engineering Highlights

1. 🛡 Production-Grade Persistence on Serverless

Deployed on GCP Cloud Run, NoteSense solves the ephemeral filesystem challenge using a custom GCS Sync Layer.

  • Strategy: Pull-on-Startup / Write-on-Sync.
  • Atomic Reliability: Uses POSIX-compliant atomic renames (os.replace) for FAISS index updates to prevent corruption during mid-write container shutdowns.

2. ⚡ 3-Layer Incremental Synchronization

Synchronizing 1000+ notes could be expensive and slow. NoteSense implements a high-performance pipeline:

  • Layer 1 (API): Filter by Notion's last_edited_time.
  • Layer 2 (Content): Block-level content hashing to skip unchanged segments.
  • Layer 3 (Vector): Only calling OpenAI Embedding for truly new or modified blocks.

3. 🤖 Anti-Hallucination & Verifiable RAG

NoteSense strictly enforces a "Cite-only" policy.

  • Threshold Guard: Calculates semantic distance via FAISS; if the closest match is below the confidence threshold, the LLM is prohibited from answering.
  • Hierarchical Citations: Every answer includes clickable citations tracing back to the specific Notion page and section.

4. 🔗 Autonomous Knowledge Link Engine

Beyond simple search, the Link Engine runs a two-stage background process (Vector Recall + LLM Verification) to discover hidden relationships between your voice notes and written pages, building an evolving knowledge graph.


🛠 Tech Stack

  • Frameworks: FastAPI (Backend), Streamlit (Frontend)
  • AI Engines: GPT-4o-mini, Text-Embedding-3-Small, Whisper ASR
  • Vector DB: FAISS (IndexIDMap with L2 Distance)
  • Infrastructure: Docker, Google Cloud Run, Google Cloud Storage (GCS)

🚀 Getting Started

Local Setup (Docker Compose)

The easiest way to run the full stack locally:

git clone https://github.com/Shanaia0805/NoteSense.git
cd NoteSense
cp .env.example .env # Fill in NOTION_API_TOKEN & OPENAI_API_KEY
docker-compose up --build

Manual Backend Setup

cd backend
python3.11 -m venv venv && source venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --host 0.0.0.0 --port 8000

📝 Author

NoteSense was designed for self-directed learners who want to turn "saved items" into "usable knowledge".

Feel free to reach out for collaboration or technical discussion!

About

No description, website, or topics provided.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages