Skip to content

Repository files navigation

Offline Local RAG Assistant

English

Project Overview & Aim

This project is the Microsoft Foundry Local Summer School final project: an offline, local Retrieval-Augmented Generation (RAG) assistant for Apple Silicon M4. It uses Microsoft Foundry Local for on-device embeddings and chat inference, SQLite for local vector storage, and Python cosine similarity for retrieval. Documents, queries, and generated answers remain on the device.

Architecture Flow

User Query
    -> Local RAG App
    -> SQLite Vector Search
    -> Retrieved Chunks
    -> Augmented Context
    -> Foundry Local LLM
    -> Generated Answer

Official References

Environment & Setup

Requires macOS ARM64 on Apple Silicon M4, Python 3.11 or newer, and Foundry Local model access. The application uses foundry-local-sdk, numpy, and Python's built-in sqlite3; no external vector database is required.

python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt

The embedding model is qwen3-embedding-0.6b. The chat model is the current catalog alias phi-3.5-mini.

Execution Guide

  1. Data ingestion: python ingestion.py
  2. Interactive CLI: python main.py
  3. Evaluation suite: python evaluate.py

Final Presentation & Demo Guide

Problem Statement: Local documents are useful but difficult to search semantically without sending private content to a remote service.

Key Features & Offline Advantage: The assistant chunks local text, embeds it on-device, searches vectors in SQLite, and gives grounded answers with source citations. It continues to work without a cloud vector database or remote inference endpoint.

Live Demo Checklist:

  • Ask an answerable question such as “How do I isolate Python project dependencies?” and show the [Source: doc2.txt] citation.
  • Ask an unanswerable question such as “What is the capital of France?” and show the exact local-knowledge-base fallback.
  • Show the SQLite database and the reported response latency.

Lessons Learned: Passage-level chunking improves the relevance of context; Python cosine search is simple and transparent for a small local corpus but will need indexing for larger collections; on-device inference provides privacy and independence at the cost of model loading and compute time.

Demo Video

Watch the project demo video


Türkçe

Proje Özeti ve Amaç

Bu proje, Apple Silicon M4 üzerinde çalışan çevrimdışı yerel Retrieval-Augmented Generation (RAG) asistanıdır ve Microsoft Foundry Local Summer School final projesi planını izler. Cihaz üzerindeki embedding ve sohbet çıkarımı için Microsoft Foundry Local, yerel vektör depolama için SQLite ve retrieval işlemi için Python cosine similarity kullanılır. Belgeler, sorgular ve üretilen yanıtlar cihaz dışına çıkmaz.

Mimari Akış

Kullanıcı Sorgusu
    -> Yerel RAG Uygulaması
    -> SQLite Vektör Araması
    -> Getirilen Chunk'lar
    -> Zenginleştirilmiş Bağlam
    -> Foundry Local LLM
    -> Üretilen Yanıt

Resmi Referanslar

Ortam ve Kurulum

Apple Silicon M4 üzerinde macOS ARM64, Python 3.11 veya daha yeni bir sürüm ve Foundry Local model erişimi gerekir. Uygulama foundry-local-sdk, numpy ve Python'ın yerleşik sqlite3 modülünü kullanır; harici vektör veritabanı gerekmez.

python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt

Embedding modeli qwen3-embedding-0.6b'dir. Sohbet modeli mevcut katalog alias'ı olan phi-3.5-mini'dir.

Çalıştırma Rehberi

  1. Veri ingestion: python ingestion.py
  2. Etkileşimli CLI: python main.py
  3. Evaluation paketi: python evaluate.py

Final Sunum ve Demo Rehberi

Problem Tanımı: Yerel belgeler değerlidir; ancak özel içerikleri uzak bir servise göndermeden anlamsal olarak aramak zordur.

Temel Özellikler ve Çevrimdışı Avantajı: Asistan yerel metinleri chunk'lara ayırır, cihaz üzerinde embedding üretir, SQLite içindeki vektörleri arar ve kaynak citation'ları içeren temellendirilmiş yanıtlar verir. Bulut vektör veritabanı veya uzak inference endpoint'i olmadan çalışır.

Canlı Demo Kontrol Listesi:

  • “Python proje bağımlılıklarını nasıl izole ederim?” gibi yanıtlanabilir bir soru sorup [Source: doc2.txt] citation'ını gösterin.
  • “Fransa'nın başkenti nedir?” gibi bilgi tabanında olmayan bir soru sorup kesin fallback mesajını gösterin.
  • SQLite veritabanını ve raporlanan yanıt gecikmesini gösterin.

Öğrenilen Dersler: Passage seviyesinde chunking bağlamın ilgililiğini artırır; Python cosine araması küçük veri kümelerinde basit ve şeffaftır ancak büyük koleksiyonlarda indeksleme gerekir; cihaz üzerinde inference gizlilik ve bağımsızlık sağlar fakat model yükleme ve hesaplama süresi maliyeti getirir.

Demo Videosu

Proje demo videosunu izleyin

🎓 Program & Mentorship

[EN] Developed as part of the 4-week intensive remote AI Project Internship cohort organized by Microsoft Turkey, under the mentorship of Barbaros Günay (Cloud Solution Architecture Manager).

  • Focus: On-device, offline-first AI architectures, Microsoft Foundry SDK implementation, local Large Language Models (LLMs), and SQLite vector retrieval.
  • Context: Project-based learning and evaluation cohort.

[TR] Microsoft Türkiye bünyesinde, CSA Manager Barbaros Günay mentorluğunda düzenlenen 4 haftalık yoğunlaştırılmış uzaktan proje stajı kapsamında geliştirilmiştir.

  • Odak: İnternet bağlantısına ihtiyaç duymadan cihaz üzerinde (offline) çalışan AI mimarileri, Microsoft Foundry SDK, yerel büyük dil modelleri (LLM) ve SQLite vektör araması entegrasyonu.
  • Kapsam: Proje tabanlı öğrenme ve değerlendirme çalışması.

About

Offline-first Local RAG assistant built during the 4-week Microsoft Turkey AI Project Internship mentored by CSA Manager Barbaros Günay. Powered by Microsoft Foundry SDK & SQLite vector search.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages