This project is the Microsoft Foundry Local Summer School final project: an offline, local Retrieval-Augmented Generation (RAG) assistant for Apple Silicon M4. It uses Microsoft Foundry Local for on-device embeddings and chat inference, SQLite for local vector storage, and Python cosine similarity for retrieval. Documents, queries, and generated answers remain on the device.
User Query
-> Local RAG App
-> SQLite Vector Search
-> Retrieved Chunks
-> Augmented Context
-> Foundry Local LLM
-> Generated Answer
- Microsoft Tech Community: Building your first local RAG application with Foundry Local
- Microsoft Learn
Requires macOS ARM64 on Apple Silicon M4, Python 3.11 or newer, and Foundry Local model access. The application uses foundry-local-sdk, numpy, and Python's built-in sqlite3; no external vector database is required.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txtThe embedding model is qwen3-embedding-0.6b. The chat model is the current catalog alias phi-3.5-mini.
- Data ingestion:
python ingestion.py - Interactive CLI:
python main.py - Evaluation suite:
python evaluate.py
Problem Statement: Local documents are useful but difficult to search semantically without sending private content to a remote service.
Key Features & Offline Advantage: The assistant chunks local text, embeds it on-device, searches vectors in SQLite, and gives grounded answers with source citations. It continues to work without a cloud vector database or remote inference endpoint.
Live Demo Checklist:
- Ask an answerable question such as “How do I isolate Python project dependencies?” and show the
[Source: doc2.txt]citation. - Ask an unanswerable question such as “What is the capital of France?” and show the exact local-knowledge-base fallback.
- Show the SQLite database and the reported response latency.
Lessons Learned: Passage-level chunking improves the relevance of context; Python cosine search is simple and transparent for a small local corpus but will need indexing for larger collections; on-device inference provides privacy and independence at the cost of model loading and compute time.
Bu proje, Apple Silicon M4 üzerinde çalışan çevrimdışı yerel Retrieval-Augmented Generation (RAG) asistanıdır ve Microsoft Foundry Local Summer School final projesi planını izler. Cihaz üzerindeki embedding ve sohbet çıkarımı için Microsoft Foundry Local, yerel vektör depolama için SQLite ve retrieval işlemi için Python cosine similarity kullanılır. Belgeler, sorgular ve üretilen yanıtlar cihaz dışına çıkmaz.
Kullanıcı Sorgusu
-> Yerel RAG Uygulaması
-> SQLite Vektör Araması
-> Getirilen Chunk'lar
-> Zenginleştirilmiş Bağlam
-> Foundry Local LLM
-> Üretilen Yanıt
- Microsoft Tech Community: Building your first local RAG application with Foundry Local
- Microsoft Learn
Apple Silicon M4 üzerinde macOS ARM64, Python 3.11 veya daha yeni bir sürüm ve Foundry Local model erişimi gerekir. Uygulama foundry-local-sdk, numpy ve Python'ın yerleşik sqlite3 modülünü kullanır; harici vektör veritabanı gerekmez.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txtEmbedding modeli qwen3-embedding-0.6b'dir. Sohbet modeli mevcut katalog alias'ı olan phi-3.5-mini'dir.
- Veri ingestion:
python ingestion.py - Etkileşimli CLI:
python main.py - Evaluation paketi:
python evaluate.py
Problem Tanımı: Yerel belgeler değerlidir; ancak özel içerikleri uzak bir servise göndermeden anlamsal olarak aramak zordur.
Temel Özellikler ve Çevrimdışı Avantajı: Asistan yerel metinleri chunk'lara ayırır, cihaz üzerinde embedding üretir, SQLite içindeki vektörleri arar ve kaynak citation'ları içeren temellendirilmiş yanıtlar verir. Bulut vektör veritabanı veya uzak inference endpoint'i olmadan çalışır.
Canlı Demo Kontrol Listesi:
- “Python proje bağımlılıklarını nasıl izole ederim?” gibi yanıtlanabilir bir soru sorup
[Source: doc2.txt]citation'ını gösterin. - “Fransa'nın başkenti nedir?” gibi bilgi tabanında olmayan bir soru sorup kesin fallback mesajını gösterin.
- SQLite veritabanını ve raporlanan yanıt gecikmesini gösterin.
Öğrenilen Dersler: Passage seviyesinde chunking bağlamın ilgililiğini artırır; Python cosine araması küçük veri kümelerinde basit ve şeffaftır ancak büyük koleksiyonlarda indeksleme gerekir; cihaz üzerinde inference gizlilik ve bağımsızlık sağlar fakat model yükleme ve hesaplama süresi maliyeti getirir.
[EN] Developed as part of the 4-week intensive remote AI Project Internship cohort organized by Microsoft Turkey, under the mentorship of Barbaros Günay (Cloud Solution Architecture Manager).
- Focus: On-device, offline-first AI architectures, Microsoft Foundry SDK implementation, local Large Language Models (LLMs), and SQLite vector retrieval.
- Context: Project-based learning and evaluation cohort.
[TR] Microsoft Türkiye bünyesinde, CSA Manager Barbaros Günay mentorluğunda düzenlenen 4 haftalık yoğunlaştırılmış uzaktan proje stajı kapsamında geliştirilmiştir.
- Odak: İnternet bağlantısına ihtiyaç duymadan cihaz üzerinde (offline) çalışan AI mimarileri, Microsoft Foundry SDK, yerel büyük dil modelleri (LLM) ve SQLite vektör araması entegrasyonu.
- Kapsam: Proje tabanlı öğrenme ve değerlendirme çalışması.