AI Engineer in progress. I build systems around LLMs: local and offline inference, retrieval-augmented generation, and applied ML tooling. I care about projects that work end to end, not just in a notebook.
A privacy-first AI assistant that runs entirely offline. Local LLM inference through Ollama, offline speech-to-text and text-to-speech (faster-whisper and Piper), RAG over your own documents using FAISS and ONNX embeddings, and a permissioned tool system where destructive actions always require explicit confirmation. Backed by 210+ unit tests.
Point at a sound in a recording and take it out. Prompt-driven audio removal built on SAM Audio.
Core AI engineering fundamentals built from scratch, step by step: tokenization with BPE, embeddings and semantic similarity, a MiniGPT implemented in PyTorch, training loops with loss curves, encoder versus decoder architectures, prompting techniques, and API and tool-use mechanics with Gemini and Ollama.
A monitoring pipeline that detects data drift by comparing incoming production data against a training-time baseline, and raises a retraining signal when the share of drifted features crosses a configurable threshold.
Python PyTorch Hugging Face Transformers Ollama FAISS Streamlit Gemini API RAG Whisper Piper
Deepening my RAG and evaluation work, and getting my projects to a standard where someone could actually deploy them: tests, documentation, and live demos, not just working code.
Open to AI Engineer opportunities. Find me on LinkedIn. CHeck my portfolio. Find me on OmJha
