I'm a Computer Science Engineering student at Madhav Institute of Technology and Science, Gwalior, focused on the intersection of computer vision, retrieval-augmented generation, and applied machine learning. I like building systems that turn raw, messy data β video, documents, sensor feeds β into something a model can reason over and a user can actually query.
Currently working as a Machine Learning Intern at Warewolf, where I'm building a YOLO-based object detection pipeline for surveillance imagery β from dataset curation and annotation to model training and evaluation.
My engineering approach leans toward product-minded ML: every model I train has to end up behind a usable interface, not just a notebook metric.
π Currently Building: Multimodal RAG pipelines for video + document intelligence
π± Currently Learning: DSPy, GEPA, and prompt/pipeline optimization for LLM systems
π§ Exploring: Agentic retrieval architectures & LLM fine-tuning
π― Open To: ML/AI Engineering roles Β· Computer Vision research Β· Open-source collaborationLanguages
AI / ML & Computer Vision
Data & Visualization
Databases & Tooling
| Domain | Proficiency | Details |
|---|---|---|
| Computer Vision | βββββ | YOLO object detection, OpenCV pipelines, Roboflow dataset annotation & curation |
| Retrieval-Augmented Generation | βββββ | LangChain, FAISS vector search, multimodal (video + document) retrieval |
| LLM Integration | βββββ | Llama-2, Whisper AI, Hugging Face model integration & fine-tuning |
| Classical ML | βββββ | Scikit-learn pipelines, regression models, TF-IDF, forecasting |
| LLM Pipeline Optimization | βββββ | DSPy, GEPA β prompt & pipeline optimization workflows |
| Data Engineering | βββββ | Data cleaning, preprocessing, deduplication, EDA at scale |
π¬ MultimodalRAG β Intelligent Video & Document Assistant
A multimodal retrieval-augmented generation system that lets you query across video and document content simultaneously β transcribing, indexing, and semantically searching both in one pipeline.
| Stack | Python, LangChain, Whisper AI, FAISS, Groq API, MoviePy, Streamlit, Sentence Transformers |
| Scale | Supports files up to 100MB across MP4/MOV/AVI and PDF/TXT/DOCX/CSV |
| Performance | Sub-second cross-modal retrieval, 40% faster video processing |
| Security | Local vector indexing, no third-party data persistence |
| Impact | 95% retrieval accuracy across mixed video/document corpora |
| Repository | GitHub |
Built with a dark-themed Streamlit interface delivering real-time, source-cited Q&A over transcribed video and parsed documents β combining Whisper AI transcription with FAISS-backed semantic search.
π€ MASTERJI β RAG-Powered Chatbot
A retrieval-augmented chatbot architected to reduce hallucination and boost factual grounding in LLM responses using document-level context.
| Stack | Python, LangChain, Hugging Face, FAISS, Llama-2, Streamlit |
| Scale | Multi-format document ingestion (PDF/TXT) |
| Performance | Sub-second query response via FAISS vector search |
| Security | Context-scoped generation, source-grounded outputs |
| Impact | 30% improvement in response accuracy over base LLM outputs |
| Repository | GitHub |
Combines LangChain orchestration with Llama-2 generation and FAISS embeddings, wrapped in a Streamlit UI supporting real-time Q&A with fully sourced citations.
π° AI Finance Assistant
An ML-driven personal finance tool that classifies expenses and forecasts budgets from transaction data.
| Stack | Python, Scikit-learn, Streamlit, Logistic/Linear Regression, TF-IDF, Exponential Smoothing |
| Scale | Multi-model pipeline across classification + forecasting tasks |
| Performance | 92% accuracy in automated expense categorization |
| Security | Local processing, no external financial data transmission |
| Impact | 15% reduction in unnecessary spending observed in test simulations |
| Repository | GitHub |
Pairs a Scikit-learn classification pipeline with time-series forecasting, surfaced through an interactive Streamlit dashboard for budget visibility.
Machine Learning Intern Β· Warewolf
May 2026 β Present
Building and refining a YOLO-based object detection pipeline for surveillance image and video analysis, working within a cross-functional engineering team.
- Curated and annotated a ~780-image surveillance dataset using Roboflow (bounding-box labeling, class organization)
- Performed data cleaning and preprocessing (filtering, deduplication, normalization) to improve training reliability
- Used OpenCV for frame extraction, image preprocessing, and detection-output visualization
Python YOLO OpenCV Roboflow Computer Vision
| Recognition | Details |
|---|---|
| π₯ National Finalist | Hack-O-Calypse, IIT Jammu |
| π₯ National Finalist & Team Leader | Code-N-Compete Hackathon, IIT Kanpur β 48-hr sprint, 25% faster integration phase |
| π Published Researcher | Second Author, "A Hybrid IoT + AI Approach for Risk Management in Algorithmic Trading", ICMCTR 2025 |
| β HackerRank 5-Star | Java |