I am a Master of Computer Science student at North Carolina State University with a background in Full-Stack Engineering and Generative AI. I specialize in building scalable cloud systems and high-performance AI architectures, ranging from multi-modal RAG pipelines to large-scale model optimization.
Whether it's reducing retrieval latency by 25% for enterprise data or pruning 100B+ parameters from an LLM in under 5 hours, I focus on engineering solutions that are both technically rigorous and impact-driven.
Java, C++, Python, JavaScript, SQL
PyTorch, TensorFlow, Scikit-learn, LangChain, Weaviate, Hugging Face, LLM Fine-tuning (Llama), Retrieval-Augmented Generation (RAG)
Node.js, Express, Spring Boot, ReactJS, AngularJS, TypeScript, RESTful APIs
Docker, Git/GitHub, Linux, Datadog APM, n8n, Jupyter
PostgreSQL, MongoDB, SQL Databases
Automated AI Phone Calls
Designed a dual-pipeline system using a Multi-Agentic Framework to automate client onboarding, eliminating 100% of manual configuration effort.
SparseGPT (LLM Pruning)
Targeted the memory wall of 175B+ parameter models, reducing VRAM requirements from 320GB to under 160GB without retraining.
Multi-Modal RAG Engine
Built a synthesis engine using CLIP embeddings and Llama 3.1 to unify visual and textual data from 100+ research papers into a single semantic space.
Real-time Music Streaming Engine
Developed a high-concurrency microservice system supporting 1,000+ concurrent updates with <100ms latency.
Enterprise Data Ingestion (Siemens)
Optimized internal dashboards to process 2M+ daily events, improving system reliability and detection time (MTTD) by 30%.
