Research Engineer @ XFlow Research
- I’m currently exploring optimizing LLM inference and autonomous multi-agent workflows.
- Reach me at mahadrehman04@gmail.com
vLLM · high-throughput and memory-efficient LLM inference engine
Research Engineer @ XFlow Research
vLLM · high-throughput and memory-efficient LLM inference engine
Forked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python
An AI-powered self-driving car model built using machine learning and neural networks to autonomously control vehicle behavior.
Python
A multi-stage evaluation pipeline for building and validating a hallucination-resistant PDF chatbot using LangChain and OpenAI.
Python
A simple yet powerful Flask-based weekly email scheduler with Gemini-generated content and Outlook integration.
Python
A RAG-based chatbot using custom mock embeddings and vector store for local document Q&A.
Python