I'm a software engineer focusing on Search, AI Infra and LLM systems.
Currently working on retrieval, inference infrastructure and engineering systems.
🎯
Focusing
Pinned Loading
-
xllm
xllm PublicForked from xLLM-AI/xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
C++
-
vllm
vllm PublicForked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python
-
milvus
milvus PublicForked from milvus-io/milvus
A cloud-native vector database, storage for next generation AI applications
Go
-
tfx
tfx PublicForked from tensorflow/tfx
TFX is an end-to-end platform for deploying production ML pipelines
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

