I work at the intersection of recommendation algorithms and inference infrastructure. The systems I build select, rank, and serve content under strict latency budgets at scale.
At BlueFocus/Domob, I own the full programmatic advertising pipeline: traffic shaping, real-time bidding, CTR/CVR prediction, and bid optimization across 30+ ad exchanges.
Search, Ads & Recommendation — CTR/CVR multi-objective modeling, retrieval-rank-rerank pipeline, bid strategy optimization, real-time bidding systems.
ML Inference Engineering — TensorFlow Serving at scale, 100ms latency SLA, feature pipelines, model calibration.
LLM Inference & Agent — KV Cache optimization, quantization & parallelism, agent orchestration, vLLM / SGLang internals.
Technical deep-dives at asterzephyr.xyz:
- GPU 推理部署学习指南:从显存计算到性能优化
- vLLM KV Cache Block Manager 深度教程
- 视频生成推理的 GPU 算力:从一道算术题说起
- 为什么强化学习训练大模型这么难
- Agent Eval 全景:怎么评、怎么设计、怎么学
Minor contributions to Apache Seata-Go and Higress (AI Native API Gateway).


