Field Application Engineer @ Taiwan AILabs
I am a Field Application Engineer targeting customer-facing AI Solutions Architect roles, with hands-on work across Generative AI (GenAI), LLM systems, and NVIDIA GPU computing. I translate customer requirements into production-ready, on-premises deployments on Linux, Kubernetes, PyTorch, CUDA, TensorRT-LLM, vLLM, and Triton—from proof of concept (PoC) through troubleshooting, acceptance, and handover.
Portfolio · PR wall · LinkedIn · CV
- AI solutions architecture: turn customer requirements and site constraints into GenAI/LLM deployment plans, acceptance criteria, and production handover.
- GPU inference: benchmark and debug serving, communication, quantization, and kernel paths across vLLM, TensorRT-LLM, SGLang, Triton, Dynamo, and FlashInfer.
- Upstream correctness: 18 merged/landed changes, with the complete live record on prs.wayne.is-a.dev.
| Project | Evidence |
|---|---|
| trtllm-triton-serving | TensorRT-LLM vs vLLM on H100; 12 controlled studies |
| tensor-core-from-scratch | 10 CUDA matmul kernels from naive to Tensor Cores |
| inference-kernel-cookbook | Flash Attention, KV cache, and paged attention from scratch |
| nccl-collectives-bench | NCCL bandwidth, latency, NVLS, and TP-decode limits |
| nim-agent-blueprint | NVIDIA NIM agentic RAG with evaluation and observability |
| llm-security-lab | Reproducible LLM attacks and defenses |



