TL

Staff Software Engineer (Full-Stack) · Transitioning to ML Inference

Tony Lee

I build full-stack products and lead architecture across cross-regional teams. I'm now transitioning into ML inference engineering through self-directed, reproducible labs — vLLM GPU serving on AWS/EKS and CUDA/Triton kernel benchmarking, measured honestly, including the negative results.

Latest writing

Read the blog