AI Infrastructure Engineer focused on inference infrastructure, ML platform runtime, and workflow orchestration.
Kubeflow member and OSS contributor. I work on backend/runtime architecture, reliability, execution semantics, and performance for ML workflow systems.
Selected merged work:
- #12023 - central driver architecture proposal based on Argo Workflows Agent
- #12010 - API server gRPC metrics and execution spec reporting optimization
- #12610 - recurring runs queue throughput optimization
- #12648 - reconciliation bug fix
- #11673 - execution-level retry fix for Argo backend
- #11585 - run retry fix for Argo
- #11925 - launcher executor input parameter fix
- #16075 - WorkflowTaskSets size reduction for large workflows
- #15440 - Configure the executor plugin at the workflow level
- #994 - Socket buffer fallback for macOS compatibility
- #1153 - Support SGLang Speculative Decoding Metrics
- #31193 - Prevents SGLang server failure mode caused by unbounded Prometheus metric cardinality
- LLM inference systems
- Inference runtime optimization
- GPU performance
- Distributed runtime reliability
- Kubernetes-native workflow systems
- Backend architecture