I build reliable and verifiable systems for large-scale model development, with a focus on distributed training, correctness, diagnosability, and resilience.
My work spans pretraining, post-training and reinforcement learning, inference and rollout, and agent and data-generation infrastructure. Recent work has supported internal dense and MoE training programs on up to 4,096 GPUs.
Previously, I was a Senior Research Engineer at Microsoft Research Asia. I was a core contributor to nnScaler (OSDI 2024), from the early development of AutoDist through later system development and model-research integrations. I also contributed real-world bug cases, design feedback, and technical support to TrainVerify (SOSP 2025).
My interests include:
- Large-scale distributed training and parallel execution
- Training correctness, diagnosis, and system resilience
- Post-training, RL, inference, and rollout systems
- Agent and data-generation infrastructure
- Model-system co-design
More details and technical writing: zyeric.github.io/cv