I am a Research Engineer at NVIDIA (Redmond, WA), working on the algorithms and systems behind efficient large-scale distributed training of large language models.
I completed my Ph.D. in Computer Science at Université de Montréal and Mila - Quebec AI Institute, advised by Ioannis Mitliagkas. My thesis, Toward Efficient and Scalable Optimization: Theoretical Insights and Practical Challenges, studied how learning rate schedules, batch sizes, and communication patterns govern the efficiency of large-scale training.
My research centers on large-scale optimization and distributed training for deep learning, with an emphasis on high-performance computing and the efficient training of large language models. I am particularly interested in semi-synchronous and large-batch training regimes, including critical batch size analysis, as well as non-Euclidean optimization and the design of scalable optimization algorithms. More broadly, I work on the co-design of optimization algorithms and the systems they run on: using the geometry of optimization to decide what to compute exactly, what to approximate, and what to communicate.
Earlier in my Ph.D., I worked on out-of-distribution generalization and confidence calibration, and investigated optimization dynamics in generative models.
During my doctoral studies I was a research intern at Meta Superintelligence Labs (Menlo Park), working on batch size scaling and optimization for foundation models; a Student Researcher at Google DeepMind (Mountain View); and a research intern at Microsoft Research (Redmond). My work has appeared at ICML, ICLR, AISTATS, and TMLR.
I am a recipient of the Masason Foundation Fellowship and the RBC Borealis Fellowship. I received my B.Eng. (2017) and M.Eng. (2019) from the Tokyo Institute of Technology, graduating as Valedictorian of the School of Computing.