Krish Agarwal

PhD Student, Carnegie Mellon University

prof_pic.jpg

I’m a second-year PhD student in InfiniAI Lab at Carnegie Mellon University, where I’m fortunate to be advised by Prof. Beidi Chen. My research is on efficient ML systems, spanning structured and sparse attention, GPU kernels, serving, and infrastructure for AI agents.

I completed my undergrad in CS at UT Austin as a Turing Scholar, where I wrote my honors thesis on Tensor Core accelerated kernels for structured masked attention and worked with Prof. Peter Stone on LLM-based planning for service robots. I’ve also interned at NVIDIA and IBM Research.

news

Sep 2026 MonarchRT was accepted to NeurIPS 2026!
Jul 2026 Released FlashRT, an agent harness that guides coding agents to turn simple reference implementations of real-time multimodal applications into optimized multi-GPU deployments.
Feb 2026 Released MonarchRT, a highly effective sparse attention parameterization for real-time video generation.
Jan 2026 Teaching assistant for 18-789: Deep Generative Modeling at CMU this spring.
Aug 2025 Started my PhD at CMU in InfiniAI Lab, advised by Prof. Beidi Chen!

selected publications

  1. MonarchRT: Efficient Attention for Real-Time Video Generation
    Krish Agarwal, Zhuoming Chen, Cheng Luo, Yongqi Chen, Haizhong Zheng, Xun Huang, Atri Rudra, and Beidi Chen
    In Advances in Neural Information Processing Systems (NeurIPS), Sydney, Australia, Dec 2026
  2. arXiv
    flashrt.png
    FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
    Krish Agarwal, Zhuoming Chen, Yanyuan Qin, Zhenyu Gu, Atri Rudra, and Beidi Chen
    Under review , 2026
  3. arXiv
    hadacore.png
    HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
    Krish Agarwal*, Rishi Astra*, Adnan Hoque, Mudhakar Srivatsa, Raghu Ganti, Less Wright, and Sijia Chen
    Preprint (featured on the PyTorch Blog) , 2024