I build GPU kernels and the CAKE kernel agent. Feel free to reach out!
Tracking all the fresh slices of CAKE served in FlashInfer — care for a taste? 🍰
I build GPU kernels and the CAKE kernel agent. Feel free to reach out!
Tracking all the fresh slices of CAKE served in FlashInfer — care for a taste? 🍰
FlashInfer: Kernel Library for LLM Serving
SGLang is a high-performance serving framework for large language models and multimodal models.
Building the Virtuous Cycle for AI-driven LLM Systems
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel