-
UC Berkeley
- Berkeley, CA
-
22:23
(UTC -07:00) - https://maoziming.github.io/
- @ziming_mao
- in/maoziming
Stars
Scalable toolkit for efficient model reinforcement
slime is an LLM post-training framework for RL Scaling.
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
htop-like TUI for real-time RDMA network monitoring.
Can LLMs Write Correct and Efficient GPU Communication Code?
The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
mKernel: fast multi-node, multi-GPU fused kernels
Ring attention implementation with flash attention
A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training
A benchmark of real-world DL kernel problems
Google Workspace CLI — one command-line tool for Drive, Gmail, Calendar, Sheets, Docs, Chat, Admin, and more. Dynamically built from Google Discovery Service. Includes AI agent skills.
Automated High-Performance GPU Kernel Generation
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
DeepGEMM: clean and efficient BLAS kernel library on GPU
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel
Research works from Tencent AI Lab regarding self-evolving agents
SysMoBench: Evaluating AI on Formally Modeling Complex Real-World Systems
Autonomous GPU Kernel Generation & Optimization via Deep Agents
Building the Virtuous Cycle for AI-driven LLM Systems
A fast communication-overlapping library for tensor/expert parallelism on GPUs.
tile-ai / tilescale
Forked from tile-ai/tilelangTile-based language built for AI computation across all scales
Tile-Based Runtime for Ultra-Low-Latency LLM Inference