Lists (1)
Sort Name ascending (A-Z)
Stars
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
torchcomms: a modern PyTorch communications API
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
[OSDI' 26] Efficient LLM Serving on Commodity GPU Clusters with Data-Reduced Cross-Instance Orchestration
A tutorial on modern GPU programming for machine learning systems
A programmable distributed training system for PyTorch
An asynchronous streaming data management module for efficient post-training.
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
MCore-Bridge: Providing Megatron-Core model definitions for state-of-the-art large models and making Megatron training as simple as Transformers — with support for 300+ large language models (Qwen3…
A unified framework for easy reinforcement learning in Flow-Matching models
[ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
A unified inference and post-training framework for accelerated video generation.
StreamDiffusion, Live Stream APP
My learning notes for ML SYS.
Transforming Video Diffusion with Temporal Sparse Attention
Distributed parallel 3D-Causal-VAE for efficient training and inference
MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the BrowseComp and BrowseComp Zh, respectively.
A framework for efficient model inference with omni-modality models
Accelerating MoE with IO and Tile-aware Optimizations
Distributed Compiler based on Triton for Parallel Systems
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
(best/better) practices of megatron on veRL and tuning guide
xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs.
Training library for Megatron-based models with bidirectional Hugging Face conversion capability
A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training