Skip to content
View duoan's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report duoan

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
10 stars written in Cuda
Clear filter

Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

Cuda 11,795 1,241 Updated Aug 17, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 10,043 1,395 Updated Aug 20, 2026

Tile primitives for speedy kernels

Cuda 3,640 318 Updated Aug 18, 2026
Cuda 513 87 Updated Dec 18, 2025

Learnings and programs related to CUDA

Cuda 439 20 Updated Jun 29, 2025

Reverse engineering NVIDIA SASS instruction dictionary, kernel audits and pattern recognition across GPU architectures.

Cuda 323 18 Updated May 18, 2026

Distributed MoE in a Single Kernel [NeurIPS '25]

Cuda 284 41 Updated May 5, 2026

🍎 One kernel a day keeps high latency away. A hands-on CUDA learning path featuring a rich collection of kernels, from the basics to peak performance, seamlessly integrated as PyTorch C++ extensions.

Cuda 211 18 Updated Aug 10, 2026