Highlights
Stars
slime is an LLM post-training framework for RL Scaling.
🧩 Hands-on SIMD Programming with C++
easydel jax kernels writen in triton for gpus and pallas for tpus
[NeurIPS 2025] Simple extension on vLLM to help you speed up reasoning model without training.
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades
Complete solutions to the Programming Massively Parallel Processors Edition 4
Glorious Engrammer keymap for Glove80 keyboard
A resource repository for machine unlearning in large language models
Playing around "Less Slow" coding practices in C++ 20, C, CUDA, PTX, & Assembly, from numerics & SIMD to coroutines, ranges, exception handling, networking and user-space IO
My learning notes for ML SYS.
Repository for training transformer _and recurrent_ language models via HuggingFace in an entirely configuration-file driven manner.
Samples for CUDA Developers which demonstrates features in CUDA Toolkit
scalable and robust tree-based speculative decoding algorithm
Faster Pytorch bitsandbytes 4bit fp4 nn.Linear ops
GPU programming related news and material links
Training materials associated with NVIDIA's CUDA Training Series (www.olcf.ornl.gov/cuda-training-series/)
🤖 A PyTorch library of curated Transformer models and their composable components
A concise but complete full-attention transformer with a set of promising experimental features from various papers
⚡ Dynamically generated stats for your github readmes
Certified Kubernetes Administrator - CKA Course
Single-document unsupervised keyword extraction