Skip to content
View demonbibi's full-sized avatar

Block or report demonbibi

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results
Python 314 38 Updated Jun 9, 2026

A kernel library written in tilelang

Python 1,674 152 Updated Apr 23, 2026

A CUDA kernel optimization toolkit for validation, benchmarking, Nsight Compute profiling, bottleneck analysis, and iterative tuning. It helps improve custom GPU operators with reproducible workflo…

Python 192 18 Updated Apr 22, 2026

Accelerating MoE with IO and Tile-aware Optimizations

Python 732 95 Updated Jul 4, 2026

A Quirky Assortment of CuTe Kernels

Python 1,074 148 Updated Jul 27, 2026

Persistent file-based planning for AI coding agents and long-running tasks. Crash-proof markdown plans, session recovery after /clear and compaction, per-turn re-injection against context rot, dete…

Python 25,788 2,162 Updated Jul 24, 2026

An agentic skills framework & software development methodology that works.

Shell 262,078 23,398 Updated Jul 24, 2026

AI agents running research on single-GPU nanochat training automatically

Python 92,166 13,164 Updated Mar 26, 2026

高性能短序列稀疏Mask Attention CUDA算子,针对<1K序列+75%稀疏度优化

Python 80 8 Updated Mar 18, 2026

CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning

Cuda 467 32 Updated Mar 30, 2026

A Next-Generation Training Engine Built for Ultra-Large MoE Models

Python 5,164 432 Updated Jul 27, 2026

KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)

Jupyter Notebook 1,160 182 Updated Mar 24, 2026

TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators

Python 137 14 Updated Jun 14, 2025

High Performance LLM Inference Operator Library

C++ 1,070 121 Updated Jul 24, 2026
C++ 189 45 Updated May 11, 2026
C++ 122 19 Updated May 16, 2025

Fused SwiGLU Triton kernels

Python 13 4 Updated Jan 25, 2024

📚200+ Tensor/CUDA Cores Kernels, ⚡️flash-attn-mma, ⚡️hgemm with WMMA, MMA and CuTe (98%~100% TFLOPS of cuBLAS/FA2 🎉🎉).

Cuda 90 8 Updated Apr 26, 2025
Python 133 11 Updated Sep 22, 2025

From Minimal GEMM to Everything

Python 229 14 Updated Jul 9, 2026

Triton implementation of FlashAttention2 that adds Custom Masks.

Python 177 16 Updated Aug 14, 2024

[DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror

C++ 540 302 Updated Jul 27, 2026

Efficient Triton Kernels for LLM Training

Python 6,537 569 Updated Jul 23, 2026

Collection of kernels written in Triton language

200 10 Updated Jan 27, 2026

DeepGEMM: clean and efficient BLAS kernel library on GPU

Cuda 7,573 1,132 Updated Jul 20, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 9,902 1,345 Updated Jul 27, 2026

Repository hosting code for "Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations" (https://arxiv.org/abs/2402.17152).

Python 1,951 404 Updated Jul 27, 2026
Next