Skip to content
View demonbibi's full-sized avatar

Block or report demonbibi

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results
Python 364 46 Updated Jun 9, 2026

A kernel library written in tilelang

Python 1,712 155 Updated Apr 23, 2026

A CUDA kernel optimization toolkit for validation, benchmarking, Nsight Compute profiling, bottleneck analysis, and iterative tuning. It helps improve custom GPU operators with reproducible workflo…

Python 195 19 Updated Apr 22, 2026

Accelerating MoE with IO and Tile-aware Optimizations

Python 737 96 Updated Jul 4, 2026

A Quirky Assortment of CuTe Kernels

Python 1,096 149 Updated Aug 8, 2026

Persistent file-based planning for AI coding agents and long-running tasks. Crash-proof markdown plans, session recovery after /clear and compaction, per-turn re-injection against context rot, dete…

Shell 26,068 2,179 Updated Aug 9, 2026

An agentic skills framework & software development methodology that works.

Shell 269,639 24,096 Updated Aug 8, 2026

AI agents running research on single-GPU nanochat training automatically

Python 93,507 13,287 Updated Mar 26, 2026

高性能短序列稀疏Mask Attention CUDA算子,针对<1K序列+75%稀疏度优化

Python 80 8 Updated Mar 18, 2026

CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning

Cuda 469 32 Updated Mar 30, 2026

A Next-Generation Training Engine Built for Ultra-Large MoE Models

Python 5,175 439 Updated Aug 7, 2026

KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)

Jupyter Notebook 1,192 187 Updated Mar 24, 2026

TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators

Python 137 15 Updated Jun 14, 2025

High Performance LLM Inference Operator Library

C++ 1,094 127 Updated Aug 6, 2026
C++ 189 45 Updated May 11, 2026
C++ 122 19 Updated May 16, 2025

Fused SwiGLU Triton kernels

Python 14 4 Updated Jan 25, 2024

📚200+ Tensor/CUDA Cores Kernels, ⚡️flash-attn-mma, ⚡️hgemm with WMMA, MMA and CuTe (98%~100% TFLOPS of cuBLAS/FA2 🎉🎉).

Cuda 91 8 Updated Apr 26, 2025
Python 135 11 Updated Sep 22, 2025

From Minimal GEMM to Everything

Python 232 15 Updated Jul 9, 2026

Triton implementation of FlashAttention2 that adds Custom Masks.

Python 177 16 Updated Aug 14, 2024

[DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror

C++ 541 303 Updated Aug 9, 2026

Efficient Triton Kernels for LLM Training

Python 6,558 578 Updated Aug 7, 2026

Collection of kernels written in Triton language

199 10 Updated Jan 27, 2026

DeepGEMM: clean and efficient BLAS kernel library on GPU

Cuda 7,638 1,159 Updated Jul 20, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 9,966 1,367 Updated Aug 5, 2026

Repository hosting code for "Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations" (https://arxiv.org/abs/2402.17152).

Python 1,958 407 Updated Aug 4, 2026
Next