Skip to content
View angry-crab's full-sized avatar
  • Tokyo, Japan

Block or report angry-crab

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Fastest kernels written from scratch

Cuda 1 Updated Sep 18, 2025

PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.

Python 222 41 Updated Dec 24, 2025

From Minimal GEMM to Everything

Python 234 17 Updated Jul 9, 2026

GPU programming related news and material links

2,273 137 Updated Jun 15, 2026

GPU documentation for humans

Python 665 84 Updated Jun 14, 2026

A fast communication-overlapping library for tensor/expert parallelism on GPUs.

C++ 1,353 113 Updated Aug 28, 2025

Distributed Compiler based on Triton for Parallel Systems

Python 1,517 165 Updated Aug 12, 2026

My learning notes for ML SYS.

HTML 6,865 477 Updated Aug 12, 2026

TensorRT inference framework for SAM2

C++ 60 11 Updated Jul 3, 2025

Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

Cuda 11,760 1,235 Updated Aug 6, 2026

A small C compiler

C 11,813 1,063 Updated Oct 30, 2023
Cuda 135 16 Updated Mar 19, 2026

CUDA Templates and Python DSLs for High-Performance Linear Algebra

C++ 10,245 2,008 Updated Aug 13, 2026

FlashInfer: Kernel Library for LLM Serving

Python 6,158 1,275 Updated Aug 13, 2026

A Easy-to-understand TensorOp Matmul Tutorial

C++ 450 55 Updated Mar 5, 2026

row-major matmul optimization

C++ 750 94 Updated May 14, 2026
C++ 5 1 Updated Mar 19, 2024

Step-by-step optimization of CUDA SGEMM

Cuda 493 64 Updated Mar 30, 2022

A list of papers, docs, codes about model quantization. This repo is aimed to provide the info for model quantization research, we are continuously improving the project. Welcome to PR the works (p…

2,424 242 Updated Jul 10, 2026

You like pytorch? You like micrograd? You love tinygrad! ❤️

Python 33,450 4,255 Updated Aug 13, 2026

A tiny scalar-valued autograd engine and a neural net library on top of it with PyTorch-like API

Jupyter Notebook 17,108 2,723 Updated Aug 3, 2026

The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.

LLVM 39,782 18,224 Updated Aug 13, 2026

how to optimize some algorithm in cuda.

Cuda 3,199 289 Updated Aug 13, 2026

A beautiful stack trace pretty printer for C++

C++ 4,300 535 Updated Apr 14, 2025

Open Machine Learning Compiler Framework

Python 13,664 3,953 Updated Aug 12, 2026

A list of awesome compiler projects and papers for tensor computation and deep learning.

2,770 329 Updated Oct 19, 2024

A playbook for systematically maximizing the performance of deep learning models.

30,282 2,420 Updated Jun 18, 2024

The Linux perf GUI for performance analysis.

C++ 5,130 288 Updated May 12, 2026

An Open Source Implementation of the Actor Model in C++

C++ 3,430 572 Updated Aug 12, 2026
Next