Stars
Ring attention implementation with flash attention
High-performance FlashAttention Implementation for Ascend NPU
🚀 Efficient implementations for emerging model architectures
A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.
Custom kernel collections using https://github.com/PTO-ISA/pto-isa
Pythonic interface and JIT compiler for https://gitcode.com/cann/pto-isa
Open-source transpiler for CUDA Tile (13.1) migration
FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kernel structure at a high level.
CUDA Tile IR is an MLIR-based intermediate representation and compiler infrastructure for CUDA kernel optimization, focusing on tile-based computation patterns and optimizations targeting NVIDIA te…
tile-ai / tilescale
Forked from tile-ai/tilelangTile-based language built for AI computation across all scales
⚡️Write HGEMM from scratch using Tensor Cores with WMMA, MMA and CuTe API, Achieve Peak⚡️ Performance.
A small OpenCL benchmark program to measure peak GPU/CPU performance.
USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference
A tool for bandwidth measurements on NVIDIA GPUs.
Puzzles for learning Triton
Exploring the scalable matrix extension of the Apple M4 processor
SGLang is a high-performance serving framework for large language models and multimodal models.
Scientific computing with Metal in C++: Matrix multiplication example
Interactive architecture diagrams for codebases
Everything we actually know about the Apple Neural Engine (ANE)