Skip to content
View fiskrt's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report fiskrt

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Ring attention implementation with flash attention

Python 1,044 101 Updated Sep 10, 2025

High-performance FlashAttention Implementation for Ascend NPU

C++ 28 23 Updated Aug 13, 2026

PTO instruction set architecture

C++ 67 46 Updated Aug 14, 2026
Python 12 3 Updated Aug 15, 2026

🚀 Efficient implementations for emerging model architectures

Python 5,564 660 Updated Aug 16, 2026

A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.

Python 239 19 Updated Jun 29, 2026

Custom kernel collections using https://github.com/PTO-ISA/pto-isa

C++ 19 14 Updated Aug 16, 2026

Pythonic interface and JIT compiler for https://gitcode.com/cann/pto-isa

Python 27 13 Updated Jun 1, 2026

Open-source transpiler for CUDA Tile (13.1) migration

TypeScript 19 4 Updated Dec 9, 2025
Jupyter Notebook 25 3 Updated Feb 13, 2026

FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kernel structure at a high level.

Python 262 108 Updated Aug 16, 2026

CUDA Tile IR is an MLIR-based intermediate representation and compiler infrastructure for CUDA kernel optimization, focusing on tile-based computation patterns and optimizations targeting NVIDIA te…

C++ 1,008 85 Updated Jul 22, 2026

Tile-based language built for AI computation across all scales

Python 185 9 Updated Aug 16, 2026

⚡️Write HGEMM from scratch using Tensor Cores with WMMA, MMA and CuTe API, Achieve Peak⚡️ Performance.

Cuda 158 10 Updated May 10, 2025

A small OpenCL benchmark program to measure peak GPU/CPU performance.

C++ 320 37 Updated Jul 14, 2026

USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference

Python 685 82 Updated May 21, 2026

A tool for bandwidth measurements on NVIDIA GPUs.

C++ 751 90 Updated Jul 28, 2026

FlashAttention (Metal Port)

Swift 612 40 Updated Sep 22, 2024

Puzzles for learning Triton

Jupyter Notebook 2,562 250 Updated Apr 1, 2026

WALDSAC (RRANSAC with SPRT)

MATLAB 6 Updated Mar 26, 2023

Exploring the scalable matrix extension of the Apple M4 processor

C 237 12 Updated Nov 7, 2024

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 31,909 7,924 Updated Aug 16, 2026

Apple GPU microarchitecture

Metal 627 30 Updated Sep 22, 2024

Scientific computing with Metal in C++: Matrix multiplication example

C++ 51 2 Updated Sep 18, 2022

Interactive architecture diagrams for codebases

Python 2,389 199 Updated Aug 16, 2026

TTS support with GGML

C++ 248 34 Updated Oct 5, 2025

Everything we actually know about the Apple Neural Engine (ANE)

2,517 97 Updated Mar 12, 2026
Next