Starred repositories
FlagGems is an operator library for large language models implemented in the Triton Language.
Synology VideoStation and DLNA FFmpeg Wrapper with AAC, DTS, EAC3 and TrueHD support via pipes (now with GStreamer support). It enables full hardware transcoding from Synology´s FFmpeg for video an…
BitBLAS is a library to support mixed-precision matrix multiplications, especially for quantized LLM deployment.
Development repository for the Triton language and compiler
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
Polyhedral Parallel Code Generation (source repository: http://repo.or.cz/ppcg.git)
Integer Set Library (source repository: http://repo.or.cz/w/isl.git)
The Torch-MLIR project aims to provide first class support from the PyTorch ecosystem to the MLIR ecosystem.
tutorials about polyhedral compilation.
LaTeX Thesis Template for the University of Chinese Academy of Sciences
An MLIR-based compiler framework bridges DSLs (domain-specific languages) to DSAs (domain-specific architectures).
[DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror
Small set of gdb commands for useful tasks in tvm
a highly-efficient library for deep neural networks based on Sunway TaihuLight supercomputer.
InfiniTensor is a high-performance inference engine tailored for GPUs and AI accelerators. Its design focuses on effective deployment and swift academic validation.
PatrickStar enables Larger, Faster, Greener Pretrained Models for NLP and democratizes AI for everyone.
A highly efficient library for GEMM operations on Sunway TaihuLight
Official repository for DistFlashAttn: Distributed Memory-efficient Attention for Long-context LLMs Training
A lightweight framework for building LLM-based agents
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…
Machine Intelligence Shader Autogen. AMDGPU ML shader code generator. (previously iGEMMgen)
ROCm / flash-attention
Forked from Dao-AILab/flash-attentionFast and memory-efficient exact attention