Skip to content
View lms-mt's full-sized avatar

Block or report lms-mt

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

vits2 backbone with multilingual-bert

Python 8,789 1,299 Updated Aug 10, 2026

FlashInfer: Kernel Library for LLM Serving

Python 6,151 1,270 Updated Aug 12, 2026

Accessible large language models via k-bit quantization for PyTorch.

Python 8,413 904 Updated Aug 12, 2026

Character Animation (AnimateAnyone, Face Reenactment)

Python 3,512 294 Updated May 31, 2024

Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.

Python 6,246 573 Updated Aug 22, 2025

LLM inference in C/C++

C++ 123,614 21,615 Updated Aug 12, 2026

MiniLLM is a minimal system for running modern LLMs on consumer-grade GPUs

Python 971 61 Updated May 15, 2023

OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.

Python 37,404 3,283 Updated Aug 17, 2024

A library for calculating the FLOPs in the forward() process based on torch.fx

Python 140 9 Updated Dec 23, 2025

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…

Python 14,366 2,658 Updated Aug 12, 2026

Count the MACs / FLOPs of your PyTorch model.

Python 5,079 535 Updated Jul 8, 2024

This is an efficient cuda implementation of 2D depthwise convolution for large kernel, it can be used in Pytorch deep learning framework.

Cuda 12 Updated Sep 28, 2023

Optimize GEMM with tensorcore step by step

40 8 Updated Dec 17, 2023

[ARCHIVED] Cooperative primitives for CUDA C++. See https://github.com/NVIDIA/cccl

Cuda 1,841 463 Updated Oct 9, 2023

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 88,861 20,584 Updated Aug 12, 2026

torch_musa is an open source repository based on PyTorch, which can make full use of the super computing power of MooreThreads graphics cards.

Python 507 40 Updated Aug 11, 2026

Fast and memory-efficient exact attention

Python 24,683 2,979 Updated Aug 12, 2026

Annotations of the interesting ML papers I read

288 28 Updated Jun 6, 2026

optimized BERT transformer inference on NVIDIA GPU. https://arxiv.org/abs/2210.03052

C++ 479 37 Updated Mar 15, 2024

Transformer related optimization, including BERT, GPT

C++ 6,445 934 Updated Mar 27, 2024

CUDA Templates and Python DSLs for High-Performance Linear Algebra

C++ 10,238 2,006 Updated Aug 8, 2026

Development repository for the Triton language and compiler

MLIR 19,926 3,102 Updated Aug 12, 2026

The C++ Core Guidelines are a set of tried-and-true guidelines, rules, and best practices about coding in C++

CSS 45,238 5,552 Updated Aug 6, 2026

PyTorch Tutorial for Deep Learning Researchers

Python 32,453 8,234 Updated Aug 15, 2023

Repository of Jupyter notebook tutorials for teaching the Deep Learning Course at the University of Amsterdam (MSc AI), Fall 2023

Jupyter Notebook 3,179 683 Updated Jun 1, 2026

State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.

Jupyter Notebook 14,846 3,409 Updated Aug 12, 2024

A list of awesome compiler projects and papers for tensor computation and deep learning.

2,770 329 Updated Oct 19, 2024