Skip to content
View ymwangg's full-sized avatar

Block or report ymwangg

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

🚀 Efficient implementations for emerging model architectures

Python 5,545 648 Updated Aug 11, 2026

high-performance linear attention kernel library built on TileLang

Python 630 66 Updated Aug 11, 2026

A PyTorch native platform for training generative AI models

Python 6 Updated Jul 2, 2026

A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.

Python 1 Updated Jun 11, 2026

NKIPy: Rapid Prototyping on Trainium

Python 30 11 Updated Aug 12, 2026
Python 69 11 Updated Jul 14, 2026

A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.

Python 921 163 Updated Aug 12, 2026
Python 19 6 Updated Aug 8, 2026

💥 Fast State-of-the-Art Tokenizers optimized for Research and Production

Rust 10,958 1,167 Updated Aug 12, 2026

Every front-end GUI client for ChatGPT, Claude, and other LLMs

3,999 275 Updated Jul 23, 2026

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

Python 5,448 431 Updated Jul 26, 2026

PipeEdge: Pipeline Parallelism for Large-Scale Model Inference on Heterogeneous Edge Devices

Python 41 28 Updated Jan 31, 2024

CUDA Templates and Python DSLs for High-Performance Linear Algebra

C++ 10,236 2,006 Updated Aug 8, 2026

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…

Python 3,426 541 Updated Aug 12, 2026

This repository contains integer operators on GPUs for PyTorch.

Python 235 55 Updated Sep 29, 2023

Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.

Python 10,677 1,091 Updated Jul 1, 2024

This repository collects papers for "A Survey on Knowledge Distillation of Large Language Models". We break down KD into Knowledge Elicitation and Distillation Algorithms, and explore the Skill & V…

1,297 72 Updated Mar 9, 2025

The official Meta Llama 3 GitHub site

Python 29,261 3,525 Updated Jan 26, 2025

FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.

Python 1,126 90 Updated Sep 4, 2024

A framework for few-shot evaluation of language models.

Python 13,598 3,477 Updated Aug 11, 2026

📰 Must-read papers and blogs on Speculative Decoding ⚡️

1,289 81 Updated Jun 27, 2026

FlashInfer: Kernel Library for LLM Serving

Python 6,146 1,266 Updated Aug 12, 2026

Measuring Massive Multitask Language Understanding | ICLR 2021

Python 1,608 117 Updated May 28, 2023

AutoAWQ implements the AWQ algorithm for 4-bit quantization with a 2x speedup during inference. Documentation:

Python 2,350 307 Updated May 11, 2025

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 31,706 7,820 Updated Aug 12, 2026

The TinyLlama project is an open endeavor to pretrain a 1.1B Llama model on 3 trillion tokens.

Python 9,019 628 Updated May 3, 2024

[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Python 3,613 318 Updated Jul 17, 2025

Numbers every LLM developer should know

4,318 141 Updated Jan 16, 2024

The simplest, fastest repository for training/finetuning medium-sized GPTs.

Python 62,047 10,691 Updated Nov 12, 2025

Fast and memory-efficient exact attention

Python 24,683 2,979 Updated Aug 11, 2026
Next