Stars
Teams-first Multi-agent orchestration for Claude Code
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo…
[EMNLP 2025] AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
Code for MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
PyTorch native quantization for training and inference
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.
An implementation of local windowed attention for language modeling
Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)
CoreNet: A library for training deep neural networks
Explorations into some recent techniques surrounding speculative decoding
[ICLR 2024] Efficient Streaming Language Models with Attention Sinks
The toolkit to test, validate, and evaluate your models and surface, curate, and prioritize the most valuable data for labeling.
Code for the AAAI 2024 Oral paper "OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models".
QLoRA: Efficient Finetuning of Quantized LLMs