Highlights
- Pro
Stars
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
A framework for few-shot evaluation of language models.
Fastest MoE/LLM inference runtime for consumer and edge Blackwell GPUs. SN74 on Gittensor.
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
SGLang is a high-performance serving framework for large language models and multimodal models.
NanoCluster: Compact & Affordable Cluster for Everyone
[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Real-Time VLAs via Future-state-aware Asynchronous Inference.
DFlash: Block Diffusion for Flash Speculative Decoding
CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
A retargetable MLIR-based machine learning compiler and runtime toolkit.
Development repository for the Triton language and compiler
TokenSpeed is a speed-of-light LLM inference engine.
FlyDSL is the Python front‑end of the project: Flexible LaYout DSL.
High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
All information and news with respect to Falcon-H1 series
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
Interactive 3D visualization of dense decoder-only LLM inference. Companion to the AI Inference Engineer 2026 course.