Stars
SGLang is a high-performance serving framework for large language models and multimodal models.
A safetensors extension to efficiently store sparse quantized tensors on disk
A high-throughput and memory-efficient inference and serving engine for LLMs
A curated list for Efficient Large Language Models