Lists (1)
Sort Name ascending (A-Z)
Stars
A workload for deploying LLM inference services on Kubernetes
【A common used C++ & Python DAG framework】 一个通用的、无三方依赖的、跨平台的、收录于awesome-cpp的、基于流图的并行计算框架。欢迎star & fork & 交流
A Datacenter Scale Distributed Inference Serving Framework
vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization
A simple, open source bilingual translation extension & Greasemonkey script (一个简约、开源的 双语对照翻译扩展 & 油猴脚本)
SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models
A framework for efficient model inference with omni-modality models
TokenSpeed is a speed-of-light LLM inference engine.
A programmable Mixture-of-Models router for heterogeneous LLM inference
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
Based on the RV32I ISA, aiming to implement the complete functions of the CPU without considering synthesis, timing, and latency.
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.
Algorithm powering the For You feed on X
a embedding infer server faster than vllm and sglang
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
FlashInfer: Kernel Library for LLM Serving
Fast and memory-efficient exact attention
The ultimate RAG for your monorepo. Query, understand, and edit multi-language codebases with the power of AI and knowledge graphs
Getting Started with Triton: A Tutorial for Python Beginners