Stars
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台…
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Synthetic Data Generation Toolkit for LLMs
Achieve state of the art inference performance with modern accelerators on Kubernetes
A high-throughput and memory-efficient inference and serving engine for LLMs
neuralmagic / nm-vllm
Forked from vllm-project/vllmA high-throughput and memory-efficient inference and serving engine for LLMs
Refine high-quality datasets and visual AI models
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing…
Sparsity-aware deep learning inference runtime for CPUs
Top-level directory for documentation and general content
ML model optimization product to accelerate inference.
Neural network model repository for highly sparse and sparse-quantized models with matching sparsification recipes
Libraries for applying sparsification recipes to neural networks with a few lines of code, enabling faster and smaller models