Stars
Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.
cuTile is a programming model for writing parallel kernels for NVIDIA GPUs
A framework for efficient model inference with omni-modality models
Gateway API Inference Extension
Achieve state of the art inference performance with modern accelerators on Kubernetes
A Datacenter Scale Distributed Inference Serving Framework
An open-source, privacy-first, self-hosted knowledge workspace where humans and AI agents work together 开源、隐私优先、自托管的知识工作空间,让人与智能体在此协作
Cost-efficient and pluggable Infrastructure components for GenAI inference
Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
[WIP] Resources for AI engineers. Also contains supporting materials for the book AI Engineering (Chip Huyen, 2025)
Optimized Agentic and LLM Bulk Processing Over Your Data
High performance self-hosted photo and video management solution.
A playbook for effectively prompting post-trained LLMs
Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.
Fast Open-Source Search & Clustering engine × for Vectors & Arbitrary Objects × in C++, C, Python, JavaScript, Rust, Java, Objective-C, Swift, C#, GoLang, and Wolfram 🔍
Adding guardrails to large language models.
DSPy: The framework for programming—not prompting—language models
An enterprise-grade AI retriever designed to streamline AI integration into your applications, ensuring cutting-edge accuracy.
[ACL'25] Official Code for LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMs
SGLang is a high-performance serving framework for large language models and multimodal models.
A simple RPC framework with protobuf service definitions
Open source, local, and self-hosted highly optimized language inference server supporting ASR/STT, TTS, and LLM across WebRTC, REST, and WS