Starred repositories
trendigger.com is a trend-tracking platform for aggregating and visualizing what people are searching on Google. It provides a hour-by-hour breakdown of popular keywords across different countries …
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…
A high-performance distributed file system designed to address the challenges of AI training and inference workloads.
Superfast AI decision making and intelligent processing of multi-modal data.
Open-source, secure environment with real-world tools for enterprise-grade agents.
MooreThreads / vllm-musa
Forked from vllm-project/vllmA high-throughput and memory-efficient inference and serving engine for LLMs
InternRobotics' open platform for building generalized navigation foundation models.
DeepGEMM: clean and efficient BLAS kernel library on GPU
Multi-Joint dynamics with Contact. A general purpose physics simulator.
Fast and memory-efficient exact attention
Fast and memory-efficient exact attention
Distributed AI Model Training and LLM Fine-Tuning on Kubernetes
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
The AI that really does things. Any OS. Any Platform. The lobster way. 🦞
Optimized primitives for collective multi-GPU communication
A CNI IPAM plugin that assigns IP addresses cluster-wide
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
A unified AI model hub for aggregation & distribution. It supports cross-converting various LLMs into OpenAI-compatible, Claude-compatible, or Gemini-compatible formats. A centralized gateway for p…
Lightweight, Modular, Kubernetes-native AI serving platform for scalable model serving.
SGLang is a high-performance serving framework for large language models and multimodal models.
A Go implementation of the Model Context Protocol (MCP), enabling seamless integration between LLM applications and external data sources and tools.
A workload for deploying LLM inference services on Kubernetes
DeepEP: an efficient expert-parallel communication library
A cert-manager webhook completing DNS01 challenge by using External DNS
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
Automatically provision and manage TLS certificates in Kubernetes