-
-
dynamo Public
Forked from ai-dynamo/dynamoA Datacenter Scale Distributed Inference Serving Framework
Rust Other UpdatedSep 23, 2026 -
aibrix Public
Forked from vllm-project/aibrixCost-efficient and pluggable Infrastructure components for GenAI inference
-
hiroute Public
Forked from higress-group/HiRouteA local-first routing and coordination engine for long-running agent work.
Rust Apache License 2.0 UpdatedSep 23, 2026 -
rbg Public
Forked from sgl-project/rbgA workload for deploying LLM inference services on Kubernetes
Go Apache License 2.0 UpdatedSep 23, 2026 -
-
-
flashinfer Public
Forked from flashinfer-ai/flashinferFlashInfer: Kernel Library for LLM Serving
Cuda Apache License 2.0 UpdatedSep 21, 2026 -
enhancements Public
Forked from kubernetes/enhancementsEnhancements tracking repo for Kubernetes
Go Apache License 2.0 UpdatedSep 20, 2026 -
agents Public
Forked from openkruise/agentsRapid and cost-effective operator and best practice for agent sandbox lifecycle management.
Go Apache License 2.0 UpdatedSep 18, 2026 -
colibri Public
Forked from JustVugg/colibriRun frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
C Apache License 2.0 UpdatedSep 18, 2026 -
kubernetes Public
Forked from kubernetes/kubernetesProduction-Grade Container Scheduling and Management
Go Apache License 2.0 UpdatedSep 15, 2026 -
llm-d-fast-model-actuation Public
Forked from llm-d-incubation/llm-d-fast-model-actuationKubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping
Go Apache License 2.0 UpdatedSep 13, 2026 -
llama.cpp Public
Forked from ggml-org/llama.cppLLM inference in C/C++
C++ MIT License UpdatedSep 12, 2026 -
dragonfly Public
Forked from dragonflyoss/dragonflyDelivers efficient, stable, and secure data distribution and acceleration powered by P2P technology, with an optional content‑addressable filesystem that accelerates OCI container launch.
Go Apache License 2.0 UpdatedSep 11, 2026 -
CUDALibrarySamples Public
Forked from NVIDIA/CUDALibrarySamplesCUDA Library Samples
Cuda Apache License 2.0 UpdatedSep 11, 2026 -
nephio Public
Forked from nephio-project/nephioNephio is a Kubernetes-based automation platform for deploying and managing highly distributed, interconnected workloads such as 5G Network Functions, and the underlying infrastructure on which tho…
Go Apache License 2.0 UpdatedSep 10, 2026 -
llm-d-router Public
Forked from llm-d/llm-d-routerllm-d Router: The intelligent entry point for inference requests
Go Apache License 2.0 UpdatedSep 9, 2026 -
-
llm-d Public
Forked from llm-d/llm-dAchieve state of the art inference performance with modern accelerators on Kubernetes
Shell Apache License 2.0 UpdatedSep 6, 2026 -
WeKnora Public
Forked from Tencent/WeKnoraOpen-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
Go Other UpdatedSep 5, 2026 -
llm-d-autoscaling Public
Forked from llm-d/llm-d-autoscalingVariant optimization autoscaler for distributed inference workloads
Go Apache License 2.0 UpdatedSep 2, 2026 -
speculators Public
Forked from vllm-project/speculatorsA unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
Python Apache License 2.0 UpdatedSep 1, 2026 -
agentic-api Public
Forked from vllm-project/agentic-apiStateful API logic for agentic applications using vLLM
Rust Apache License 2.0 UpdatedSep 1, 2026 -
vLLM-SGLang-book Public
Forked from thu/vLLM-SGLang-bookvLLM 与 SGLang: 大模型高效推理双引擎实战配套代码
Python MIT License UpdatedAug 26, 2026 -
gpu-perf-engineering-resources Public
Forked from wafer-ai/gpu-perf-engineering-resourcesA curated resource list for learning AI performance engineering, from GPU fundamentals to production inference.
Python UpdatedAug 23, 2026 -
DeepTutor Public
Forked from HKUDS/DeepTutorDeepTutor: Lifelong Personalized Tutoring. https://deeptutor.info/.
Python Apache License 2.0 UpdatedAug 21, 2026 -
comet Public
Forked from rpamis/cometComet: OpenSpec + Superpowers dual-star development workflow
JavaScript MIT License UpdatedAug 21, 2026 -
pocketInfer Public
Forked from pocketInfer/pocketInferTurn massive model configs into memory-bounded test models that preserve key architecture and inference paths. (把跑不动的超大模型配置,缩成 保留关键架构与推理路径的“可调试测试模型“)
Python Other UpdatedAug 21, 2026 -
diagram-design Public
Forked from cathrynlavery/diagram-design38 editorial diagram types for Claude Code, Codex, and Pi. Self-contained HTML + SVG. No shadows. No Mermaid slop.
HTML MIT License UpdatedAug 20, 2026