Lists (19)
Sort Name ascending (A-Z)
Agent Skills & Tooling
Claude Skills 生态、agent harness 工具AI Agent Frameworks
Claude/LangChain/LangGraph/MCP/Agent SDKAI Infra Learning (中文)
中文 AI infra/LLM 教程与 awesome listContainer Runtimes / Wasm
containerd/runc/kata/youki/wasmCS Fundamentals / Interview
算法、系统设计、面试题、经典书Go & Rust Foundations
Go/Rust 语言学习、基础库、Web 框架K8s × AI Serving
K8s 上跑 LLM/AI 工作负载的平台K8s Core & Controllers
K8s 主线、controllers、operator SDK、CRD、kubectl、client-goK8s GPU & Device Plugins
GPU/异构、device-plugin、DRA、vGPUK8s Multi-Cluster & Edge
多集群、虚拟集群、边缘K8s Networking & CNI
CNI、Service Mesh、Envoy、GatewayK8s Schedulers & Batch
调度器、队列、batch、Volcano/Kueue/Karpenter/KoordinatorLLM Inference Engines
vLLM/SGLang/LMDeploy/Ollama/TensorRT/DynamoLLM Training & Fine-tuning
训练、微调、量化Misc / Personal Tools
其余工具与个人项目My / Friends' Projects
自己仓库 + 学习圈仓库Observability & Tracing
Prom/Grafana/OTel/Loki/Pyroscope/eBPF observabilityQuant / Finance
量化交易、金融数据Storage on K8s
CSI、本地存储、缓存、数据加速Starred repositories
A local-first routing and coordination engine for long-running agent work.
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding,…
GPT-2-style LLM built from scratch in C/CUDA with hand-written backprop, BPE tokenizer, FlashAttention, pretraining, and SFT.
Generate text, images, video, speech, and music by MiniMax.
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.
Agent skill that removes signs of AI-generated writing from text
《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码
vLLM plugin for attention-ffn disaggregation support
《深入理解 AI Infra:量化分析与系统设计》(李博杰 著)开源书稿:从硬件约束和模型架构出发,量化推导 LLM 推理与训练系统设计。含全书正文、PDF、配套计算工具与实验
An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: https://arxiv.org/abs/2606.09682
A transparent, in-container GPU resource controller that enforces memory and compute limits by intercepting CUDA calls without application or driver changes.
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.
Optimized primitives for collective multi-GPU communication
Ecosystem of libraries and tools for writing and executing fast GPU code fully in Rust.
A modern replacement for Redis and Memcached
Nephio is a Kubernetes-based automation platform for deploying and managing highly distributed, interconnected workloads such as 5G Network Functions, and the underlying infrastructure on which tho…
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Go bindings to systemd socket activation, journal, D-Bus, and unit files
Kubernetes controllers for fast model actuation using vLLM sleep/wake and launcher-based model swapping
how to optimize some algorithm in cuda.