Stars
A tutorial on modern GPU programming for machine learning systems
AI 基础知识 - GPU 架构、CUDA 编程、大模型基础及AI Agent 相关知识。
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等
[TPAMI 2026] 3D and 4D World Modeling: A Survey
Proxy that captures and visualizes in-flight Claude Code requests and conversations.
强化学习中文教程(蘑菇书🍄),在线阅读地址:https://datawhalechina.github.io/easy-rl/
CUDA Templates and Python DSLs for High-Performance Linear Algebra
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
AI-powered, vision-driven UI automation for every platform.
Midscene connector for pc,include local pc and remote pc server. Supports windows/linux/macos.基于midscene的跨平台PC桌面端自动化操作代理。同时支持本地和远程桌面操作方式。
Paper list in the survey: A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
[ACL 2026] Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
SpotServe: Serving Generative Large Language Models on Preemptible Instances
Persist and reuse KV Cache to speedup your LLM.
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
Offline optimization of your disaggregated Dynamo graph
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
A Simplified PyTorch Implementation of Vision Transformer (ViT)
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton