- Beijing, China
Lists (3)
Sort Name ascending (A-Z)
Stars
A rule-based tunnel for Android.
TokenSpeed is a speed-of-light LLM inference engine.
便宜机场,一元机场,性价比机场,白嫖机场,免费机场,机场推荐,github加速,github文件加速,机场订阅加速,2025年最新科学上网,vpn机场推荐
Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and full-stack tuning.
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
Alibaba Cloud's high-performance KVCache system for LLM inference, with components for global cache management, inference simulation(HiSim), and more.
🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm Maps )
mimalloc is a compact general purpose allocator with excellent performance.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
[CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
[NeurIPS 2025 D&B] 🚀 SWE-bench Goes Live!
Fast Hadamard transform in CUDA, with a PyTorch interface
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
how to optimize some algorithm in cuda.
DeepXTrace is a lightweight tool for precisely diagnosing slow ranks in DeepEP-based environments.
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
AIInfra(AI 基础设施)指AI系统从底层芯片等硬件,到上层软件栈支持AI大模型训练和推理。
A Datacenter Scale Distributed Inference Serving Framework
🚀 Efficient implementations for emerging model architectures
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.