Stars
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
Here are my personal paper reading notes (including machine learning systems, AI infrastructure, and other interesting stuffs).
Pre-built wheels that erase Flash Attention 3 installation headaches.
Neovim 🤝 OpenCode in the flow that you already know.
TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained GPUs.
null-ls.nvim reloaded / Use Neovim as a language server to inject LSP diagnostics, code actions, and more via Lua.
Distributed Task Queue (development branch)
🥢像老乡鸡🐔那样做饭。已添加2026年发布的《老乡鸡菜品溯源报告 2.0中新出现的菜品。主要部分于2024年完工,非老乡鸡官方仓库。文字来自《老乡鸡菜品溯源报告》,并做归纳、编辑与整理。CookLikeHOC.
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
⚡️SwanLab - an open-source, modern-design AI training tracking and visualization tool. Supports Cloud / Self-hosted use. Integrated with PyTorch / Transformers / verl / LLaMA Factory / ms-swift / U…
Compare different hardware platforms via the Roofline Model for LLM inference tasks.
CPM.cu is a lightweight, high-performance CUDA implementation for LLMs, optimized for end-device inference and featuring cutting-edge techniques in sparse architecture, speculative sampling and qua…
Distributed Compiler based on Triton for Parallel Systems
Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without …
A markup-based typesetting system that is powerful and easy to learn.
A fast communication-overlapping library for tensor/expert parallelism on GPUs.
[ACL 2025 main] FR-Spec: Frequency-Ranked Speculative Sampling
Fast and Memory-Efficient Exact Attention for Large Headdim, 1.5x~6x speedup over PyTorch SDPA.
DeepEP: an efficient expert-parallel communication library