Lists (4)
Sort Name ascending (A-Z)
Starred repositories
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
🥷 Engineering habits you already know, turned into skills Claude can run.
A rewrite of Tachiyomi for the Desktop
SGLang is a high-performance serving framework for large language models and multimodal models.
Kernel Design Agents (KDA) is a agent-centric workflow to write high-performance CUDA Kernels.
A tutorial on modern GPU programming for machine learning systems
An AI-native distributed data plane built in Rust that supports high performance RPC, KV Cache, Message Queue, and File & Object Acceleration
DeepGEMM: clean and efficient BLAS kernel library on GPU
Skills for Real Engineers. Straight from my .agents directory.
Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
PiKV: KV Cache Management System for Mixture of Experts [Efficient ML System]
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
A framework for efficient model inference with omni-modality models
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
High-performance RL post-training infrastructure. Designed to achieve bitwise operator-level train-inference consistency across heterogeneous engines and extreme memory efficiency for GRPO, PPO, etc.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
An AI-powered IDE for long-form fiction writing, combining software engineering workflows, modern storytelling methodologies, and multi-agent systems.
Compile SillyTavern character cards into pi-native, event-driven interactive narrative runtimes.
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI