Highlights
- Pro
Lists (5)
Sort Name ascending (A-Z)
Starred repositories
AgentENV (AENV) is a distributed platform for running agent environments at scale.
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST.…
Production-ready MoE load balancing via real-time expert replication
JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
Block-GTQ: RoPE-aware bit allocation for KV-cache quantization
An open toolkit and public dataset hub for collecting, sanitizing, analyzing, and visualizing coding agent traces.
Vortex: Programmable Sparse Attention for Agents as Algorithm Designers
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
mKernel: fast multi-node, multi-GPU fused kernels
TokenSpeed is a speed-of-light LLM inference engine.
A lightweight inference engine supporting speculative speculative decoding (SSD).
MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the BrowseComp and BrowseComp Zh, respectively.
Distributed MoE in a Single Kernel [NeurIPS '25]
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Efficient Long-context Language Model Training by Core Attention Disaggregation
Accelerating MoE with IO and Tile-aware Optimizations
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
一个基于nano banana pro🍌的原生AI PPT生成应用,迈向"Vibe PPT"; 支持上传任意模板图片,上传任意素材&智能解析,一句话/大纲/页面描述自动生成PPT,口头修改指定区域、一键导出可编辑ppt - An AI-native slides generator based on nano banana pro🍌
SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling
Genai-bench is a powerful benchmark tool designed for comprehensive token-level performance evaluation of large language model (LLM) serving systems.