Skip to content
View Tom-CaoZH's full-sized avatar
👋
Focusing
👋
Focusing

Highlights

  • Pro

Block or report Tom-CaoZH

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results
Go 97 9 Updated Sep 15, 2025

AgentENV (AENV) is a distributed platform for running agent environments at scale.

Rust 3,167 266 Updated Aug 12, 2026

MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts

Python 1,068 118 Updated Aug 7, 2026

Open Frontier Intelligence

8,388 648 Updated Aug 6, 2026

FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST.…

C++ 503 61 Updated Aug 12, 2026

Production-ready MoE load balancing via real-time expert replication

Cuda 233 11 Updated Jul 17, 2026

Language model tokenization at GB/s

Rust 3,968 208 Updated Aug 6, 2026

A distributed framework for LLM agents

Python 607 26 Updated Aug 6, 2026

Scalable RL for Any Agent and Sandbox.

Python 580 66 Updated Aug 7, 2026

JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting

Python 176 6 Updated Aug 9, 2026

DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms

Python 6,932 647 Updated Jul 9, 2026

Block-GTQ: RoPE-aware bit allocation for KV-cache quantization

Python 13 1 Updated Jun 24, 2026

An open toolkit and public dataset hub for collecting, sanitizing, analyzing, and visualizing coding agent traces.

Python 89 13 Updated Jul 24, 2026

Vortex: Programmable Sparse Attention for Agents as Algorithm Designers

Python 68 13 Updated Aug 11, 2026

UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

C++ 1,488 168 Updated Aug 12, 2026

mKernel: fast multi-node, multi-GPU fused kernels

Cuda 264 24 Updated Aug 12, 2026
Python 9 1 Updated May 12, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,854 234 Updated Aug 12, 2026

Fast Polar Decomposition for Muon

Python 174 14 Updated Jul 2, 2026

A lightweight inference engine supporting speculative speculative decoding (SSD).

Python 986 78 Updated May 10, 2026

MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the BrowseComp and BrowseComp Zh, respectively.

Python 8,370 644 Updated Jul 6, 2026

Distributed MoE in a Single Kernel [NeurIPS '25]

Cuda 281 40 Updated May 5, 2026

OpenTinker is an RL-as-a-Service infrastructure for foundation models

Python 677 63 Updated Mar 21, 2026

Efficient Long-context Language Model Training by Core Attention Disaggregation

Python 105 7 Updated Apr 7, 2026

Accelerating MoE with IO and Tile-aware Optimizations

Python 739 95 Updated Jul 4, 2026

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

Python 4,739 785 Updated May 17, 2026

一个基于nano banana pro🍌的原生AI PPT生成应用,迈向"Vibe PPT"; 支持上传任意模板图片,上传任意素材&智能解析,一句话/大纲/页面描述自动生成PPT,口头修改指定区域、一键导出可编辑ppt - An AI-native slides generator based on nano banana pro🍌

TypeScript 15,438 1,773 Updated Aug 10, 2026

SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling

Python 16 3 Updated Aug 14, 2025

Genai-bench is a powerful benchmark tool designed for comprehensive token-level performance evaluation of large language model (LLM) serving systems.

Python 318 56 Updated Jul 29, 2026
Next