Skip to content
View whn09's full-sized avatar
  • AWS
  • Beijing, China

Block or report whn09

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

FlashKDA: high-performance Kimi Delta Attention kernels

Cuda 867 87 Updated Jul 29, 2026

MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts

Python 813 82 Updated Jul 28, 2026
Python 1 Updated Jul 21, 2026

OpenClaw-RL: Train any agent simply by talking

Python 5,612 608 Updated May 23, 2026

An LLM post-training framework with vLLM for RL Scaling

Python 390 69 Updated Jul 28, 2026

slime is an LLM post-training framework for RL Scaling.

Python 7,684 1,103 Updated Jul 24, 2026

A safetensors extension to efficiently store sparse quantized tensors on disk

Python 308 107 Updated Jul 28, 2026

MSCCL++: A GPU-driven communication stack for scalable AI applications

C++ 5 1 Updated Feb 6, 2026

The Agent Harness for AI-Human Collaboration, inspired by the AI-DLC (AI-Driven Development Lifecycle)

TypeScript 1,104 98 Updated Jul 27, 2026

Tile-based language built for AI computation across all scales

C++ 176 7 Updated Jul 29, 2026

High-performance LLM operator library built on TileLang.

Python 165 53 Updated Jul 29, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 7,018 661 Updated Jul 29, 2026

Tile-Based Runtime for Ultra-Low-Latency LLM Inference

Python 1,597 112 Updated Jul 14, 2026

A PyTorch native platform for training generative AI models

Python 6 Updated Jul 2, 2026

Source Han Sans | 思源黑体 | 思源黑體 | 思源黑體 香港 | 源ノ角ゴシック | 본고딕

Python 16,980 1,377 Updated Jun 25, 2025

The official repo for STCast (CVPR2026 Highlight).

Python 24 3 Updated Apr 10, 2026

MixFormer models for recommender systerm

Python 4 3 Updated Mar 15, 2026

Common recipes to run vLLM

JavaScript 936 348 Updated Jul 29, 2026

🚀 Efficient implementations for emerging model architectures

Python 5,468 615 Updated Jul 27, 2026

NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmer…

C++ 2 Updated May 14, 2026

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Python 22,240 2,709 Updated Jul 14, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 1 Updated May 5, 2026

NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmer…

C++ 9 2 Updated Jul 24, 2026

Optimized primitives for collective multi-GPU communication

C++ 6 5 Updated Jul 27, 2026

MSCCL++: A GPU-driven communication stack for scalable AI applications

C++ 544 102 Updated Jul 29, 2026
Next