Skip to content
View whn09's full-sized avatar
  • AWS
  • Beijing, China

Block or report whn09

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Cuda 3,640 478 Updated Jan 17, 2026

SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.

Python 805 334 Updated Aug 14, 2026

Lightweight coding agent that runs in your terminal

Rust 105,935 16,081 Updated Aug 14, 2026

🎥 Make videos programmatically with React

TypeScript 56,307 4,209 Updated Aug 14, 2026

USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference

Python 685 82 Updated May 21, 2026

AI video skill for Claude Code & Codex — cinematic product videos with Remotion: 152 shot recipe cards, 209 motion previews, a production-ready template

TypeScript 4,999 429 Updated Aug 14, 2026

Train speculative decoding models effortlessly and port them smoothly to SGLang serving.

Python 1,077 315 Updated Aug 14, 2026

Community maintained hardware plugin for vLLM on AWS Neuron

Python 3 Updated Jul 31, 2026

Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and full-stack tuning.

Python 1,808 254 Updated Aug 4, 2026

FlashKDA: high-performance Kimi Delta Attention kernels

Cuda 1,211 115 Updated Jul 30, 2026

MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts

Python 1,078 118 Updated Aug 13, 2026
Python 1 Updated Jul 21, 2026

OpenClaw-RL: Train any agent simply by talking

Python 5,633 606 Updated May 23, 2026

An LLM post-training framework with vLLM for RL Scaling

Python 419 77 Updated Aug 3, 2026

slime is an LLM post-training framework for RL Scaling.

Python 7,960 1,142 Updated Aug 14, 2026

A safetensors extension to efficiently store sparse quantized tensors on disk

Python 313 112 Updated Aug 14, 2026

MSCCL++: A GPU-driven communication stack for scalable AI applications

C++ 5 1 Updated Feb 6, 2026

The Agent Harness for AI-Human Collaboration, inspired by the AI-DLC (AI-Driven Development Lifecycle)

TypeScript 1,130 100 Updated Aug 14, 2026

Tile-based language built for AI computation across all scales

Python 185 9 Updated Aug 14, 2026

High-performance LLM operator library built on TileLang.

Python 168 55 Updated Aug 14, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 7,215 695 Updated Aug 14, 2026

Tile-Based Runtime for Ultra-Low-Latency LLM Inference

Python 1,681 120 Updated Aug 13, 2026

A PyTorch native platform for training generative AI models

Python 6 Updated Jul 2, 2026

Source Han Sans | 思源黑体 | 思源黑體 | 思源黑體 香港 | 源ノ角ゴシック | 본고딕

Python 17,059 1,380 Updated Jun 25, 2025

The official repo for STCast (CVPR2026 Highlight).

Python 26 3 Updated Apr 10, 2026
Next