Skip to content
View imShZh's full-sized avatar
🎯
Focusing
🎯
Focusing

Organizations

@ShZh-Playground @ShZh-libraries @ShZh-websites

Block or report imShZh

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

The official Lark/飞书 CLI tool, maintained by the larksuite team — built for humans and AI Agents. Covers core business domains including Messenger, Docs, Base, Sheets, Calendar, Mail, Tasks, Meetin…

Go 16,190 1,285 Updated Aug 5, 2026

A framework for efficient model inference with omni-modality models

Python 5,877 1,411 Updated Aug 5, 2026

Agentic RL Training at Scale

Python 1,804 380 Updated Aug 5, 2026

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

Python 1,898 346 Updated Aug 5, 2026

Perplexity open source garden for inference technology

Rust 611 66 Updated May 27, 2026

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

Python 9,884 997 Updated Jul 14, 2026

An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models

Python 3,346 304 Updated Aug 5, 2026

NVIDIA GPU metrics exporter for Prometheus leveraging DCGM

Go 1,820 320 Updated Jul 25, 2026

DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

Python 42,868 4,920 Updated Aug 5, 2026

The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.

Python 5,642 573 Updated Aug 5, 2026

slime is an LLM post-training framework for RL Scaling.

Python 7,772 1,118 Updated Aug 5, 2026

A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology

C 1,402 193 Updated Jul 14, 2026

My learning notes for ML SYS.

HTML 6,825 473 Updated Aug 5, 2026

Optimized primitives for collective multi-GPU communication

C++ 4,938 1,361 Updated Aug 4, 2026

A PyTorch native platform for training generative AI models

Python 5,591 929 Updated Aug 5, 2026

NCCL Tests

Cuda 1,612 396 Updated Aug 3, 2026

What would you do with 1000 H100s...

Jupyter Notebook 1,187 72 Updated Jan 10, 2024

Efficient Triton Kernels for LLM Training

Python 6,553 572 Updated Aug 3, 2026

Puzzles for learning Triton, play it with minimal environment configuration!

Python 743 99 Updated Mar 17, 2026

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++ 6,174 1,053 Updated Aug 5, 2026

CUDA Python: Performance meets Productivity

Cython 3,331 317 Updated Aug 5, 2026

FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.

Python 1,120 90 Updated Sep 4, 2024
C++ 97 9 Updated Mar 26, 2025

A minimal implementation of vllm.

Cuda 73 Updated Jul 27, 2024

FlashInfer: Kernel Library for LLM Serving

Python 6,107 1,232 Updated Aug 5, 2026

text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)

Python 12,937 1,324 Updated Nov 4, 2025

Virtual whiteboard for sketching hand-drawn like diagrams

TypeScript 129,033 14,710 Updated Aug 5, 2026

Tile primitives for speedy kernels

Cuda 3,604 315 Updated Jul 13, 2026

Material for gpu-mode lectures

Jupyter Notebook 6,398 639 Updated Jun 15, 2026
Next