Skip to content
View YongjunHe's full-sized avatar

Organizations

@sfu-db @DS3Lab @sfu-dis @llm-db

Block or report YongjunHe

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A PyTorch-based framework for Quantum Classical Simulation, Quantum Machine Learning, Quantum Neural Networks, Parameterized Quantum Circuits with support for easy deployments on real quantum compu…

Jupyter Notebook 1,654 260 Updated Jul 6, 2026

C++ and Python support for the CUDA Quantum programming model for heterogeneous quantum-classical workflows

C++ 1,102 433 Updated Aug 4, 2026

Microsoft Quantum Development Kit, including the Q# programming language, resource estimator, and Quantum Katas

Rust 980 205 Updated Aug 3, 2026

Communication patterns for AI, built on top of NCCL device and host APIs

Cuda 25 6 Updated Aug 3, 2026

Accurate, large-scale, and extensible simulator for LLM inference Systems

Python 652 119 Updated Jul 25, 2025

High-performance RL post-training infrastructure. Designed to achieve bitwise operator-level train-inference consistency across heterogeneous engines and extreme memory efficiency for GRPO, PPO, etc.

Python 220 62 Updated Aug 3, 2026

UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

C++ 1,481 167 Updated Aug 4, 2026

InfiniCCL is a unified, cross-platform collective communication library designed for heterogeneous accelerator environments.

C++ 18 5 Updated Aug 4, 2026

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 31,233 7,624 Updated Aug 4, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 9,940 1,361 Updated Aug 4, 2026

ETHZ Heterogeneous Accelerated Compute Cluster.

42 4 Updated Jun 12, 2026

MLSys competition for the best MOE NKI kernels

Python 48 15 Updated May 29, 2026

Google Research

Jupyter Notebook 38,469 8,463 Updated Jul 30, 2026

LLM Inference analyzer for different hardware platforms

Jupyter Notebook 122 24 Updated Jul 30, 2026

Distributed Evolutionary Algorithms in Python

Python 6,430 1,165 Updated Apr 17, 2026

common in-memory tensor structure

C++ 1,235 168 Updated Jun 19, 2026

Open ABI and FFI for Machine Learning Systems

C++ 443 90 Updated Aug 3, 2026

Google's Operations Research tools:

C++ 13,848 2,453 Updated Aug 3, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 7,116 673 Updated Aug 4, 2026

Extremely fast Query Engine for DataFrames, written in Rust

Rust 39,240 2,991 Updated Aug 4, 2026

Development repository for the Triton language and compiler

MLIR 19,847 3,072 Updated Aug 4, 2026

cuDF - GPU DataFrame Library

C++ 9,723 1,086 Updated Aug 4, 2026

NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmer…

C++ 567 98 Updated Jul 29, 2026

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

Python 9,880 994 Updated Jul 14, 2026

Low-Latency Transaction Scheduling via Userspace Interrupts: Why Wait or Yield When You Can Preempt? (SIGMOD 2025 Best Paper Award)

C++ 81 11 Updated Dec 30, 2025

Pytorch domain library for recommendation systems

Python 2,592 677 Updated Aug 4, 2026

Graph Neural Network Library for PyTorch

Python 23,978 4,032 Updated Jul 31, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,781 4,336 Updated Aug 4, 2026

PyTorch native post-training library

Python 5,795 742 Updated Aug 3, 2026

Fast and memory-efficient exact attention

Python 24,614 2,960 Updated Aug 4, 2026
Next