Skip to content
View Irvingwangjr's full-sized avatar

Block or report Irvingwangjr

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Ulysses sequence-parallel all-to-all as a torch custom op, moved by the GPU copy engines over NVSHMEM symmetric memory. Zero SM usage; 1.7-2.2x over torch.distributed on NVLink.

Python 34 4 Updated Aug 7, 2026

Agent Substrate: the core system

Go 1,175 223 Updated Aug 8, 2026

K8S on FoundationDB

Go 45 4 Updated Dec 7, 2025

Tensors and Dynamic neural networks in Python with strong GPU acceleration

Python 102,294 28,784 Updated Aug 9, 2026
Rust 26 13 Updated Jun 29, 2026

High Performance LLM Inference Operator Library

C++ 1,094 127 Updated Aug 6, 2026

LoRAFusion: Efficient LoRA Fine-Tuning for LLMs

Python 28 3 Updated Jul 2, 2026

Supercharge Your Model Training

Python 5,493 464 Updated Apr 29, 2026

Accelerating MoE with IO and Tile-aware Optimizations

Python 737 96 Updated Jul 4, 2026

Training API and CLI

Python 689 77 Updated Aug 8, 2026

An experimental implementation of compiler-driven automatic sharding of models across a given device mesh.

Python 94 25 Updated Aug 7, 2026

A storage solution for PyTorch tensors with distributed tensor support.

Python 83 16 Updated Jul 17, 2026

[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.

Python 512 52 Updated Mar 30, 2026

https://github.com/eunomia-bpf homepage, documents and blogs

TypeScript 229 40 Updated Aug 9, 2026

Perplexity GPU Kernels

C++ 600 99 Updated Nov 7, 2025

slime is an LLM post-training framework for RL Scaling.

Python 7,818 1,129 Updated Aug 7, 2026

Cosmos-RL is a flexible and scalable Reinforcement Learning framework specialized for Physical AI applications.

Python 467 69 Updated Jul 31, 2026

Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"

Python 1,070 140 Updated May 30, 2026

Exploring context retrieval strategies for code editing

Python 13 3 Updated Sep 23, 2024

A std::execution style runtime context and High Performance RPC Transport for using OpenUCX. Including CUDA/ROCM/... devices with RDMA.

C++ 33 5 Updated May 26, 2026

Thread-safe asyncio-aware queue for Python

Python 964 54 Updated Aug 3, 2026

Scalable toolkit for efficient model reinforcement

Python 1,891 501 Updated Aug 9, 2026

[DEPRECATED] Moved to ROCm/rocm-libraries repo

Python 260 167 Updated Aug 4, 2026

The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (V…

Python 37,057 5,187 Updated Aug 8, 2026

TensorDict is a pytorch dedicated tensor container.

Python 1,034 116 Updated Aug 9, 2026

Research and development for optimizing transformers

Python 132 16 Updated Feb 16, 2021

Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.

Python 6,243 573 Updated Aug 22, 2025

Fast Multi-dimensional Sparse Attention

C++ 781 67 Updated Jul 26, 2026

🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.

Python 34,270 7,229 Updated Aug 9, 2026
Next