Skip to content
View Irvingwangjr's full-sized avatar

Block or report Irvingwangjr

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Ulysses sequence-parallel all-to-all as a torch custom op, moved by the GPU copy engines into torch symmetric memory. Zero SM usage; 1.66-2.17x over torch.distributed on NVLink.

Python 41 5 Updated Aug 11, 2026

Agent Substrate: the core system

Go 1,217 229 Updated Aug 13, 2026

K8S on FoundationDB

Go 45 4 Updated Dec 7, 2025

Tensors and Dynamic neural networks in Python with strong GPU acceleration

Python 102,355 28,862 Updated Aug 13, 2026
Rust 26 13 Updated Jun 29, 2026

High Performance LLM Inference Operator Library

C++ 1,107 132 Updated Aug 6, 2026

LoRAFusion: Efficient LoRA Fine-Tuning for LLMs

Python 30 3 Updated Jul 2, 2026

Supercharge Your Model Training

Python 5,494 465 Updated Apr 29, 2026

Accelerating MoE with IO and Tile-aware Optimizations

Python 739 95 Updated Jul 4, 2026

Training API and CLI

Python 693 78 Updated Aug 8, 2026

An experimental implementation of compiler-driven automatic sharding of models across a given device mesh.

Python 94 25 Updated Aug 11, 2026

A storage solution for PyTorch tensors with distributed tensor support.

Python 83 16 Updated Jul 17, 2026

[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.

Python 514 52 Updated Mar 30, 2026

https://github.com/eunomia-bpf homepage, documents and blogs

TypeScript 231 39 Updated Aug 13, 2026

Perplexity GPU Kernels

C++ 600 99 Updated Nov 7, 2025

slime is an LLM post-training framework for RL Scaling.

Python 7,885 1,136 Updated Aug 12, 2026

Cosmos-RL is a flexible and scalable Reinforcement Learning framework specialized for Physical AI applications.

Python 465 69 Updated Aug 13, 2026

Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"

Python 1,072 141 Updated May 30, 2026

Exploring context retrieval strategies for code editing

Python 13 3 Updated Sep 23, 2024

A std::execution style runtime context and High Performance RPC Transport for using OpenUCX. Including CUDA/ROCM/... devices with RDMA.

C++ 33 5 Updated May 26, 2026

Thread-safe asyncio-aware queue for Python

Python 964 54 Updated Aug 3, 2026

Scalable toolkit for efficient model reinforcement

Python 1,903 512 Updated Aug 13, 2026

[DEPRECATED] Moved to ROCm/rocm-libraries repo

Python 260 167 Updated Aug 10, 2026

The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (V…

Python 37,063 5,189 Updated Aug 11, 2026

TensorDict is a pytorch dedicated tensor container.

Python 1,034 116 Updated Aug 13, 2026

Research and development for optimizing transformers

Python 132 16 Updated Feb 16, 2021

Simple and efficient pytorch-native transformer text generation in <1000 LOC of python.

Python 6,245 573 Updated Aug 22, 2025

Fast Multi-dimensional Sparse Attention

C++ 782 68 Updated Jul 26, 2026

🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.

Python 34,303 7,237 Updated Aug 13, 2026
Next