Skip to content
View JiwenJ's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report JiwenJ

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

实现 mHC: Manifold-Constrained Hyper-Connections

Python 17 Updated Jan 2, 2026

implementations and experimentation on mHC by deepseek - https://arxiv.org/abs/2512.24880

Shell 372 34 Updated Feb 17, 2026

Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory

Python 134 3 Updated Jul 27, 2026

A self-improving RLM agent for coding workflows and long-running autonomous tasks.

TypeScript 16,243 1,745 Updated Aug 15, 2026

Piecewise-Taylor Attention

Python 41 2 Updated Aug 4, 2026

NVSentinel is a cross-platform fault remediation service designed to rapidly remediate runtime node-level issues in GPU-accelerated computing environments

Go 371 107 Updated Aug 14, 2026

Embodied AI Operating System (EAIOS)

Rust 328 50 Updated Aug 15, 2026

A tool for recording RL trajectories.

Python 128 19 Updated Jul 30, 2026

Helpful kernel tutorials, examples and SKILLs for tile-based GPU programming

Python 796 82 Updated Aug 13, 2026

Mixture-of-experts (MoE) training megakernel for NVL72s

Python 535 60 Updated Aug 14, 2026

MAGI-2-preview: Scaling Video Generation Models Efficiently

Python 519 12 Updated Aug 6, 2026

DeepStack: Facilitating Co-Design Exploration of 3D DRAM-Stacked Accelerators for Distributed LLM Inference. Includes the MICRO 2026 AE artifact.

Python 38 Updated Jul 24, 2026
Python 37 1 Updated Aug 4, 2026

20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.

Python 13,616 1,487 Updated Aug 13, 2026

Making large AI models cheaper, faster and more accessible

Python 41,437 4,502 Updated Aug 10, 2026

Everything you need to know about LLM inference

TypeScript 384 37 Updated Aug 14, 2026

Official repository for paper Routing-Free Mixture-of-Experts.

Python 17 1 Updated Apr 19, 2026

[TMLR 2026] LibMoE: A LIBRARY FOR COMPREHENSIVE BENCHMARKING MIXTURE OF EXPERTS IN LARGE LANGUAGE MODELS

Jupyter Notebook 52 Updated May 26, 2026

Reference implementation and examples of the CuTe Layout representation and algebra.

Python 266 25 Updated Aug 6, 2026

Kanana: Compute-efficient Bilingual Language Models

286 15 Updated Jul 23, 2025

🍎 One kernel a day keeps high latency away. A hands-on CUDA learning path featuring a rich collection of kernels, from the basics to peak performance, seamlessly integrated as PyTorch C++ extensions.

Cuda 208 18 Updated Aug 10, 2026

large language model internal-medicine monitor toolbox

Python 14 11 Updated Aug 13, 2026

General purpose GPU compute framework built on Vulkan to support 1000s of cross vendor graphics cards (AMD, Qualcomm, NVIDIA & friends). Blazing fast, mobile-enabled, asynchronous and optimized for…

C++ 2,554 198 Updated Aug 15, 2026

GPU documentation for humans

Python 666 85 Updated Jun 14, 2026

High-performance GPU kernels written in TIRx.

Python 87 7 Updated Aug 15, 2026
Python 73 9 Updated Jul 27, 2026

Open Frontier Intelligence

8,461 659 Updated Aug 6, 2026

MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts

Python 1,079 119 Updated Aug 13, 2026

AgentENV (AENV) is a distributed platform for running agent environments at scale.

Rust 3,200 268 Updated Aug 15, 2026

Root Mean Square Layer Normalization

Python 285 19 Updated Mar 28, 2023
Next