-
Institute of Automation, Chinese Academy of Sciences
- Beijing
-
03:16
(UTC +08:00) - https://jiwenj.github.io/
- https://www.zhihu.com/people/JiwenJ
Lists (8)
Sort Name ascending (A-Z)
Starred repositories
implementations and experimentation on mHC by deepseek - https://arxiv.org/abs/2512.24880
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
A self-improving RLM agent for coding workflows and long-running autonomous tasks.
NVSentinel is a cross-platform fault remediation service designed to rapidly remediate runtime node-level issues in GPU-accelerated computing environments
A tool for recording RL trajectories.
Helpful kernel tutorials, examples and SKILLs for tile-based GPU programming
Mixture-of-experts (MoE) training megakernel for NVL72s
MAGI-2-preview: Scaling Video Generation Models Efficiently
DeepStack: Facilitating Co-Design Exploration of 3D DRAM-Stacked Accelerators for Distributed LLM Inference. Includes the MICRO 2026 AE artifact.
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Making large AI models cheaper, faster and more accessible
Everything you need to know about LLM inference
Official repository for paper Routing-Free Mixture-of-Experts.
[TMLR 2026] LibMoE: A LIBRARY FOR COMPREHENSIVE BENCHMARKING MIXTURE OF EXPERTS IN LARGE LANGUAGE MODELS
Reference implementation and examples of the CuTe Layout representation and algebra.
🍎 One kernel a day keeps high latency away. A hands-on CUDA learning path featuring a rich collection of kernels, from the basics to peak performance, seamlessly integrated as PyTorch C++ extensions.
large language model internal-medicine monitor toolbox
General purpose GPU compute framework built on Vulkan to support 1000s of cross vendor graphics cards (AMD, Qualcomm, NVIDIA & friends). Blazing fast, mobile-enabled, asynchronous and optimized for…
High-performance GPU kernels written in TIRx.
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
AgentENV (AENV) is a distributed platform for running agent environments at scale.