-
Peking University
- https://fxmeng.github.io
Lists (3)
Sort Name ascending (A-Z)
Stars
One-pass, O(L) memory computation of attention KL loss
🚀 Efficient implementations for emerging model architectures
slime is an LLM post-training framework for RL Scaling.
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
A kernel library written in tilelang
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Minimal AI coding agent (~1,000 lines of Python) inspired by Claude Code. Works with any LLM. Think NanoGPT for coding agents. Formerly NanoCoder.
LegalOne: A Family of Foundation Models for Reliable Legal Reasoning
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
The repo for SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
Byted PyTorch Distributed for Hyperscale Training of LLMs and RLs
🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"
Youtu-RAG: Next-Generation Agentic Intelligent Retrieval-Augmented Generation System
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)
Youtu-Tip: Tap for Intelligence, Keep on Device.
The local UI to run and train text and diffusion models, including Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, FLUX and more.
Design hardware-friendly model architectures and migrate existing LLMs with minimal performance loss
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
A repository aimed at pruning DeepSeek V3, R1 and R1-zero to a usable size
[ICLR 2026] Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning
A plug-and-play library for parameter-efficient-tuning (Delta Tuning)