Skip to content
View fxmeng's full-sized avatar

Block or report fxmeng

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results
Python 402 50 Updated Jul 30, 2026

One-pass, O(L) memory computation of attention KL loss

Python 7 Updated Jun 11, 2026

🚀 Efficient implementations for emerging model architectures

Python 5,531 644 Updated Aug 9, 2026

slime is an LLM post-training framework for RL Scaling.

Python 7,822 1,130 Updated Aug 7, 2026

An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

Python 566 137 Updated Aug 10, 2026

A kernel library written in tilelang

Python 1,712 155 Updated Apr 23, 2026

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++ 6,210 1,066 Updated Aug 10, 2026

Minimal AI coding agent (~1,000 lines of Python) inspired by Claude Code. Works with any LLM. Think NanoGPT for coding agents. Formerly NanoCoder.

Python 1,638 380 Updated Aug 4, 2026

LegalOne: A Family of Foundation Models for Reliable Legal Reasoning

71 7 Updated Feb 3, 2026

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …

Python 15,096 1,581 Updated Aug 9, 2026

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,016 109,231 Updated Aug 6, 2026

The repo for SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass

Jupyter Notebook 97 6 Updated May 23, 2026

IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse

132 11 Updated Mar 14, 2026
Swift 11 Updated Mar 7, 2026

Byted PyTorch Distributed for Hyperscale Training of LLMs and RLs

Python 1,036 64 Updated Mar 3, 2026

🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"

Python 1,017 53 Updated Feb 5, 2026

Youtu-RAG: Next-Generation Agentic Intelligent Retrieval-Augmented Generation System

Python 274 31 Updated Apr 1, 2026

VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

Python 2,135 247 Updated Aug 7, 2026

MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)

Python 1,662 91 Updated Feb 14, 2026

Youtu-Tip: Tap for Intelligence, Keep on Device.

Python 593 67 Updated Feb 27, 2026

The local UI to run and train text and diffusion models, including Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, FLUX and more.

Python 69,774 6,298 Updated Aug 10, 2026

Design hardware-friendly model architectures and migrate existing LLMs with minimal performance loss

Python 497 32 Updated Jul 14, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 7,166 684 Updated Aug 9, 2026

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

C++ 1,515 276 Updated Aug 10, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,888 4,365 Updated Aug 10, 2026

A repository aimed at pruning DeepSeek V3, R1 and R1-zero to a usable size

Python 87 9 Updated Sep 5, 2025

[ICLR 2026] Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning

Python 1,236 183 Updated Feb 26, 2026

Nano vLLM

Python 14,924 2,439 Updated Apr 26, 2026

A plug-and-play library for parameter-efficient-tuning (Delta Tuning)

Python 1,046 83 Updated Sep 19, 2024
Next