-
Sun Yat-Sen University
- Alibaba, Hangzhou, China
- https://www.zhihu.com/people/jian-xin-15-96
Stars
Qwen-AgentWorld: Language World Models for General Agents
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
AI agents running research on single-GPU nanochat training automatically
🦉 OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
"OpenHarness: Open Agent Harness with a Built-in Personal Agent--Ohmo!"
A benchmark for LLMs on complicated tasks in the terminal
Official code of "StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs".
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
Pytorch Implementation of "Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language Models", AAAI 2025
Pytorch Implementation of "Sinkhorn Distance Minimization for Knowledge Distillation", COLING 2024 and TNNLS 2024
Unleashing the Power of Reinforcement Learning for Math and Code Reasoners
Super-Efficient RLHF Training of LLMs with Parameter Reallocation
Official Repo for Open-Reasoner-Zero
Scalable RL solution for advanced reasoning of language models
Efficient Triton Kernels for LLM Training
[ICLR 2024]EMO: Earth Mover Distance Optimization for Auto-Regressive Language Modeling(https://arxiv.org/abs/2310.04691)
OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
[ICLR2023] PLOT: Prompt Learning with Optimal Transport for Vision-Language Models
Implementation of Sinkhorn algorithms in Torch.
code for paper "BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation"
An Open Source Toolkit For LLM Distillation
llm deploy project based mnn. This project has merged into MNN.
[CVPR 2023] DepGraph: Towards Any Structural Pruning; LLMs, Vision Foundation Models, etc.