-
Peking University
- Beijing, China
-
20:05
(UTC -12:00) - phython96.github.io
- https://scholar.google.com/citations?user=MZXDSSUAAAAJ&hl=zh-CN
Lists (2)
Sort Name ascending (A-Z)
Stars
DROL is a one-step offline RL actor trained with top-1 dynamic routing.
Decoupling Manifold Modeling and Value Maximization for Offline Policy Extraction
Official Repo for paper: Scaling Behavior Cloning Improves Causal Reasoning: An Open Model for Real-Time Video Game Playing
[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.
A simple yet powerful agent framework that delivers with open-source models
AndroidWorld is an environment and benchmark for autonomous agents
🔥 A minimal training framework for scaling FLA models
Code release for paper "Test-Time Training Done Right"
A Dockerized Android emulator supporting multiple CPU architectures (x86 and arm64) with native performance and seamless ADB & Scrcpy Web access.
Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
Android in docker solution with noVNC supported and video recording
Model Context Protocol Server for Mobile Automation and Scraping (iOS, Android, Emulators, Simulators and Real Devices)
Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
A cross-platform GUI automation Python module for human beings. Used to programmatically control the mouse & keyboard.
Agent framework and applications built upon Qwen>=3.0, featuring Function Calling, MCP, Code Interpreter, RAG, Chrome extension, etc.
🌐 Make websites accessible for AI agents. Automate tasks online with ease.
A repo for open research on building large reasoning models
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention
RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RN…
Lets make video diffusion practical!
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
[AAAI'26 Oral] DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping
Paper list in the survey: A Survey on Vision-Language-Action Models: An Action Tokenization Perspective