-
SJTU
- Shanghai, China
- https://sunzey.github.io/
- @sunzeyi6
- https://www.cnblogs.com/zery-blog/
Highlights
- Pro
Lists (3)
Sort Name ascending (A-Z)
Starred repositories
Lightweight coding agent that runs in your terminal
DeepSeek Harness: Everything is a Plugin.
Qwen-CUA: Native Computer Use for (Almost) Everything — a screenshot-driven agent that operates computers with keyboard and mouse, jointly developed by the Qwen Team and XLang Lab.
Make any agent harness multimodal-native.
The unified framework for sim & real robot teleoperation
Open-source framework for computer use agents: VeriGen verifiable task synthesis, online RL training (AgentRL), and OSWorld/ScienceBoard evaluation.
MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research · 浏览器里运行的安卓模拟器 · Browser-hosted Android Simulator · Verifiable Evaluation · Scalable Online RL Training
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Qwen-AgentWorld: Language World Models for General Agents
Post-training with Tinker
Official repository for the survey "Reinforcement Learning for Visual Generation: A Model-Oriented Survey".
HY-Embodied-0.5-X: An Enhanced Embodied Foundation Model for Real-World Agents
AI-powered, vision-driven UI automation for every platform.
A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io
Multimodal RL training framework for diffusion & omni models
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System
[ECCV 2026] Generative Refinement Networks for Visual Synthesis (Support C2I & T2I & T2V)
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
MiMo Code: Where Models and Agents Co-Evolve
Official codebase for Fast-WAM: Do World Action Models Need Test-time Future Imagination?
GPT Image 2 prompt gallery, image prompt library, agentic skill, and CLI for OpenAI image generation/editing