-
Peking University
- Beijing, China
- https://zhwang4ai.github.io/
Stars
OpenGame: Open Agentic Coding for Games
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
Turn Claude Code into a full game dev studio — 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.
[ICLR 2026] LLM/VLM gaming agents and model evaluation through games.
Benchmark environment for evaluating vision-language models (VLMs) on popular video games!
Repo for Paper "OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft"
[ACL 2026] OxyGent: Making Multi-Agent Systems Modular, Observable, and Evolvable via Oxy Abstraction https://arxiv.org/abs/2604.25602
A lightweight, local-first, and free experiment tracking library from Hugging Face 🤗
Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving state-of-the-art performance on 38 out of 60 public benchmarks.
Official repository for the paper "LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code"
☁️ KUMO: Generative Evaluation of Complex Reasoning in Large Language Models
Official Implementation of "JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse"
[ICCV 2025] Official implementation of Open-World Skill Discovery from Unsegmented Demonstration Videos
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
Fully open reproduction of DeepSeek-R1
⚡⚡ Lightning Fast (~300TPS) Reinforcement Learning environment on latest Minecraft 🏝️
🏡 GitHub Pages template for personal academic homepage
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
DeepSeek-VL: Towards Real-World Vision-Language Understanding
A suite of image and video neural tokenizers
Official implementation of paper "ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting" (CVPR'25)
A collection of guides and examples for the Gemma open models from Google.
[IROS'25 Oral & NeurIPSw'24] Official implementation of "MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World Control "