Hi there! I am a senior undergraduate at Huazhong University of Science and Technology (B.S. in Intelligent Science and Technology, 2023 - 2027), originally from Taipei, Taiwan. I am currently a research intern at Tencent Hunyuan with Tianyu Pang, and at VisionX Lab, NYU Courant with Prof. Saining Xie, where I spent Fall 2025 as an exchange student. Previously I worked at OpenDILab, Shanghai AI Lab with Yazhe Niu, and at CAS-SIAT with Prof. Min Yang.
I am applying for CS PhD programs starting Fall 2027, focused on vision-centric multimodal reasoning and visual agents. I am also open to long-term research internships and collaborations in the meantime — if you are working on any of the above, I would like to hear from you.
I work on the gap between what a model sees and what it understands. Today’s VLMs can name what is visible yet miss what matters: implication, spatial state, and how that state changes over time. My work traces this gap from benchmarks that expose it to visual reinforcement learning that begins to close it.
I believe vision still lacks what code became for language models: a training ground where long-horizon work is grounded in real scenes, outcomes are verifiable, and skills transfer. Building it requires agents that preserve and revise visual state over time, use tools to inspect the world, and learn from their actions. But perception and memory are not enough. Long-horizon intelligence also needs an experience-shaped value signal that guides action before deliberation is complete. Humans have such a signal: emotion. My goal is to build vision-centric agents that see, remember, reason, and make decisions in the real world.
Wuhan, China
Supervised by Tianyu Pang
Focus: Multimodal evaluation, Multimodal RL post-training, Long-horizon visual agent harnesses, Video generation.
Supervised by Prof. Saining Xie
Focus: Thinking with images, Spatial intelligence, Explicit visual states, Visual state tracking and memory, Multi-step visual inference.
Supervised by Yazhe Niu, Chaochao Lu and Prof. Hongsheng Li
Focus: Multimodal RL post-training, Agentic RL for long-horizon trajectories, Latent world models, Open-source RL infrastructure.
Supervised by Minghuan Tan, Prof. Shiwen Ni and Prof. Min Yang
Focus: Multimodal benchmarks, Image implication understanding, Human-centered evaluation of language models.
Supervised by Prof. Hong Huang and Prof. Hai Jin
Focus: Graph understanding and reasoning, Instruction tuning for large language models.