Yutao Cui
Yutao Cui

Yutao Cui 崔玉涛

Senior Researcher, Tencent Hunyuan

About Me

I am a Senior Researcher at Tencent Hunyuan. My work focuses on the pre-training and post-training of multi-modal foundation models, spanning image, video, and audio generation as well as native unified multi-modal architectures.

As a core contributor, I have been involved in the development and release of HunyuanImage 3.0, HunyuanVideo, and related multi-modal systems. I also lead the projects of MixGRPO and the HYDRA / HYDRA-X series.

I received my Ph.D. in Computer Science and Technology from Nanjing University in 2024, advised by Prof. Limin Wang.

News

Selected Publications

* indicates equal contribution. Full list on Google Scholar.

Technical Reports

HunyuanImage 3.0

HunyuanImage 3.0 Technical Report

Siyu Cao, Hangting Chen, Peng Chen, Yiji Cheng, Yutao Cui, Xinchi Deng, et al.

An industrial-scale native multi-modal model for high-quality image generation.

Tech Report, 2025 · arXiv · Code

HunyuanVideo: A Systematic Framework for Large Video Generative Models

Weijie Kong, Qi Tian, Zijian Zhang, et al. (incl. Yutao Cui)

A systematic framework for large-scale video generation from Tencent Hunyuan.

Tech Report, 2024 · arXiv

HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Sizhe Shan, Qiulin Li, Yutao Cui, Miles Yang, Yuehai Wang, Qun Yang, Jin Zhou, Zhao Zhong

High-fidelity video-to-audio generation with multimodal diffusion and representation alignment.

Tech Report, 2025 · arXiv · Code

Papers

MixGRPO method overview

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Junzhe Li*, Yutao Cui*, Tao Huang*, Yinping Ma, Chun Fan, Miles Yang, Zhao Zhong

An efficient RL method that mixes ODE/SDE sampling to accelerate flow-based preference alignment.

ECCV 2026 · arXiv · Code

HYDRA architecture

HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization

Xuerui Qiu*, Yutao Cui*, Guozhen Zhang*, Junzhe Li, JiaKui Hu, Xiao Zhang, et al.

A unified multi-modal architecture that harmonizes representations for joint generation and understanding.

arXiv 2026 · arXiv

HYDRA-X architecture

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

Guozhen Zhang*, Xuerui Qiu*, Yutao Cui*, Tianhui Song, Changlin Li, Junzhe Li, et al.

Native unified multimodal models with holistic visual tokenizers for both images and videos.

arXiv 2026 · arXiv

MixFormer framework

MixFormer: End-to-End Tracking with Iterative Mixed Attention

Yutao Cui, Cheng Jiang, Gangshan Wu, Limin Wang

An end-to-end transformer tracker with iterative mixed attention for joint feature extraction and target integration.

IEEE TPAMI 2024 (Featured Article) · IEEE · Code

StableDrag: Stable Dragging for Point-based Image Editing

Yutao Cui, Xiaotong Zhao, Guozhen Zhang, Sizhe Cao, Kai Ma, Limin Wang

A stable point-based image editing method that improves dragging control and visual consistency.

ECCV 2024 · arXiv

MixFormerV2 model

MixFormerV2: Efficient Fully Transformer Tracking

Yutao Cui, Tianhui Song, Gangshan Wu, Limin Wang

A fully transformer tracker designed for high efficiency and real-time deployment.

NeurIPS 2023 · arXiv · Code

SportsMOT dataset overview

SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes

Yutao Cui, Chenkai Zeng, Xiaoyu Zhao, Yichun Yang, Gangshan Wu, Limin Wang

A large multi-object tracking dataset and benchmark focused on challenging sports scenes.

ICCV 2023 · arXiv · Code

MixFormer framework

MixFormer: End-to-End Tracking with Iterative Mixed Attention

Yutao Cui, Cheng Jiang, Limin Wang, Gangshan Wu

CVPR 2022 (Oral) · PDF · arXiv · Code

Fully Convolutional Online Tracking

Yutao Cui, Cheng Jiang, Limin Wang, Gangshan Wu

Computer Vision and Image Understanding (CVIU) 2022 · arXiv

Target Transformed Regression for Accurate Tracking

Yutao Cui, Cheng Jiang, Limin Wang, Gangshan Wu

arXiv 2021 · arXiv

Awards & Honors