๐Ÿ‘€ About me

I build real-time, always-on video AI that perceives, reasons, and responds on the fly.

I am Zhenyu Yang (ๆจๆŒฏๅฎ‡), a fourth-year Ph.D. student (2022-2027) at the State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences, advised by Prof. Changsheng Xu. Previously, I earned my Bachelorโ€™s degree from Beijing University of Posts and Telecommunications in 2022.

My research interests include 1. Streaming Video Understanding, 2. Multimodal Large Language Models, 3. Multimodal Retrieval. I have previously worked as a research intern with the Tencent Hunyuan team, the Kuaishou Keye team, and the 360 AI Department. I welcome collaboration and am always open to discussing research opportunitiesโ€”feel free to reach out via email!!!

๐Ÿ”ฅ News

๐Ÿ“š Survey and Awesome List

A Survey on Streaming Video Understanding โ€” a comprehensive review of papers, models, and datasets for real-time, always-on, interactive video AI, covering proactive decision-making (when to act) and efficient long-context resource management (how to sustain).

I also maintain Awesome-Streaming-Video-Understanding, a continuously updated, curated list accompanying the survey.

GitHub stars

๐Ÿ“ Publications

NeurIPS 2025
sym

๐Ÿ”ฅ [NeurIPSโ€™2025] Poster

LiveStar: Live Streaming Assistant for Real-World Online Video Understanding

Zhenyu Yang, Kairui Zhang, Yuhang Hu, Bing Wang, Shengsheng Qian, Bin Wen, Fan Yang, Tingting Gao, Weiming Dong, Changsheng Xu

[Code] [Project] [Paper] [ไธญๆ–‡่งฃ่ฏป]

ICLR 2025
sym

๐Ÿš€ [ICLRโ€™2025] Spotlight

SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding

Zhenyu Yang, Yuhang Hu, Zemin Du, Dizhan Xue, Shengsheng Qian, Jiahong Wu, Fan Yang, Weiming Dong, Changsheng Xu

[Code] [Project] [Paper] [Dataset] [Model] [Leaderboard] [Submission] [ไธญๆ–‡่งฃ่ฏป]

SIGIR 2024
sym

๐Ÿ† [SIGIRโ€™2024] Best Paper Honorable Mention

LDRE: LLM-based Divergent Reasoning and Ensemble for Zero-Shot Composed Image Retrieval

Zhenyu Yang, Dizhan Xue, Shengsheng Qian, Weiming Dong, Changsheng Xu

[Code] [Paper] [Video]

ACM MM 2024
sym

๐ŸŽ‰ [ACM MMโ€™2024] Poster

Semantic Editing Increment Benefits Zero-Shot Composed Image Retrieval

Zhenyu Yang, Shengsheng Qian, Dizhan Xue, Jiahong Wu, Fan Yang, Weiming Dong, Changsheng Xu

[Code] [Paper]

ECCV 2026
sym

๐ŸŽ‰ [ECCVโ€™2026] Poster

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

Xumin Yu*, Zuyan Liu*, Zhenyu Yang*, Yuhao Dong, Shengsheng Qian, Jiwen Lu, Han Hu, Yongming Rao

[Code] [Paper] [Model]

* Equal contribution

ICLR 2026
sym

๐ŸŽ‰ [ICLRโ€™2026] Poster

Querystream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response

Kairui Zhang, Zhenyu Yang, Bing Wang, Shengsheng Qian, Changsheng Xu

[Code] [Paper]

ACM MM 2025
sym

๐ŸŽ‰ [ACM MMโ€™2025] Poster

StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA

Yuhang Hu, Zhenyu Yang, Shihan Wang, Shengsheng Qian, Bin Wen, Fan Yang, Tingting Gao, Changsheng Xu

[Code] [Paper]

ICIG 2025
sym

๐Ÿ† [ICIGโ€™2025] Best Paper Award

Multi-View Captioning with Semantic Delta Re-Ranking for Zero-Shot Composed Video Retrieval

Zhixiang Ding, Lilong Liu, Zhenyu Yang, Shengsheng Qian

[Code] [Project] [Paper]

๐ŸŽ– Honors and Awards

  • Best Paper Honorable Mention (5/791), SIGIR, 2024
  • Best Paper Award, ICIG, 2025
  • Spotlight Paper (~3.27%), ICLR, 2025
  • CIE-Tencent Doctoral Research Incentive Project / ๆททๅ…ƒๅญฆ่€… (ไธญๅ›ฝ็”ตๅญๅญฆไผš-่…พ่ฎฏๅšๅฃซ็”Ÿ็ง‘็ ”ๆฟ€ๅŠฑ่ฎกๅˆ’), 2025
  • National Scholarship, Ministry of Education, China, 2024
  • Outstanding Graduate, Beijing, 2022
  • Outstanding Graduate, Beijing University of Posts and Telecommunications, 2022
  • First-Class Scholarship, Beijing University of Posts and Telecommunications, 2020/2021
  • First Prize in American Mathematical Contest in Modeling (MCM), Top 6.7% Globally, 2020

๐Ÿ“– Education

  • 2022.09 - 2027.06: Ph.D, State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing. Major: Computer Applied Technology.
  • 2018.09 - 2022.06: Undergraduate, School of Computer Science (National Pilot Software Engineering School), Beijing University of Posts and Telecommunications, Beijing. Major: Intelligent Science and Technology.

๐Ÿ™‹ Services

  • Conference Reviewer: AAAI 2027, NeurIPS 2026, ECCV 2026, CVPR 2026, ICLR 2026, AAAI 2026, NeurIPS 2025, ICCV 2025, ACML 2025, etc.
  • Journal Reviewer: IEEE Transactions on Image Processing (TIP), ACM Transactions on Multimedia Computing Communications and Applications (TOMM), ACM Transactions on Information Systems (TOIS), Neurocomputing, Pattern Recognition.