๐ About me
I build real-time, always-on video AI that perceives, reasons, and responds on the fly.
I am Zhenyu Yang (ๆจๆฏๅฎ), a fourth-year Ph.D. student (2022-2027) at the State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences, advised by Prof. Changsheng Xu. Previously, I earned my Bachelorโs degree from Beijing University of Posts and Telecommunications in 2022.
My research interests include 1. Streaming Video Understanding, 2. Multimodal Large Language Models, 3. Multimodal Retrieval. I have previously worked as a research intern with the Tencent Hunyuan team, the Kuaishou Keye team, and the 360 AI Department. I welcome collaboration and am always open to discussing research opportunitiesโfeel free to reach out via email!!!
๐ฅ News
- 2026.07: ย ๐๐ Two of our papers have been accepted to ACM MM 2026!
- 2026.06: ย ๐๐ Our paper โViQ: Text-Aligned Visual Quantized Representations at Any Resolutionโ has been accepted to ECCV 2026!
- 2026.06: ย โจโจ We released our โTowards Online Interactors: A Comprehensive Survey on Streaming Video Understandingโ and an Awesome-Streaming-Video-Understanding list โ the most comprehensive collection of papers, code, and datasets for real-time, always-on video AI. Check it out and give us a โญ!
- 2026.02: ย ๐๐ Our paper โQuerystream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Responseโ has been accepted to ICLR 2026!
- 2025.11: ย ๐๐ Congratulations to our โMulti-View Captioning with Semantic Delta Re-Ranking for Zero-Shot Composed Video Retrievalโ for winning the Best Paper Award at ICIG 2025!
- 2025.09: ย ๐๐ Our paper โLiveStar: Live Streaming Assistant for Real-World Online Video Understandingโ about streaming Video-LLMs has been accepted to NeurIPS 2025!
- 2025.08: ย ๐๐ Our paper โStreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQAโ has been accepted to ACM MM 2025 Datasets!
- 2025.07: ย ๐๐ I was awarded the CIE-Tencent Doctoral Research Incentive Project, a competitive grant awarded to only 23 recipients nationwide, along with a research fund of 100,000 RMB.
๐ Survey and Awesome List
A Survey on Streaming Video Understanding โ a comprehensive review of papers, models, and datasets for real-time, always-on, interactive video AI, covering proactive decision-making (when to act) and efficient long-context resource management (how to sustain).
I also maintain Awesome-Streaming-Video-Understanding, a continuously updated, curated list accompanying the survey.
๐ Publications
๐ฅ [NeurIPSโ2025] Poster
LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
Zhenyu Yang, Kairui Zhang, Yuhang Hu, Bing Wang, Shengsheng Qian, Bin Wen, Fan Yang, Tingting Gao, Weiming Dong, Changsheng Xu
๐ [ICLRโ2025] Spotlight
SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
Zhenyu Yang, Yuhang Hu, Zemin Du, Dizhan Xue, Shengsheng Qian, Jiahong Wu, Fan Yang, Weiming Dong, Changsheng Xu
[Code] [Project] [Paper] [Dataset] [Model] [Leaderboard] [Submission] [ไธญๆ่งฃ่ฏป]
๐ [SIGIRโ2024] Best Paper Honorable Mention
LDRE: LLM-based Divergent Reasoning and Ensemble for Zero-Shot Composed Image Retrieval
Zhenyu Yang, Dizhan Xue, Shengsheng Qian, Weiming Dong, Changsheng Xu
๐ [ACM MMโ2024] Poster
Semantic Editing Increment Benefits Zero-Shot Composed Image Retrieval
Zhenyu Yang, Shengsheng Qian, Dizhan Xue, Jiahong Wu, Fan Yang, Weiming Dong, Changsheng Xu
๐ [ECCVโ2026] Poster
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
Xumin Yu*, Zuyan Liu*, Zhenyu Yang*, Yuhao Dong, Shengsheng Qian, Jiwen Lu, Han Hu, Yongming Rao
* Equal contribution
๐ [ICLRโ2026] Poster
Querystream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response
Kairui Zhang, Zhenyu Yang, Bing Wang, Shengsheng Qian, Changsheng Xu
๐ [ACM MMโ2025] Poster
Yuhang Hu, Zhenyu Yang, Shihan Wang, Shengsheng Qian, Bin Wen, Fan Yang, Tingting Gao, Changsheng Xu
๐ [ICIGโ2025] Best Paper Award
Multi-View Captioning with Semantic Delta Re-Ranking for Zero-Shot Composed Video Retrieval
Zhixiang Ding, Lilong Liu, Zhenyu Yang, Shengsheng Qian
๐ Honors and Awards
- Best Paper Honorable Mention (5/791), SIGIR, 2024
- Best Paper Award, ICIG, 2025
- Spotlight Paper (~3.27%), ICLR, 2025
- CIE-Tencent Doctoral Research Incentive Project / ๆททๅ ๅญฆ่ (ไธญๅฝ็ตๅญๅญฆไผ-่ พ่ฎฏๅๅฃซ็็ง็ ๆฟๅฑ่ฎกๅ), 2025
- National Scholarship, Ministry of Education, China, 2024
- Outstanding Graduate, Beijing, 2022
- Outstanding Graduate, Beijing University of Posts and Telecommunications, 2022
- First-Class Scholarship, Beijing University of Posts and Telecommunications, 2020/2021
- First Prize in American Mathematical Contest in Modeling (MCM), Top 6.7% Globally, 2020
๐ Education
- 2022.09 - 2027.06: Ph.D, State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing. Major: Computer Applied Technology.
- 2018.09 - 2022.06: Undergraduate, School of Computer Science (National Pilot Software Engineering School), Beijing University of Posts and Telecommunications, Beijing. Major: Intelligent Science and Technology.
๐ Services
- Conference Reviewer: AAAI 2027, NeurIPS 2026, ECCV 2026, CVPR 2026, ICLR 2026, AAAI 2026, NeurIPS 2025, ICCV 2025, ACML 2025, etc.
- Journal Reviewer: IEEE Transactions on Image Processing (TIP), ACM Transactions on Multimedia Computing Communications and Applications (TOMM), ACM Transactions on Information Systems (TOIS), Neurocomputing, Pattern Recognition.