|
ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images
Yunfeng Wu, Hongying Cheng, Zihao He, Songhua Liu
ECCV 2026
Paper /
Code
Naively fine-tuning by single LoRA with high-resolution images introduces noise and artifacts. We introduce Relay-LoRA, a two-stage fine-tuning method that reduces noise and enhances visual detail.
|
|
FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation
Yunfeng Wu, Jiayi Song, Zhenxiong Tan, Zihao He, Songhua Liu
ECCV 2026
Paper /
Code
We identify the root cause of degradation at high resolutions and propose an efficient Flex-Attention-based interpolation window masking mechanism for seamless 4K video generation.
|
|
FiT: Flexible Image Transformer with Degradation Awareness
Zihao He, Yunfeng Wu, Xinchao Wang, Songhua Liu
ICML 2026
Paper
We present Flexible Image Transformer (FIT) that explicitly models degradation awareness across the entire pipeline, from patch sampling to pixel reconstruction.
|
|
MitPose: Multi-Granularity Guided Vision Transformer for Human Pose Estimation
Yunfeng Wu, Qizhong Gao, Yize Liu, Jun Sun, Zhuozhi Li, Yuhao Jin, Yong Yue, Xiaohui Zhu
INDIN 2025
Paper /
Code
We introduce an innovative over-parameterized convolution and global-attention complementary mechanism for multi-granularity feature representation, achieving SOTA performance on COCO and MPII benchmarks.
|
|
Alibaba Group
Research Intern | Supervised by Xiangxiang Chu
Research Direction: World Model
|
|
Shanghai Jiao Tong University
Research Intern | Supervised by Songhua Liu
Research Direction: Video Generation
|
|
Xi'an Jiaotong-Liverpool University
Research Assistant | Supervised by Yong Yue
Research Direction: Human Pose Estimation
|
Feel free to steal this website's source code. Do not scrape the HTML from this page itself, as it includes analytics tags that you do not want on your own website — use the github code instead. Also, consider using Leonid Keselman's Jekyll fork of this page.
|
|