Welcome to
Computer Vision and Learning Group.

...
...
...
...
...

Our group conducts research in Computer Vision, focusing on perceiving and modeling humans.

We study computational models that enable machines to perceive and analyze human activities from visual input. We leverage machine learning and optimization techniques to build statistical models of humans and their behaviors. Our goal is to advance algorithmic foundations of scalable and reliable human digitalization, enabling a broad class of real-world applications. Our group is part of the Institute for Visual Computing (IVC) at the Department of Computer Science of ETH Zurich.

Featured Projects

In-depth look at our work.

SmoothMotionVectors: Optimizing Your Content for Video Codecs in Free View Video Compression

ConferenceSIGGRAPH 2026 Conference Track

Authors:Mingyang SongYang ZhangSiyu TangTunc Ozan Aydin

We show that Dynamic Gaussian Splatting can be aggressively compressed by combining quantization-aware training with carefully structured motion vectors. The principle is borrowed from conventional video codecs: smoother content is cheaper to encode.

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

ConferenceEuropean Conference on Computer Vision (ECCV 2026)

Authors:Xiaozhong Lyu*Gen Li*Zhiyin QianXucong ZhangMarc PollefeysSiyu Tang (*equal contribution; order interchangeable)

ReViV reconstructs viewer-centric human motion (body, hand, and gaze) and view-centric scene geometry (camera and depth) from a single egocentric RGB video in a unified feed-forward model.

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

ConferenceSIGGRAPH 2026 Journal Track

Authors:Kaifeng ZhaoMathis PetrovichHaotian ZhangTingwu WangSiyu TangDavis Rempe

ARDY is an autoregressive diffusion model for interactive human motion generation that supports online text prompting and flexible long-horizon kinematic constraints with real-time responsiveness.

GrowFields: Compositional 4D Neural Fields for Topology-Changing Plant Growth

ConferenceEuropean Conference on Computer Vision (ECCV 2026)

Authors:Joaquin GajardoMichele VolpiMarko MihajlovicSiyu TangLukas RothSergey Prokudin

GrowFields models 4D plant growth by decomposing a plant into organs and evolving them with a shared, latent-conditioned neural velocity field that learns cross-organ growth priors while handling changing topology.

NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control

ConferenceEuropean Conference on Computer Vision (ECCV 2026)

Authors:Chia-Wen Chen, Yan WuKorrawe KarunratanakulSiyu Tang

NaP-Control uses reinforcement learning to navigate the latent noise of a task-agnostic diffusion policy prior for fast, robust, and versatile physics-based character control.

Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation

ConferenceEuropean Conference on Computer Vision (ECCV 2026)

Authors:Rui WangQuentin LohmeyerSiyu TangMirko Meboldt

Multi4D enables high-quality, efficient dynamic scene reconstruction via competitive multi-level specialization, and compact, high-accuracy 4D segmentation with fast inference.

MATCH: Feed-forward Gaussian Registration for Head Avatar Creation and Editing

ConferenceConference on Computer Vision and Pattern Recognition (CVPR 2026)

Authors:Malte PrinzlerPaulo GotardoSiyu TangTimo Bolkart

Given calibrated multi-view images of human heads, MATCH infers static Gaussian splat textures in dense semantic correspondence.

BulletTime: Decoupled Control of Time and Camera Pose for Video Generation

ConferenceConference on Computer Vision and Pattern Recognition (CVPR 2026)

Authors:Yiming WangQihang ZhangShengqu CaiTong WuJan AckermannZhengfei KuangYang ZhengFrano RajičSiyu TangGordon Wetzstein

Time- and camera-controlled 4D video generation that enables decoupled control over world time and camera pose from a single input video.

Latest News

Here’s what we've been up to recently.