Starred repositories
Official repository for “HelloWorld: Enabling Socially Interactive Characters in Video World Models”
Full-stack open-source interactive long-horizon world model.
[CVPR 2026] UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking
Provide with pre-build flash-attention 2 and 3 package wheels on Linux and Windows using GitHub Actions
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice…
[TPAMI 2026] Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views (with Visual Imitation Learning for Robots)
FloodDiffusion: Tailored Diffusion Forcing for Streaming Motion Generation
LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing
[NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning
Official PyTorch implementation of the paper: UniGaze: Towards Universal Gaze Estimation via Large-scale Pre-Training.
A complete, cross-platform solution to record, convert and stream audio and video.
A project page template for academic papers. Demo at https://eliahuhorwitz.github.io/Academic-project-page-template/
Github Pages template based upon HTML and Markdown for personal, portfolio-based websites.
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.
Official Code for CVPR 2024 paper "Complementing Event Streams and RGB Frames for Hand Mesh Reconstruction"
3D Pose Estimation of Two Interacting Hands from a Monocular Event Camera [3DV'24]
HaMeR: Reconstructing Hands in 3D with Transformers
Face alignment & Anime FFHQ alignment
[ICCV 2023] BlendFace: Re-designing Identity Encoders for Face-Swapping https://arxiv.org/abs/2307.10854
Official implementation of Würstchen: Efficient Pretraining of Text-to-Image Models
[ICCV 2023, Oral] Iterative Prompt Learning for Unsupervised Backlit Image Enhancement
Research code for CVPR 2021 paper "End-to-End Human Pose and Mesh Reconstruction with Transformers"