-
BS@UoL; RA@SJTU; Intern@Alibaba; Incoming MSc@HKU
- Beijing, China
-
14:05
(UTC +08:00) - https://willwu111.github.io
Stars
Official implementation for paper "MaskFlow: Precise, Consistent and Seamless Regional Image Editing". Code and Dataset are publicly available.
Benchmarking Knowledge Transfer in Lifelong Robot Learning
We propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos.
Context Forcing: Consistent Autoregressive Video Generation with Long Context [ICML26]
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
[CVPR 2026] [Best Paper Finalist] [Oral] Official repository of Vision Test-Time Training
JoyAI-Echo: Pushing the Frontier of Long Audio-Visual Generation
A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
Xetrieval: Mechanistically Explaining Dense Retrieval
A Minimalist, Batteries-included Repository for Advancing World Model Science.
Code for Fast Training of Diffusion Models with Masked Transformers
aaaaxiaoxu / biztrack
Forked from sptin2002/biztrackBizTrack is a simple web application for small businesses, providing an intuitive order tracking system to manage finances, calculate total order income and expenses, and track profits or losses se…
[ICLR 2026] Official implementation of JavisDiT and JavisDiT++ series.
[ECCV2026] ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images
[ICCV 2025] Official implementation of the paper: REPA-E: Unlocking VAE for End-to-End Tuning of Latent Diffusion Transformers
hammershock / moffee
Forked from wbopan/moffeemoffee: Make Markdown Ready to Present
A Lightweight, Configuration-Driven, Flexible Fine-Tuning Framework for 🤗 Diffusers
DreamX-World: A General-Purpose Interactive World Model
A platform for reproducible world model research and evaluation
TempoFit: Plug-and-Play Layer-Wise Temporal KV Memory for Long-Horizon Vision-Language-Action Manipulation
A comprehensive benchmark specifically designed to evaluate the interactive response capabilities of world models in 4D settings.
4-steps distilled version of Wan2.2-TI2V-5B
[AAAI-2026]FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
Lets make video diffusion practical!
MOVA: Towards Scalable and Synchronized Video–Audio Generation
[ICML 2026] Official repository for the paper "Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention"