-
Tsinghua University
- Beijing, China
-
08:26
(UTC +08:00) - thuwzy.github.io
Starred repositories
A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
[ECCV 2026] Official code of GEM: Generative Supervision Helps Embodied Intelligence
[ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
Muon is an optimizer for hidden layers in neural networks
[ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
Qwen-Image-Lightning: Speed up Qwen-Image model with distillation
Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference
Hunyuan3D-Omni: A Unified Framework for Controllable Generation of 3D Assets
ViPE: Video Pose Engine for Geometric 3D Perception
A curated collection of fun and creative examples generated with Nano Banana & Nano Banana Pro🍌, Gemini-2.5-flash-image based model. We also release Nano-consistent-150K openly to support the commu…
Voyager is an interactive RGBD video generation model conditioned on camera input, and supports real-time 3D reconstruction.
4DNeX: Feed-Forward 4D Generative Modeling Made Easy
Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
Generate large-scale explorable 3D scenes with high-quality panorama videos from a single image or text prompt.
Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
[NeurIPS 2024 & TPAMI 2026] Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
PhysX: Physical-Grounded 3D Asset Generation (NeurIPS 2025, Spotlight)
[CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
Towards a Generative 3D World Engine for Embodied Intelligence
[ICLR'26] Topology-Preserved Auto-regressive Mesh Generation in the Manner of Weaving Silk