-
xAI; NVIDIA Research; CUHK MMLab; Zhejiang University
- Seattle, WA
- https://alvinliu0.github.io/
- https://scholar.google.com/citations?hl=en&user=QHx5ncgAAAAJ
- @AlvinLiu27
- in/xian-liu-9840b52a3
Stars
[IEEE TPAMI 2026] Simulating the Real World: Survey & Resources, which contains our survey "Simulating the Real World: A Unified Survey of Multimodal Generative Models" (IEEE TPAMI, 2026) and Aweso…
[CVPR 2025] HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation
Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the form of video.
Cosmos-Predict2 is a collection of general-purpose world foundation models for Physical AI that can be fine-tuned into customized world models for downstream applications.
[ICLR’26] Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
Cosmos-Predict1 is a collection of general-purpose world foundation models for Physical AI that can be fine-tuned into customized world models for downstream applications.
Cosmos-Transfer1 is a world-to-world transfer model designed to bridge the perceptual divide between simulated and real-world environments.
NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.
[ICLR'25] 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation
[ICLR 2025] EdgeRunner: Auto-regressive Auto-encoder for Efficient Mesh Generation
[ICLR 2025] Implementation of Accelerating Auto-regressive Text-to-Image Generation with Training-free Speculative Jacobi Decoding
A suite of image and video neural tokenizers
TC4D: Trajectory-Conditioned Text-to-4D Generation
[ECCV 2024] The official implementation of paper "BrushNet: A Plug-and-Play Image Inpainting Model with Decomposed Dual-Branch Diffusion"
[ICCV 2023] The official implementation of paper "HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation"
4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling
[CVPR 2024 Highlight] Code for "HumanGaussian: Text-Driven 3D Human Generation with Gaussian Splatting"
[ICLR 2024] Github Repo for "HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion"
[CVPR'2023] Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation
The repository for paper Unsupervised Volumetric Animation
Elucidating the Design Space of Diffusion-Based Generative Models (EDM)
Text-to-3D & Image-to-3D & Mesh Exportation with NeRF + Diffusion.
Machine Learning Interviews from FAANG, Snapchat, LinkedIn. I have offers from Snapchat, Coupang, Stitchfix etc. Blog: mlengineer.io.
Official PyTorch Implementation of "Learning to Learn with Generative Models of Neural Network Checkpoints"
Official implementation of Cold-Diffusion for different transformations in pytorch.
Awesome Lists for Tenure-Track Assistant Professors and PhD students. (助理教授/博士生生存指南)
[CVPR 2022] Code for "Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation"
Tackling the Generative Learning Trilemma with Denoising Diffusion GANs https://arxiv.org/abs/2112.07804