-
University of Rochester
- Rochester, NY, US
Highlights
- Pro
Stars
[CVPR 2025] VideoWorld is a simple generative model that learns purely from unlabeled videos—much like how babies learn by observing their environment.
Code release for "PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop" (ICML 2025)
Code release for https://kovenyu.com/WonderWorld/
Simulation platform for general-purpose robotics & embodied AI learning.
Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).
[CVPR 2025 Highlight] Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
[NeurIPS'2024]: DiffGS: Functional Gaussian Splatting Diffusion
[arXiv 2023] DreamGaussian4D: Generative 4D Gaussian Splatting
[CVPR 2024] 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838
[ICLR'25 Oral] No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images
[SIGGRAPH 2024] Motion I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling
Official inference repo for FLUX.1 models
This repo contains the code for 1D tokenizer and generator
Official PyTorch Implementation of "SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers"
This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
Open-Sora: Democratizing Efficient Video Production for All
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation [TMLR 2024]
[WIP] Layer Diffusion for WebUI (via Forge)
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Aligning LMMs with Factually Augmented RLHF
PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
LVDM: Latent Video Diffusion Models for High-Fidelity Long Video Generation
General technology for enabling AI capabilities w/ LLMs and MLLMs
AlignProp uses direct reward backpropogation for the alignment of large-scale text-to-image diffusion models. Our method is 25x more sample and compute efficient than reinforcement learning methods…
[CVPR 2024] Code for the paper "Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model"
🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.