Skip to content
View pkuanjie's full-sized avatar
  • University of Rochester
  • Rochester, NY, US

Highlights

  • Pro

Block or report pkuanjie

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[CVPR 2025] VideoWorld is a simple generative model that learns purely from unlabeled videos—much like how babies learn by observing their environment.

Python 794 40 Updated Feb 25, 2026

Code release for "PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop" (ICML 2025)

Jupyter Notebook 59 3 Updated May 8, 2025

Code release for https://kovenyu.com/WonderWorld/

Python 741 36 Updated Apr 14, 2025

Simulation platform for general-purpose robotics & embodied AI learning.

Python 29,758 2,837 Updated Aug 17, 2026

Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).

Python 13,451 1,313 Updated Jun 26, 2026

[CVPR 2025 Highlight] Align3R: Aligned Monocular Depth Estimation for Dynamic Videos

Python 461 30 Updated Apr 4, 2025
Python 1,301 96 Updated Aug 2, 2025

[NeurIPS'2024]: DiffGS: Functional Gaussian Splatting Diffusion

C++ 289 24 Updated Apr 10, 2025

[arXiv 2023] DreamGaussian4D: Generative 4D Gaussian Splatting

Python 620 37 Updated Jun 10, 2024

[CVPR 2024] 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering

Jupyter Notebook 3,880 385 Updated Oct 27, 2024

PyTorch implementation of MAR+DiffLoss https://arxiv.org/abs/2406.11838

Python 1,947 125 Updated Feb 20, 2026

[ICLR'25 Oral] No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images

Python 979 54 Updated Feb 25, 2026

[SIGGRAPH 2024] Motion I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling

Python 192 16 Updated Sep 27, 2024

Official inference repo for FLUX.1 models

Python 25,892 1,912 Updated Jul 31, 2025
Python 30 5 Updated Oct 24, 2023

This repo contains the code for 1D tokenizer and generator

Jupyter Notebook 1,173 69 Updated Mar 20, 2025

Official PyTorch Implementation of "SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers"

Python 1,203 80 Updated Dec 22, 2025

This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.

Python 12,153 1,061 Updated Mar 8, 2026

Open-Sora: Democratizing Efficient Video Production for All

Python 29,278 3,005 Updated Apr 9, 2026

ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation [TMLR 2024]

Python 260 14 Updated Jul 1, 2024

[WIP] Layer Diffusion for WebUI (via Forge)

Python 4,116 352 Updated Aug 30, 2024

PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Python 1,937 99 Updated Oct 31, 2024

Aligning LMMs with Factually Augmented RLHF

Python 398 31 Updated Nov 1, 2023

PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Python 3,303 202 Updated Oct 31, 2024

LVDM: Latent Video Diffusion Models for High-Fidelity Long Video Generation

Python 503 23 Updated Nov 16, 2024

General technology for enabling AI capabilities w/ LLMs and MLLMs

Python 4,464 376 Updated Jul 25, 2026

AlignProp uses direct reward backpropogation for the alignment of large-scale text-to-image diffusion models. Our method is 25x more sample and compute efficient than reinforcement learning methods…

Python 326 11 Updated Nov 1, 2024

[CVPR 2024] Code for the paper "Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model"

Python 244 17 Updated Apr 6, 2024
Python 648 35 Updated Feb 15, 2024

🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.

3,266 150 Updated Aug 12, 2026
Next