Lists (1)
Sort Name ascending (A-Z)
Starred repositories
[CVPR Oral 2022] PyTorch Implementation for "Learning to Deblur using Light Field Generated and Real Defocused Images"
TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction
Edit Banana: A framework for converting statistical formats into editable.
official github code for "SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing"
[CVPR2026] Exploring Spatial Intelligence from a Generative Perspective
Official Implementation of "AssemLM: Spatial Reasoning Multimodal Large Language Models for Robotic Assembly"
[ECCV2026] Anchor Forcing is a cache-centric framework for interactive streaming video generation that preserves visual quality and coherent motion across prompt switches
[ECCV 2026 Oral] DreamID-V: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer
[ACL'26] EvoToken-DLM (Beyond Hard Masks: Progressive Token Evolution for Diffusion Language)
This repository collects and organises state‑of‑the‑art papers on spatial reasoning for Multimodal Vision–Language Models (MVLMs).
[ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process
Enjoy the magic of Diffusion models!
[CVPR 2025] A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
[NeurIPS 2025 Spotlight] A Generalist Diffusion Model for Vision Perception
Frechet Video Distance metric implemented on PyTorch
Microsoft PowerToys is a collection of utilities that supercharge productivity and customization on Windows
One-shot and Few-shot 3D Editing without Per-Scene Optimization
A curated list of awesome papers for reconstructing 4D spatial intelligence from video. (arXiv 2507.21045)
Official implementation of the paper "Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content".
[ICLR2026] Any-to-Bokeh is a novel one-step video bokeh framework that converts arbitrary input videos into temporally coherent, depth-aware bokeh effects.
Official repo for paper "Structured 3D Latents for Scalable and Versatile 3D Generation" (CVPR'25 Spotlight).
[Arxiv] Discrete Diffusion in Large Language and Multimodal Models: A Survey
[3DV 2026] Revisiting Depth Representations for Feed-Forward 3D Gaussian Splatting
[ICML2026] ACTIVE-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
[NeurIPS 2025] Official Repo of Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
Python client for Baidu Yun (Personal Cloud Storage) 百度云/百度网盘Python客户端