Stars
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
A Curated List of Awesome Video World Models with AR Diffusion: Covering Algorithms, Applications, and Infrastructure, Aimed at Serving as a Comprehensive Resource for Researchers, Practitioners, a…
The official implementation of “MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction”
[WIP] Your personal AI-powered arXiv digest — fetch, score, summarize, and browse daily papers effortlessly.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
A local AI assistant running on your device. It turns your files into actionable memory.
[NeurIPS 2025] Controllable Human-centric Keyframe Interpolation with Generative Prior
[CVPR 2026 Highlight] MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator
[ICCV 2025] FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing
Lets make video diffusion practical!
TradingAgents: Multi-Agents LLM Financial Trading Framework
[ICML 2025] Official PyTorch Implementation of "History-Guided Video Diffusion"
code for "Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion"
Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation [Siggraph Asian 2025]
[CVPR'26] ObjectClear: Precise Object and Effect Removal with Adaptive Target-Aware Attention
MAGI-1: Autoregressive Video Generation at Scale
Official Implementation of paper "MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion"
[CVPR 2025] MatAnyone: Stable Video Matting with Consistent Memory Propagation
The official implementation of CVPR'25 Oral paper "Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise"
[NeurIPS 2024] Neural Localizer Fields for Continuous 3D Human Pose and Shape Estimation
[NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling
Stable Virtual Camera: Generative View Synthesis with Diffusion Models
[ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
A unified inference and post-training framework for accelerated video generation.