Stars
MOVA: Towards Scalable and Synchronized Video–Audio Generation
A unified and fully open-source framework for instruction-guided and reference-guided video editing using natural language.
"CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/
The ultimate training toolkit for finetuning diffusion models
[CVPR'26 Demo] Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
MAGI-1: Autoregressive Video Generation at Scale
[CVPR 2024] VidToMe: Video Token Merging for Zero-Shot Video Editing
Di♪♪Rhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
[ICML 2025] SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation
[CVPR2025 Highlight] Video Generation Foundation Models: https://saiyan-world.github.io/goku/
YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
[NeurIPS 2024] SHMT: Self-supervised Hierarchical Makeup Transfer via Latent Diffusion Models
[CVPR 2025] Official implementation of "AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models"
Official pytorch implementation of paper "Portrait Eyeglasses and Shadow Removal by Leveraging 3D Synthetic Data" (CVPR 2022).
Official repository of In-Context LoRA for Diffusion Transformers
victorchall / genmoai-smol
Forked from genmoai/mochiThe best OSS video generation models
[NeurIPS 2024] Depth Anything V2. A More Capable Foundation Model for Monocular Depth Estimation
[Patterns (Cell subsidiary journal)] The official code for "UltraLight VM-UNet: Parallel Vision Mamba Significantly Reduces Parameters for Skin Lesion Segmentation".
reproduction of AnimateAnyone
SUPIR aims at developing Practical Algorithms for Photo-Realistic Image Restoration In the Wild. Our new online demo is also released at suppixel.ai.
APISR: Anime Production Inspired Real-World Anime Super-Resolution (CVPR 2024)
Emote Portrait Alive: Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions
[SIGGRAPH ASIA 2024 TCS] AnimateLCM: Computation-Efficient Personalized Style Video Generation without Personalized Video Data
Unofficial Implementation of Animate Anyone