-
Zhejiang University, Harbin Institute of Technology
- Shanghai
- https://orcid.org/0009-0005-3732-3035
Stars
Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
[ECCV 2026 Oral] DreamID-V: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer
[CVPR 2026] Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
Official inference repo for FLUX.2 models
Qwen-Image-Lightning: Speed up Qwen-Image model with distillation
Qwen-Image text to image lora trainer
Face-MakeUp (SD1.5): Multimodal Facial Prompts for Text-to-Image Generation (ECAI-2025)
[CVPR 2025 Highlight🔥] Identity-Preserving Text-to-Video Generation by Frequency Decomposition
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
[ICLR 2026] Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning
deepbeepmeep / Wan2GP
Forked from Wan-Video/Wan2.1A fast AI Video Generator for the GPU Poor. Supports Wan 2.1/2.2, LTX-2, Qwen Image, Hunyuan Video, LTX Video and Flux.
Unlimited-length talking video generation that supports image-to-video and video-to-video generation
Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset
HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation
[ArXiv 2025] A survey about controllable video generation: This repo is the official awesome of "Controllable video generation: A survey"
The minimal opencv for Android, iOS, ARM Linux, Windows, Linux, MacOS, HarmonyOS, WebAssembly, watchOS, tvOS, visionOS
Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
[CVPR2026 🎉] Stand-In is a lightweight, plug-and-play framework for identity-preserving video generation.
Official inference repo for FLUX.1 models
Wan: Open and Advanced Large-Scale Video Generative Models
Enjoy the magic of Diffusion models!
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
[NeurIPS 2025 D&B🔥] OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment