Stars
ViPE: Video Pose Engine for Geometric 3D Perception
[NeurIPS 2025 D&B🔥] OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment
Combining Teacache with xDiT to Accelerate Visual Generation Models
This repository is an implementation that recreates the SketchGuidance feature of "ToonCrafter".
中文nlp解决方案(大模型、数据、模型、训练、推理)
Unofficial implementation of the paper "The Chosen One: Consistent Characters in Text-to-Image Diffusion Models"
[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)