Stars
FireRed-Image-Edit is a powerful image editing foundation model achieving open-source state-of-the-art performance with precise instruction following, high-fidelity generation, superior identity co…
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenario…
This is the official implementation of our paper: "MiniMax-Remover: Taming Bad Noise Helps Video Object Removal"
ComfyUI wrapper for segment anything 3
A ComfyUI custom node designed for advanced image background removal and object, face, clothes, and fashion segmentation, utilizing multiple models including RMBG-2.0, INSPYRENET, BEN, BEN2, BiRefN…
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.
基于AI的图片/视频硬字幕去除、文本水印去除,无损分辨率生成去字幕、去水印后的图片/视频文件。无需申请第三方API,本地实现。AI-based tool for removing hard-coded subtitles and text-like watermarks from videos or Pictures.
[CVPR 2026 Highlight] MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator
[CVPR 2025] MatAnyone: Stable Video Matting with Consistent Memory Propagation
A voice conversion extension node for ComfyUI based on FreeVC, enabling high-quality voice conversion capabilities within the ComfyUI framework.
FreeVC: Towards High-Quality Text-Free One-Shot Voice Conversion
Easily train a good VC model with voice data <= 10 mins!
[TMLR] Memory-Guided Diffusion for Expressive Talking Video Generation
Official implementation of "Sonic: Shifting Focus to Global Audio Perception in Portrait Animation"
real time face swap and one-click video deepfake with only a single image
Hackable and optimized Transformers building blocks, supporting a composable construction.
CUDA integration for Python, plus shiny features
[CVPR 2026] PersonaLive! : Expressive Portrait Image Animation for Live Streaming
SoftVC VITS Singing Voice Conversion
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
Clone a voice in 5 seconds to generate arbitrary speech in real-time
AIGCPanel 是一个简单易用的一站式AI数字人系统,支持视频合成、声音合成、声音克隆,简化本地模型管理、一键导入和使用AI模型。
Enhanced the comfyui savevideo node to support previewing and saving videos containing alpha channels.