Starred repositories
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
转换网易云音乐 ncm 到 mp3 / flac. Convert Netease Cloud Music ncm files to mp3/flac files.
🎵 The Ultimate Open Source Suno Alternative - Professional UI for ACE-Step 1.5 AI Music Generation. Free, local, unlimited. Stop paying for Suno!
FL AceStep Training - LoRA training nodes for ACE-Step 1.5 music generation in ComfyUI
An node for ComfyUI that implements AceStep 1.5 SFT (Supervised Fine-Tuning), a high-quality music generation model. This node replicates the full functionality of the official Gradio pipeline, off…
[ICLR 2026] Taming large-scale few-step training with self-adversarial flows! 👏🏻
[Tutorial] Few-Step Distillation for Text-to-Image Generation: A Practical Guide
ACE-Step: A Step Towards Music Generation Foundation Model
Fun-CosyVoice3-0.5B-2512 语音合成服务的简化部署方案,以及快速测试和部署提供应用调用
Chinese voice corpus. 中文语音语料,语音更加清晰自然,包含8个开源数据集,3200个说话人,900小时语音,1300万字。
Python runtime for WeTextProcessing (does not depend on Pynini)
FlashCosyVoice: A lightweight vLLM implementation built from scratch for CosyVoice.
Text-audio foundation model from Boson AI
Wan: Open and Advanced Large-Scale Video Generative Models
Code for ICML 2025 Paper "Highly Compressed Tokenizer Can Generate Without Training"
A Massive Contextual Speech Recognition Benchmark.
[NeurIPS 2025] PyTorch implementation of [ThinkSound], a unified framework for generating audio from any modality, guided by Chain-of-Thought (CoT) reasoning.
[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Generative models for conditional audio generation
A family of state-of-the-art Transformer-based audio codecs for low-bitrate high-quality audio coding.
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
Ming - facilitating advanced multimodal understanding and generation capabilities built upon the Ling LLM.
Codec for paper: LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis
Codebase for 'Scaling Rich Style-Prompted Text-to-Speech Datasets'
A collection of datasets for the purpose of emotion recognition/detection in speech.