Skip to content
View TythonLee's full-sized avatar

Block or report TythonLee

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents

JavaScript 2,178 177 Updated Aug 18, 2026

转换网易云音乐 ncm 到 mp3 / flac. Convert Netease Cloud Music ncm files to mp3/flac files.

C++ 4,346 509 Updated Oct 5, 2025

🎵 The Ultimate Open Source Suno Alternative - Professional UI for ACE-Step 1.5 AI Music Generation. Free, local, unlimited. Stop paying for Suno!

JavaScript 4,763 737 Updated Jun 27, 2026

FL AceStep Training - LoRA training nodes for ACE-Step 1.5 music generation in ComfyUI

Python 143 15 Updated Apr 25, 2026

An node for ComfyUI that implements AceStep 1.5 SFT (Supervised Fine-Tuning), a high-quality music generation model. This node replicates the full functionality of the official Gradio pipeline, off…

Python 55 12 Updated May 11, 2026

[ICLR 2026] Taming large-scale few-step training with self-adversarial flows! 👏🏻

Python 537 27 Updated Feb 24, 2026

[Tutorial] Few-Step Distillation for Text-to-Image Generation: A Practical Guide

Python 370 24 Updated Dec 31, 2025

ACE-Step: A Step Towards Music Generation Foundation Model

Python 4,777 613 Updated Feb 15, 2026

Fun-CosyVoice3-0.5B-2512 语音合成服务的简化部署方案,以及快速测试和部署提供应用调用

Python 101 19 Updated Dec 24, 2025

Chinese voice corpus. 中文语音语料,语音更加清晰自然,包含8个开源数据集,3200个说话人,900小时语音,1300万字。

752 125 Updated Jun 12, 2020

Python runtime for WeTextProcessing (does not depend on Pynini)

Python 54 10 Updated Jul 29, 2026

FlashCosyVoice: A lightweight vLLM implementation built from scratch for CosyVoice.

Python 250 25 Updated Feb 25, 2026

Text-audio foundation model from Boson AI

Python 8,319 642 Updated Jun 5, 2026

Wan: Open and Advanced Large-Scale Video Generative Models

Python 17,190 2,173 Updated Mar 17, 2026

Code for ICML 2025 Paper "Highly Compressed Tokenizer Can Generate Without Training"

Jupyter Notebook 206 12 Updated Jun 10, 2025

A Massive Contextual Speech Recognition Benchmark.

Python 109 3 Updated Aug 6, 2025

[NeurIPS 2025] PyTorch implementation of [ThinkSound], a unified framework for generating audio from any modality, guided by Chain-of-Thought (CoT) reasoning.

Python 1,377 82 Updated Apr 3, 2026

[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Python 2,259 265 Updated Feb 23, 2026

Generative models for conditional audio generation

Python 3,842 478 Updated Aug 9, 2026

A family of state-of-the-art Transformer-based audio codecs for low-bitrate high-quality audio coding.

Python 442 30 Updated Jul 17, 2026

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Python 23,153 2,794 Updated Aug 18, 2026

YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open

Python 6,388 758 Updated Jun 4, 2025

Ming - facilitating advanced multimodal understanding and generation capabilities built upon the Ling LLM.

Jupyter Notebook 667 59 Updated Jul 27, 2026

Codec for paper: LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis

Python 363 56 Updated Jun 25, 2026

Spark-TTS Inference Code

Python 11,003 1,158 Updated Apr 9, 2025

使用vllm加速cosyvoice2的推理

Jupyter Notebook 498 66 Updated Apr 26, 2025

Codebase for 'Scaling Rich Style-Prompted Text-to-Speech Datasets'

Python 167 11 Updated Mar 26, 2026

A collection of datasets for the purpose of emotion recognition/detection in speech.

HTML 421 48 Updated Sep 30, 2024
Next