Stars
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
A native-PyTorch library for large scale M-LLM (text/audio) training with tp/cp/dp.
Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
Super-Efficient RLHF Training of LLMs with Parameter Reallocation
[NAACL 2025] WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching
A simple screen parsing tool towards pure vision based GUI agent
Making large AI models cheaper, faster and more accessible
Reverse Engineering of Supervised Semantic Speech Tokenizer (S3Tokenizer) proposed in CosyVoice
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷
A Framework for Speech, Language, Audio, Music Processing with Large Language Model
✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.
This project provides a way to create a Docker image based on an official Ubuntu Image with an SSH server (SSHD) enabled
Implementation of E2-TTS, "Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS", in Pytorch
Approaching (Almost) Any Machine Learning Problem
Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI
OpenTAD is an open-source temporal action detection (TAD) toolbox based on PyTorch.
This repo implements VQVAE on mnist and as well as colored version of mnist images. It also implements simple LSTM for generating sample numbers using the encoder outputs of trained VQVAE
Implementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in Pytorch
Character Animation (AnimateAnyone, Face Reenactment)
小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 百度贴吧评论回复爬虫 | 知乎问答文章|评论爬虫
Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.