Skip to content
View yaguanghu's full-sized avatar

Block or report yaguanghu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.

Python 1,414 187 Updated May 21, 2026

Inworld TTS

Python 744 80 Updated Aug 10, 2026

A native-PyTorch library for large scale M-LLM (text/audio) training with tp/cp/dp.

Python 233 30 Updated Jul 2, 2026

Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Python 1,037 142 Updated Dec 2, 2025

The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.

Python 5,667 574 Updated Aug 15, 2026

Super-Efficient RLHF Training of LLMs with Parameter Reallocation

Python 336 22 Updated Apr 24, 2025

Notes and commented code for RLHF (PPO)

Python 136 38 Updated Feb 27, 2024

[NAACL 2025] WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching

Python 133 12 Updated Apr 8, 2026

Democratizing Reinforcement Learning for LLMs

Python 5,784 606 Updated Aug 15, 2026

Reproduce R1 Zero on Logic Puzzle

Python 2,449 163 Updated Mar 20, 2025

A simple screen parsing tool towards pure vision based GUI agent

Jupyter Notebook 25,258 2,225 Updated Jul 20, 2026

Making large AI models cheaper, faster and more accessible

Python 41,437 4,502 Updated Aug 10, 2026

本项目用于大模型数学解题能力方面的数据集合成,模型训练及评测,相关文章记录。

Python 104 9 Updated Sep 14, 2024

Reverse Engineering of Supervised Semantic Speech Tokenizer (S3Tokenizer) proposed in CosyVoice

Python 527 69 Updated Dec 22, 2025

Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷

Python 6,887 404 Updated Aug 13, 2026

A Framework for Speech, Language, Audio, Music Processing with Large Language Model

Python 1,054 117 Updated Jan 15, 2026

✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Python 2,527 182 Updated Mar 28, 2025

An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.

Python 4,409 361 Updated Aug 14, 2025

This project provides a way to create a Docker image based on an official Ubuntu Image with an SSH server (SSHD) enabled

Shell 53 30 Updated Sep 19, 2024

Implementation of E2-TTS, "Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS", in Pytorch

Python 517 52 Updated Dec 20, 2025

Approaching (Almost) Any Machine Learning Problem

8,370 1,122 Updated Mar 25, 2023

Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI

Python 1,390 74 Updated Aug 14, 2026

OpenTAD is an open-source temporal action detection (TAD) toolbox based on PyTorch.

Python 344 25 Updated Jul 14, 2026

This repo implements VQVAE on mnist and as well as colored version of mnist images. It also implements simple LSTM for generating sample numbers using the encoder outputs of trained VQVAE

Python 62 10 Updated Feb 6, 2024

Implementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in Pytorch

Python 2,627 278 Updated Jan 12, 2025

Character Animation (AnimateAnyone, Face Reenactment)

Python 3,512 295 Updated May 31, 2024

小红书笔记 | 评论爬虫、抖音视频 | 评论爬虫、快手视频 | 评论爬虫、B 站视频 | 评论爬虫、微博帖子 | 评论爬虫、百度贴吧帖子 | 百度贴吧评论回复爬虫 | 知乎问答文章|评论爬虫

Python 62,477 12,174 Updated Aug 14, 2026

Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.

Python 72,007 6,492 Updated Aug 15, 2026
Jupyter Notebook 93 11 Updated Sep 28, 2024
Next