Skip to content
View abnerwang's full-sized avatar
💭
I may be slow to respond.
💭
I may be slow to respond.

Block or report abnerwang

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents

JavaScript 2,199 178 Updated Aug 20, 2026
Python 1,243 129 Updated Aug 17, 2026

从零开始玩转OpenClaw:最全面的中文教程,涵盖安装、配置、实战案例和避坑指南(github版)

Shell 4,552 684 Updated Jul 6, 2026

Bridge Claude Code / Codex to IM platforms — chat with AI coding agents from Telegram, Discord, or Feishu/Lark.

TypeScript 2,863 317 Updated Mar 23, 2026

Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice…

Python 13,025 1,688 Updated Mar 17, 2026

Ultimate Vocal Remover Inference CLI

Python 123 12 Updated Feb 27, 2026

GUI for a Vocal Remover that uses Deep Neural Networks.

Python 25,824 1,947 Updated Mar 13, 2025

Easily train a good VC model with voice data <= 10 mins!

Python 37,676 5,228 Updated Aug 4, 2026

CNN-based audio segmentation toolkit. Allows to detect speech, music, noise and speaker gender. Has been designed for large scale gender equality studies based on speech time per gender.

Python 910 153 Updated Mar 12, 2026

Official Pytorch Implementation of "Diff-HierVC: Diffusion-based Hierarchical Voice Conversion with Robust Pitch Generation and Masked Prior for Zero-shot Speaker Adaptation"

Python 237 20 Updated Jul 3, 2024

Code for the paper Hybrid Spectrogram and Waveform Source Separation

Python 10,356 1,573 Updated Apr 24, 2024

FreeVC: Towards High-Quality Text-Free One-Shot Voice Conversion

Python 716 128 Updated Jan 19, 2025

SoulX-Podcast is an inference codebase by the Soul AI team for generating high-fidelity podcasts from text.

Python 3,524 458 Updated Dec 11, 2025

High-quality speech synthesis with LoRA fine-tuning on index-tts, enhancing prosody and naturalness for single and multi-speaker voices.

Python 334 30 Updated Mar 12, 2026

Contexts Optical Compression

Python 23,831 2,201 Updated Jan 27, 2026

Turn detection for full-duplex dialogue communication

Python 607 41 Updated Dec 26, 2025

Silero VAD: pre-trained enterprise-grade Voice Activity Detector

Python 10,017 827 Updated Aug 18, 2026

An implementation for Frame-level Speech Signal-to-Noise Ratio Estimation using deep learning

Jupyter Notebook 43 6 Updated Mar 23, 2022

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Python 23,229 2,801 Updated Aug 18, 2026

An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.

Python 4,421 363 Updated Aug 14, 2025

The official repository for ERNIE 4.5 and ERNIEKit – its industrial-grade development toolkit based on PaddlePaddle.

Python 7,736 1,445 Updated Jul 24, 2026

Vital tracker implemented using PyTorch

Python 40 9 Updated Nov 2, 2020

A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

Python 26,194 2,051 Updated Aug 12, 2026

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …

Python 15,306 1,617 Updated Aug 20, 2026

The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.

Python 2,098 169 Updated Apr 21, 2025

Chinese text normalization for speech processing

Python 738 151 Updated Mar 18, 2023

Mamba SSM architecture

Python 18,755 1,802 Updated Jul 22, 2026

JARVIS, a system to connect LLMs with ML community. Paper: https://arxiv.org/pdf/2303.17580.pdf

Python 25,187 2,199 Updated Jul 29, 2025
Next