Stars
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
从零开始玩转OpenClaw:最全面的中文教程,涵盖安装、配置、实战案例和避坑指南(github版)
Bridge Claude Code / Codex to IM platforms — chat with AI coding agents from Telegram, Discord, or Feishu/Lark.
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice…
GUI for a Vocal Remover that uses Deep Neural Networks.
Easily train a good VC model with voice data <= 10 mins!
CNN-based audio segmentation toolkit. Allows to detect speech, music, noise and speaker gender. Has been designed for large scale gender equality studies based on speech time per gender.
Official Pytorch Implementation of "Diff-HierVC: Diffusion-based Hierarchical Voice Conversion with Robust Pitch Generation and Masked Prior for Zero-shot Speaker Adaptation"
Code for the paper Hybrid Spectrogram and Waveform Source Separation
FreeVC: Towards High-Quality Text-Free One-Shot Voice Conversion
SoulX-Podcast is an inference codebase by the Soul AI team for generating high-fidelity podcasts from text.
High-quality speech synthesis with LoRA fine-tuning on index-tts, enhancing prosody and naturalness for single and multi-speaker voices.
Turn detection for full-duplex dialogue communication
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
An implementation for Frame-level Speech Signal-to-Noise Ratio Estimation using deep learning
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.
The official repository for ERNIE 4.5 and ERNIEKit – its industrial-grade development toolkit based on PaddlePaddle.
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.
Chinese text normalization for speech processing
JARVIS, a system to connect LLMs with ML community. Paper: https://arxiv.org/pdf/2303.17580.pdf