Stars
A SOTA Industrial-Grade Voice Activity Detection & Audio Event Detection, supporting 100+ languages, outperforming Silero-VAD, TEN-VAD, FunASR-VAD and WebRTC-VAD
Easy to use stem (e.g. instrumental/vocals) separation from CLI or as a python package, using a variety of amazing pre-trained models (primarily from UVR)
Toolbox for Evaluation of AEC/AES Systems
Production First and Production Ready End-to-End Speech Recognition Toolkit
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
A training code template for DNN-based speech enhancement.
使用Tensorflow实现声纹识别
This Repostory contains the pretrained DTLN-aec model for real-time acoustic echo cancellation.
A minimum unofficial implementation of the "A Convolutional Recurrent Neural Network for Real-Time Speech Enhancement" (CRN) using PyTorch
Noise supression using deep filtering
implementation of "DCCRN-Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement" by pytorch
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation
Automatic Speech Recognition (ASR), Speaker Verification, Speech Synthesis, Text-to-Speech (TTS), Language Modelling, Singing Voice Synthesis (SVS), Voice Conversion (VC)
Towards hot directions in industrial end to end speech recognition
An unofficial implementation of DeepVQE proposed by Microsoft Corp.
Unofficial PyTorch implementation of Google AI's VoiceFilter system
Text-audio foundation model from Boson AI