Stars
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
The official repo of BridgeVoC, which explores using the Schrödinger Bridge framework for neural vocoding.
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Andr…
a fully open-source implementation of a GPT-4o-like speech-to-speech video understanding model.
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
[ICLR2026] FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
[AAAI 2026 oral] KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
SlamKit is an open source tool kit for efficient training of SpeechLMs. It was used for "Slamming: Training a Speech Language Model on One GPU in a Day"
Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
LLM-based ASR recipe with Zipformer encoder and Qwen LLM
Multi-speaker separation, identification, diarization ALL-IN-ONE. It can isolate the target speaker from a conversation audio and do ASR.
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Wo…
Persian Kaldi profile for Rhasspy built from open speech data
Unofficial WIP LoRa Finetuning repository for VibeVoice
[EMNLP 2025 Findings] Official code for EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
A toolkit for speaker diarization.
IndexTTS Fine-tuning notebooks
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Zero-shot voice cloning text-to-speech (TTS) with explicit emotion class conditioning built on F5-TTS