Stars
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singi…
A Fully Self-Hosted Solution for Full-Duplex Voice Interaction
Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
A native-PyTorch library for large scale M-LLM (text/audio) training with tp/cp/dp.
a lightweight speech processing toolkit based on Pytorch and (Py)Kaldi
State-of-the-art audio codec with 90x compression factor. Supports 44.1kHz, 24kHz, and 16kHz mono/stereo audio.
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
FSA/FST algorithms, differentiable, with PyTorch compatibility.
kaldi-asr/kaldi is the official location of the Kaldi project.
Tools for handling multimodal data in machine learning projects.
C++ implementation of LSTM (Long Short Term Memory), in Kaldi's nnet1 framework. Used for automatic speech recognition, possibly language modeling etc, the training can be switched between CPU and …