Skip to content
View ndkgit339's full-sized avatar

Block or report ndkgit339

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Converts text to speech in realtime

Python 4,010 401 Updated Aug 2, 2026

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

Python 60,881 6,613 Updated Jul 22, 2026

gentle forced aligner

Python 1,706 307 Updated Jul 24, 2026

日本語音声に対して音素ラベルをアラインメントするためのツールです

C++ 40 5 Updated Aug 19, 2025

FACodec: Speech Codec with Attribute Factorization used for NaturalSpeech 3

Python 255 24 Updated Apr 20, 2024

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

Python 23,579 2,380 Updated Jul 13, 2026

StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.

Python 1,282 104 Updated Jun 29, 2025

An Open Source text-to-speech system built by inverting Whisper.

Jupyter Notebook 4,629 274 Updated Dec 14, 2025

Whisper realtime streaming for long speech-to-text transcription and translation

Python 3,662 409 Updated Nov 12, 2025

Real Time Speech Enhancement in the Waveform Domain (Interspeech 2020)We provide a PyTorch implementation of the paper Real Time Speech Enhancement in the Waveform Domain. In which, we present a ca…

Python 1,900 319 Updated Mar 14, 2023

PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html

Python 2,214 329 Updated Sep 10, 2025

An unofficial PyTorch implementation of the audio LM VALL-E

Python 2,978 398 Updated May 10, 2023

In defence of metric learning for speaker recognition

Python 1,172 290 Updated Apr 22, 2026

An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/

Python 7,934 782 Updated Feb 11, 2024

An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io

Python 69 7 Updated Sep 21, 2023

Analysis of gutenberg dataset

Jupyter Notebook 44 12 Updated Dec 22, 2018

Pipeline to generate the Standardized Project Gutenberg Corpus

Python 221 56 Updated Jun 22, 2026

Provides training, inference and voice conversion recipes for RADTTS and RADTTS++: Flow-based TTS models with Robust Alignment Learning, Diverse Synthesis, and Generative Modeling and Fine-Grained …

Roff 291 46 Updated Apr 6, 2023

Tacotron 2 - PyTorch implementation with faster-than-realtime inference

Jupyter Notebook 5,296 1,410 Updated Jun 12, 2024

An implementation of Microsoft's "FastSpeech 2: Fast and High-Quality End-to-End Text to Speech"

Python 2,185 614 Updated Oct 27, 2023

An implementation of Microsoft's "FastSpeech 2: Fast and High-Quality End-to-End Text to Speech"

Python 29 6 Updated Mar 28, 2024

VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Python 7,889 1,382 Updated Dec 6, 2023

A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.

Python 2,310 220 Updated Apr 13, 2026

This repository contains a set of codes to run (i.e., train, perform inference with, evaluate) a diarization method called EEND-vector-clustering.

Python 81 18 Updated Oct 18, 2022

An open source dataset for source separation

Python 501 83 Updated Feb 9, 2024

Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Python 1,150 133 Updated Aug 7, 2024