-
The University of Tokyo
- https://sites.google.com/view/ymatsunaga/home
- @YutaMResearch
- in/yuta-matsunaga-86766a223
Lists (4)
Sort Name ascending (A-Z)
Stars
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
FACodec: Speech Codec with Attribute Factorization used for NaturalSpeech 3
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
An Open Source text-to-speech system built by inverting Whisper.
Whisper realtime streaming for long speech-to-text transcription and translation
Real Time Speech Enhancement in the Waveform Domain (Interspeech 2020)We provide a PyTorch implementation of the paper Real Time Speech Enhancement in the Waveform Domain. In which, we present a ca…
PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html
An unofficial PyTorch implementation of the audio LM VALL-E
In defence of metric learning for speaker recognition
An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/
An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io
Analysis of gutenberg dataset
Pipeline to generate the Standardized Project Gutenberg Corpus
Provides training, inference and voice conversion recipes for RADTTS and RADTTS++: Flow-based TTS models with Robust Alignment Learning, Diverse Synthesis, and Generative Modeling and Fine-Grained …
Tacotron 2 - PyTorch implementation with faster-than-realtime inference
An implementation of Microsoft's "FastSpeech 2: Fast and High-Quality End-to-End Text to Speech"
Wataru-Nakata / FastSpeech2-JSUT
Forked from ming024/FastSpeech2An implementation of Microsoft's "FastSpeech 2: Fast and High-Quality End-to-End Text to Speech"
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.
This repository contains a set of codes to run (i.e., train, perform inference with, evaluate) a diarization method called EEND-vector-clustering.
Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis