🎯
Focusing. I may be slow to reply.
Focusing on multimodal synthesis (speech/audio/sing), speech translation, and self-supervised learning.
-
Facebook AI Research (FAIR)
- Menlo Park
- rongjiehuang.github.io
Pinned Loading
-
AIGC-Audio/AudioGPT
AIGC-Audio/AudioGPT PublicAudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
-
Alpha-VLLM/Lumina-T2X
Alpha-VLLM/Lumina-T2X PublicLumina-T2X is a unified framework for Text to Any Modality Generation
-
Text-to-Audio/Make-An-Audio
Text-to-Audio/Make-An-Audio PublicPyTorch Implementation of Make-An-Audio (ICML'23) with a Text-to-Audio Generative Model
-
TranSpeech
TranSpeech PublicPyTorch Implementation of TranSpeech (ICLR'23): Textless NAR Speech-to-Speech Translation with Bilateral Perturbation
-
yangdongchao/AcademiCodec
yangdongchao/AcademiCodec PublicAcademiCodec: An Open Source Audio Codec Model for Academic Research
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.