Stars
AudioLDM: Generate speech, sound effects, music and beyond, with text.
An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/
Automatically create prompts and make them fight each other to know which is the best
Code for robust monocular depth estimation described in "Ranftl et. al., Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer, TPAMI 2022"
High accuracy RAG for answering questions from scientific documents with citations
Real-time end-to-end singing voice conversion system based on DDSP (Differentiable Digital Signal Processing)
Singing Voice Conversion via diffusion model
PyTorch code for BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
swayducky / auto
Forked from Significant-Gravitas/AutoGPTAn experimental open-source attempt to make GPT-4 fully autonomous.
An unofficial PyTorch implementation of the audio LM VALL-E