Stars
A survey of spoken dialogue models (SDMs) with speech input and speech output. Focus on their Intermediate Representation and Generation Pattern
State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
EVAR ~ Evaluation package for Audio Representations
🔊 Repository for our NAACL-HLT 2019 paper: AudioCaps
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
Code for the paper "Do Audio Language Models Understand Linguistic Variations"?
The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how t…
This repository aims to collect Transformer-based sound event detection (SED) algorithms.
A benchmark for evaluating audio encoders on various audio tasks.
PyTorch code and models for VJEPA2 self-supervised learning from video.
🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.
Voice Activity Detector (VAD) : low-latency, high-performance and lightweight
This is a list of speech tasks and datasets, which can provide training data for Generative AI, AIGC, AI model training, intelligent speech tool development, and speech applications.
Interactively inspect module inputs, outputs, parameters, and gradients.
Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.
Awesome Speech Dataset, including download links and a brief explanation for each resource. These datasets provide diverse and high-quality speech data covering various domains such as conversation…
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
Code for DeSTA2.5-Audio, general-purpose LALM
Ke-Omni-R is an advanced audio reasoning model and achieved SOTA on MMAU
Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models
Codebase for 'Scaling Rich Style-Prompted Text-to-Speech Datasets'