Stars
An open source library designed to provide community examples of Joint Embedding Predictive Architectures (JEPAs). It contains code and examples for learning representations from images, video, and…
A toolkit for speaker diarization.
A Conversational Speech Generation Model
Multi-speaker diarization from video using SyncNet’s cross-modal embedding space to match multiple face tracks to corresponding audio tracks.
The visual communication layer between humans and AI agents. Capture, annotate, render diagrams, and organize with AI — powered by Electron and Ollama. macOS & Linux.
Official Pytorch implementation of "Large Language Models are Strong Audio-Visual Speech Recognition Learners" [ICASSP 2025] and "Mitigating Attention Sinks and Massive Activations in Audio-Visual …
Baseline system for CNVSRC2023 (Chinese Continuous Visual Speech Recognition Challenge 2023)
Foundational Models for State-of-the-Art Speech and Text Translation
Faster Whisper transcription with CTranslate2
ICASSP 2023-2024 Papers: A complete collection of influential and exciting research papers from the ICASSP 2023-24 conferences. Explore the latest advancements in acoustics, speech and signal proce…
CVPR 2023-2024 Papers: Dive into advanced research presented at the leading computer vision conference. Keep up to date with the latest developments in computer vision and deep learning. Code inclu…
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
GeneFace: Generalized and High-Fidelity 3D Talking Face Synthesis; ICLR 2023; Official code
AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
PyTorch implementation of "Distinguishing Homophenes using Multi-Head Visual-Audio Memory" (AAAI2022)
INTERSPEECH 2023-2024 Papers: A complete collection of influential and exciting research papers from the INTERSPEECH 2023-24 conference. Explore the latest advances in speech and language processin…
MultiMAE: Multi-modal Multi-task Masked Autoencoders, ECCV 2022
A High-Performance Pytorch Implementation of face detection models, including RetinaFace and DSFD
MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation
Code and Pretrained Models for ICLR 2023 Paper "Contrastive Audio-Visual Masked Autoencoder".
🤖 AgentVerse 🪐 is designed to facilitate the deployment of multiple LLM-based agents in various applications, which primarily provides two frameworks: task-solving and simulation
Research repository for LipLearner: Customizable Silent Speech Interactions on Mobile Devices (CHI 2023).
Audio-Visual Corruption Modeling of our paper "Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability Scoring" in CVPR23
ImageBind One Embedding Space to Bind Them All
Official implementation of RAVEn (ICLR 2023) and BRAVEn (ICASSP 2024)
Zero-1-to-3: Zero-shot One Image to 3D Object (ICCV 2023)
Supplementary materials for paper MegaPortraits [ACMM22]