Skip to content
View mpc001's full-sized avatar

Block or report mpc001

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

An open source library designed to provide community examples of Joint Embedding Predictive Architectures (JEPAs). It contains code and examples for learning representations from images, video, and…

Python 755 99 Updated Jul 17, 2026

A toolkit for speaker diarization.

Jupyter Notebook 524 62 Updated Aug 4, 2026

A Conversational Speech Generation Model

Python 14,719 1,478 Updated May 27, 2025

Multi-speaker diarization from video using SyncNet’s cross-modal embedding space to match multiple face tracks to corresponding audio tracks.

Python 3 Updated Oct 20, 2025

The visual communication layer between humans and AI agents. Capture, annotate, render diagrams, and organize with AI — powered by Electron and Ollama. macOS & Linux.

JavaScript 280 24 Updated May 7, 2026
Python 2 Updated Apr 13, 2026

Official Pytorch implementation of "Large Language Models are Strong Audio-Visual Speech Recognition Learners" [ICASSP 2025] and "Mitigating Attention Sinks and Massive Activations in Audio-Visual …

Python 64 8 Updated Jan 18, 2026

Baseline system for CNVSRC2023 (Chinese Continuous Visual Speech Recognition Challenge 2023)

Python 23 5 Updated Apr 27, 2024
Swift 4 1 Updated Feb 18, 2023

Foundational Models for State-of-the-Art Speech and Text Translation

Jupyter Notebook 11,838 1,175 Updated Jul 28, 2026

Faster Whisper transcription with CTranslate2

Python 24,905 2,021 Updated Nov 19, 2025

ICASSP 2023-2024 Papers: A complete collection of influential and exciting research papers from the ICASSP 2023-24 conferences. Explore the latest advancements in acoustics, speech and signal proce…

Python 525 23 Updated May 5, 2025

CVPR 2023-2024 Papers: Dive into advanced research presented at the leading computer vision conference. Keep up to date with the latest developments in computer vision and deep learning. Code inclu…

Python 453 28 Updated Jul 15, 2024

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

Python 19,835 1,988 Updated Aug 14, 2026

GeneFace: Generalized and High-Fidelity 3D Talking Face Synthesis; ICLR 2023; Official code

Python 2,657 291 Updated Oct 18, 2024

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Python 10,172 850 Updated Jul 6, 2024

PyTorch implementation of "Distinguishing Homophenes using Multi-Head Visual-Audio Memory" (AAAI2022)

Python 27 5 Updated Mar 9, 2024

INTERSPEECH 2023-2024 Papers: A complete collection of influential and exciting research papers from the INTERSPEECH 2023-24 conference. Explore the latest advances in speech and language processin…

685 43 Updated Dec 25, 2024

MultiMAE: Multi-modal Multi-task Masked Autoencoders, ECCV 2022

Python 635 75 Updated Dec 13, 2022

A High-Performance Pytorch Implementation of face detection models, including RetinaFace and DSFD

Python 231 61 Updated Jun 18, 2025

MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

Python 404 37 Updated Sep 11, 2023

Code and Pretrained Models for ICLR 2023 Paper "Contrastive Audio-Visual Masked Autoencoder".

Python 292 24 Updated Mar 20, 2024

🤖 AgentVerse 🪐 is designed to facilitate the deployment of multiple LLM-based agents in various applications, which primarily provides two frameworks: task-solving and simulation

JavaScript 5,105 511 Updated Sep 9, 2024

Research repository for LipLearner: Customizable Silent Speech Interactions on Mobile Devices (CHI 2023).

Swift 63 7 Updated Nov 8, 2023

Audio-Visual Corruption Modeling of our paper "Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability Scoring" in CVPR23

Python 35 2 Updated Jun 20, 2023

ImageBind One Embedding Space to Bind Them All

Python 9,065 845 Updated Nov 21, 2025

Official implementation of RAVEn (ICLR 2023) and BRAVEn (ICASSP 2024)

Python 82 9 Updated Feb 27, 2025

Zero-1-to-3: Zero-shot One Image to 3D Object (ICCV 2023)

Python 3,056 221 Updated Dec 5, 2023

Supplementary materials for paper MegaPortraits [ACMM22]

265 18 Updated Oct 16, 2023
Next