Highlights
- Pro
Stars
An agentic skills framework & software development methodology that works.
The official implementation of 'LAMB: LLM-based Audio Captioning with Modality Gap Bridging via Cauchy-Schwarz Divergence' (ICASSP2026)
LightlyStudio - The Unified Data Platform for Multimodal ML
[NeurIPS 2025] AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
Implementation of a single layer of the MMDiT, proposed in Stable Diffusion 3, in Pytorch
Interspeech_ablation_study (Boundary-Conscious Pruning: Hard Set-Aware Model Compression for Efficient Speaker Recognition)
Techniques and tools for optimizing how AI coding assistants understand your codebase, with a focus on cost reduction and efficiency.
[INTERSPEECH 2025] Official code for "SEED: Speaker Embedding Enhancement Diffusion Model"
[INTERSPEECH 2024] Official pytorch code for the paper "Disentangled Representation Learning for Environment-agnostic Speaker Recognition"
Official Pytorch Implementation of 'LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport' (ICASSP2025)
Official Python SDK for the Agent2Agent (A2A) Protocol
Samples using the Agent2Agent (A2A) Protocol
A TTS model capable of generating ultra-realistic dialogue in one pass.
High accuracy RAG for answering questions from scientific documents with citations
Agent Laboratory is an end-to-end autonomous research workflow meant to assist you as the human researcher toward implementing your research ideas
Code and Pretrained Models for ICLR 2023 Paper "Contrastive Audio-Visual Masked Autoencoder".
Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
CAIRI Supervised, Semi- and Self-Supervised Visual Representation Learning Toolbox and Benchmark
ZiyuGuo99 / Awesome-MIM-1
Forked from Lupin1998/Awesome-MIMAwesome List of Masked Image Modeling (MIM) Papers for Self-supervised Visual Representation Learning
[INTERSPEECH 2024] Official pytorch code for the paper "Disentangled Representation Learning for Environment-agnostic Speaker Recognition"
PyTorch implementation of MAE https//arxiv.org/abs/2111.06377
The official repo of Qwen-Audio (通义千问-Audio) chat & pretrained large audio language model proposed by Alibaba Cloud.
ICASSP 2023: 'Speaker recognition with two-step multi-modal deep cleansing'
INTERSPEECH2023: Target Active Speaker Detection with Audio-visual Cues