Skip to content
View BongkiLee's full-sized avatar

Block or report BongkiLee

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A survey of spoken dialogue models (SDMs) with speech input and speech output. Focus on their Intermediate Representation and Generation Pattern

32 Updated Mar 24, 2026

State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!

Jupyter Notebook 2,346 160 Updated Apr 13, 2026
Python 41 9 Updated Feb 18, 2026

OpenFLAM: Framewise Language Audio Model

Python 111 7 Updated Jun 4, 2026

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

Python 319 33 Updated Jan 19, 2026

open-vocabulary sound event detection

Python 54 6 Updated Dec 17, 2025

EVAR ~ Evaluation package for Audio Representations

Jupyter Notebook 81 5 Updated Feb 19, 2026
Python 15 Updated Jan 24, 2025

🔊 Repository for our NAACL-HLT 2019 paper: AudioCaps

Python 216 24 Updated Oct 6, 2025

Masked Modeling Duo: Towards a Universal Audio Pre-training Framework

Jupyter Notebook 163 10 Updated Feb 23, 2026

Code for the paper "Do Audio Language Models Understand Linguistic Variations"?

3 Updated Jul 11, 2025

The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how t…

Python 3,595 326 Updated May 26, 2026

This repository aims to collect Transformer-based sound event detection (SED) algorithms.

Jupyter Notebook 105 8 Updated Feb 10, 2026

A benchmark for evaluating audio encoders on various audio tasks.

Python 56 9 Updated Apr 27, 2026

PyTorch code and models for VJEPA2 self-supervised learning from video.

Python 4,479 551 Updated Mar 23, 2026

🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.

3,266 150 Updated Aug 12, 2026

Voice Activity Detector (VAD) : low-latency, high-performance and lightweight

C 2,239 176 Updated Feb 2, 2026

This is a list of speech tasks and datasets, which can provide training data for Generative AI, AIGC, AI model training, intelligent speech tool development, and speech applications.

83 6 Updated Jun 7, 2024

Interactively inspect module inputs, outputs, parameters, and gradients.

Python 360 21 Updated Dec 22, 2025

Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.

Python 1,495 112 Updated Mar 16, 2026

The LLM Evaluation Framework

Python 17,641 1,818 Updated Aug 17, 2026

Awesome Speech Dataset, including download links and a brief explanation for each resource. These datasets provide diverse and high-quality speech data covering various domains such as conversation…

28 Updated Jul 4, 2025

Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper

Jupyter Notebook 5,624 502 Updated Aug 15, 2026

Vox-Profile Benchmark

Python 90 13 Updated Feb 16, 2026

Code for DeSTA2.5-Audio, general-purpose LALM

Python 141 7 Updated Feb 4, 2026

Ke-Omni-R is an advanced audio reasoning model and achieved SOTA on MMAU

Python 61 1 Updated Jun 11, 2025

Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation

Python 4,720 375 Updated Jun 21, 2025

PyTorch implementation of Audio Flamingo: Series of Advanced Audio Understanding Language Models

1,177 102 Updated Dec 15, 2025

Codebase for 'Scaling Rich Style-Prompted Text-to-Speech Datasets'

Python 167 11 Updated Mar 26, 2026
Next