Skip to content
View DevKiHyun's full-sized avatar

Highlights

  • Pro

Block or report DevKiHyun

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

An agentic skills framework & software development methodology that works.

Shell 273,987 24,527 Updated Aug 13, 2026

Speaker Language Model

Python 10 2 Updated Aug 11, 2025

The official implementation of 'LAMB: LLM-based Audio Captioning with Modality Gap Bridging via Cauchy-Schwarz Divergence' (ICASSP2026)

5 Updated Jan 9, 2026
Python 9 2 Updated Nov 6, 2025

LightlyStudio - The Unified Data Platform for Multimodal ML

Python 877 33 Updated Aug 19, 2026

[NeurIPS 2025] AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding

Python 27 1 Updated Nov 3, 2025

Implementation of a single layer of the MMDiT, proposed in Stable Diffusion 3, in Pytorch

Python 555 17 Updated Jan 18, 2026

Interspeech_ablation_study (Boundary-Conscious Pruning: Hard Set-Aware Model Compression for Efficient Speaker Recognition)

Python 1 Updated Feb 25, 2025

Techniques and tools for optimizing how AI coding assistants understand your codebase, with a focus on cost reduction and efficiency.

JavaScript 25 Updated Jul 1, 2025

[INTERSPEECH 2025] Official code for "SEED: Speaker Embedding Enhancement Diffusion Model"

Python 61 3 Updated Nov 3, 2025

[INTERSPEECH 2024] Official pytorch code for the paper "Disentangled Representation Learning for Environment-agnostic Speaker Recognition"

2 Updated Jun 20, 2024

Official Pytorch Implementation of 'LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport' (ICASSP2025)

Python 10 1 Updated Apr 14, 2025

Official Python SDK for the Agent2Agent (A2A) Protocol

Python 2,090 478 Updated Aug 19, 2026

Samples using the Agent2Agent (A2A) Protocol

Jupyter Notebook 1,737 734 Updated Aug 17, 2026

Official code of Diffusion -bridge

Python 12 2 Updated Oct 29, 2025

A TTS model capable of generating ultra-realistic dialogue in one pass.

Python 19,369 1,693 Updated Nov 19, 2025

High accuracy RAG for answering questions from scientific documents with citations

Python 9,057 907 Updated Aug 12, 2026

Agent Laboratory is an end-to-end autonomous research workflow meant to assist you as the human researcher toward implementing your research ideas

Python 5,800 806 Updated Aug 20, 2025

Code and Pretrained Models for ICLR 2023 Paper "Contrastive Audio-Visual Masked Autoencoder".

Python 292 24 Updated Mar 20, 2024

Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Jupyter Notebook 177 10 Updated Sep 26, 2022

Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities

Python 22,192 2,703 Updated Jan 23, 2026

CAIRI Supervised, Semi- and Self-Supervised Visual Representation Learning Toolbox and Benchmark

Python 658 62 Updated Oct 15, 2025

Awesome List of Masked Image Modeling (MIM) Papers for Self-supervised Visual Representation Learning

12 3 Updated Jun 4, 2023

[INTERSPEECH 2024] Official pytorch code for the paper "Disentangled Representation Learning for Environment-agnostic Speaker Recognition"

Python 19 3 Updated Jul 23, 2024

PyTorch implementation of MAE https//arxiv.org/abs/2111.06377

Python 8,371 1,348 Updated Jul 23, 2024

The official repo of Qwen-Audio (通义千问-Audio) chat & pretrained large audio language model proposed by Alibaba Cloud.

Python 1,941 147 Updated Jul 5, 2024

ICASSP 2023: 'Speaker recognition with two-step multi-modal deep cleansing'

Python 44 6 Updated Oct 31, 2022

Code repository for FreGrad

Python 52 5 Updated May 19, 2024

INTERSPEECH2023: Target Active Speaker Detection with Audio-visual Cues

Python 61 4 Updated May 29, 2023
Next