Skip to content
View alufia's full-sized avatar
  • Yonsei University

Block or report alufia

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.

Jupyter Notebook 3,922 279 Updated Apr 23, 2026

[NeurIPS' 25] Benchmark for evaluating TTS models on complex prosodic, expressiveness, and linguistic challenges.

Python 226 16 Updated Dec 9, 2025

Audio Normalization for Python/ffmpeg

HTML 1,522 128 Updated Jul 10, 2026

A project for tri-modal LLM benchmarking and instruction tuning.

Python 61 8 Updated Mar 27, 2025

Leaderboard and code for "Speech-IFEval", Interspeech 2025

Python 24 2 Updated May 27, 2025

[NeurIPS 2025] Benchmark data and code for MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

Python 214 5 Updated Feb 25, 2026

Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation

Python 4,702 372 Updated Jun 21, 2025

Kimi-VL: Mixture-of-Experts Vision-Language Model for Multimodal Reasoning, Long-Context Understanding, and Strong Agent Capabilities

1,218 99 Updated Jul 15, 2025

Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.

Jupyter Notebook 4,056 325 Updated Jun 12, 2025

Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/TTS/Audio Codec/audio LLM evaluation

Python 311 24 Updated Jul 29, 2026

[TACL'26] VoiceBench: Benchmarking LLM-Based Voice Assistants

Python 378 28 Updated Jun 11, 2026

AudioBench: A Universal Benchmark for Audio Large Language Models

Python 319 15 Updated May 29, 2026

Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Python 223 15 Updated Feb 28, 2025

Official repository for KoMT-Bench built by LG AI Research

Python 73 3 Updated Aug 8, 2024

An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. The goal of this repo is to provide the si…

TypeScript 19,449 1,984 Updated Apr 11, 2026

Unofficial implementation of Titans, SOTA memory for transformers, in Pytorch

Python 1,972 206 Updated Jul 13, 2026

Unified automatic quality assessment for speech, music, and sound.

Python 749 53 Updated Jun 5, 2025

Open Data Platform for analysts, quants and AI agents.

Python 71,193 7,263 Updated Jul 30, 2026

Ola: Pushing the Frontiers of Omni-Modal Language Model

Python 395 16 Updated Jun 13, 2025

The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery 🧑‍🔬

Jupyter Notebook 14,319 2,033 Updated Dec 19, 2025

Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!

Python 9,369 783 Updated Jul 30, 2026

[EMNLP 2025] OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking

Python 487 63 Updated Aug 23, 2025

AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Python 133 5 Updated Dec 9, 2024

Official inference repo for FLUX.1 models

Python 25,832 1,904 Updated Jul 31, 2025

[NeurIPS 2024 Spotlight] The official implement of research paper "MotionBooth: Motion-Aware Customized Text-to-Video Generation"

Python 138 10 Updated Oct 8, 2024

The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.

Python 2,098 168 Updated Apr 21, 2025

A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.

Python 2,303 220 Updated Apr 13, 2026
Next