Lists (1)
Sort Name ascending (A-Z)
Stars
Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
[NeurIPS' 25] Benchmark for evaluating TTS models on complex prosodic, expressiveness, and linguistic challenges.
A project for tri-modal LLM benchmarking and instruction tuning.
Leaderboard and code for "Speech-IFEval", Interspeech 2025
[NeurIPS 2025] Benchmark data and code for MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
Kimi-VL: Mixture-of-Experts Vision-Language Model for Multimodal Reasoning, Long-Context Understanding, and Strong Agent Capabilities
Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.
Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/TTS/Audio Codec/audio LLM evaluation
[TACL'26] VoiceBench: Benchmarking LLM-Based Voice Assistants
AudioBench: A Universal Benchmark for Audio Large Language Models
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
Official repository for KoMT-Bench built by LG AI Research
An AI-powered research assistant that performs iterative, deep research on any topic by combining search engines, web scraping, and large language models. The goal of this repo is to provide the si…
Unofficial implementation of Titans, SOTA memory for transformers, in Pytorch
Unified automatic quality assessment for speech, music, and sound.
Open Data Platform for analysts, quants and AI agents.
Ola: Pushing the Frontiers of Omni-Modal Language Model
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery 🧑🔬
Easily fine-tune, evaluate and deploy Qwen, Gemma, or any open weight LLM!
[EMNLP 2025] OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
Official inference repo for FLUX.1 models
[NeurIPS 2024 Spotlight] The official implement of research paper "MotionBooth: Motion-Aware Customized Text-to-Video Generation"
The official repo of Qwen2-Audio chat & pretrained large audio language model proposed by Alibaba Cloud.
A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.