- Shanghai,China
- blog.csdn.net/zhulinniao
Stars
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
Fast, Sharp & Reliable Agentic Intelligence
Plug-and-play streaming semantic VAD for real-time full-duplex spoken dialogue systems.
The Full-Duplex Interaction Track of the ICASSP 2026 Human-like Spoken Dialogue Systems Challenge aims to advance the evaluation of full-duplex dialogue systems by in- troducing a dual-channel dial…
MOSS-Audio-Tokenizer is a Causal Transformer-based audio tokenizer built on the CAT architecture. Trained on 3M hours of diverse audio, it supports streaming and variable bitrates, delivering SOTA …
MOSS-TTS-Nano is an open-source multilingual tiny speech generation model from MOSI.AI and the OpenMOSS team. With only 0.1B parameters, it is designed for realtime speech generation, can run direc…
State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.
An Open-source Streaming High-fidelity Neural Audio Codec
AcademiCodec: An Open Source Audio Codec Model for Academic Research
High-Quality Voice Cloning TTS for 600+ Languages
A high-quality rapid TTS voice cloning model that reaches speeds of 150x realtime.
Fast CosyVoice3 inference with tensorRT and tensorRT-LLM
Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (LLM).
Qwen3-ASR speech-to-text for llama.cpp — patch, GGUF models, and benchmarks
Implementation of Qwen3-ASR-0.6B in GGML
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singi…
LightTTS is a lightweight TTS inference framework optimized for CosyVoice2 and CosyVoice3, enabling fast and scalable speech synthesis in Python and supports stream and bistream modes.
Official Python toolkit for the Qwen3-ASR API. Parallel high‑throughput calls, robust long‑audio transcription, multi‑sample‑rate support.
Qwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilingual speech/music/song recognition, language detection and timestamp prediction.
Worlds first open-source real-time end-to-end spoken dialogue model with personalized voice cloning.
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice…
Liquid Audio - Speech-to-Speech audio models by Liquid AI