Skip to content
View whaozl's full-sized avatar

Block or report whaozl

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

C 21,107 1,905 Updated Aug 9, 2026
Python 693 50 Updated Apr 29, 2026

Fast, Sharp & Reliable Agentic Intelligence

C++ 2,092 85 Updated Apr 3, 2026

Plug-and-play streaming semantic VAD for real-time full-duplex spoken dialogue systems.

Python 288 38 Updated Jul 17, 2026

The Full-Duplex Interaction Track of the ICASSP 2026 Human-like Spoken Dialogue Systems Challenge aims to advance the evaluation of full-duplex dialogue systems by in- troducing a dual-channel dial…

Python 38 Updated Apr 27, 2026

MOSS-Audio-Tokenizer is a Causal Transformer-based audio tokenizer built on the CAT architecture. Trained on 3M hours of diverse audio, it supports streaming and variable bitrates, delivering SOTA …

Python 250 18 Updated Jun 16, 2026

MOSS-TTS-Nano is an open-source multilingual tiny speech generation model from MOSI.AI and the OpenMOSS team. With only 0.1B parameters, it is designed for realtime speech generation, can run direc…

Python 4,108 523 Updated Jul 26, 2026
Python 287 60 Updated Aug 7, 2026

State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.

Python 4,013 362 Updated Jan 4, 2024

An Open-source Streaming High-fidelity Neural Audio Codec

Python 512 32 Updated Mar 4, 2025

AcademiCodec: An Open Source Audio Codec Model for Academic Research

Python 674 84 Updated Dec 27, 2023

High-Quality Voice Cloning TTS for 600+ Languages

Python 8,910 1,467 Updated Aug 5, 2026

State-of-the-art TTS model under 25MB 😻

Python 15,311 883 Updated Jun 11, 2026

A high-quality rapid TTS voice cloning model that reaches speeds of 150x realtime.

Python 4,916 636 Updated Jun 5, 2026

Fast CosyVoice3 inference with tensorRT and tensorRT-LLM

Python 77 17 Updated Mar 7, 2026

Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (LLM).

Python 683 75 Updated Jul 31, 2026

Qwen3-ASR speech-to-text for llama.cpp — patch, GGUF models, and benchmarks

Python 15 2 Updated Feb 2, 2026

Implementation of Qwen3-ASR-0.6B in GGML

C++ 107 42 Updated Jul 28, 2026

A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singi…

Python 631 45 Updated Jun 2, 2026

LightTTS is a lightweight TTS inference framework optimized for CosyVoice2 and CosyVoice3, enabling fast and scalable speech synthesis in Python and supports stream and bistream modes.

Python 48 8 Updated Apr 14, 2026

Official Python toolkit for the Qwen3-ASR API. Parallel high‑throughput calls, robust long‑audio transcription, multi‑sample‑rate support.

Python 993 94 Updated Feb 5, 2026

Qwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilingual speech/music/song recognition, language detection and timestamp prediction.

Python 3,341 333 Updated Jun 26, 2026

Worlds first open-source real-time end-to-end spoken dialogue model with personalized voice cloning.

Jupyter Notebook 550 60 Updated Apr 17, 2026

The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.

Python 948 49 Updated Aug 10, 2026

Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice…

Python 12,893 1,668 Updated Mar 17, 2026
Python 1,616 218 Updated Jun 17, 2026

Liquid Audio - Speech-to-Speech audio models by Liquid AI

Python 558 96 Updated Jun 5, 2026

A fast, local neural text to speech system

C++ 11,277 1,051 Updated Aug 26, 2025

使用python进行语音识别

Python 170 531 Updated Feb 16, 2022
Next