Skip to content
View ABexit's full-sized avatar
💭
I may be slow to respond.
💭
I may be slow to respond.
  • University of Chinese Academy of Sciences
  • BeiJing

Block or report ABexit

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

The open-source AI voice studio. Clone, dictate, create.

TypeScript 50,497 6,260 Updated Aug 9, 2026

Open Source framework for voice agents, multimodal apps, and realtime AI. Maintained by Daily and the community.

Python 14,133 2,463 Updated Aug 15, 2026

High-Quality Voice Cloning TTS for 600+ Languages

Python 9,141 1,502 Updated Aug 10, 2026

[ACL 2026 Main] FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pre-training

Python 36 2 Updated Apr 20, 2026

VITA-QINYU: Expressive Spoken Language Model for Role-Playing and Singing

Python 123 7 Updated Jul 14, 2026
Python 447 34 Updated Mar 25, 2026

The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how t…

Python 3,592 326 Updated May 26, 2026

GitHub Repository for the AudSemThinker Model and the AudSem Dataset

Python 14 2 Updated Jun 4, 2025

A framework for efficient model inference with omni-modality models

Python 6,132 1,470 Updated Aug 15, 2026

Send a phone call from AI agent, in an API call. Or, directly call the bot from the configured phone number!

Python 6,556 774 Updated Jul 1, 2026

Open-source framework for conversational voice AI agents

Python 11,051 1,349 Updated Aug 14, 2026

[ACL 2025] OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching

Jupyter Notebook 45 6 Updated Feb 9, 2025

The first medical SpeechLM, open-sourced with weight, data, and code of training, inference, and evaluation.

Python 10 Updated Apr 23, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 386,405 81,214 Updated Aug 16, 2026
Python 291 61 Updated Aug 7, 2026

Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice…

Python 12,966 1,681 Updated Mar 17, 2026

PersonaPlex code.

Python 10,365 1,444 Updated Mar 2, 2026

A Fully Self-Hosted Solution for Full-Duplex Voice Interaction

Python 580 46 Updated Sep 28, 2025

GPT-4o-level, real-time spoken dialogue system.

Python 375 33 Updated Jan 27, 2025

An voice activity detection and audio segmentation tool

Python 858 100 Updated Jul 28, 2026

Silero VAD: pre-trained enterprise-grade Voice Activity Detector

Python 9,954 825 Updated Jul 16, 2026

FlowMirror-HydraVox — A natively accelerated multi-head autoregressive TTS system derived from CosyVoice 3.0. It predicts multiple tokens per step for faster, high-quality speech synthesis, featuri…

Python 49 4 Updated Feb 17, 2026

Added vLLM support to IndexTTS for faster inference.

Python 1,227 175 Updated Apr 13, 2026

Multilingual TTS model with voice cloning and duration control, based on T5Gemma encoder-decoder LLM

Python 312 31 Updated Apr 3, 2026

SoTA open-source TTS

Python 25,997 3,481 Updated Jul 21, 2026

💖🧸 Self hosted, you-owned Grok Companion, a container of souls of waifu, cyber livings to bring them into our worlds, wishing to achieve Neuro-sama's altitude. Capable of realtime voice chat, Minec…

TypeScript 47,926 4,750 Updated Aug 15, 2026

GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Python 2,366 177 Updated Jul 21, 2026

The official repo for "Vidi: Large Multimodal Models for Video Understanding and Editing"

Python 651 44 Updated Aug 3, 2026

Controllable and fast Text-to-Speech for over 7000 languages!

Python 2,210 316 Updated Jan 25, 2026
Next