Stars
First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come …
A framework for efficient model inference with omni-modality models
MCP server and Claude Code skill for Excalidraw — programmatic canvas toolkit to create, edit, and export diagrams via AI agents with real-time canvas sync.
List of open-source TTS, voice cloning, and music generation models
The agent that grows with you
Generation of diagrams like flowcharts or sequence diagrams from text in a similar manner as markdown
React component for 2D, 3D, VR and AR force directed graphs
Omnilingual ASR Open-Source Multilingual SpeechRecognition for 1600+ Languages
Liquid Audio - Speech-to-Speech audio models by Liquid AI
A PyTorch native platform for training generative AI models
😎 Awesome lists about all kinds of interesting topics
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal is…
Kyutai's Speech-To-Text and Text-To-Speech models based on the Delayed Streams Modeling framework.
zero-shot voice conversion & singing voice conversion, with real-time support
python bindings for symphonia/opus - read various audio formats from python and write opus files
Continuation of the "Official Python Client for the Discogs API"
Awesome speech/audio LLMs, representation learning, and codec models
A Conversational Speech Generation Model
Fast and memory-efficient exact attention
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis
Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities。
open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
Learning Vim and Vimscript doesn't have to be hard. This is the guide that you're looking for 📖
Verbatim Automatic Speech Recognition with improved word-level timestamps and filler detection
Transcription, forced alignment, and audio indexing with OpenAI's Whisper