VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
-
Updated
Sep 24, 2026 - Python
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required. Meetily (Meetly Ai - https://meetily.ai) is the #1 Self-hosted, Open-source Ai meeting note taker for macOS & Windows. Understand How to write meeting minutes
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
A lightweight yet powerful audio-to-MIDI converter with pitch bend detection
On-device subtitle generation that connects directly to DaVinci Resolve, Premiere, and After Effects.
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
Legacy Swift Hex app. Try the Rust rewrite at hex.kitlangton.com; new source at github.com/anomalyco/hex.
Open-source meeting transcription API for Google Meet, Microsoft Teams & Zoom. Auto-join bots, real-time WebSocket transcripts, MCP server for AI agents. Self-host or use hosted SaaS.
🔊 Awesome list for Whisper — an open-source AI-powered speech recognition system developed by OpenAI
Instant, controllable, local pre-trained AI models in Rust
Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
Cutting edge AI technology for automated audio transcription. A nice GUI for OpenAIs Whisper and pyannote (speaker identification)
A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness
A python package to build AI-powered real-time audio applications
To associate your repository with the transcription topic, visit your repo's landing page and select "manage topics."