Skip to content
View lalimili6's full-sized avatar

Block or report lalimili6

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.

Swift 13,689 1,506 Updated Jul 24, 2026
Python 192 20 Updated Feb 27, 2026

The official repo of BridgeVoC, which explores using the Schrödinger Bridge framework for neural vocoding.

Python 143 13 Updated Nov 20, 2025

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Andr…

C++ 14,221 1,632 Updated Aug 13, 2026

a fully open-source implementation of a GPT-4o-like speech-to-speech video understanding model.

Python 38 4 Updated Apr 7, 2025
Python 54 7 Updated Mar 28, 2026

Build your custom LLM

Python 3 Updated Oct 2, 2025

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

Python 23,613 2,385 Updated Jul 13, 2026

[ICLR2026] FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates

Python 51 6 Updated Aug 6, 2026

[AAAI 2026 oral] KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction

Python 43 9 Updated Sep 25, 2025

SlamKit is an open source tool kit for efficient training of SpeechLMs. It was used for "Slamming: Training a Speech Language Model on One GPU in a Day"

Python 228 14 Updated Mar 14, 2026
Python 3,599 435 Updated Aug 13, 2026

superfast text to speech in any voice

Python 63 6 Updated Feb 16, 2026

Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Python 1,038 142 Updated Dec 2, 2025
Python 30 1 Updated Apr 29, 2026

On-device TTS model by Neuphonic

Python 6,237 656 Updated Jul 30, 2026

LLM-based ASR recipe with Zipformer encoder and Qwen LLM

Python 35 7 Updated Sep 25, 2025

Multi-speaker separation, identification, diarization ALL-IN-ONE. It can isolate the target speaker from a conversation audio and do ASR.

Python 96 12 Updated Oct 13, 2025

Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Wo…

Python 9,438 792 Updated Aug 17, 2026

Persian Kaldi profile for Rhasspy built from open speech data

Shell 18 3 Updated Oct 13, 2021

Unofficial WIP LoRa Finetuning repository for VibeVoice

Python 373 96 Updated Sep 24, 2025

Train any LLM to support speech input

Python 3 1 Updated Sep 10, 2025

[EMNLP 2025 Findings] Official code for EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

Python 43 11 Updated Sep 9, 2025

A toolkit for speaker diarization.

Jupyter Notebook 525 62 Updated Aug 4, 2026

Sisyphus recipies for ASR

Python 19 25 Updated Aug 10, 2026

IndexTTS Fine-tuning notebooks

Jupyter Notebook 139 23 Updated Jun 17, 2025

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Python 23,091 2,793 Updated Aug 13, 2026

Zero-shot voice cloning text-to-speech (TTS) with explicit emotion class conditioning built on F5-TTS

Python 40 6 Updated Mar 3, 2026

Very fast, accurate speaker diarization

Python 291 28 Updated Aug 5, 2026
Next