Stars
An interface library for RL post training with environments.
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice…
JATTS: A modern, research-oriented Japanese Text-to-speech Open-sourced Toolkit
A Datacenter Scale Distributed Inference Serving Framework
Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement the code.
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang is a high-performance serving framework for large language models and multimodal models.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…
tsukumijima / pyopenjtalk-plus
Forked from r9y9/pyopenjtalkpyopenjtalk-plus: A Python wrapper for OpenJTalk with additional improvements
[ICASSP 2024] TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Official implementation of the paper "BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec"
Controllable and fast Text-to-Speech for over 7000 languages!
Inference and training library for high-quality TTS models.
Implementation of Band Split Roformer, SOTA Attention network for music source separation out of ByteDance AI Labs
Nendo is an open source platform for AI-driven audio management, intelligence, and generation.
Starter-kit to build constrained agents with Nextjs, FastAPI and Langchain
🔊 Text-Prompted Generative Audio Model
Foundational model for human-like, expressive TTS
[ICASSP 2024] This is the official code for "VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching"
Easily use and train state of the art late-interaction retrieval methods (ColBERT) in any RAG pipeline. Designed for modularity and ease-of-use, backed by research.