Skip to content
View qqsantaclaus's full-sized avatar

Block or report qqsantaclaus

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

Python 46 4 Updated Nov 18, 2025

Faster Whisper transcription with CTranslate2

Python 24,613 1,999 Updated Nov 19, 2025

Streamlit component allowing to record audio from the user's microphone and/or perform speech to text easily

Python 178 29 Updated Mar 6, 2024

Official PyTorch Implementation of "Scalable Diffusion Models with Transformers"

Python 8,693 807 Updated May 31, 2024

StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation

Python 256 33 Updated Sep 13, 2024

Use the Asana API to export Asana tasks and comments to HTML

PHP 7 3 Updated May 3, 2023

A small demonstration of using WebDataset with ImageNet and PyTorch Lightning

Python 76 6 Updated Dec 19, 2023

Developer friendly Natural Language Processing ✨

JavaScript 1,384 63 Updated May 27, 2026

Box SDK for Python

Python 458 221 Updated Jul 28, 2026

A command line interface for interacting with the Box API.

JavaScript 280 66 Updated Jul 28, 2026

Large, modern dataset for speech recognition

Shell 731 67 Updated Feb 26, 2024

🔊 Text-prompted Generative Audio Model - With the ability to clone voices

Jupyter Notebook 3,338 443 Updated Aug 24, 2025

The code for the bark-voicecloning model. Training and inference.

Python 710 115 Updated Sep 13, 2023

🔊 Text-Prompted Generative Audio Model

Jupyter Notebook 39,215 4,673 Updated Aug 19, 2024

A lightweight implementation of Beam Search for sequence models in PyTorch.

Python 61 7 Updated Jul 15, 2024

Unofficial implementation of miipher

Python 137 20 Updated Apr 19, 2024

De-essing software to reduce sibilance in speech

C 24 6 Updated Aug 8, 2020

This repo provides the processed samples of the manuscript "MossFormer: Pushing the Performance Limit of Monaural Speech Separation using Gated Single-head Transformer with Convolution-augmented Jo…

107 10 Updated Nov 28, 2024

An implementation of Performer, a linear attention-based transformer, in Pytorch

Python 1,179 150 Updated Feb 2, 2022

Implementation of Rotary Embeddings, from the Roformer paper, in Pytorch

Python 819 66 Updated Jun 20, 2026

Library for Textless Spoken Language Processing

Python 559 57 Updated Aug 29, 2023

A GPU-optional modular synthesizer in pytorch, 16200x faster than realtime, for audio ML researchers.

Python 383 16 Updated Feb 16, 2026

Contrastive Language-Audio Pretraining

Python 2,233 214 Updated May 15, 2025

Official Jax Implementation of MaskGIT

Jupyter Notebook 562 54 Updated Nov 18, 2022

A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation

Python 574 72 Updated Apr 2, 2023

Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable…

Jupyter Notebook 23,527 2,671 Updated Mar 3, 2026

Implementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind, in Pytorch

Python 1,544 94 Updated Apr 24, 2025

FMA: A Dataset For Music Analysis

Jupyter Notebook 2,644 457 Updated Jan 5, 2023

Object-oriented handling of audio data, with GPU-powered augmentations, and more.

Python 350 79 Updated Apr 1, 2025

Olive: Simplify ML Model Finetuning, Conversion, Quantization, and Optimization for CPUs, GPUs and NPUs.

Python 2,366 306 Updated Jul 29, 2026
Next