Lists (1)
Sort Name ascending (A-Z)
Stars
Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.
Noise supression using deep filtering
A comprehensive audio, image, video, CSV, and JSONL viewer extension for VSCode and Cursor.
Open-Source Turn-Taking Detection Model and Dataset for Full-Duplex Spoken Dialogue Systems
Some comprehensive papers about speaker diarization
Python Kalman filtering and optimal estimation library. Implements Kalman filter, particle filter, Extended Kalman filter, Unscented Kalman filter, g-h (alpha-beta), least squares, H Infinity, smoo…
Unofficial implementation of Titans, SOTA memory for transformers, in Pytorch
Official repository of SepReformer for speech separation
The official Pytorch implementation of "Frame-wise streaming end-to-end speaker diarization with non-autoregressive self-attention-based attractors". [ICASSP 2024] and "LS-EEND: long-form streaming…
Official implementation for our paper "Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations"
wsj0-{2, 3, 4, 5} mix generation scripts, in Python.
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
Kalman Filter in Python (파이썬으로 구현하는 칼만 필터)
Collection of papers on state-space models
Two-stage progressive neural network for acoustic echo cancellation
Unofficial implementation of SCP-GAN
The official implementation of GTCRN, an ultra-lightweight SE model.
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
21 Lessons, Get Started Building with Generative AI
This repo contains the official PyTorch implementation of "A Systematic Comparison of Phonetic Aware Techniques for Speech Enhancement" (Interspeech 2022)
Implementation of Mega, the Single-head Attention with Multi-headed EMA architecture that currently holds SOTA on Long Range Arena
A simple way to keep track of an Exponential Moving Average (EMA) version of your Pytorch model
ADAPTING SELF-SUPERVISED MODELS TO MULTI-TALKER SPEECH RECOGNITION USING SPEAKER EMBEDDINGS
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities