Stars
WEFE: The Word Embeddings Fairness Evaluation Framework. WEFE is a framework that standardizes the bias measurement and mitigation in Word Embeddings models. Please feel welcome to open an issue in…
INTERSPEECH 2023-2024 Papers: A complete collection of influential and exciting research papers from the INTERSPEECH 2023-24 conference. Explore the latest advances in speech and language processin…
A list of awesome AI in libraries, archives, and museum collections from around the world 🕶️
A collection of datasets for the purpose of emotion recognition/detection in speech.
AI Audio Datasets (AI-ADS) 🎵, including Speech, Music, and Sound Effects, which can provide training data for Generative AI, AIGC, AI model training, intelligent audio tool development, and audio a…
Must-read Papers on Knowledge Editing for Large Language Models.
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Unofficial reimplementation of ECAPA-TDNN for speaker recognition (EER=0.86 for Vox1_O when train only in Vox2)
Matlab and Python libraries for an unsupervised method for robust voice activity detection (rVAD), as in the paper rVAD: An Unsupervised Segment-Based Robust Voice Activity Detection Method.
Fast audio data augmentation in PyTorch. Inspired by audiomentations. Useful for deep learning.
INA's library with pretrained models for gender and age prediction from faces.
PolyglotDB is a package for phonetic corpus storage and analysis
biblatex is a sophisticated bibliography system for LaTeX users. It has considerably more features than traditional bibtex and supports UTF-8
clean CV LaTex template with GitHub Actions that compile and publish new changes
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
Official repository for MixFaceNets: Extremely Efficient Face Recognition Networks
Tensorflow-Keras Model Profiler: Tells you model's memory requirement, no. of parameters, flops etc.
CNN-based audio segmentation toolkit. Allows to detect speech, music, noise and speaker gender. Has been designed for large scale gender equality studies based on speech time per gender.
Fully-Convolutional Network for Pitch Estimation of Speech Signals
Robust Speech Recognition via Large-Scale Weak Supervision
Neighborhood Attention Transformer, arxiv 2022 / CVPR 2023. Dilated Neighborhood Attention Transformer, arxiv 2022
SPINOS: A Dataset of Subtle Polarity and Intensity Opinion Shifts
Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
🎥 Python and OpenCV-based scene cut/transition detection program & library.