-
Sony CSL
- France
- @howariou
Stars
alsa-scarlett-gui is a Gtk4 GUI for the ALSA controls presented by the Linux kernel Focusrite USB Drivers
WhisperDRZ: a Whisper-large-v3 fine-tune that transcribes, diarizes (who said what), predicts word-level timestamps, and tags non-speech events. Inference only.
The Free Software Media System - Server Backend & API
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Code for the paper “Automatic Music Sample Identification with Multi-Track Contrastive Learning”.
A Pytorch (support batch and channel) implementation of GoogleBrain's SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Self-supervised key estimation model that matches performance with supervised state-of-the-art model.
C++ polyphonic pitch/time library (GitHub mirror)
Simple python library for pitch shifting and time stretching. Wrapper of Signalsmith Stretch C++ Library
Time-stretch audio clips quickly with PyTorch (CUDA supported)! Additional utilities for searching efficient transformations are included.
An automatic sample identification (ASID) system using a contrastively trained GNN encoder.
A Tiktok bot built in python that can send a lot of reports automatically in any Tiktok Video.
Repository for the ISMIR 2024 Paper "STONE: Self-supervised Tonality Estimator".
Pytorch code for "Improving Self-Supervised Learning by Characterizing Idealized Representations"
Code release for "Improved baselines for vision-language pre-training"
PyTorch Lightning + Hydra. A very user-friendly template for ML experimentation. ⚡🔥⚡
State-of-the-art pretrained music models for training, evaluation, inference
Localized watermarking for AI-generated speech audios, with SOTA on robustness and very fast detector
Supplementary material for the ISMIR 2019 paper entitled "Multi-Task Learning of Tempo and Beat: Learning One to Improve the Other"
Post-processing for CREPE to turn f0 pitch estimates into discrete notes e.g. MIDI
Comparison of Python audio resampling implementations
Hackable and optimized Transformers building blocks, supporting a composable construction.
Fast and memory-efficient exact attention
Mustango: Toward Controllable Text-to-Music Generation
ISMIR 2023 Papers: A complete collection of influential and exciting research papers from the ISMIR 2023 conference.