-
ITMO University
Highlights
- Pro
Lists (9)
Sort Name ascending (A-Z)
Stars
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
U-MusT: A Unified Framework for Cross-modal Translation of Score Images, Symbolic Music, and Performance Audio
Create acoustic diffusers with custom images!
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python depend…
YingMusic-Singer-Plus: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
Run the full 2.78-trillion-parameter Kimi K3 model, DeepSeek V4.1 Flash or GLM-5.3-Flash beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C infe…
SyMuPe: Affective and Controllable Symbolic Music Performance (ACM MM '25, Outstanding Paper Award)
A multi-instrument music transcription model developed by Kyutai and Mirelo.
RapidIn: Scalable Influence Estimation for Large Language Models (LLMs). The implementation for paper "Token-wise Influential Training Data Retrieval for Large Language Models" (Accepted on ACL 2024).
TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
Towards a general language-audio model for computational paralinguistic tasks
DEMON: Diffusion Engine for Musical Orchestrated Noise
Codec for paper: LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis
Official Release of CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction
Move files and folders to the trash
SA3 medium audio inpainter — MLX SAME-L decoder + FastAPI + vanilla Svelte UI
An open-source model for music captioning, lyrics transcription, structural analysis, and musical question answering
an architecture for neural network inference in real-time audio applications
Implementation of Multiscreen proposed by Ken Nakanishi for "Screening is Enough"
Implementation of the Hierarchical Latent Action Model, proposed by Hanjung Kim et al. of Yonsei University
ICASSP2026 - Code for "Joint Estimation of Piano Dynamics and Metrical Structure". Estimate Piano dynamics markings from audio, with Bark-scale specific loudness feature extractor in PyTorch.
Implementation of Flow Matching model for MNIST to understand how it works