Lists (3)
Sort Name ascending (A-Z)
Stars
AI Image Detector - Check if the Image is AI
Vidu S1: A Real-Time Interactive Video Generation Model
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling
Open-source humanize text toolkit. Documents 4 humanization methodologies with reference implementations, plus a production pipeline combining LLM rewriting with a cross-engine translation chain. P…
MultiModal Audio Generation in Raw Waveform Space.
YAML-native agent workflow execution engine, written in Rust
Official Cortex development plugin with native coding tools, workflow skills, project analysis, and git-aware automation.
Enterprise-ready Spring AI platform for RAG, tool calling, async ingestion, JWT/RBAC security, and observability.
Omni2Sound — Your Multimodal Audio Generation Codebase (CVPR 2026 Highlight)
Being-H is BeingBeyond's family of human-centric embodied foundation models.
[RSS26'] Welcome to Psi-Zero, a Humanoid VLA towards Universal Humanoid Intelligence.
MiMo-Audio: Audio Language Models are Few-Shot Learners
The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how t…
Official repository of Myna: Masking-Based Contrastive Learning of Musical Representations
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
Vector (and Scalar) Quantization, in Pytorch
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
[ICLR 2025] Official PyTorch implementation of our paper for general continual learning "Advancing Prompt-Based Methods for Replay-Independent General Continual Learning".
Official PyTorch implementation of our CVPR 2025 paper, "LoRA Subtraction for Drift-Resistant Space in Exemplar-Free Continual Learning."
Awesome Incremental Learning
[ICLR 2025] Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
A library built for easier audio self-supervised training, downstream tasks evaluation
A Comprehensive Survey on Continual Learning in Generative Models.
This is the official repository of the papers "Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers" [IEEE MLSP 2024] and "Efficient Fine-tuning of Audio Spectrogram Transformers…