Starred repositories
🥢像老乡鸡🐔那样做饭。已添加2026年发布的《老乡鸡菜品溯源报告 2.0中新出现的菜品。主要部分于2024年完工,非老乡鸡官方仓库。文字来自《老乡鸡菜品溯源报告》,并做归纳、编辑与整理。CookLikeHOC.
Programmer's guide about how to cook at home.
Inference for the STFT-VAE continuous audio codec (24kHz, 3.125Hz latent)
KVAE-Audio: a continuous full-band audio waveform autoencoder
Self-supervised learning (SSL) audio embedding framework with PyTorch Lightning + Hydra supporting Audio-JEPA, RQA-JEPA, BEST-RQ (ViT based), and BEST-RQ-2.
Official implementation of "USAD: Universal Speech and Audio Representation via Distillation"
[ECCV 2026] Towards Scalable Pre-training of Visual Tokenizers for Generation
Official code release for the paper "One-Step Generative Modeling via Wasserstein Gradient Flows"
Official Repo of "Flow-OPD: On-Policy Distillation for Flow Matching Models"
MultiModal Audio Generation in Raw Waveform Space.
[CVPR 2026 Findings] V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
[CVPR 2026] Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation
[KDD 2026] Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
A dual-rate LLM architecture bridging DSP and NLP. Decouples semantic planning from lexical synthesis to solve O(N2) bottlenecks.
Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI
Scaled diffusion transformer for text-to-speech synthesis (DiT + T5Gemma2 conditioning, TorchTitan & Megatron backends, tested up to 1024 GPUs)
The agent that grows with you
CVPR 2026 (Oral)-Differentiable Vector Quantization for Rate-Distortion Optimization of Generative Image Compression
Single-stage End-to-End Training for Tokenization and Generation
DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick
A Large-scale Wu Dialect Speech Corpus with Multi-dimensional Annotations
Qwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilingual speech/music/song recognition, language detection and timestamp prediction.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞