Skip to content
View dukGuo's full-sized avatar
  • Northwestern Polytechnical University
  • China

Block or report dukGuo

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

🥢像老乡鸡🐔那样做饭。已添加2026年发布的《老乡鸡菜品溯源报告 2.0中新出现的菜品。主要部分于2024年完工,非老乡鸡官方仓库。文字来自《老乡鸡菜品溯源报告》,并做归纳、编辑与整理。CookLikeHOC.

Dockerfile 23,741 2,357 Updated Jul 19, 2026

🍲 好的,今天我们来做菜!OK, Let's Cook!

TypeScript 6,456 425 Updated Apr 12, 2026

Programmer's guide about how to cook at home.

101,458 11,051 Updated Jul 24, 2026

Inference for the STFT-VAE continuous audio codec (24kHz, 3.125Hz latent)

Python 43 2 Updated Jul 12, 2026
Python 65 Updated Jul 1, 2026

KVAE-Audio: a continuous full-band audio waveform autoencoder

Python 102 6 Updated Jul 23, 2026

Self-supervised learning (SSL) audio embedding framework with PyTorch Lightning + Hydra supporting Audio-JEPA, RQA-JEPA, BEST-RQ (ViT based), and BEST-RQ-2.

Python 7 Updated Feb 20, 2026
Python 4 2 Updated Jun 19, 2026

美股指南

6,199 920 Updated Jul 22, 2026

Official implementation of "USAD: Universal Speech and Audio Representation via Distillation"

Python 11 1 Updated Jun 7, 2026
Python 997 107 Updated Jul 10, 2026

[ECCV 2026] Towards Scalable Pre-training of Visual Tokenizers for Generation

Python 496 16 Updated Apr 15, 2026

Official code release for the paper "One-Step Generative Modeling via Wasserstein Gradient Flows"

Python 71 5 Updated Jun 9, 2026

Official Repo of "Flow-OPD: On-Policy Distillation for Flow Matching Models"

Python 266 4 Updated Jun 24, 2026

MultiModal Audio Generation in Raw Waveform Space.

Python 154 10 Updated May 26, 2026

[CVPR 2026 Findings] V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think

Python 56 2 Updated Apr 28, 2026

[CVPR 2026] Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation

Python 100 4 Updated Apr 26, 2026

[KDD 2026] Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe

Python 43 4 Updated Aug 10, 2025
Python 49 2 Updated May 2, 2026

A dual-rate LLM architecture bridging DSP and NLP. Decouples semantic planning from lexical synthesis to solve O(N2) bottlenecks.

Python 7 Updated Apr 11, 2026

Pytorch implementation of Transfusion, "Predict the Next Token and Diffuse Images with One Multi-Modal Model", from MetaAI

Python 1,385 73 Updated Jan 27, 2026

Scaled diffusion transformer for text-to-speech synthesis (DiT + T5Gemma2 conditioning, TorchTitan & Megatron backends, tested up to 1024 GPUs)

Python 24 Updated Mar 29, 2026

The agent that grows with you

Python 222,175 42,567 Updated Jul 29, 2026

CVPR 2026 (Oral)-Differentiable Vector Quantization for Rate-Distortion Optimization of Generative Image Compression

Python 52 Updated Jun 16, 2026

Single-stage End-to-End Training for Tokenization and Generation

Python 117 1 Updated Mar 24, 2026

DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick

Python 11 2 Updated May 12, 2026

A Large-scale Wu Dialect Speech Corpus with Multi-dimensional Annotations

Python 171 4 Updated Feb 6, 2026

Qwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilingual speech/music/song recognition, language detection and timestamp prediction.

Python 3,246 330 Updated Jun 26, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 384,453 80,797 Updated Jul 29, 2026
Next