Skip to content
View ZeyueT's full-sized avatar

Block or report ZeyueT

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

TypeScript 285 24 Updated Aug 3, 2026

On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

Python 512 36 Updated Aug 2, 2026

A collection of multimodal reasoning papers, codes, datasets, benchmarks and resources.

37 1 Updated Jul 31, 2026

Boogu-Image-0.1 is an Apache-2.0 open-source image generation and editing model family that delivers near-closed-source performance with an order of magnitude less data.

Python 939 54 Updated Jul 23, 2026

🚀 Fastest Anything-to-Audio Gen for conditioned sound and music creation.

Python 230 33 Updated Jun 22, 2026

MMAE: A Massive Multitask Audio Editing Benchmark

Python 102 4 Updated Jun 8, 2026

Official implementation of the CVPR 2026 paper "SonoWorld: From One Image to a 3D Audio-Visual Scene."

Python 41 2 Updated Jul 6, 2026

[SIGGRAPH 2026] Repository of Audio-Omni

Python 400 35 Updated Jun 10, 2026

Official Implementation of SAGE-GRPO:Manifold-Aware Exploration for Reinforcement Learning in Video Generation

Python 126 3 Updated Apr 2, 2026

[ICLR 2026] A two-stage evaluation benchmark for Text-to-Audio generation, assessing category, count, ordering, and timestamp accuracy with Gemini 2.5 Pro.

Python 3 Updated Mar 22, 2026

Allow your 🦞 bot to Shout, Speak, with "human" vibe

Python 525 77 Updated May 7, 2026
Python 203 Updated Feb 27, 2026

语音方向实验室/公司/资源/实习等,欢迎推荐或自荐

609 68 Updated Nov 13, 2024

A comprehensive list of papers for the definition of World Models and using World Models for General Video Generation, Embodied AI, and Autonomous Driving, including papers, codes, and related webs…

Python 1,950 68 Updated Jul 31, 2026

Advancing Open-source World Models

Python 4,341 396 Updated Jul 9, 2026
Python 195 35 Updated Nov 19, 2025

Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

Python 206 23 Updated May 29, 2024

ICML 2024 "From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation"

10 Updated Oct 13, 2024
Python 31 3 Updated Feb 4, 2021

[CVPR 2025] Pytorch implementation of the paper "Hearing Anywhere in Any Environment"

Python 34 1 Updated Sep 18, 2025

[NeurIPS 2024] Code, Dataset, Samples for the VATT paper “ Tell What You Hear From What You See - Video to Audio Generation Through Text”

Python 38 2 Updated Jul 24, 2025

[ICLR 2026 Oral] ScaleCUA is the open-sourced computer use agents that can operate on cross-platform environments (Windows, macOS, Ubuntu, Android).

Python 1,129 80 Updated Jan 7, 2026

[ICLR2026] WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction

Python 71 3 Updated Sep 3, 2025

A curated list of Vision (video/image) to Audio Generation

107 5 Updated Feb 10, 2026

[CVPR2024] Make Your Dream A Vlog

Python 435 47 Updated May 19, 2025

[ICLR2026] Video-GPT via Next Clip Diffusion.

Python 45 1 Updated Jun 2, 2025

[ICLR 2026] Repository of AudioX

Python 1,545 142 Updated Mar 10, 2026

[Neurips 2025 NextVid Workshop Oral✨] Official Implementation of VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention

Python 63 3 Updated Sep 22, 2025

Latest Advances on System-2 Reasoning

Python 1,352 80 Updated Jun 8, 2025
Next