Stars
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
A collection of multimodal reasoning papers, codes, datasets, benchmarks and resources.
Boogu-Image-0.1 is an Apache-2.0 open-source image generation and editing model family that delivers near-closed-source performance with an order of magnitude less data.
🚀 Fastest Anything-to-Audio Gen for conditioned sound and music creation.
Official implementation of the CVPR 2026 paper "SonoWorld: From One Image to a 3D Audio-Visual Scene."
Official Implementation of SAGE-GRPO:Manifold-Aware Exploration for Reinforcement Learning in Video Generation
[ICLR 2026] A two-stage evaluation benchmark for Text-to-Audio generation, assessing category, count, ordering, and timestamp accuracy with Gemini 2.5 Pro.
Allow your 🦞 bot to Shout, Speak, with "human" vibe
A comprehensive list of papers for the definition of World Models and using World Models for General Video Generation, Embodied AI, and Autonomous Driving, including papers, codes, and related webs…
Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models
ICML 2024 "From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation"
[CVPR 2025] Pytorch implementation of the paper "Hearing Anywhere in Any Environment"
[NeurIPS 2024] Code, Dataset, Samples for the VATT paper “ Tell What You Hear From What You See - Video to Audio Generation Through Text”
[ICLR 2026 Oral] ScaleCUA is the open-sourced computer use agents that can operate on cross-platform environments (Windows, macOS, Ubuntu, Android).
[ICLR2026] WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
A curated list of Vision (video/image) to Audio Generation
[ICLR2026] Video-GPT via Next Clip Diffusion.
[Neurips 2025 NextVid Workshop Oral✨] Official Implementation of VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
Latest Advances on System-2 Reasoning