Stars
Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
Boogu-Image-0.1 is an Apache-2.0 open-source image generation and editing model family that delivers near-closed-source performance with an order of magnitude less data.
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video proโฆ
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects
Official Python inference and LoRA trainer package for the LTX-2 audioโvideo generative model.
A digital data-generation pipeline that synthesizes humanoid loco-manipulation data from 3D assets and video priors.
Various retargeting optimizers to translate human hand motion to robot hand motion.
[CVPR 2026] Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking
[NeurIPS 2025 Spotlight] SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
[ICRA 2026] VITRA: Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
Official implementation of Forge4D: Feed-Forward 4D Human Reconstruction and Interpolation from Uncalibrated Sparse Videos
Foundation Models and Data for Human-Human and Human-AI interactions.
๐น A more flexible framework that can generate videos at any resolution and creates videos from images.
The code of paper "LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning" accepted by ICLR'25
Generate ARKit expression from audio in realtime
Lets make video diffusion practical!
[SIGGRAPH 2025] LAM: Large Avatar Model for One-shot Animatable Gaussian Head
[CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
An autonomous agent that conducts deep research on any data using any LLM providers
๐ Explore Egocentric Vision: research, data, challenges, real-world apps. Stay updated & contribute to our dynamic repository! Work-in-progress; join us!
A curated list of state-of-the-art research in embodied AI, focusing on vision-language-action (VLA) models, vision-language navigation (VLN), and related multimodal learning approaches.
[RSS 2023] Diffusion Policy Visuomotor Policy Learning via Action Diffusion
Simulation platform for general-purpose robotics & embodied AI learning.
A JavaScript library like PyTorch, with GPU acceleration.
๐ง๐ปโโ๏ธA list of papers curated for you to dive into the Awesome Radiance Field-based 3D Editing.