Stars
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts Interested in World Modeling.
This is a collection of recent papers on reasoning in video generation models.
Public code for XFactor: Introduces the first geometry-free model to achieve true self-supervised / pose-free Novel View Synthesis (NVS) by learning transferable latent camera pose representations.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
A collection of papers on semantic correspondence, organized by year.
[ICLR'26] IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
[Awesome-Spatial-VLMs] This repository is the official, community-maintained resource for the survey paper: Spatial Intelligence in Vision-Language Models: A Comprehensive Survey;
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
Official implementation of paper "Unified World Models: Memory-Augmented Planning and Foresight for Visual Navigation"
[ECCV 2026🔥] SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT
Official Repository for “CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models" [CVPR2025]
InteriorGS: 3D Gaussian Splatting Dataset of Semantically Labeled Indoor Scenes
RAGEN leverages reinforcement learning to train LLM reasoning agents in interactive, stochastic environments.
World model reasoning RL for multi-turn VLM agents
The first collection of academic iKUN papers in the world
Github repository for "Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas" (ICML 2025)
Official code for Paper "Mantis: Multi-Image Instruction Tuning" [TMLR 2024 Best Paper]
A paper list for spatial reasoning
TheaterGen: Character Management with LLM for Consistent Multi-turn Image Generation