Highlights
- Pro
Stars
Pytorch implementation of MeanFlow on ImageNet and CIFAR10
[ICLR 2026] SoFlow: Solution Flow Models for One-Step Generative Modeling
BASTION: Budget-Aware Speculative Decoding with Tree-structured Block Diffusion Drafting
Cambrian-P: Pose-Grounded Video Understanding
[CVPR 2026]SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
Official codebase for the paper "How to build a consistency model: Learning flow maps via self-distillation" (NeurIPS 2025).
Official implementation of V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising (ECCV 2026)
Official implementation of ReDi: Rectified Discrete Flow (NeurIPS 2025)
PyTorch implementation of JiT https://arxiv.org/abs/2511.13720
Official Codebase For paper "One-step Language Modeling via Continuous Denoising"
[ICLR 2026] MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
An open-source evaluation toolkit to evaluate MLLMs on Spatial Intelligence using the EASI protocol
[CVPR 2026] Scaling Spatial Intelligence with Multimodal Foundation Models
[CVPR 2026] G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
A paper list for spatial reasoning
ViPE: Video Pose Engine for Geometric 3D Perception
Cambrian-S: Towards Spatial Supersensing in Video
Official repo and evaluation implementation of VSI-Bench
[CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
[CVPR 2024 & NeurIPS 2024] EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI
SpaceR: The first MLLM empowered by SG-RLVR for video spatial reasoning
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
[NeurIPS 2025] 3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding