Stars
👀「大模型」2小时从0训练65M参数的视觉多模态VLM!Train a 65M-parameter VLM from scratch in just 2h!
UFM: A Unified Dense Image Correspondence Estimator for both Optical Flow & Wide Baseline Matching Tasks. Matches any pair of images. (NeurIPS 2025)
One framework to evaluate any VLA model on any robot simulation benchmark.
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
[ICML 2026] ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
1K resolution vision transformers pretrained on 1B human images.
INF Tech's open-source MLLMs for SOTA visual-language understanding and advanced document intelligence.
Allen Institute for AI: WildDet3D: Scaling Promptable 3D Detection in the Wild
SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments
This is a repository for listing papers on scene graph generation and application.
This repository collects and organises state‑of‑the‑art papers on spatial reasoning for Multimodal Vision–Language Models (MVLMs).
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
Build scalable data pipelines on YTsaurus with automatic stage management, local development simulation, and more.
NVIDIA Alpamayo 1.5 Nano is an open 10B reasoning VLA model for autonomous vehicles with reinforcement-learning enhanced reasoning, navigation guidance, and visual question answering.
[CVPR26] LEAD: Minimizing Learner–Expert Asymmetry in End-to-End Driving
NVIDIA Alpamayo 1 Nano is an open 10B reasoning VLA model for autonomous vehicles that pairs driving trajectories with Chain-of-Causation reasoning.
AlpaSim is an open-source autonomous vehicle simulation platform designed for development and testing of end-to-end AV policies
Official implementation of Kimodo, a kinematic motion diffusion model for high-quality human(oid) motion generation.
Seoul World Model: Grounding World Simulation Models in a Real-World Metropolis
The ultimate collection of high-fidelity Seedance 2.0 prompts and Seedance AI resources. Discover Seedance 2.0 how to use for cinematic film, anime, UGC, social media, meme and advertising. Include…
[ICML 2026 Oral] Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
[CVPR 2026] Offical implementation of the paper "HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images".
[ECCV 2026] Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
[ICML'26] Official repository of Utonia: Toward One Encoder for All Point Clouds
REFLEX Dataset: A Multimodal Dataset of Human Reactions to Robot Failures and Explanations
Official code for CVPR 2026 paper: VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection