Stars
[CVPR2026]🚀🚀🚀Official code for the paper "YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection." *(YOLO = You Only Look Once)* 🔥🔥🔥
🔥 Quo Vadis, World Modeling? Towards Interactive World Proxies for Continually Improving Agents
Code for SM4RT: Learning Structured Motion Geometry for 4D Reconstruction
fabio-sim / LightGlue-ONNX
Forked from cvg/LightGlueONNX-compatible LightGlue: Local Feature Matching at Light Speed. Supports TensorRT, OpenVINO
Official PyTorch implementation of "Let RGB Be the Language of Vision".
BoQ: A Place is Worth a Bag of learnable Queries (CVPR 2024)
Self-supervised learning for spatial perception
Non-invasive decoding of typed sentences from MEG and EEG brain recordings using a convolutional encoder, transformer, and character-level language model.
Official Implementation of "Multi-View Point Tracking as Geometric Supervision for 4D Video Generation"
[ICLR 25, TPAMI 26, CVPR 26] Track-On: Online Point Tracking with Memory
[CVPR'26 Highlight] Official Code for “V²-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence”
[CVPR26] MuM's a pretty good feature extractor for 3D tasks, probably the best.
[NeurIPS'23] Emergent Correspondence from Image Diffusion
[CVPR 2026 Highlight] MV-RoMa: From Pairwise Matching into Multi-View Track Reconstruction
GLUEMAP: Global Structure-from-Motion Meets Feedforward Reconstruction
Official implementation of Déjà View: Looping Transformers for Multi-View 3D Reconstruction
Deep functional map codebase for 3D shape matching with a 33× faster batched solver, multiple method implementations, and evaluation metrics.
TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction
Official implementation of "GARD: Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction"
A curated, continuously updated reading list, paper blogs, and resources for World Action Models (WAMs) in embodied AI.
A theoretical reconstruction of the Claude Mythos architecture, built from first principles using the available research literature.
[CVPR 2026 (Highlight)] Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
Allen Institute for AI: WildDet3D: Scaling Promptable 3D Detection in the Wild
Mastering Diverse Domains through World Models
pySLAM is a hybrid Python/C++ Visual SLAM pipeline supporting monocular, stereo, and RGB-D cameras. It provides a broad set of modern local and global feature extractors, multiple loop-closure stra…