Starred repositories
Official code for CVPR 2026 paper: VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
CVPR 2026 - TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detection
[CVPR 2026] R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection
[ICCV 2025] Detect Anything 3D in the Wild
OpenPCDet Toolbox for LiDAR-based 3D Object Detection.
OpenMMLab's next-generation platform for general 3D object detection.
[CVPR 2026] Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments
Official implementation of "AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison"
CADAM is the open source text-to-CAD web application
A library of agent skills for CAD, CAE and CAM
[ECCV 2026] RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
Multi-Joint dynamics with Contact. A general purpose physics simulator.
VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer (CVPR 2026)
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
[CVPR2025] AnomalyNCD: Towards Novel Anomaly Class Discovery in Industrial Scenarios. Paper is available at https://arxiv.org/abs/2410.14379
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
[CVPR 2025] Towards Training-free Anomaly Detection with Vision and Language Foundation Models
This is the official repository for our recent paper "Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models".
[AAAI 2024 Oral] AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language Models
Open Multi-Agent Interactive Classroom — Get an immersive, multi-agent learning experience in just one click
Normal-Abnormal Guided Generalist Anomaly Detection (NeurIPS 2025)
Official repository for the paper "MGPC: Multimodal Network for Generalizable Point Cloud Completion With Modality Dropout and Progressive Decoding"
[CORL 2025 Oral]One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation.
[CVPR2025] Hand-held Object Reconstruction from RGB Video with Dynamic Interaction
3D BBox refinement interface used in LabelAny3D (NeurIPS 2025)
[NeurIPS 2025] LabelAny3D: Label Any Object 3D in the Wild
[CVPR 2025] Official Implementation of "Dinomaly: The Less Is More Philosophy in Multi-Class Unsupervised Anomaly Detection". The first multi-class UAD model that can compete with single-class SOTAs
(ICCV 2025) DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup