Stars
[AAAI 2026] OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
[ICRA 2026] RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
OpenMMLab Pre-training Toolbox and Benchmark
[ICCV 2023] Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object Prediction
[ICCV 2023 & TPAMI 2026] SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera Videos
[CVPR 2023] VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
Code release for our CVPR 2023 paper "Detecting Everything in the Open World: Towards Universal Object Detection".
An open source implementation of CLIP.
[ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
cvpr2024/cvpr2023/cvpr2022/cvpr2021/cvpr2020/cvpr2019/cvpr2018/cvpr2017 论文/代码/解读/直播合集,极市团队整理
PromptDet: Towards Open-vocabulary Detection using Uncurated Images, ECCV2022
[CVPR 2022] Official code for "RegionCLIP: Region-based Language-Image Pretraining"
Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone
Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
Improving the Transferability of Adversarial Examples with Resized-Diverse-Inputs, Diversity-Ensemble and Region Fitting
Object detection, 3D detection, and pose estimation using center point detection:
Crack LeetCode, not only how, but also why.
FCOS: Fully Convolutional One-Stage Object Detection (ICCV'19)
SOLO and SOLOv2 for instance segmentation, ECCV 2020 & NeurIPS 2020.
Code for "Deep Snake for Real-Time Instance Segmentation" CVPR 2020 oral