Skip to content
View RogerZhangzz's full-sized avatar
  • Sydney, Au

Highlights

  • Pro

Block or report RogerZhangzz

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that sho…

Python 11,298 1,705 Updated Jul 31, 2026

Official codes for paper: 3DGS-DET: Empower 3D Gaussian Splatting with Boundary Guidance and Box-Focused Sampling for Indoor 3D Object Detection

165 5 Updated Mar 16, 2026

Build resilient agents.

Python 39,590 6,643 Updated Aug 12, 2026

😎 Awesome lists of papers and codes about Large Vision-Language Models

13 Updated Apr 1, 2024

😎 Awesome lists of papers and codes about open-vocabulary perception, including both 3D and 2D

64 4 Updated Jul 27, 2025

Awesome Data-Driven Autonomous Driving Solutions. Also the official repository of our survey paper: Data-Centric Evolution in Autonomous Driving: A Comprehensive Survey of Big Data System, Data Min…

173 12 Updated Mar 20, 2024

Emu Series: Generative Multimodal Models from BAAI

Python 1,777 83 Updated Jan 12, 2026

[CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.

Python 3,345 268 Updated Jul 17, 2026

(TPAMI 2024) A Survey on Open Vocabulary Learning

1,002 43 Updated May 12, 2026

Official code for NeurIPS2023 paper CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection and TPAMI 2025 paper CoDAv2

Jupyter Notebook 224 16 Updated May 28, 2026

[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.

Python 24,977 2,776 Updated Aug 12, 2024

🧀 Code and models for the ICML 2023 paper "Grounding Language Models to Images for Multimodal Inputs and Outputs".

Jupyter Notebook 484 37 Updated Oct 30, 2023
Python 54 3 Updated Oct 11, 2023

[ECCV 2024] Official implementation of the paper "Semantic-SAM: Segment and Recognize Anything at Any Granularity"

Python 2,855 146 Updated Jul 10, 2025
Python 816 48 Updated Jul 8, 2024

(ECCVW 2025)GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Python 555 28 Updated Jun 3, 2025

This is the official code for MobileSAM project that makes SAM lightweight for mobile applications and beyond!

Jupyter Notebook 5,844 586 Updated May 5, 2026

A collection of resources and papers on Diffusion Models

HTML 12,366 1,010 Updated Aug 1, 2024

Segment Anything in High Quality [NeurIPS 2023]

Jupyter Notebook 4,253 264 Updated Sep 12, 2025

A collection of papers on the topic of ``Computer Vision in the Wild (CVinW)''

1,373 57 Updated Mar 14, 2024
Python 7,805 523 Updated Apr 14, 2024

A curated list of foundation models for vision and language tasks

1,172 62 Updated Apr 20, 2026

[ICRA 2022] Towards Scale Consistent Monocular Visual Odometry by Learning from the Virtual World

Python 31 2 Updated Apr 19, 2023

[IJCV 2022] Information-Theoretic Odometry Learning

Python 17 2 Updated Apr 19, 2023

MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.

Python 954 38 Updated Mar 19, 2025

The official repo for [AAAI 2024] "SimDistill: Simulated Multi-modal Distillation for BEV 3D Object Detection""

Python 43 3 Updated May 16, 2024

The official repo for [TPAMI'23] "Vision Transformer with Quadrangle Attention"

Python 239 10 Updated Sep 25, 2025

Official implementation of "Composer: Creative and Controllable Image Synthesis with Composable Conditions"

1,559 47 Updated Dec 26, 2023
Python 25 1 Updated Mar 9, 2023

Github Pages template based upon HTML and Markdown for personal, portfolio-based websites.

SCSS 17,436 8,401 Updated Aug 3, 2026
Next