Stars
Code & Dataset repository for the paper "Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation"
DINO-X: The World's Top-Performing Vision Model for Open-World Object Detection and Understanding
Official Repo For OMG-LLaVA and OMG-Seg codebase [CVPR-24 and NeurIPS-24]
RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RN…
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use th…
Llama中文社区,实时汇总最新Llama学习资料,构建最好的中文Llama大模型开源生态,完全开源可商用
A Framework of Small-scale Large Multimodal Models
[CVPR 2024] PixelLM is an effective and efficient LMM for pixel-level reasoning and understanding.
We proposed a large-scale benchmark for traffic accidents detection from video surveillance
✨✨Latest Advances on Multimodal Large Language Models
Project Page for "LISA: Reasoning Segmentation via Large Language Model"
GLM-4 series: Open Multilingual Multimodal Chat LMs | 开源多语言多模态对话模型
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
This repo contains the code for our paper Towards Open-Ended Visual Recognition with Large Language Model
The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
Open-Sora: Democratizing Efficient Video Production for All
This project aim to reproduce Sora (Open AI T2V model), we wish the open source community contribute to this project.
A curated list of recent diffusion models for video generation, editing, and various other applications.
The repo for "Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image" and "Metric3Dv2: A Versatile Monocular Geometric Foundation Model..."
A curated publication list on open vocabulary semantic segmentation and related area (e.g. zero-shot semantic segmentation) resources..
we want to create a repo to illustrate usage of transformers in chinese
EntitySeg Toolbox: Towards Open-World and High-Quality Image Segmentation
PytorchAutoDrive: Segmentation models (ERFNet, ENet, DeepLab, FCN...) and Lane detection models (SCNN, RESA, LSTR, LaneATT, BézierLaneNet...) based on PyTorch with fast training, visualization, ben…
❄️ ChatGPT Desktop Application (Mac, Windows and Linux)
[CVPR 2023] Official repository of Generative Semantic Segmentation
Official implementation of the CVPR 2022 paper "DETReg: Unsupervised Pretraining with Region Priors for Object Detection".