Stars
[ICLR 2025 Spotlight] OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
[AAAI 2024] EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question Answering
Official code repository for NeurIPS 2022 paper "SatMAE: Pretraining Transformers for Temporal and Multi-Spectral Satellite Imagery"
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
An open source implementation of CLIP.
[CVPR 2023] Official Implementation of X-Decoder for generalized decoding for pixel, image and language
SAM (Segment Anything Model) for generating rotated bounding boxes with MMRotate, which is a comparison method of H2RBox-v2.
The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery, ISPRS. Also, including other vision transformers and CNNs for satellite, aerial image …
A playbook for systematically maximizing the performance of deep learning models.
chongzhou96 / MaskCLIP
Forked from open-mmlab/mmsegmentationOfficial PyTorch implementation of "Extract Free Dense Labels from CLIP" (ECCV 22 Oral)
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
Awesome Remote Sensing Toolkit based on PaddlePaddle.
MDPI Journal: Remote Sensing Track 2023
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Attention Branch Network (CIFAR100, ImageNet models)
The implementations of "Unsupervised Domain Adaptation for Remote Sensing Semantic Segmentation with Transformer"
The official repo for [TGRS'22] "Advancing Plain Vision Transformer Towards Remote Sensing Foundation Model"
Implementation of Denoising Diffusion Probabilistic Model in Pytorch
Code of SwinSTFM: Remote Sensing Spatiotemporal Fusion using Swin Transformer
Building Extraction from remote sensing image using Vision Transformer, IEEE Transactions on Geoscience and Remote Sensing, 2022
[IGARSS 2022] CapFormer: Pure transformer for remote sensing image caption
Danfeng Hong, Zhu Han, Jing Yao, Lianru Gao, Bing Zhang, Antonio Plaza, Jocelyn Chanussot. Spectralformer: Rethinking hyperspectral image classification with transformers, IEEE Transactions on Geos…