Stars
A list of resouces for multispectral pedestrian detection,including the datasets, methods, annotations and tools.
An AI prompt optimizer for writing better prompts and getting better AI results.
Expert Systems With Applications:Shooting condition insensitive unmanned aerial vehicle object detection
mujianyu / TwoStream_Yolov8
Forked from ultralytics/ultralyticsNEW - YOLOv8 🚀 in PyTorch > ONNX > OpenVINO > CoreML > TFLite
M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision (ICCV 2025)
[ACM MM 2023 Oral] Multispectral Object Detection via Cross-Modal Conflict-Aware Learning.
This is the official repository for “Drone-based RGB-Infrared Cross-Modality Vehicle Detection via Uncertainty-Aware Learning” (IEEE T-CSVT 2022).
Drone-based RGB-Infrared Cross-Modality Vehicle Detection via Uncertainty-Aware Learning
End-to-End CLIP-driven Mamba Model for Multi-modal Fusion
MP-HSIR: A Multi-Prompt Framework for Universal Hyperspectral Image Restoration
Official Repo for MICCAI 25 Paper: Geometry-Guided Local Alignment for Multi-View Visual Language Pre-Training in Mammography
Use visible and infrared images to train the network. This method is better to face the dark environment.
GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding
YOLOv11-RGBT: Towards a Comprehensive Single-Stage Multispectral Object Detection Framework(Supports RGBT detection for all YOLO series from YOLOv3 to YOLOv13, as well as RTDETR. 【Ultralytics YOLOv…
Lightweight modal-guided cross-attention fusion network for visible-infrared object detection (Pattern Recognition, 2026)
[ICML 2024] Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Reference PyTorch implementation and models for DINOv3
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
Code of Paper OmniFuse: Composite Degradation-Robust Image Fusion with Language-Driven Semantics.
Official Code of Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion (CVPR 2024)