[ICCVW 2025] LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
-
Updated
Aug 8, 2025 - Python
[ICCVW 2025] LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
[CVPR'25-Demo] Official repository of "TryOffDiff: Virtual-Try-Off via High-Fidelity Garment Reconstruction using Diffusion Models".
[NeurIPS 2024] AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
[ICLR 2026] The implementation of the paper Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
[ICLR 2025] - Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
Official PyTorch implementation of the WACV 2025 Oral paper "Composed Image Retrieval for Training-FREE DOMain Conversion".
MODA: open fashion retrieval benchmark and models by Hopit AI. MODA (203M, open source), MODA Pro Lite (213M, open weights), MODA Pro (hosted). Full-corpus benchmarks vs FashionSigLIP, SigLIP-SO400M and ZooClaw — one harness, losses shown. #1 open model on LookBench.
Code for Post-hoc Probabilistic Vision-Language Models
Chitrarth: Bridging Vision and Language for a Billion People
Meme search and discovery engine using CLIP and BLIP
An AI powered Video Serach Engine with google's SigLIP and Qdrant. It allows to search objects or key moments in videos just using natural language.
CLIP & SigLIP model training from scratch
256M-param document VLM: SigLIP + SmolLM2. LoRA fine-tuning, ONNX export, runs on CPU. Apache 2.0.
Open reproducible benchmarks for food-image recognition models and APIs.
Generalized Referring Expression Segmentation on Aerial Photos with Aerial-D, a 37,288-image dataset with 1.52M referring expressions covering instances, groups, and semantic regions across 21 categories.
PIN Architecture: a six-layer forensic system for detecting AI-generated and manipulated imagery. Fifteen independent analysis pins run concurrently: C2PA provenance, compression forensics, CLIP/SigLIP2/frequency detectors, Grad-CAM explainability, LLM adjudication and a calibrated XGBoost ensemble. Signed provenance outranks statistical inference.
Visual Embedding Reduction and Space Exploration — Clustering-guided Insights for Training Data Enhancement in V-rDu
Este proyecto presenta una solución de Computer Vision para la detección y clasificación de objetos en imágenes, las cuales son extraídas como frames de vídeos. Utiliza el modelo FastSAM para la detección de objetos, y para la clasificación, emplea embeddings que pueden ser generados mediante dos modelos distintos: CLIP o SigLIP.
Identify mountain peaks in your photos using AI—zero-shot retrieval, landmark re-ranking, and geospatial priors.
A custom Vision-Language Model (VLM) built from scratch, using SigLip for contrastive learning and a ViT-based encoder to generate meaningful image captions and semantic descriptions.
To associate your repository with the siglip topic, visit your repo's landing page and select "manage topics."