Projects based on SigLIP (Zhai et. al, 2023) and Hugging Face transformers integration 🤗
-
Updated
Feb 21, 2025 - Jupyter Notebook
Projects based on SigLIP (Zhai et. al, 2023) and Hugging Face transformers integration 🤗
Inference and fine-tuning examples for vision models from 🤗 Transformers
[ICCVW 2025] LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
[CVPR'25-Demo] Official repository of "TryOffDiff: Virtual-Try-Off via High-Fidelity Garment Reconstruction using Diffusion Models".
Spotlight-style local search for everything you saved and forgot: GitHub stars, local files, images, and bookmarks. Privacy-first, no full-disk scanning, fully on-device.
[NeurIPS 2024] AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
Local-first cinematic still archive for filmmakers. Search your frames by look, mood, technique or director with SigLIP semantic search; ingest video and URLs; build moodboards. Free desktop app for Win/Mac/Linux, Pro unlocks studio tools. Your library stays on your disk.
[ICLR 2026] The implementation of the paper Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
本项目以应用为主出发,结合了从基础的机器学习、深度学习到目标检测以及目前最新的大模型,采用目前成熟的 第三方库、开源预训练模型以及相关论文的最新技术,目的是记录学习的过程同时也进行分享以供更多人可以直接进行使用。
[ICLR 2025] - Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
Fine-Tuning SigLIP 2 for Single/Multi-Label Image Classification. Image classification vision-language encoder model fine-tuned for Image Classification Tasks
Low-latency ONNX and TensorRT based zero-shot classification and detection with contrastive language-image pre-training based prompts
Download flickr8k, flickr30k image caption datasets
Official PyTorch implementation of the WACV 2025 Oral paper "Composed Image Retrieval for Training-FREE DOMain Conversion".
MODA: open fashion retrieval benchmark and models by Hopit AI. MODA (203M, open source), MODA Pro Lite (213M, open weights), MODA Pro (hosted). Full-corpus benchmarks vs FashionSigLIP, SigLIP-SO400M and ZooClaw — one harness, losses shown. #1 open model on LookBench.
A minimal, but effective implementation of CLIP (Contrastive Language-Image Pretraining) in PyTorch
Code for Post-hoc Probabilistic Vision-Language Models
Chitrarth: Bridging Vision and Language for a Billion People
Meme search and discovery engine using CLIP and BLIP
High quality data curation for training AI models for robotics
To associate your repository with the siglip topic, visit your repo's landing page and select "manage topics."