Stars
🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), ga…
[ICLR 2025] Official code repository for "TULIP: Token-length Upgraded CLIP"
Sync and Async Object-oriented Python SDK for the 3x-ui API.
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
[CVPR 2025] DEIM: DETR with Improved Matching for Fast Convergence
Next generation frontend tooling. It's fast!
[CVPR2023] Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning (https://arxiv.org/abs/2212.04500)
[RAL 2022] Topo-boundary: A Benchmark Dataset on Topological Road-boundary Detection Using Aerial Images for Autonomous Driving
[CVPR 2025] Multiple Object Tracking as ID Prediction
Official implementation of "Local All-Pair Correspondence for Point Tracking" (ECCV 2024)
Associate Everything Detected: Facilitating Tracking-by-Detection to the Unknown
[CVPR2023] MOTRv2: Bootstrapping End-to-End Multi-Object Tracking by Pretrained Object Detectors
Official repository of "SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory"
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
[CVPR 2023] VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
Active Learning in the era of Foundation Models
InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation 🔥
Official Implementation for "Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models" (SIGGRAPH 2023)
Quality-Aware Image-Text Alignment for Opinion-Unaware Image Quality Assessment
Measures and metrics for image2image tasks. PyTorch.
Official PyTorch implementation of "Scaling Up Personalized Image Aesthetic Assessment via Task Vector Customization" (ECCV 2024)
CLIP+MLP Aesthetic Score Predictor
[AAAI 2023] Exploring CLIP for Assessing the Look and Feel of Images
PyTorch implementation of CLIP Maximum Mean Discrepancy (CMMD) for evaluating image generation models.
[NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
[NeurIPS2023] DatasetDM:Synthesizing Data with Perception Annotations Using Diffusion Models
[ICCV2023] DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models
sagieppel / fine-tune-train_segment_anything_2_in_60_lines_of_code
Forked from facebookresearch/sam2The repository provides code for training/fine tune the Meta Segment Anything Model 2 (SAM 2)