Stars
Domain-agnostic multi-agent software design and evolution harness
Cosmos-Reason2 models understand the physical common sense and generate appropriate embodied decisions in natural language through long chain-of-thought reasoning processes.
This repository contains the code for the paper - "Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision" (CVPR 2026)
EfficientSAM3 compresses SAM3 into lightweight, edge-friendly models via progressive knowledge distillation for fast promptable concept segmentation and tracking.
Stereo4D dataset and processing code
Unofficial DynaDUSt3R reimplementation trained on Stereo4D (research only).
NextFlow🚀: Unified Sequential Modeling Activates Multimodal Understanding and Generation
[ICLR 2026] Training Visual Reasoners with Multimodal Verifiers
State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!
[arXiv 2023] Set-of-Mark Prompting for GPT-4V and LMMs
[NeurIPS 2023] Official implementation of the paper "Segment Everything Everywhere All at Once"
Official repository of paper "Subobject-level Image Tokenization" (ICML-25)
Official inference framework for 1-bit LLMs
Repair malformed JSON from LLMs, APIs, logs, and user input in Python.
Everything about the SmolLM and SmolVLM family of models
A suite of image and video neural tokenizers
Hello AI World guide to deploying deep-learning inference networks and deep vision primitives with TensorRT and NVIDIA Jetson.
Densely Captioned Images (DCI) dataset repository.
Grounded SAM 2: Ground and Track Anything in Videos with Grounding DINO, Florence-2 and SAM 2
Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
Caption-Anything is a versatile tool combining image segmentation, visual captioning, and ChatGPT, generating tailored captions with diverse controls for user preferences. https://huggingface.co/sp…
Data release for the ImageInWords (IIW) paper.
[IJCV 2026] Multimodal Referring Segmentation
Project Page for "LISA: Reasoning Segmentation via Large Language Model"
Reference PyTorch implementation and models for DINOv3