Stars
Hundreds of models & providers. One command to find what runs on your hardware.
Official implementaiton of RefAM: Attention Magnets for Zero-Shot Referral Segmentaiton
Transform arXiv papers into a single LaTeX source that can be used as a prompt for asking LLMs questions about the paper.
This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!
VisualOverload (CVPR 2026) is a VQA benchmark for image understanding in dense, high-resolution scenes.
The first Large Audio Language Model that enables native in-depth thinking, which is trained on large-scale audio Chain-of-Thought data.
This repo holds the implementation of PAVE: Patching and Adapting Video Large Language Models (CVPR2025)
[CVPR24] Official Implementation of GEM (Grounding Everything Module)
Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".
Generic PyTorch dataset implementation to load and augment VIDEOS for deep learning training loops.
Original PyTorch implementation of the code for the paper "Straight to the Point: Fast-forwarding Videos via Reinforcement Learning Using Textual Data" at the IEEE/CVF Conference on Computer Vision…
Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch