Skip to content
View edsonroteia's full-sized avatar

Block or report edsonroteia

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Hundreds of models & providers. One command to find what runs on your hardware.

Rust 37,044 2,361 Updated Sep 21, 2026

Official implementaiton of RefAM: Attention Magnets for Zero-Shot Referral Segmentaiton

Jupyter Notebook 17 1 Updated Feb 6, 2026

Transform arXiv papers into a single LaTeX source that can be used as a prompt for asking LLMs questions about the paper.

Python 169 10 Updated Aug 24, 2026

This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!

1,443 71 Updated Aug 2, 2026

VisualOverload (CVPR 2026) is a VQA benchmark for image understanding in dense, high-resolution scenes.

Python 18 2 Updated May 31, 2026

The first Large Audio Language Model that enables native in-depth thinking, which is trained on large-scale audio Chain-of-Thought data.

Python 299 24 Updated May 15, 2025

This repo holds the implementation of PAVE: Patching and Adapting Video Large Language Models (CVPR2025)

Python 28 3 Updated Sep 6, 2025

[CVPR24] Official Implementation of GEM (Grounding Everything Module)

Python 140 7 Updated Apr 10, 2025
Python 5 Updated Oct 7, 2024

Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".

Python 479 41 Updated Apr 24, 2024

CLIP-It! Language-Guided Video Summarization

75 2 Updated Jun 21, 2021

Generic PyTorch dataset implementation to load and augment VIDEOS for deep learning training loops.

Python 472 45 Updated Jan 18, 2023

Original PyTorch implementation of the code for the paper "Straight to the Point: Fast-forwarding Videos via Reinforcement Learning Using Textual Data" at the IEEE/CVF Conference on Computer Vision…

Python 8 2 Updated Mar 26, 2022

Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch

Python 25,522 3,495 Updated Sep 20, 2026