Stars
Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction
[ICLR 2026] Official implementation of the paper "Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs"
[CVPR 2026] Official Pytorch Code for ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
[CVPR 2026] Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
🔥🔥🔥[AAAI 2026 Oral] Official Implementation of Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…
Code for "Agentic Very Long Video Understanding" (EGAgent) [ACL 2026 Main]
AI agents running research on single-GPU nanochat training automatically
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
[NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning
Paper list for Efficient Reasoning.
📖 This is a repository for organizing papers, codes, and other resources related to Latent Reasoning.
Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Video-R1: Reinforcing Video Reasoning in MLLMs [🔥the first paper to explore R1 for video]
Official code for "Rethinking Chain-of-Thought Reasoning for Videos"
🔥An open-source survey of the latest video reasoning tasks, paradigms, and benchmarks.
The development and future prospects of large multimodal reasoning models.
Collect the awesome works evolved around reasoning models like O1/R1 in visual domain
This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!
[TMLR 2025] Efficient Reasoning Models: A Survey
[ICML 2025] Official repository for paper "Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation"
A curated list of awesome Multimodal studies.
[TMLR 2026] Survey: https://arxiv.org/pdf/2507.20198
[ICCV 2025] Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
[MICCAI 2025 Early Accept] PRETI: Patient-Aware Retinal Foundation Model via Metadata-Guided Representation Learning
[CVPR 2025] Official Pytorch Code for Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
A paper list of some recent works about Token Compress for Vit and VLM