Skip to content
View dusruddl2's full-sized avatar

Highlights

  • Pro

Block or report dusruddl2

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

Python 24 Updated Jun 12, 2026
Python 42 Updated Nov 8, 2024

[ICLR 2026] Official implementation of the paper "Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs"

Python 25 3 Updated Mar 3, 2026

[CVPR 2026] Official Pytorch Code for ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting

Python 11 1 Updated May 11, 2026

A paper list of Awesome Latent Space.

947 40 Updated Jul 13, 2026

[CVPR 2026] Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO

Python 119 1 Updated Feb 28, 2026

🔥🔥🔥[AAAI 2026 Oral] Official Implementation of Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding

Python 530 17 Updated Jan 20, 2026

ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…

Python 13,792 1,237 Updated Jul 22, 2026

Code for "Agentic Very Long Video Understanding" (EGAgent) [ACL 2026 Main]

Python 49 5 Updated Jul 1, 2026

AI agents running research on single-GPU nanochat training automatically

Python 91,942 13,142 Updated Mar 26, 2026

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

JavaScript 232,669 35,464 Updated Jul 24, 2026

[NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning

Python 268 10 Updated Oct 18, 2025

Paper list for Efficient Reasoning.

900 46 Updated May 29, 2026

📖 This is a repository for organizing papers, codes, and other resources related to Latent Reasoning.

405 9 Updated Nov 5, 2025

Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Python 491 33 Updated May 20, 2026

Video-R1: Reinforcing Video Reasoning in MLLMs [🔥the first paper to explore R1 for video]

Python 882 46 Updated Dec 14, 2025

Official code for "Rethinking Chain-of-Thought Reasoning for Videos"

21 Updated Dec 14, 2025

🔥An open-source survey of the latest video reasoning tasks, paradigms, and benchmarks.

189 8 Updated Jun 14, 2026

The development and future prospects of large multimodal reasoning models.

614 22 Updated Jan 9, 2026

Collect the awesome works evolved around reasoning models like O1/R1 in visual domain

55 2 Updated Jul 21, 2025

This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!

1,435 64 Updated May 11, 2026

[TMLR 2025] Efficient Reasoning Models: A Survey

Python 315 23 Updated Jun 26, 2026

[ICML 2025] Official repository for paper "Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation"

Python 194 36 Updated Sep 23, 2025

A curated list of awesome Multimodal studies.

341 25 Updated Jul 20, 2026

[TMLR 2026] Survey: https://arxiv.org/pdf/2507.20198

371 23 Updated May 29, 2026

[ICCV 2025] Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs

Python 61 2 Updated Feb 2, 2026

[MICCAI 2025 Early Accept] PRETI: Patient-Aware Retinal Foundation Model via Metadata-Guided Representation Learning

Python 12 2 Updated Aug 1, 2025

[CVPR 2025] Official Pytorch Code for Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis

Jupyter Notebook 15 3 Updated Jun 21, 2025

A paper list of some recent works about Token Compress for Vit and VLM

944 43 Updated Jul 20, 2026
Next