Skip to content
View shipengai's full-sized avatar

Block or report shipengai

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[CVPR 2026] TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs

Python 169 13 Updated Jul 21, 2026

[ICML2026]Efficient Post-Training of MLLMs for Temporal Video Grounding via On-Policy Distillation

Python 11 Updated Jun 29, 2026

qqr is an RL training framework for open-ended agents.

Python 276 22 Updated Aug 5, 2026

Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.

Python 23,002 2,349 Updated Jul 29, 2026

https://diadem-captioner.github.io/

Python 6 Updated Jan 31, 2026

D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning

Python 15 1 Updated Aug 4, 2026

Scalable toolkit for efficient model reinforcement

Python 1,891 501 Updated Aug 9, 2026

Multimodal RL training framework for diffusion & omni models

Python 763 128 Updated Aug 8, 2026

https://avocado-captioner.github.io/

Python 38 1 Updated Oct 16, 2025

[ICML 2026] Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Python 52 1 Updated Jun 29, 2026

[ICLR 2026] An official implementation of "CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning"

Python 227 8 Updated Jun 23, 2026

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Python 905 61 Updated Jun 29, 2026
Python 13 2 Updated Mar 23, 2026

Official Repo for SvS: A Self-play with Variational Problem Synthesis strategy for RLVR training

Python 55 5 Updated Dec 13, 2025

FuseAI Project

Python 600 37 Updated Jan 25, 2025

Codebase for Merging Language Models (ICML 2024)

Python 872 52 Updated May 5, 2024

Official Implementation of OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning

Python 887 27 Updated Apr 11, 2026
Python 18 Updated Mar 16, 2026

πŸ”₯ OneThinker: All-in-one Reasoning Model for Image and Video [CVPR 2026]

Python 463 34 Updated Feb 28, 2026
Python 78 4 Updated Apr 9, 2026

[CVPR 2026] Official repo for "VideoSSR: Video Self-Supervised Reinforcement Learning"

Python 41 2 Updated Nov 11, 2025

[CVPR 2026] Boosting Reasoning in Large Multimodal Models via Activation Replay

Python 24 Updated May 7, 2026

Official code for "Rethinking Chain-of-Thought Reasoning for Videos"

21 Updated Dec 14, 2025

[CVPR2026] VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice

Python 89 5 Updated Feb 27, 2026

[CVPRF 2026] Official PyTorch code of "Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning".

Jupyter Notebook 11 Updated Aug 5, 2026

Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)

Python 726 30 Updated Sep 24, 2025

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Python 19,209 1,513 Updated Aug 8, 2026

[ICLR2026] VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Python 526 20 Updated Jul 19, 2026

GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Python 2,363 176 Updated Jul 21, 2026

RoboBrain 2.5: Advanced version of RoboBrain. Depth in Sight, Time in Mind. πŸŽ‰πŸŽ‰πŸŽ‰

Python 1,122 113 Updated Feb 28, 2026
Next