-
ShanghaiTech University
- Shanghai
- https://scholar.google.com/citations?user=j_8OPwwAAAAJ&hl=en
Stars
[ICML 2026 Oral] Agent-native Mid-training for Software Engineering
CVPR 2026 Accepted Paper WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
[ICLR2026] The official repository for the CodeGym project: "Generalizable End-to-End Tool-Use RL with Synthetic CodeGym"
ICLR 2026 Accepted Paper Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
An Open-Source Asynchronous Coding Agent
Multimodal Retrieval-augmented Generation Framework Built by Tongyi Lab, Alibaba Group.
Official Implementation for TMLR25 DA-DPO: Cost-efficient Difficulty-aware Preference Optimization for Reducing MLLM Hallucinations
Youtu-Tip: Tap for Intelligence, Keep on Device.
[ICML 2026] Official resources of "Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning".
[EMNLP 2025] Official implementation for paper "MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval"
Official Repository of MMLONGBENCH-DOC: Benchmarking Long-context Document Understanding with Visualizations
NeurIPS 2025 Accepted Paper NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
Official Code for "Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search"
SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
[NeurIPS 2025] NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
✨✨ [ICLR 2026] R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
[CVPR 2025 (Oral)] Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
The Next Step Forward in Multimodal LLM Alignment
[CVPR 2025] Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
MME-CoT: Benchmarking Chain-of-Thought in LMMs for Reasoning Quality, Robustness, and Efficiency
[ICLR2026] This is the first paper to explore how to effectively use R1-like RL for MLLMs and introduce Vision-R1, a reasoning MLLM that leverages cold-start initialization and RL training to incen…
Repo for Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent