Biography

I am an Adjunct Ph.D. Supervisor at the School of AI, Shanghai Jiao Tong University. From August 2022 to July 2026, I served as a Research Scientist at Shanghai AI Lab. My work focuses on the R&D of General-Purpose Foundation Models, with extensive experience in Multimodal Understanding, Document Parsing, and Data-Centric AI.

I believe that true innovation stems from deep diving, and more importantly, from the relentless refinement and bold reshaping of existing technologies. Refusing to settle for the status quo, my goal is to deliver research that is not only scientifically rigorous but also practically transformative—tackling the “hard problems” others cannot, to provide unique solutions for the industry’s most critical challenges.

Guided by this philosophy, I led the R&D of MinerU, an open-source toolkit for high-quality document parsing. The project has garnered over 50k GitHub stars in just 1.5 years, frequently topping GitHub Trending charts. It is widely adopted by both academia and industry, serving as a mainstream solution for enterprises and developers building high-quality LLM and RAG corpora. Additionally, I have published over 50 papers in top-tier conferences such as CVPR, ICCV, NeurIPS, and ICLR, with over 6,000 Google Scholar citations.


我是上海交通大学人工智能学院兼职博士生导师,入选上海市东方英才拔尖人才项目。2022年8月至2026年7月,我曾在上海人工智能实验室(Shanghai AI Lab)担任青年科学家。我专注于通用基础大模型的研发;在多模态理解智能文档解析以数据为中心的人工智能(Data-Centric AI)等方向有长期积累。

我相信真正的创新源于深耕,更源于对现有技术的极致打磨与勇敢重塑。我不囿于既有的技术边界,而是致力于产出既具备科学严谨性,又具有变革意义的研究——通过攻克那些别人做不到的难题,为行业最关键的挑战提供独一无二的解决方案。秉持这一理念,我主导研发了开源文档解析工具 MinerU。该项目在一年半内斩获 50k+ GitHub Stars,多次登顶 GitHub Trending 全球榜单,不仅在学术界广受好评,更被产业界广泛采用,成为众多企业与开发者构建高质量大模型语料及 RAG 语料库的主流选择。我在AI相关领域发表高水平论文50余篇,包含CVPR, ICCV, NeurIPS, ICLR 等顶级会议,谷歌学术引用超 6000 次。

📧 Email: ictwangbin@gmail.com

🔥 News

2026:

  • 2026.04:  🔥🔥🔥 MinerU2.5-Pro is released! Pushing the limits of data-centric document parsing — achieves 95.69 on OmniDocBench v1.6, surpassing models with 200× more parameters (Gemini 3 Pro, GPT-5.2, Qwen3-VL-235B). [Paper] [GitHub]
  • 2026.03:  🔥🔥🔥 MinerU-Diffusion is released! Rethinking Document OCR as inverse rendering via diffusion decoding — up to 3.26× faster than MinerU2.5 with near-lossless accuracy. [Paper] [GitHub]
  • 2026.02:  🎉🎉 UniMERNet, TRivia, OmniDocLayout and ARM-Thinker are accepted by CVPR 2026.
  • 2026.02:  🎉🎉 MoDora is accepted by SIGMOD 2026.

2025:

  • 2025.09:  🎉🎉 MinerU 2.5 is released! A 1.2B vision-language model for document parsing. [Tech Report] [Hugging Face Model] [GitHub]
    • SOTA Performance: Surpasses general models (Gemini 2.5-Pro, GPT-4o, etc.) and specialized tools (MonkeyOCR, PP-StructureV3).
    • High Efficiency: Achieves top accuracy with significantly greater speed than large-model solutions.
  • 2025.06:  🎉🎉 OHR, LEGION and Chimera are accepted by ICCV 2025.
  • 2025.02:  🎉🎉 OmniDocBench and CDM are accepted by CVPR 2025.
  • 2025.01:  🎉🎉 GeoX and OmniCorpus are accepted by ICLR 2025.

2024:

  • 2024.09:  🎉🎉 InternLM-XComposer2-4KHD is accepted by NeurIPS 2024.
  • 2024.07:  🔥🔥🔥 has received 3500+ GitHub stars within one month.
  • 2024.07:  🔥🔥🔥 has received 4200+ GitHub stars and ranked #1 on the GitHub Trending list.
  • 2024.07:  🎉🎉 CLIP-Parrot-Bias is accepted by ECCV 2024 (Oral).
  • 2024.02:  🎉🎉 OPERA is accepted by CVPR 2024.
  • 2023.12:  🎉🎉 VIGC is accepted by AAAI 2024.
  • 2023.12:  🎉🎉 One paper is accepted by IJAEOG 2024.
  • 2023.08:  🎉🎉 DropQueries is accepted by TMM 2023.
  • 2023.08:  🎉🎉 V3Det is accepted by ICCV 2023 (Oral).

🚀 Project

New 🔥
sym

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale (Project Lead) | Models(Hugging Face) | Models(ModelScope) | Github

  • Current SOTA on OmniDocBench v1.6, scoring 95.69 overall — surpassing models with 200× more parameters (Gemini 3 Pro, GPT-5.2, Qwen3-VL-235B).
  • Key insight: data quality is the real ceiling, not model size. We maintain the same 1.2B model and unlock its full potential through data engineering: diversity-aware sampling (65.5M samples), cross-model verification for reliable annotations, and iterative refinement for hard samples.
  • Three-stage progressive training: large-scale pre-training → hard sample fine-tuning → GRPO alignment.
New 🔥
sym

MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding (Project Lead) | Github

  • A novel paradigm that reframes document OCR as inverse rendering via diffusion decoding, replacing autoregressive generation with block-wise parallel processing.
  • Achieves up to 3.26× faster throughput compared to MinerU2.5, with 2.12× speedup at 99.9% relative accuracy.
Github Repo
sym

PDF-Extract-Kit: A Comprehensive Toolkit for High-Quality PDF Content Extraction (Project Lead)

Models(Hugging Face) | Models(ModelScope) | Github

📝 Publications

ECCV 2024
sym

Parrot Captions Teach CLIP to Spot Text

Yiqi Lin*, Conghui He*, Alex Jinpeng Wang*, Bin Wang*, Weijia Li, Mike Zheng Shou

ECCV 2024 Oral, | Project | Github

AAAI 2024
sym

VIGC: Visual Instruction Generation and Correction

Bin Wang, Fan Wu, Xiao Han, Jiahui Peng, Huaping Zhong, Pan Zhang, Xiaoyi Dong, Weijia Li, Wei Li, Jiaqi Wang, Conghui He

AAAI 2024, | Project | Github

TMM 2023
sym

DropQueries: A Simple Way to Discover Comprehensive Segment Representations

Haojie Ding, Bin Wang, Guoliang Kang, Weijia Li, Conghui He, Yao Zhao, and Yunchao Wei

TMM 2023

ICCV 2023
sym

V3Det: Vast Vocabulary Visual Detection Dataset

Jiaqi Wang, Pan Zhang, Tao Chu, Yuhang Cao, Yujie Zhou, Tong Wu, Bin Wang, Conghui He, and Dahua Lin

ICCV 2023 Oral, | Project | Github

IJCAI 2019
sym

Boundary perception guidance: A scribble-supervised semantic segmentation approach

Bin Wang, Guojun Qi, Sheng Tang, Tianzhu Zhang, Yunchao Wei, Linghui Li, and Yongdong Zhang

IJCAI 2019

MICCAI 2019
sym

Spatiotemporal Breast Mass Detection Network(MD-Net) in 4D DCE-MRI Images

Lixi Deng, Sheng Tang, Huazhu Fu, Bin Wang, and Yongdong Zhang

MICCAI 2019

MICCAI 2018
sym

Automated pulmonary nodule detection: High sensitivity with few candidates

Bin Wang, Guojun Qi, Sheng Tang, Liheng Zhang, Lixi Deng, and Yongdong Zhang

MICCAI 2018

🎖 Honors and Awards

  • 2020.06, Zhu Li Yuehua Outstanding Ph.D. student Scholarship, Chinese Academy of Sciences (CAS).
  • 2016.09, Won 3rd place in the ILSVRC 2016 VID task (Object Detection from Video).

🏢 Work Experience

  • 2022.08 - 2026.07, Research Scientist, Shanghai AI Lab, Shanghai, China.
  • 2020.07 - 2022.08, Researcher, SenseTime, Shenzhen, China.

📖 Education

  • 2015.09 - 2020.06, Ph.D., University of Chinese Academy of Sciences, Beijing, China.
  • 2013.09 - 2015.06, M.S., Beijing Jiaotong University, Beijing, China.