I am an AI researcher and engineer at the Beijing Academy of Artificial Intelligence (BAAI), with previous experience at ByteDance and Meituan. My work has evolved from computer vision and OCR to large language models, with a current focus on domain adaptation, post-training, agentic systems, retrieval-augmented generation, and multimodal reasoning.
我目前在北京智源人工智能研究院从事 AI 研究与工程工作,曾就职于字节跳动和美团。我的研究关注如何让大模型真正进入专业领域:从高质量数据、监督微调和强化学习,到 Agent、RAG 与多模态推理,并尽可能将论文对应的代码、模型和数据开放出来。
- Domain LLMs & post-training: data construction, supervised fine-tuning, reinforcement learning, knowledge adaptation, and capability retention
- AI agents & retrieval: academic search, scientific survey generation, agentic RAG, and long-horizon knowledge workflows
- Multimodal reasoning: visual-language models, technical drawing understanding, cross-chart reasoning, and OCR
- Open data & models: multilingual industry corpora, instruction datasets, domain models, benchmarks, and reproducible evaluation
Multilingual, multi-industry pre-training corpora with accompanying data-rating and classification models
A 2.7M-sample multilingual, multi-industry instruction collection with domain-adapted models
Open pre-training corpora spanning finance, medicine, law, education, technology, and other industries
MechVL, CareBot, industry-specific LLMs, MechVQA, SPARBench, and related research releases
An end-to-end natural-scene Chinese text detection and recognition pipeline built with CTPN, CRNN, and CTC
Image-to-LaTeX recognition for printed and handwritten mathematical formulas
A Deep Knowledge Tracing implementation for adaptive learning
Research discussions and open-source collaboration are welcome.