Skip to content
View rongkunxue's full-sized avatar
🎯
Focusing
🎯
Focusing
  • China, Shanghai
  • 22:17 (UTC +08:00)

Highlights

  • Pro

Block or report rongkunxue

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
rongkunxue/README.md

Rongkun Xue

Agent Post-training · Reinforcement Learning · Foundation Model Reasoning

Email

About

I currently work on agent post-training at Alibaba Cloud. Previously, I worked with ByteDance Seed and Shanghai AI Laboratory. I received both my bachelor's and master's degrees in Automation from Xi'an Jiaotong University.

My research focuses on reinforcement learning for foundation models, agent post-training, reasoning and tool use, generative policies, and multimodal intelligence.

Selected Publications

2026

Yufei Zhan, Ziheng Wu, Yousong Zhu, Rongkun Xue, Guanghao Zhou, Ruipu Luo, Zhenghao Chen, Can Zhang, Yifan Li, Zhentao He, Zheming Yang, Ming Tang, Minghui Qiu, Jinqiao Wang
CVPR 2026 · Paper · arXiv

Introduces cue-guided visual rethinking to help multimodal language models revise misleading visual interpretations during complex reasoning.

2025

Qingbin Li, Rongkun Xue, Jie Wang, Ming Zhou, Zhi Li, Xiaofeng Ji, Yongqi Wang, Miao Liu, Zheming Yang, Minghui Qiu, Jing Yang
arXiv 2025 · Paper · Code

Branches trajectories at high-entropy critical tokens to sustain exploration and mitigate entropy collapse in reinforcement learning with verifiable rewards.

Rongkun Xue, Jinouwen Zhang, Yazhe Niu, Dazhong Shen, Bingqi Ma, Yu Liu, Jing Yang
ICCV 2025 · Paper · Project · Code

Reuses pretrained continuous generative models as general-purpose feature extractors by reversing their generation process.

Rongkun Xue, Yazhe Niu, Shuai Hu, Zixin Yin, Yongqiang Yao, Jing Yang
ICML 2025 Tokenshop · Paper · OpenReview · Code

Learns highly compressed discrete speech representations designed for efficient spoken language modeling.

2024

Jinouwen Zhang, Rongkun Xue, Yazhe Niu, Yun Chen, Jing Yang, Hongsheng Li, Yu Liu
arXiv 2024 · Paper · OpenReview · Code

Unifies and simplifies generative-policy optimization through GMPO and GMPG, with a standardized experimental framework for offline reinforcement learning.

Rongkun Xue, Jing Yang, Yuyang Jiang, Yiming Feng, Zi Yang
IEEE Intelligent Vehicles Symposium (IV) 2024 · Paper · arXiv

Combines motion-dynamic RRT, fluid-field modeling, and PPO for real-time long-range terrain-following and terrain-avoidance planning.

Experience

  • Alibaba Cloud — Agent Post-training
  • ByteDance Seed — Fast and Slow Reasoning for Foundation Models
  • Shanghai AI Laboratory — Multimodal Foundation Model Training

Contact

For research collaboration, feel free to reach me at rongkunxue@outlook.com.

Pinned Loading

  1. XJTU-Automation-share XJTU-Automation-share Public

    本项目为西安交通大学自动化专业课程资料。我是西安交通大学自动化专业19级学生,本项目包含我从大一到大四的所搜集整理的课程资料以及实验资料,包括但不限于往年题,实验代码以及实验报告,复习提纲,课后习题答案等。希望该项目在学业上对学弟学妹有所帮助,记得留下star.

    C 116 8

  2. opendilab/DI-engine opendilab/DI-engine Public

    OpenDILab Decision AI Engine. The Most Comprehensive Reinforcement Learning Framework B.P.

    Python 3.6k 437

  3. opendilab/PRG opendilab/PRG Public

    [ICCV 2025] Pretrained Reversible Generation as Unsupervised Visual Representation Learning

    Python 47 1

  4. opendilab/awesome-ui-agents opendilab/awesome-ui-agents Public

    A curated list of of awesome UI agents resources, encompassing Web, App, OS, and beyond (continually updated)

    314 36

  5. opendilab/HH-Codec opendilab/HH-Codec Public

    [ICML 2025 Tokenization Workshop] HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling

    Python 106 4