Skip to content
View Rongfeng-Guo's full-sized avatar

Block or report Rongfeng-Guo

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Vision-OPD is a regional-to-global on-policy self-distillation framework that transfers a model's own privileged crop-conditioned perception to its full-image policy, enabling fine-grained visual u…

Python 285 11 Updated Jul 17, 2026

This repo implements the unification of diffusion model and recified flow through variational bridge.

Jupyter Notebook 1 Updated Jul 25, 2026

World model reasoning RL for multi-turn VLM agents

Python 493 60 Updated Aug 15, 2026

SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Python 944 73 Updated May 17, 2026

[ICML 2026] XSkill: Continual Learning from Experience and Skills in Multimodal Agents

Python 254 34 Updated May 13, 2026

The codebase of Cola DLM

Python 280 18 Updated Jul 20, 2026
Python 243 29 Updated Apr 23, 2024

A curated collection of papers and resources on On-Policy Distillation for Large Language Models.

Python 520 10 Updated Aug 12, 2026

[CVPR2026] Official codebase for the paper "Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space"

Python 88 5 Updated May 12, 2026
Python 30 4 Updated Apr 7, 2026

Pretraining and inference code for a large-scale depth-recurrent language model

Python 911 81 Updated Dec 29, 2025

PyTorch-based open-source code for paper "SOD: Step-wise On-policy Distillation for Small Language Model Agents"

Python 156 10 Updated May 22, 2026

On Policy Distillation Build on top of Verl

Python 95 9 Updated May 25, 2026
Python 15 11 Updated Apr 14, 2026
Python 11 4 Updated Apr 14, 2026

Repository hosting code for "Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations" (https://arxiv.org/abs/2402.17152).

Python 1,962 408 Updated Aug 13, 2026
Python 278 42 Updated Oct 26, 2025

Minimal reproduction of OneRec

Python 1,759 256 Updated May 14, 2026
Python 76 2 Updated Oct 2, 2024

We propose the first Tri-party LLM-agent Recommendation framework (TriRec) that explicitly coordinates user utility, item exposure, and platform-level fairness.

Python 6 Updated Jun 1, 2026

[ACL 2025] iAgent: LLM Agent as a Shield between User and Recommender Systems

Python 36 5 Updated May 23, 2025
Python 8 Updated Mar 10, 2026
Python 5 Updated May 30, 2026

2026TAAC腾讯广告算法大赛-KDDCUP方案,best score:0.832321,rank:51

Python 45 10 Updated Jul 30, 2026

参考 [参赛队伍] TAAC 2026 腾讯广告算法大赛 X KDD 2026

Python 5 Updated Apr 6, 2026

Fixing GRPO training collapse in long-horizon multi-tool agents. A lightweight PRM-Lite + LATA joint approach achieves +37% over vanilla GRPO on τ-bench airline (50-task, multi-turn).

Python 198 13 Updated Jun 27, 2026

[ACM-RecSys26] Baseline code for music-crs challenge

Python 18 9 Updated Jun 23, 2026
Next