Lists (10)
Sort Name ascending (A-Z)
Stars
Vision-OPD is a regional-to-global on-policy self-distillation framework that transfers a model's own privileged crop-conditioned perception to its full-image policy, enabling fine-grained visual u…
This repo implements the unification of diffusion model and recified flow through variational bridge.
World model reasoning RL for multi-turn VLM agents
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
[ICML 2026] XSkill: Continual Learning from Experience and Skills in Multimodal Agents
A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
[CVPR2026] Official codebase for the paper "Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space"
Pretraining and inference code for a large-scale depth-recurrent language model
PyTorch-based open-source code for paper "SOD: Step-wise On-policy Distillation for Small Language Model Agents"
On Policy Distillation Build on top of Verl
Repository hosting code for "Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations" (https://arxiv.org/abs/2402.17152).
We propose the first Tri-party LLM-agent Recommendation framework (TriRec) that explicitly coordinates user utility, item exposure, and platform-level fairness.
[ACL 2025] iAgent: LLM Agent as a Shield between User and Recommender Systems
2026TAAC腾讯广告算法大赛-KDDCUP方案,best score:0.832321,rank:51
LC1332 / TAAC_2026-ref1
Forked from Puiching-Memory/TAAC_2026参考 [参赛队伍] TAAC 2026 腾讯广告算法大赛 X KDD 2026
Fixing GRPO training collapse in long-horizon multi-tool agents. A lightweight PRM-Lite + LATA joint approach achieves +37% over vanilla GRPO on τ-bench airline (50-task, multi-turn).
[ACM-RecSys26] Baseline code for music-crs challenge