Highlights
- Pro
Stars
The paper list of the survey "The Past Frames the Future: Memory Mechanisms for Autoregressive Video Generation"
Official code of Motus: A Unified Latent Action World Model
EvoPolicyGym is infrastructure for evaluating coding agents and generating training experience through Autonomous Policy Evolution.
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
[ECAI2025] The Official Implementation for "Region-aware Compositional Context Prompting for Zero-Shot Anomaly Detection".
[ECCV2026] The Official Implementation for ''CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection''
[NeurIPS 2025] Official codebase for T2DA: Offline Meta-RL from Natural Language Supervision
[arxiv 2606] Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning
Beyond SFT-to-RL: Pre-alignment via Black-BoxOn-Policy Distillation for Multimodal RL
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents
Official repository for SceneCode, a framework that turns natural language prompts into executable, editable indoor world programs with articulated objects.
Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Keep tabs on your tabs. Turn your "New tabs" page into a mission control, so you can close them easily. Built for people who open too many tabs and never close them.
OpenClaw-RL: Train any agent simply by talking
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
A unified framework for vision-language environments with Gymnasium-compatible interface
We propose Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale community signals as supervision, and formulate scientific taste learning as a preference…
✨✨[ICML 2026] Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
[ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process
[KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML 2026]
[ICLR 2026] The official repository for the paper "AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning".
✨✨ [ICLR 2026] MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
Advances and Frontiers of LLM-based Issue Resolution in Software Engineering A Comprehensive Survey
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
slime is an LLM post-training framework for RL Scaling.