Skip to content
View huzican's full-sized avatar
  • Nanjing University
  • Nanjing

Highlights

  • Pro

Block or report huzican

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

The paper list of the survey "The Past Frames the Future: Memory Mechanisms for Autoregressive Video Generation"

38 1 Updated Sep 24, 2026
Python 24 2 Updated Sep 17, 2026

Official code of Motus: A Unified Latent Action World Model

Python 1,293 75 Updated Jan 5, 2026
Python 34 2 Updated Aug 14, 2026

EvoPolicyGym is infrastructure for evaluating coding agents and generating training experience through Autonomous Policy Evolution.

Python 176 10 Updated Sep 18, 2026

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

Python 22 1 Updated Jul 8, 2026

[ECAI2025] The Official Implementation for "Region-aware Compositional Context Prompting for Zero-Shot Anomaly Detection".

3 Updated Oct 13, 2025

[ECCV2026] The Official Implementation for ''CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection''

Python 6 Updated Jul 3, 2026

[NeurIPS 2025] Official codebase for T2DA: Offline Meta-RL from Natural Language Supervision

Python 16 1 Updated Jun 1, 2025

[arxiv 2606] Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning

Python 61 1 Updated Jul 3, 2026

Beyond SFT-to-RL: Pre-alignment via Black-BoxOn-Policy Distillation for Multimodal RL

Python 101 2 Updated May 6, 2026

A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models

843 35 Updated Aug 30, 2026

A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents

Python 21 2 Updated Jun 9, 2026

Official repository for SceneCode, a framework that turns natural language prompts into executable, editable indoor world programs with articulated objects.

Python 61 3 Updated Jun 18, 2026

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

Python 57 1 Updated Jun 10, 2026

Keep tabs on your tabs. Turn your "New tabs" page into a mission control, so you can close them easily. Built for people who open too many tabs and never close them.

JavaScript 1,800 514 Updated Apr 14, 2026

OpenClaw-RL: Train any agent simply by talking

Python 5,698 609 Updated May 23, 2026

verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"

Python 2,335 225 Updated Jun 9, 2026

A unified framework for vision-language environments with Gymnasium-compatible interface

Python 38 1 Updated Mar 17, 2026

We propose Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale community signals as supervision, and formulate scientific taste learning as a preference…

432 11 Updated Sep 22, 2026

✨✨[ICML 2026] Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

Python 155 7 Updated Mar 12, 2026

[ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process

Python 1,073 45 Updated Feb 10, 2026

[KernelGYM & Dr. Kernel] A distributed GPU environment and a collection of RL training methods to support RL for Kernel Generations [ICML 2026]

Python 210 33 Updated Mar 29, 2026

[ICLR 2026] The official repository for the paper "AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning".

Jupyter Notebook 84 5 Updated Aug 11, 2026

✨✨ [ICLR 2026] MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models

Python 43 4 Updated Apr 10, 2025

VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model

Python 2,313 208 Updated Mar 19, 2026

Advances and Frontiers of LLM-based Issue Resolution in Software Engineering A Comprehensive Survey

Python 88 4 Updated Sep 23, 2026

A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.

Python 825 37 Updated Aug 6, 2026

slime is an LLM post-training framework for RL Scaling.

Python 8,526 1,268 Updated Sep 24, 2026
Next